Synthesia is the market leader for AI avatar video — the category of tool that turns a text script into a presenter-led video without any camera, crew, or studio. In 2026 the output quality for training videos, product explainers, and internal communications has crossed the threshold where most viewers accept it without flagging it as AI. Here is how to use it well.
What changed in 2026
- Expressive avatars V3 launched with significantly improved body language and facial expression range, reducing the "reading a script" stiffness that defined earlier versions.
- Custom avatar quality improved; a 5-minute recorded sample now produces a reliable custom avatar, down from 30+ minutes required in earlier iterations.
- Translate feature now covers 140+ languages with lip-sync; one video can become a multilingual asset library without re-scripting.
- Pricing shifted to seat-based annual plans (~$30–$90/user/month depending on tier) with video export limits at entry tiers and unlimited at enterprise.
Choosing an avatar
| Avatar type |
Setup required |
Best for |
| Stock avatar (230+) |
None — pick and use |
Training, marketing, explainers |
| Custom avatar |
5-min recording session |
Brand-consistent presenter |
| Expressive avatar |
Select in avatar panel |
Emotional content, storytelling |
| Personal avatar (your face) |
Consent form + recording |
Executive comms, authentic brand |
For most use cases, a well-chosen stock avatar is indistinguishable from a custom one at this quality level. Use a custom avatar when brand recognition or a specific "face" matters.
Scripting for AI delivery
This is the single highest-leverage skill in Synthesia. Poor scripts produce poor videos regardless of avatar quality.
Rules for AI-readable scripts:
- Write in short sentences of 12–20 words.
- Use commas and periods to control breathing pauses; add
[pause] tags for deliberate beats.
- Avoid tongue twisters, acronym clusters, and words with unusual pronunciation.
- Write what you would say, not what you would write. "We're" sounds better than "We are" in a spoken context.
- Read it aloud before submitting. If you stumble, the avatar will sound unnatural there.
Example of weak script: The integration facilitates bi-directional synchronization between enterprise CRM systems and downstream analytics infrastructure.
Improved: This integration syncs your CRM and your analytics tools automatically. Changes in either system update the other in real time.
Building a multi-scene project
For training videos and onboarding content, use the multi-scene editor:
- Create one scene per concept — aim for 45–90 seconds per scene.
- Add text overlays, screen recordings, or slides alongside the avatar.
- Use the scene dividers to add chapter titles — Synthesia generates a table of contents automatically.
- Set a consistent background across scenes for visual coherence.
- Export as a single video or as an interactive player with chapter navigation.
Multi-scene projects with chapters dramatically improve completion rates for training content.
Custom avatar setup
- Request a recording link from Synthesia (available on paid plans).
- Film yourself in a quiet, well-lit room — solid background, camera at eye level.
- Read the provided script naturally for 4–6 minutes; the system needs voice and facial data.
- Submit for processing (~24–48 hours).
- Test the avatar on short scripts before your first production video.
Common recording mistakes: poor lighting (even soft shadows cause quality issues), background noise, glasses glare, and reading too fast.
How to pick the right Synthesia workflow
- Quick explainer or update? Single scene, stock avatar, ~90-second script.
- Training module or course content? Multi-scene with chapters, screen recording integration.
- Multilingual asset? Script once, use Translate — saves full re-production for each language.
- High-stakes executive communication? Personal avatar or custom avatar, review delivery on each key phrase.
- Social video clips? Use the 9:16 aspect ratio template and keep scripts under 45 seconds.
Common mistakes
Using written language instead of spoken language. This is the most common failure. Run your script through the voice preview before committing to production.
Overloading a single scene. More than 2 minutes of continuous avatar without a visual change loses audiences. Break it up with slides, screen recordings, or scene cuts.
Skipping the subtitle review. Auto-generated subtitles are ~95% accurate. Review and correct before exporting — errors in subtitles damage credibility more than they would in spoken audio.
Recording the custom avatar in bad light. Synthesia needs clean facial data. Invest 20 minutes in proper lighting before your recording session.
What to skip
- Asking Synthesia to handle complex emotional scenes. Expressive avatars handle moderate emotion well; they break on grief, panic, or complex humor. Plan around this.
- Dense, technical scripts without simplification. Synthesia is not a reading service; it is a communication tool. If the content is dense, the script needs restructuring, not better AI.
- Exporting without preview. Always preview the full video in the player before downloading — AI-generated pauses and emphasis sometimes need manual correction with SSML tags.
FAQ
How does Synthesia compare to HeyGen in 2026?
Both are strong. Synthesia leads on enterprise features, LMS integrations, and multilingual lip-sync. HeyGen leads on ultra-realistic avatar quality for marketing videos. For training at scale, Synthesia's team features and integrations give it an edge.
Can viewers tell the video is AI-generated?
With expressive avatars and a well-written script, most viewers do not flag it unprompted. Disclosure requirements vary by context — check your regional guidelines and platform policies.
What is the minimum plan to get custom avatars?
Custom avatar creation is available on the Starter plan and above. Personal avatar (your own likeness) requires the Enterprise plan with an additional consent verification process.
How long does video export take?
Short videos (under 3 minutes) typically export in 2–5 minutes. Long multi-scene projects (20+ minutes) can take 15–30 minutes.
Where to go next
See How to use Descript in 2026, AI for podcast producers in 2026, and AI for course creators in 2026.