An explainer video needs three things most teams do not have on hand at the same time: a presenter, a recording day, and a voice that reads the script the way it was written. AI collapses those three into one pass. You write the script, pick an avatar and a voice, and get a lip-synced narrated video back in the time it takes to make coffee, with none of the usual scheduling around a camera, a studio or a talent's availability.
The classic explainer video pipeline is slow because every step waits on a person: booking a presenter, renting space, recording multiple takes, then handing the footage to an editor. A script rewrite after the shoot means reshooting. AI narration removes that dependency chain. The presenter is a model, the voice is generated from text, and a script change is a re-run rather than a re-book.
Picasso IA runs heygen/avatar-v for the presenter layer, taking a typed script plus a chosen voice and returning a lip-synced talking-head video at up to 4K, in 16:9 or 9:16. Pair it with a text-to-speech model such as elevenlabs/v3 or minimax/speech-2.8-hd for the voiceover, and the two together cover what a studio shoot used to require: a face, a voice and a script that stays in sync. New accounts get free credits, enough to test a script on more than one avatar before committing to a final cut.
The presenter choice sets the tone before a single word of the script is heard. A studio avatar in a blazer reads as corporate; a casual, seated avatar reads as a product walkthrough; no avatar at all, just narration over slides, reads as a tutorial. Match the style to where the video will actually be watched.
| Presenter style | How it is made | Fits best |
|---|
| Studio avatar | heygen/avatar-v with a formal voice, 16:9, 1080p or 4K | Product launches, investor updates |
| Casual seated avatar | heygen/avatar-v, relaxed voice, slower pace | Onboarding, internal training |
| Photo-based presenter | A real photo animated into a talking video | Founder-led updates, personal branding |
| Vertical social avatar | heygen/avatar-v, 9:16, burned-in captions | Short-form ads, social explainers |
| Faceless narration | Voiceover only, paired with slides or b-roll | Tutorials, documentation walkthroughs |
An avatar is not mandatory. Plenty of effective explainer videos never show a face at all, running a generated voiceover over screen recordings, product shots or AI b-roll instead. Choose the avatar when a presenter builds trust, and skip it when the product itself is the more convincing thing to look at.
A script written to be read silently and a script written to be spoken out loud are not the same document. Avatar-v reads exactly what you type, punctuation included, so a script that looks fine on the page can come out stiff on screen if it was written for the eye rather than the ear.
- Short sentences: a script written in short, spoken sentences avoids the run-on delivery that long written clauses produce.
- Read it aloud first: reading your own draft out loud catches the phrases that look fine on paper and sound awkward once spoken.
- One idea per line: breaking the script into single-idea lines makes it easier to trim or reorder without rewriting the whole thing.
- Spell out tricky terms: writing an acronym or a brand name the way it should sound helps the voice model land the pronunciation.
- Keep it under the limit: avatar-v accepts up to 5,000 characters per run, so a long script needs splitting into segments rather than one giant take.
Pro tip: generate the voiceover on its own first and listen to the pacing before you commit to the avatar render. A sentence that reads fine silently can land at an odd speed once spoken, and it is far cheaper to fix a script line than to redo a finished video.
This walkthrough goes from a written script to a finished narrated video, using the avatar and voice models in the catalog. Nothing here needs a camera or a studio.
- Write the script in short, spoken sentences. Keep it under 5,000 characters per segment, and read it aloud once before moving on.
- Generate the voiceover. Run the script through elevenlabs/v3 or minimax/speech-2.8-hd, pick a voice that fits the tone, and listen back for pacing.
- Generate the avatar video. Feed the script and a voice ID into heygen/avatar-v, choose an avatar, set the resolution and the aspect ratio, 16:9 for a website or 9:16 for social.
- Add supporting visuals. Generate slide graphics or product shots with a model like FLUX or nano banana to cut away to between avatar segments.
- Caption and export. Turn on burned-in captions in avatar-v, or run the finished audio through the Toolkit for a styled subtitle track, then export the file.
A single unbroken shot of a talking avatar works for a short clip, but a two-minute explainer benefits from cutting away to something else every twenty or thirty seconds, the same reason a human-hosted video reaches for a screen share or a product shot. Generate supporting stills or short clips separately and edit them in as cutaways around the avatar footage.
Captions matter for the same reason they matter on any video: a meaningful share of viewers watch muted, on a phone, in a place where sound is not an option. Avatar-v can burn captions directly into the render, or you can pull a timed transcript through automatic video subtitles and style it separately if you want more control over the look.
For teams publishing the same explainer in several markets, AI video dubbing translates the finished narration into another language and resyncs the lip movement, which is faster than writing a second script and rendering a second avatar pass from scratch. It is worth comparing the full pipeline against a single-purpose tool; the Picasso IA vs HeyGen page walks through what changes when the avatar sits inside a broader catalog instead of standing alone.
An honest page says where a tool stops being useful. For AI-narrated explainer videos, a handful of limits show up often enough to plan around.
Hand gestures stay limited. The avatar moves naturally at the shoulders and face, but it will not point at an on-screen chart or mime an action the way a live presenter can, so pair it with visual cutaways instead of relying on gesture to carry a point. Long scripts need splitting too, since avatar-v caps input at 5,000 characters per run, which means a longer video is several segments stitched together rather than one continuous take.
Pronunciation of unusual names is worth double-checking: a brand name, a person's name or a technical term can come out slightly off, so review the voiceover before the final render rather than after. The emotional range also stays even rather than dramatic, composed and professional in a way that suits a product explainer but underdelivers a script that needs real emphasis or surprise. And the model reads a script, it does not write one. It narrates exactly what you give it, so a weak or unclear script produces a clean but unconvincing video regardless of how good the avatar looks.
None of this rules the tool out for explainer work. It rules out treating the avatar as a replacement for a well-written script, which was always the harder half of the job.
What do I need to make an AI explainer video?
A written script and an idea of the tone you want, corporate, casual or somewhere between. Everything else, the presenter, the voice and the lip-sync, comes from the models themselves. You do not need a camera, a microphone or an actor. A first draft script and ten minutes is enough to see whether the direction works before refining it further.
Which model makes the talking avatar?
heygen/avatar-v generates the presenter, taking a typed script and a chosen voice ID and returning a lip-synced video at 720p, 1080p or 4K, in 16:9 or 9:16. It supports scripts up to 5,000 characters per run and can burn captions directly into the output, so a single generation covers the presenter, the sync and the on-screen text together.
Can I use my own photo instead of a stock avatar?
Yes, through a separate photo-to-talking-video model that animates a real photo instead of a preset character. This suits a founder-led update or a personal brand where the point is that it is actually you speaking, even though the delivery is generated. The result reads more personal than a studio avatar but works best with a clear, front-facing photo.
How do I get natural-sounding narration?
Generate the voiceover separately before the avatar render, using elevenlabs/v3 or minimax/speech-2.8-hd, and listen to the pacing on its own first. Short, spoken sentences read more naturally than long written clauses, and a quick read-aloud pass on your own script before generating usually catches the lines that will sound stiff.
Can the video be in a different language than the script?
Yes. Write and render the explainer once, then run the finished narration through AI video dubbing to translate the audio and resync the lip movement to the new language, rather than rewriting the script and rendering a second avatar pass from the start. This is the faster route for teams publishing the same video across several markets.
Does the avatar look and sound realistic?
Close enough that most viewers register it as a video presenter rather than obviously synthetic media, especially in a well-lit studio-style avatar reading a clean script. It is not indistinguishable from a filmed actor in every case, particularly on fast hand motion or a script with heavy emotional swings, and it is honest practice to disclose AI narration where the audience would reasonably expect a live presenter.
How long can an AI explainer video be?
heygen/avatar-v accepts up to 5,000 characters of script per generation, roughly five to seven minutes of spoken narration depending on pacing. A longer video is built from several segments stitched together in an editor rather than one continuous render, which also gives you a natural place to cut to supporting visuals.
Can I add captions automatically?
Yes, two ways. Avatar-v can burn captions directly into the render as part of the same generation, or you can pull a timed transcript separately through automatic video subtitles and style the look yourself. The first is faster; the second gives more control over font, placement and timing.
What does an AI explainer video cost to make?
Generations are paid in credits, and the cost depends on the resolution, length and the specific models you run, so a fixed number here would be wrong within a month. New accounts start with free credits, enough to test a script on the avatar and voice models before committing further, and current plan pricing lives on the pricing page, the only place worth trusting for numbers.
How is this different from just recording a voiceover over slides?
A voiceover over slides is a faster, cheaper route and works well for tutorials and documentation, where the product on screen carries more weight than a face does. An avatar adds a presence that helps with trust and retention, particularly for a launch or an onboarding video meant to feel like a person is walking you through something rather than reading at you.
Write the script, then let a presenter say it. Open the Picasso IA toolkit and turn today's draft into a narrated video before the meeting where you would have pitched a shoot instead.