Photo to talking avatar: make any picture speak with AI
One still image, one script or voice recording, and you get a lip-synced video of up to 28 seconds. No camera, no studio, no editing timeline.
Make my photo speak40 free credits at sign-up Β· talking clips from β¬4.99
How a still photo becomes a speaking presenter
A photo contains a face frozen in one position. Speech, on the other hand, is a continuous sequence of shapes: the lips round for an "o", close entirely for a "b", spread for an "ee", and the jaw drops and rises between them. Making a picture talk means generating every one of those intermediate positions from a single frame, and keeping them aligned with the audio to within a fraction of a second β because the human eye notices desynchronised lips long before it can explain why a video feels wrong.
That is the work the tool does for you. It locates the face in your image, maps its geometry β the corners of the mouth, the jawline, the eyes, the head angle β then reads the audio track and animates that geometry frame by frame to match what is being said. The rest of the face moves too: small blinks, subtle brow motion, the micro shifts that separate a speaking person from a puppet with a hinged jaw. The background and the identity of the subject stay untouched.
Two consequences are worth knowing before you start. First, the source does not have to be a photograph at all: a drawn character, a mascot or a cartoon avatar animates on exactly the same principle, since the model works from facial structure rather than from realism. Second, the audio is what drives the timing, so the length of your clip is the length of your recording β capped at 28 seconds, which is where accurate lip-sync stays reliable.
Three steps, no editing software
The whole process happens in the browser, from upload to finished video file:
1. Pick the photo
A portrait, a headshot, an illustrated character or a cartoon avatar you generated on ToonZap. The face is detected automatically.
2. Supply the words
Type the script, record straight from the browser, or upload an audio file you already have. Text and voice both work as the driver.
3. Collect the video
The tool returns a video where the mouth and expressions follow the audio, up to 28 seconds long, ready to publish or embed.
Where a talking photo earns its place
Twenty-eight seconds sounds short until you count what fits into it β an intro, an announcement, a lesson opener, a personal message:
Presentations & pitches
Open a deck with a short spoken intro instead of a title slide, or attach a thirty-second walkthrough to a proposal. A presenter who speaks holds attention in a way a bullet list never does β and you record nothing on camera.
Training & onboarding
Give each module of a course a presenter reading the key points. Updating a lesson means editing the script and regenerating, not booking a studio and re-shooting an entire take to fix one sentence.
Social media
Reels, Shorts and TikToks where a character speaks for you. The 28-second ceiling is well matched to short-form feeds, where the first three seconds decide everything and brevity is a feature.
Choosing the right source photo
Most disappointing results trace back to the input, not the model. Three habits fix almost everything:
Frame the head generously
Head and shoulders, roughly centred, with a little room above the hair. A full-body shot leaves too few pixels on the face for the animation to read expressions cleanly.
One face per photo
If two people share the frame, it becomes ambiguous which one should speak. Crop to the person you want animated before uploading.
Match voice to face
A calm portrait paired with an energetic delivery reads as odd. Choosing a voice whose pace and warmth fit the expression in the photo makes the result far more believable.
Writing for 28 seconds
A spoken script is not a written paragraph read aloud. At a comfortable pace, 28 seconds holds roughly sixty to eighty words, so the discipline is to say one thing well rather than three things quickly. Open with the point instead of building up to it, keep sentences short enough to speak in a single breath, and end on the action you want β visit, register, watch the next module. Read your draft out loud once before generating: anything you stumble over, the voice will stumble over too.
For the voice itself you have three options β record in the browser, upload an existing audio file, or generate the narration from your text with one of the 21 AI voices available in English and many other languages (with a subscription, from the Essential plan at β¬9.99/month; your own recording works on every plan). Browse the full library on the AI voice generator page. Whichever route you take, the lip-sync is generated against that exact audio, so the delivery you choose is the delivery you see.
No photo you want to use? Generate the face first
Plenty of people would rather not put their own face in a video. Create an original character with the avatar generator, then bring it back here and give it a voice.
Create a free AI avatar firstPhoto to talking avatar β frequently asked questions
What kind of photo works best for a talking avatar?
A front-facing picture where the face is clearly visible, reasonably large in the frame and evenly lit. The mouth must not be hidden by a hand, a microphone or a scarf, and hard shadows across one half of the face tend to confuse the animation. Beyond that the source is flexible: a phone selfie, a studio headshot, a cartoon avatar generated on ToonZap or an illustrated character all animate the same way, because the model works from facial geometry rather than from photographic realism.
How long can the talking avatar video be?
Up to 28 seconds with accurate lip-sync. Written out, that is roughly sixty to eighty spoken words at a natural pace, which is enough for an introduction, a product announcement, a lesson intro or a personal video message. If your script is longer, the practical approach is to cut it into consecutive segments and generate one clip per segment, then assemble them in the order you want.
Can the avatar speak with my own voice?
Yes, and there are three routes. You can record your voice directly in the browser, upload an audio file you already have β a voice memo, a podcast excerpt, a voice-over exported from your editing software β or type a script and pick one of the 21 natural AI voices available in English and many other languages (AI voices come with a subscription, from the Essential plan at β¬9.99/month β a credit pack alone does not unlock them; your own recording works on every plan). Whichever you choose, the mouth movements and facial expressions are generated to match that specific audio track.
Do I need video editing skills to make a photo speak?
None at all. There is no timeline, no keyframing and no lip-sync work to do by hand: you supply the photo and the audio or script, and the tool returns a finished video file. Every new ToonZap account gets 40 welcome credits with no credit card β enough to test photo-to-cartoon three times. A talking clip costs 60 credits minimum, so producing one takes the Mini pack (β¬4.99, 150 credits, no subscription): two 5-second clips with your own voice.
Give one of your photos a voice
Create your free account (40 credits), then a β¬4.99 Mini pack produces your first two talking clips. All you need is one clear portrait and a couple of sentences.
More AI tools to explore
ToonZap is not affiliated with, endorsed by or sponsored by Instagram, TikTok, YouTube or any animation studio. Platform names are mentioned only to describe where a generated video can be published. Only upload photos of people who have agreed to appear in a generated video.