Recording a training video that the engine will accept
Complete recipe for recording a training video the avatar engine will accept first try: framing, clothing, lighting, audio, and the common rejection causes we
Home › Help Center › AI Avatars › Recording a training video that the engine will accept
Training a video-based AI avatar takes 3–4 hours and burns 4,120 credits per attempt. Getting the recording right the first time saves you both. This is the recipe we'd hand a friend before they hit record.
~60 seconds
30s talking · 30s silent
Landscape 16:9
Camera sideways
30 fps
Not 60, not 24
Head & shoulders
Cut at mid-chest
The shape you're aiming for: head-and-shoulders, defined collar edge, soft even light, plain background.
Why it matters
Our training pipeline is good — but it's not magic. It learns your face, your voice, your micro-expressions from one 60-second clip. If anything in that clip is ambiguous (where your shoulders end, where your face turns, what your skin tone is in shadow), the model trains on the ambiguity too. When the trained avatar then fails the engine's quality check, the entire 3–4 hour run is thrown away.
The good news: every rejection we've seen traces back to one of a small number of recording issues. Fix those upfront and your success rate jumps dramatically.
Two-phase recording: the first ~30 seconds you speak naturally — read a short script, introduce yourself, tell a story. The second ~30 seconds you stop talking and sit silently, looking at the lens with a relaxed neutral expression. Both halves are critical: the engine needs to learn how you move when speaking AND how you look at rest.
Quick checklist (read first)
- Landscape orientation — turn your phone or camera sideways.
- 30 fps — not 60, not 24. Check your camera settings.
- ~60 seconds total — 30s talking, 30s sitting still and silent.
- Head-and-shoulders framing. Top of head near the top of the frame; cut off mid-chest.
- Top with a clear collar or neckline. A defined edge between your shoulder and the background.
- Soft, even light on your whole face — no harsh window behind you, no half-shadow.
- Plain background — a wall, a curtain, a calm room. No other faces, no posters, no busy patterns.
- Quiet room — no music, no air-con hum, no echo.
- Sit still during the silent half. Don't look around, don't sway, don't fidget.
Orientation, frame rate, length
The engine is calibrated for landscape 16:9 at 30 fps. It can technically ingest portrait or 60 fps, but the trained model is consistently lower quality and rejection rates are noticeably higher.
✓ Do this
- Landscape 16:9 — camera sideways
- 30 fps capture
- 60-second take, give or take
- Same distance from camera throughout
✗ Avoid
- Portrait phone video (vertical)
- 60 fps "smooth" mode
- Cinematic 24 fps
- Leaning forward / back as you talk
Framing and distance
Imagine a passport photo, then zoom out about 30%. You want:
- Top of your head near the top of the frame (don't crop hair off)
- Frame ending around mid-chest
- Your face centred horizontally and slightly above the vertical midpoint
- Eyes looking directly into the lens — not above it, not below it
- The camera at roughly eye level (laptop webcams are usually too low — stack books under the laptop)
Eyes on the lens, head-and-shoulders framing, a single soft front light — exactly what the engine wants to see.
Stay the same distance from the camera for the whole 60 seconds. Don't lean forward when you make a point and back when you finish. The engine reads "distance change" as "physical movement to model" and overcorrects.
Clothing — the rule everyone gets wrong
Wear a top with a visible collar, neckline, or seam at the shoulders. The engine uses the edge between your shoulders and the background to figure out where your body ends. If you wear a halter top, a strapless dress, a tank top with thin straps — or a top with strong patterns or in the same colour as the background — the engine can't draw that edge cleanly. The trained model ends up with artifacts around the neck and shoulders, and the quality check rejects it.
Collared top — sharp shoulder edge
Bare shoulders — no edge
Stripes + busy background
✓ Wear this
- Collared shirt (contrasting with background)
- Crew-neck T-shirt
- Mock-neck / turtle-neck top
- Sweater or hoodie with hood down
- Blazer or jacket
✗ Don't wear
- Halter, strapless, off-the-shoulder
- Tank top with thin straps
- Same colour as your background
- Strong patterns (stripes, plaid, busy florals)
- Reflective fabric (sequins, satin)
Most common rejection cause we see: the engine reports "Quality issue detected" but the real reason is shoulder-edge detection failing because of clothing. If you've been rejected for "quality" and your video looks fine to you, your clothing is the first thing to change.
Lighting
The model learns your skin tones, your highlights, the shape of your face from light. Get this wrong and it learns the wrong face.
✓ Good lighting
- Light source in front of you
- Diffuse (north window, ring light through softbox)
- One source, not three
- Warm-white LED if night-shooting
✗ Bad lighting
- Bright window behind you (silhouette)
- Direct sun or bare bulb (harsh shadows)
- Mixed colour temperatures on your face
- Flickering cheap LED bulbs
Background
Keep it calm. The engine doesn't fail on a "busy" background, but face-tracking accuracy drops the more there is going on behind you — and that drop can be enough to push the trained model below the quality threshold.
A busy background like this is risky — text, vehicles, and high-contrast shapes behind your head can confuse face-tracking. Move to a plain wall.
- Plain wall, curtain, or bookshelf at least 1 metre behind you. Distance prevents your shadow from landing on the background.
- No other faces in the frame — including photos on the wall, posters, screens playing video, mirrors that might reflect someone.
- No high-contrast patterns directly behind your head. Avoid anything with strong stripes, sharp geometric prints, or text.
- Not too busy. A messy bookshelf is fine; a kitchen with appliances, plants, mail, and movement is not.
Audio
Audio quality matters for the cloned voice. A quiet recording trains a quiet clone — there's no "louder" knob after training.
✓ Good audio
- Speak at normal volume — mic meter in the green
- Headset, lavalier, or USB condenser mic if you have one
- Small room with soft furnishings (bedroom)
- Phone on aeroplane mode, notifications off
- One language only
✗ Bad audio
- Whispering or trailing off mid-sentence
- Laptop mic in an echo-y kitchen or bathroom
- Background music, AC hum, traffic noise
- Notification dings, fridge cycling, neighbours
- Switching languages mid-take
Speaking segment (first 30 seconds)
You have ~30 seconds before the silent half begins. The content doesn't matter — the engine doesn't transcribe what you say, it learns how you move and sound. But you want continuous natural speech, not pauses.
Speaking — mouth open, looking at the lens, relaxed
Silent — neutral expression, calm posture, eyes on lens
A simple script that works:
"Hi, my name is [your first name] and I'm recording this short clip so an AI can learn how I look and sound on camera. I'm looking straight into the lens and speaking in my normal everyday voice. My shoulders are relaxed and I'm trying not to move my head too much. In a moment I'll stop talking and sit completely still and silent so the model can capture my resting expression."
Read it once at normal pace — it's deliberately tuned to ~30 seconds. Don't whisper, don't shout, don't perform — speak the way you'd speak to a colleague you've known for years.
Silent segment (next 30 seconds)
This is where most people accidentally sabotage their take. After the speaking script ends, you need to:
- Stop talking entirely. Don't mouth words, don't hum.
- Keep looking at the lens. Don't look away, don't read your script again, don't check the recording.
- Relax your face into a neutral expression. Not smiling, not frowning — the face you make when listening to someone politely.
- Sit still. Don't shift in your chair, don't rebalance, don't scratch.
- Blink normally. Don't try to suppress blinks — that looks alien on the trained avatar.
This half is what makes your avatar feel alive at rest. It's why a trained avatar feels different from a stock one — the stock ones don't have your neutral expression.
Full-body or 3/4-body replicas (optional)
The default training takes a head-and-shoulders crop — the only part of you that ends up in the trained avatar. If you want a 3/4-body or full-body avatar (one that shows your torso, arms, and posture during video generation), the recording rules tighten:
3/4-body: full torso visible, clean outline against a plain background, arms relaxed.
- Step further back from the camera so your full upper body is in frame, with some headroom and some space below the waist.
- Use a plain backdrop — a single-colour wall, a roller backdrop, or a green/blue screen. The engine needs a clean silhouette around your whole body, not just your head.
- Wear clothing that has a defined edge along the body — the same shoulder-edge rule, but extended to your arms and waist. Loose, drapey clothing makes the silhouette ambiguous.
- Stand or sit still — even more important than the head-and-shoulders take. Any large body movement makes the trained model less stable.
- Make sure the floor / ground is consistent — same colour and texture for the whole take. Don't film with feet half on a rug, half on bare floor.
After you submit
3–4 hoursTypical training time. Leave the page — we'll update My Avatars.
Live progress barHH:MM:SS ticking on your avatar card.
Auto-refund on failureIf training fails, your credits come straight back.
Talk live + generateOnce ready, your avatar is yours to use.
Common rejection causes (and how to fix them)
- "Quality issue detected" — usually clothing (no clear shoulder edge), lighting (face in shadow), or the engine being unable to track your face for parts of the clip. Re-record with a collared top, soft front-lighting, and minimal head movement.
- Avatar comes out blurry around the mouth — usually the camera was out of focus, the source video was over-compressed, or you were moving too much while speaking. Re-record closer to the camera, slightly slower delivery.
- Avatar's eyes look wrong — usually you weren't looking directly into the lens. Tape a small arrow above your laptop camera as a target.
- Voice clone sounds too quiet — your source audio was below ~-22 dB RMS. The recorder shows you a mic meter; if it stays in the yellow, you're too quiet to clone well.
- Avatar's skin tone drifts between scenes when you talk to it later — you had multiple coloured light sources during the take. Re-record with one light source only.
When you should NOT use the video-trained path
The video path is for "I want an AI version of me, with my exact face and my exact voice." It's high effort and not always the right answer.
Generic professional face?
Use a
stock replica — ready in seconds, free, 100+ to choose from.
One good photo only?
Photo-trained path — single headshot + voice from catalog. Same 3–4 hours but easier setup.
Two video attempts failed already?
Don't try a third with the same setup. Change clothing, room, camera — OR fall back to stock / photo-trained.
TL;DR
Landscape · 30 fps · 60 seconds (30 talking, 30 silent) · head-and-shoulders · top with a visible collar · soft front lighting · plain quiet room · sit still during the silent half. Get those right and the engine is overwhelmingly likely to accept the first take.
Related articles in AI Avatars
Help Center · Learn · Privacy Policy · Terms of Service · Contact