Recording a training video that the engine will accept

Complete recipe for recording a training video the avatar engine will accept first try: framing, clothing, lighting, audio, and the common rejection causes we

HomeHelp CenterAI Avatars › Recording a training video that the engine will accept

Training a video-based AI avatar takes 3–4 hours and burns 4,120 credits per attempt. Getting the recording right the first time saves you both. This is the recipe we'd hand a friend before they hit record.

~60 seconds
30s talking · 30s silent Landscape 16:9
Camera sideways 30 fps
Not 60, not 24 Head & shoulders
Cut at mid-chest The shape you're aiming for: head-and-shoulders, defined collar edge, soft even light, plain background.

Why it matters

Our training pipeline is good — but it's not magic. It learns your face, your voice, your micro-expressions from one 60-second clip. If anything in that clip is ambiguous (where your shoulders end, where your face turns, what your skin tone is in shadow), the model trains on the ambiguity too. When the trained avatar then fails the engine's quality check, the entire 3–4 hour run is thrown away.

The good news: every rejection we've seen traces back to one of a small number of recording issues. Fix those upfront and your success rate jumps dramatically.

Two-phase recording: the first ~30 seconds you speak naturally — read a short script, introduce yourself, tell a story. The second ~30 seconds you stop talking and sit silently, looking at the lens with a relaxed neutral expression. Both halves are critical: the engine needs to learn how you move when speaking AND how you look at rest.

Quick checklist (read first)

Orientation, frame rate, length

The engine is calibrated for landscape 16:9 at 30 fps. It can technically ingest portrait or 60 fps, but the trained model is consistently lower quality and rejection rates are noticeably higher.

✓ Do this

✗ Avoid

Framing and distance

Imagine a passport photo, then zoom out about 30%. You want:

Eyes on the lens, head-and-shoulders framing, a single soft front light — exactly what the engine wants to see.

Stay the same distance from the camera for the whole 60 seconds. Don't lean forward when you make a point and back when you finish. The engine reads "distance change" as "physical movement to model" and overcorrects.

Clothing — the rule everyone gets wrong

Wear a top with a visible collar, neckline, or seam at the shoulders. The engine uses the edge between your shoulders and the background to figure out where your body ends. If you wear a halter top, a strapless dress, a tank top with thin straps — or a top with strong patterns or in the same colour as the background — the engine can't draw that edge cleanly. The trained model ends up with artifacts around the neck and shoulders, and the quality check rejects it.

Collared top — sharp shoulder edge Bare shoulders — no edge Stripes + busy background

✓ Wear this

✗ Don't wear

Most common rejection cause we see: the engine reports "Quality issue detected" but the real reason is shoulder-edge detection failing because of clothing. If you've been rejected for "quality" and your video looks fine to you, your clothing is the first thing to change.

Lighting

The model learns your skin tones, your highlights, the shape of your face from light. Get this wrong and it learns the wrong face.

✓ Good lighting

✗ Bad lighting

Background

Keep it calm. The engine doesn't fail on a "busy" background, but face-tracking accuracy drops the more there is going on behind you — and that drop can be enough to push the trained model below the quality threshold.

A busy background like this is risky — text, vehicles, and high-contrast shapes behind your head can confuse face-tracking. Move to a plain wall.

Audio

Audio quality matters for the cloned voice. A quiet recording trains a quiet clone — there's no "louder" knob after training.

✓ Good audio

✗ Bad audio

Speaking segment (first 30 seconds)

You have ~30 seconds before the silent half begins. The content doesn't matter — the engine doesn't transcribe what you say, it learns how you move and sound. But you want continuous natural speech, not pauses.

Speaking — mouth open, looking at the lens, relaxed Silent — neutral expression, calm posture, eyes on lens A simple script that works:
"Hi, my name is [your first name] and I'm recording this short clip so an AI can learn how I look and sound on camera. I'm looking straight into the lens and speaking in my normal everyday voice. My shoulders are relaxed and I'm trying not to move my head too much. In a moment I'll stop talking and sit completely still and silent so the model can capture my resting expression."

Read it once at normal pace — it's deliberately tuned to ~30 seconds. Don't whisper, don't shout, don't perform — speak the way you'd speak to a colleague you've known for years.

Silent segment (next 30 seconds)

This is where most people accidentally sabotage their take. After the speaking script ends, you need to:

This half is what makes your avatar feel alive at rest. It's why a trained avatar feels different from a stock one — the stock ones don't have your neutral expression.

Full-body or 3/4-body replicas (optional)

The default training takes a head-and-shoulders crop — the only part of you that ends up in the trained avatar. If you want a 3/4-body or full-body avatar (one that shows your torso, arms, and posture during video generation), the recording rules tighten:

3/4-body: full torso visible, clean outline against a plain background, arms relaxed.

After you submit

3–4 hours
Typical training time. Leave the page — we'll update My Avatars. Live progress bar
HH:MM:SS ticking on your avatar card. Auto-refund on failure
If training fails, your credits come straight back. Talk live + generate
Once ready, your avatar is yours to use.

Common rejection causes (and how to fix them)

When you should NOT use the video-trained path

The video path is for "I want an AI version of me, with my exact face and my exact voice." It's high effort and not always the right answer.

Generic professional face?
Use a stock replica — ready in seconds, free, 100+ to choose from. One good photo only?
Photo-trained path — single headshot + voice from catalog. Same 3–4 hours but easier setup. Two video attempts failed already?
Don't try a third with the same setup. Change clothing, room, camera — OR fall back to stock / photo-trained. TL;DR Landscape · 30 fps · 60 seconds (30 talking, 30 silent) · head-and-shoulders · top with a visible collar · soft front lighting · plain quiet room · sit still during the silent half. Get those right and the engine is overwhelmingly likely to accept the first take.

Related articles in AI Avatars

Help Center · Learn · Privacy Policy · Terms of Service · Contact