Getting the most out of Avatar V comes down to three inputs: your motion recording, your voice clone, and your base look. This guide covers how to optimize each one for the best possible results.
New to Avatar V? Start with Avatar V guide first, then come back here to level up your results.
🛎️ Quick facts about Avatar V 🛎️
Avatar V and Avatar IV use per-Look credit rates:
Avatar V: Video Look = 48 credits per minute (not available for photo-based looks)
Avatar IV: Photo Look = 16 credits per minute; Video Look = 31 credits per minute
Avatar V works best for real human avatars, while Avatar IV works better for virtual characters, including 2D, 3D, and non-human avatars.
Avatar V is only available for video-based looks. For photo-based looks, use Avatar IV.
In Video Agent, Avatar V is applied automatically only for real human avatars, while non-human or virtual avatars will default to Avatar IV.
Avatar V is available in Studio, Home Page Shortcut, & Video Agent.
Motion recording
Your 15-second video is the most important input in the entire process. Avatar V learns your gestures, expressions, and mannerisms from this single clip.
Key principles:
Be extremely expressive - more than feels natural
Vary your tone and use your hands
Make direct eye contact with the camera
Bring energy even if it feels like you're overdoing it
The energy you put in is the energy you get out. A flat recording produces a stiff avatar. An expressive recording produces one that feels alive.
Example:
Recording style | Output |
Calm, reserved | Stiff, robotic motion |
Expressive, energetic | Natural, dynamic motion |
Voice clone
A dedicated standalone voice clone produces noticeably more realistic results than using audio from your base video.
Key principles:
Record a standalone voice clone rather than relying on your base video audio
Be expressive and vary your delivery when recording
Iterate until it genuinely sounds like you at your best
Base look
Your base look is your avatar's identity. Every outfit, setting, and scene you generate is built from this single photo - so choosing the right one is critical.
Key principles:
Use a close-up or half-body shot with your face clearly visible
Keep your expression subtle - avoid big smiles or unusual angles
Avoid accessories that obscure your face
Choose a photo where you feel you look your best
Example:
Base look quality | Output |
Unclear face, accessories, unusual angle | Inconsistent, lower quality generated looks |
Clear close-up, subtle expression, no accessories | Consistently strong generated looks |
⚠️ Face too small = lip sync degrades: if your face takes up too little of the frame (common in zoomed-out shots), the AI can't see enough of your mouth to sync speech accurately. Use a close-up or half-body shot where your face is large and clearly visible.
Test a few options before committing. The right base look will produce noticeably better outputs across all your prompts.
⚠️ Match what you give to what you want back: whatever the AI doesn't have, it invents — this is called hallucinating. A face-only clip means it guesses your body. A closed-mouth photo means it guesses your teeth. Guesses are where things look off. Show it what you want it to get right.
💡 You have to like it: you won't like the output more than the input. Pick a photo where you genuinely like how your face looks — if it doesn't look good to you in the source, it won't look better coming out.
💡 Accessories logic: skip hats and glasses in your base look unless you want them to appear in every generated look. It's easier to add a hat later than remove one, because the AI already knows your hair from the video reference and can add accessories cleanly.
⚠️ Important: Always use the same base look when generating new Looks. Switching your base look mid-project causes slight visual inconsistency across your Look library — your avatar may appear subtly different from one Look to the next.
Refining your looks
Use the Edit feature to make targeted adjustments to any generated look without starting from scratch. This is useful for tweaking outfit details, swapping backgrounds, or refining specific elements.
Motion reference
If you've uploaded multiple video looks to your avatar group, you can select a specific one as your motion reference in Advanced Settings. Your motion reference shapes the style and energy of movement in your output.
Key principles:
Match your motion reference to the tone of the video you're making
Use a high-energy recording for dynamic, expressive output
Use a calm recording for natural, subtle output
Use a side-angle recording for better results with side-angle photos
💡 Pro Tip: Match your motion reference video angle to your photo's angle. If your Look photo is a side-angle shot, use a side-angle recording for your motion reference. Mismatched angles produce noticeably less consistent results — the avatar's pose and generated Look won't align as naturally.
Common issues and solutions
Avatar looks stiff or robotic Your motion recording may lack expressiveness. Re-record with more energy, gestures, and facial expression.
Voice doesn't sound like me Use a standalone voice clone rather than your base video audio. Record expressively and iterate until satisfied.
Generated looks are inconsistent Your base look may not be strong enough. Try a clearer close-up with a subtle expression and no accessories.
Output doesn't match the photo angle Select a motion reference that matches the angle of your photo. For side-angle photos, use a side-angle motion reference.
My generated Looks look inconsistent from one to the next — what can I do?
This is usually a base look quality issue. Here's the most effective fix:
Generate 50 Looks from your current base look
Identify the 10 most consistent and accurate ones
Retrain your avatar using only those 10 as the reference set
Repeat if needed — each iteration typically improves consistency
This approach uses the model's own output to refine itself, progressively narrowing in on your most stable likeness.
