Wan 2.1 Mocha

End-to-end video character replacement system that swaps a character in a source video with a new character from reference images. No need for explicit structural guidance like pose or depth maps.

Cost: 60 credits

Input

Upload a clear character photo

Upload a clear character photo (PNG/JPG recommended, avoid WEBP). The composition, camera position, and human pose should match the video for best results.

Upload the motion clip

Upload the video containing the motion/expressions. Keep the same aspect ratio as the image for best results. Maximum 120 seconds supported.

Try:

Optional. Brief rules to guide the character replacement, such as preserving outfit, natural expressions, or avoiding background changes.

Choose output video resolution. Higher resolution increases generation cost.

Set a specific seed for reproducible results. Use -1 for random generation.

Output

Generated content will appear here

Example Results

Reference character image

Character Replacement:

Reference character image

Character Replacement:

Motion video

Character Replacement:

Motion video

Frequently Asked Questions

What is Wan 2.1 Mocha?
MoCha is an end-to-end video character replacement system designed to swap a character in a source video with a new character (provided via reference images), without relying on explicit structural guidance (e.g., pose/depth maps) for every frame.
How are credits calculated?
Duration- and resolution-based: 480p ≈ 12 credits/s, 720p ≈ 24 credits/s. Minimum charge is 60 credits (5 seconds at 480p). Maximum video length is 120 seconds.
What are the use cases for character replacement?
Creating personalized content, character animation for games/films, translation/localization with different actors, creating demo content, replacing actors in existing footage, and generating character variations for storytelling.
What are the key features of MoCha?
End-to-end character replacement without pose maps, natural expression transfer, holistic movement replication, maintains character appearance consistency, supports various resolutions (480p/720p), and preserves motion dynamics from source video.
How do I get best results?
Match composition & pose between reference image and video, keep aspect ratios consistent, use clear reference photos (PNG/JPG), ensure camera position and human pose align, and use optional prompts to preserve specific elements like outfit or expression.
What inputs are required?
Character image (PNG/JPG recommended) and motion video are required. Optional: prompt for guidance rules, resolution selection (480p/720p), and seed for reproducibility.
Can I control which character gets replaced?
MoCha automatically identifies and replaces characters based on the reference image. For better targeting, ensure your reference image and video share similar composition, camera angle, and pose.
What video formats are supported?
Input: MP4 or WebM. Output: MP4 format. Maximum input video length is 120 seconds.
How does MoCha differ from other character animation tools?
MoCha is specialized for character replacement in existing videos, not generating new motion. It transfers a character's appearance to another character's motion without requiring pose estimation or depth maps, making it efficient for production workflows.