AI Applications
Choose an AI model to get started
Video Generation
Image Generation
Audio Generation
Video Effects
Wan 2.6 Image to Video Flash
Transform images into dynamic videos up to 15 seconds with Alibaba Wan 2.6 Flash. Supports 720p/1080p resolution, single or multi-shot modes, and optional audio generation.
Input
Output
Generated content will appear here
Example Results

Video Generation:
Upload your image to animate
Video Generation:
zoom out and she is dancing
Frequently Asked Questions
- What inputs are required?
- An input image and a text prompt are required. You can optionally adjust resolution, duration, shot type, audio settings, and other parameters.
- What is the difference between single and multi shot types?
- Single shot produces a continuous, smooth video. Multi shot creates dynamic content with scene transitions, better for action-packed or storytelling content.
- What audio formats are supported?
- WAV and MP3 formats are supported. You can upload custom audio to sync with your video, or enable audio generation for AI-created sound.
- What is the maximum video duration?
- The maximum duration is 15 seconds. You can select any duration from 2 to 15 seconds using the slider.
- Any tips for better results?
- Use high-quality, clear images. Write detailed prompts describing motion direction, speed, and atmosphere. Use negative prompts to avoid unwanted artifacts. Single shot mode works better for smooth, continuous motion.
You Might Also Like
View AllWan 2.2 Image to Video Fast
A very fast and cheap optimized version of Wan 2.2 A14B image-to-video model. Transform static images into dynamic video sequences with smooth transitions and high-quality motion generation.
Wan 2.5 Image to Video Fast
Generate high-quality videos from a single image with Alibaba Wan 2.5-fast. Choose 720p/1080p, 5s/10s, optionally guide with audio, and enable prompt expansion.
Seedance 2.0 API
ByteDance's next-generation AI video model featuring 2K cinema-grade output, native audio-video sync, multi-modal input with up to 12 reference assets, multi-shot storytelling, and 8+ language lip-sync — all from a single prompt.