Ovi Model Integration Guide for ComfyUI
Posted on October 15, 2025 - Comfyui

Introduction: Bringing Synchronized Audio-Video Generation to Your ComfyUI Workflow
Character.AI's Ovi model represents a significant advancement in generative AI, offering the ability to simultaneously create synchronized video and audio content from a single prompt. This powerful model can interpret either text-only or a combination of text and image inputs to produce cohesive multimedia clips. This guide provides a comprehensive walkthrough for integrating Ovi into the popular and flexible ComfyUI framework using the ComfyUI-Ovi custom node package, a tool designed to streamline the entire process.
The Ovi model is distinguished by several powerful features that set it apart in the generative media landscape:
- Video+Audio Generation: Ovi's core capability is the simultaneous and synchronized creation of both video and audio streams, ensuring that sound and motion are perfectly aligned.
- High-Quality Audio: The model incorporates a bespoke 5B audio branch, pretrained on high-quality internal datasets to deliver rich and clear sound.
- Flexible Input: It supports both text-only prompts for generating entirely new scenes and text-plus-image conditioning to guide the generation from a starting visual.
- Video Specifications: By default, Ovi generates 5-second videos at 24 frames per second (FPS), with a training resolution of 720x720 pixels.
- High-Resolution Upscaling: Despite its training resolution, Ovi can naturally generate content at higher resolutions, such as 960x960, and in various aspect ratios while maintaining spatial and temporal consistency.
The ComfyUI-Ovi custom node set serves as the bridge between this advanced model and the ComfyUI ecosystem. It provides the preferred integration method by delivering streamlined setup, selectable precision (BF16/FP8), fine-grained attention-backend control, and per-node device targeting for multi-GPU rigs. This document is a practical guide for developers, creators, and technical artists, designed to walk you through the complete process from initial installation to advanced performance tuning.
With the proper setup, you can harness Ovi's unique audio-video capabilities directly within your existing ComfyUI workflows. We will begin by outlining the necessary prerequisites to ensure a smooth installation.
1. System Prerequisites and Requirements
Verifying that your system meets the necessary requirements is the critical first step for a successful installation and optimal performance. Running a sophisticated model like Ovi requires specific hardware and software configurations to handle its computational demands. This section outlines the dependencies needed to run Ovi effectively within ComfyUI.
Hardware and Software Requirements
| Category | Component | Requirement |
|---|---|---|
| GPU VRAM | FP8 Model | 16–24 GB (with CPU offload) |
| BF16 Model | >32 GB+ (without CPU offload) | |
| CUDA Stack | PyTorch | 2.4+ |
| Driver/Runtime | CUDA 12.x |
Component Dependencies
The ComfyUI-Ovi nodes are designed for component reuse, leveraging model files that may already be present in your ComfyUI environment. These files are not downloaded automatically and are expected to be present from an existing Wan 2.2 installation. If your setup is new, you will need to source these files manually.
wan2.2_vae.safetensors- One of the following UMT5 text encoders:
umt5-xxl-enc-bf16.safetensors(for BF16 precision)umt5-xxl-enc-fp8_e4m3fn.safetensors(for FP8 precision)
Once you have confirmed that your system meets these prerequisites, you can proceed with the custom node installation.
2. Installation and Weight Management
The installation process is made remarkably straightforward by the automation features built into the ComfyUI-Ovi package. This section covers the installation of the custom nodes and explains how the required model weights are managed, both automatically by the loader and manually if needed.
Step 1: Install the Custom Node
Execute the following commands in your terminal, starting from your main ComfyUI directory:
# Navigate to the custom_nodes directory
cd custom_nodes
# Clone the ComfyUI-Ovi repository
git clone https://github.com/snicolast/ComfyUI-Ovi.git
# Navigate into the newly cloned directory
cd ComfyUI-Ovi
# Install the required Python packages
pip install -r requirements.txt
After the installation commands are complete, you must restart ComfyUI for the new nodes to be recognized.
Step 2: Understanding Automatic Weight Handling
The Ovi Engine Loader node includes a self-bootstrapping loader feature that automates much of the setup. When you first use this node, it will automatically download the necessary MMAudio assets and the selected Ovi fusion weights (Ovi-11B-bf16.safetensors or Ovi-11B-fp8.safetensors) to the correct directories.
Note that this automatic process intentionally skips the VAE and text encoder, as the system is designed to reuse these common components from existing Wan 2.2 installations to avoid data duplication.
Step 3: Verifying Model Directory Structure
While the loader handles the Ovi and MMAudio weights, it relies on the pre-existence of the Wan 2.2 VAE and UMT5 text encoder. Proper file placement is crucial for the system to function correctly. The expected directory structure within your ComfyUI installation is as follows:
ComfyUI/
├── models/
│ ├── diffusion_models/
│ │ ├── Ovi-11B-bf16.safetensors
│ │ └── Ovi-11B-fp8.safetensors
│ ├── text_encoders/
│ │ └── umt5-xxl-enc-bf16.safetensors # or the fp8 version
│ └── vae/
│ └── wan2.2_vae.safetensors
└── custom_nodes/
└── ComfyUI-Ovi/
└── ckpts/
└── MMAudio/
└── ext_weights/...
With the nodes installed and the model weights correctly placed, you are ready to build your first audio-video generation workflow.
3. Core Workflow: Your First Audio-Video Generation
This section provides a practical, step-by-step walkthrough of a complete generation workflow using the core Ovi nodes. By following these steps, you will learn the function of each component and see how they connect to produce a final, synchronized audio-video output. All Ovi-specific nodes can be found under the Ovi category in the ComfyUI node search dialog.
Node-by-Node Breakdown
- Ovi Engine Loader: Responsible for downloading any missing weights, building the core generation engine, and caching it for subsequent runs. Also provides settings for precision (BF16/FP8), CPU offload, and GPU targeting.
- Ovi Video Generator: The generation core. Takes the engine and a text prompt as input, producing the raw video and audio data in latent form. Accepts an optional image input for text-plus-image conditioning.
- Ovi Latent Decoder: The final step. Decodes the video and audio latents into tangible media: image frames and an audio stream.
The Six-Step Generation Process
- Load Engine: Add the Ovi Engine Loader node and configure precision and CPU offload as needed.
- (Optional) Connect Wan Components: Use the Ovi Wan Component Loader node if your files are in custom directories.
- (Optional) Tune Attention: Add the Ovi Attention Selector node to manually set the attention backend.
- Generate Latents: Connect the OVI_ENGINE output from the loader to the Ovi Video Generator and enter your prompt.
- Decode to Media: Connect the outputs to the Ovi Latent Decoder node.
- Save Output: Connect to “Save Image” and “Save Audio File” nodes to export the final results.
Crafting Effective Prompts
Ovi uses special tags within the text prompt to differentiate between speech and environmental sounds:
- Speech:
<S>and<E>tags - Audio Description:
<AUDCAP>and<ENDAUDCAP>tags
4. Advanced Configuration and Performance Optimization
Optimizing your Ovi workflow is a balancing act between VRAM, speed, and quality.
Managing VRAM and Precision
- BF16: Higher precision, requires 32 GB+ VRAM, best quality.
- FP8: Lower precision, works on 16–24 GB VRAM, slightly lower quality.
Leveraging CPU Offload
Enabling CPU offload moves model modules from GPU VRAM to system RAM, reducing VRAM use but increasing runtime (~+20 seconds).
Multi-GPU Device Targeting
- Single-GPU systems: Hidden by default.
- Multi-GPU systems: Dropdown appears on the loader node to select which GPU (e.g.,
cuda:0,cuda:1) to use.
Optimizing with Attention Backends
The Ovi Attention Selector node allows manual backend control. Options include:
autoFlashAttentionSDPAxFormersnative
5. Troubleshooting and Best Practices
Common Issues and Solutions
- High VRAM After a Run: Use ComfyUI’s Unload Models feature to clear memory.
- Missing Weights: Manually place missing files in the correct directory.
- Switching Precision: Switch between BF16/FP8 via dropdown without restarting ComfyUI.
- Backend Errors: Automatically falls back to native if backend dependencies are missing—check console logs for details.
6. Conclusion
By leveraging the ComfyUI-Ovi custom nodes, integrating Character.AI's Ovi model into your creative workflow becomes straightforward and powerful. With streamlined setup, automated weight management, and flexible performance controls, synchronized audio-video generation is now accessible to a wide range of systems.
Whether aiming for the highest quality or balancing limited resources, these tools provide the flexibility you need. Experiment with precision settings, attention backends, and prompt styles to unlock Ovi’s creative potential.
We extend our thanks to Character.AI for developing the Ovi model, and to the maintainers of Wan 2.2 VAE, MMAudio, and UMT5 for their foundational work.
Related Posts

ComfyUI Luma AI API: Una guía sencilla
Aprende a crear vídeos impresionantes sin esfuerzo. Nuestra guía te muestra cómo integrar, instalar y usar la API Luma AI Dream Machine en ComfyUI. ¡Transforma hoy mismo tu proceso de creación de vídeos!

ComfyUI Snapshot Manager: Gestión de nodos personalizados y entornos
¿Luchas por mantener organizados tus nodos personalizados y entornos de ComfyUI? El ComfyUI Snapshot Manager está ahí para apoyarte.

Atajos de teclado de ComfyUI: ¡Potencia tu flujo de trabajo ahora!
¿Cansado de flujos de trabajo lentos? Estos revolucionarios atajos de teclado de ComfyUI harán que tu productividad se eleve vertiginosamente.

Comfyui VisualQueryTemplate Node for Precision Control
Este nodo, denominado "VisualQueryTemplateNode", está diseñado para realizar tareas de Respuesta a Preguntas Visuales (VQA, por sus siglas en inglés) en imágenes.