Vaani 1
Vaani is our flagship audio generation model.
IFR
57.32
CR
63.67
EMR
8.20
Omni and Novel Architecture
We propose Thinker-Talker architecture, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. We propose a novel position embedding, named TMRoPE (Time-aligned Multimodal RoPE), to synchronize the timestamps of video inputs with audio.
Real-Time Voice and Video Chat
Architecture designed for fully real-time interactions, supporting chunked input and immediate output.
Natural and Robust Speech Generation
Surpassing many existing streaming and non-streaming alternatives, demonstrating superior robustness and naturalness in speech generation.