The Research Company of Pune

Vaani 1

Vaani is our flagship audio generation model.

IFR 57.32
CR 63.67
EMR 8.20

Omni and Novel Architecture

We propose Thinker-Talker architecture, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. We propose a novel position embedding, named TMRoPE (Time-aligned Multimodal RoPE), to synchronize the timestamps of video inputs with audio.

Real-Time Voice and Video Chat

Architecture designed for fully real-time interactions, supporting chunked input and immediate output.

Natural and Robust Speech Generation

Surpassing many existing streaming and non-streaming alternatives, demonstrating superior robustness and naturalness in speech generation.