vllm-project/vllm-omni

[New Model] Add TADA (Hume AI) TTS model support

Open

#2,001 opened on Mar 19, 2026

View on GitHub
 (8 comments) (0 reactions) (1 assignee)Python (1,067 forks)github user discovery
good first issuehelp wantednew model

Repository metrics

Stars
 (4,990 stars)
PR merge metrics
 (PR metrics pending)

Description

Model

TADA by Hume AI — a speech-language model that unifies speech synthesis with text generation using 1:1 text-speech token alignment built on Llama 3.2.

Architecture

  • Encoder (HumeAI/tada-codec): encodes reference audio into aligned token sequences
  • LLM (TadaForCausalLM): autoregressive generation, available in 1B and 3B-ML (multilingual) variants
  • Generates the complete speech segment per text token in a single AR step (dynamic duration/prosody)

This would likely be a 2-stage pipeline in vllm-omni (AR generation → codec decode), similar to Qwen3 TTS / Fish Speech.

References

cc @ashok-arora — would you be interested in taking a look at this?

Contributor guide