OmniVoice
Local zero-shot TTS covering 600+ languages, cloning from a short sample.
OmniVoice (k2-fsa) is a diffusion-LM zero-shot TTS model with 600+ language coverage. It clones a voice from a 3–10 second reference sample and understands inline non-verbal cues.
<laugh> <sigh> <surprise> <confirm>Setup
Add the service
Manage Services → + Add Services → OmniVoice → Add. The model downloads on first use.
Give it a reference sample
3 to 10 seconds of clean speech is enough to clone from.
Precision (float16, bfloat16, float32), diffusion steps and guidance scale trade speed against quality.
Needs roughly 5–7 GB of VRAM, wants flash-attn 2 for full speed, and does not truly stream — long replies are chunked. The model is very new.