Voxta docs

OmniVoice

Local zero-shot TTS covering 600+ languages, cloning from a short sample.

OmniVoice (k2-fsa) is a diffusion-LM zero-shot TTS model with 600+ language coverage. It clones a voice from a 3–10 second reference sample and understands inline non-verbal cues.

<laugh>  <sigh>  <surprise>  <confirm>

Setup

Add the service

Manage Services → + Add Services → OmniVoice → Add. The model downloads on first use.

Give it a reference sample

3 to 10 seconds of clean speech is enough to clone from.

Precision (float16, bfloat16, float32), diffusion steps and guidance scale trade speed against quality.

Needs roughly 5–7 GB of VRAM, wants flash-attn 2 for full speed, and does not truly stream — long replies are chunked. The model is very new.

On this page