Voxta docs

SenseVoice

Local speech-to-text that also detects emotion and audio events.

SenseVoice is an open-source neural STT that reports more than words: it detects emotion and audio events — laughter, music, applause. It runs locally.

Setup

Add the service

Manage Services → + Add Services → SenseVoice → Add. The model downloads on first use.

Pick a model

ModelCovers
SenseVoiceSmall5 languages, plus emotions and audio events.
Fun-ASR NanoChinese, English, Japanese and dialects.
Fun-ASR MLT Nano31 languages.

Set the language

Leave it on auto-detect, or pin it if you always speak the same one — pinning is more reliable.

Handles one transcription at a time, and needs some GPU memory.

On this page