Voxta docs
DevelopersModules

Service types

Every service type a Voxta module can implement, from TextGen and TextToSpeech to SpeechToText and ChatAugmentations, and how the score ranks them.

When you register a module, you declare what kind(s) of service it provides via the Supports dictionary. This page lists every supported value and what to use it for.

The Supports dictionary

builder.Register(new ServiceDefinition
{
    // ...
    Supports = new()
    {
        ChatAugmentations = new(ServiceDefinitionCategoryScore.High),
    }
});

Each entry says: "I can act as this kind of service, with this confidence level." A single module can support multiple service types — for example, a vision provider that does both image generation and image understanding declares both.

The score

ServiceDefinitionCategoryScore is a relative weight Voxta uses to rank candidate services for a slot. Values: Low, Medium, High. Use High if your module is purpose-built for this service type, Low if it's a side capability.

Catalog

ServiceTypes valueWhat it doesTypical use caseRegistration
TextGenGenerates assistant replies from prompts.OpenAI, Anthropic, llama.cpp, KoboldCpp, Ollama.AddTextGenService
ActionInferenceDecides which semantic action the user is triggering (run alongside TextGen).Usually delegated to a fast LLM. Configure separately from main chat LLM for cost / speed.AddActionInferenceService
DecisionsPicks one of a fixed set of options and returns it with a probability; writes no text.Who speaks next, argument-less actions and tool call review, in place of ActionInference. Declared with DecisionsSupport; set RunsOnActionInference when it answers with the model your ActionInference service loaded. See Decisions.AddDecisionsService
SummarizationCondenses long chat histories into memory.A cheap LLM endpoint kept separate from the main chat LLM.AddSummarizationService
TextToSpeechRenders text to audio.ElevenLabs, Azure TTS, Coqui, Piper, SAPI.AddTextToSpeechService
SpeechToTextTranscribes microphone audio.Vosk, Whisper, Azure Speech, Google STT.AddSpeechToTextService
AudioInputRaw microphone capture.Custom audio backends, virtual mics, network audio.AddAudioInputService
AudioOutputSpeaker playback.Custom output routing, network audio.AddAudioOutputService
AudioPipelineAudio processing (noise suppression, gain, format conversion).DSP plugins, sample-rate converters.AddAudioPipelineService
WakeWordListens for a hotword to start a chat turn.Picovoice Porcupine, openWakeWord.AddWakeWordService
VisionCaptureCaptures images from a source (screen, webcam, game window).Screen grabbers, OBS bridges.AddVisionCaptureService
ComputerVisionUnderstands images — describes scenes, identifies objects, reads text.GPT-4o vision, Claude vision, local vision models.AddComputerVisionService
ChatAugmentationsInjects context into the chat and exposes semantic actions the LLM can trigger.In-process game and app integrations — anything you can drive from .NET without needing another runtime. Elite Dangerous COVAS is the reference example.AddChatAugmentationsService
MemoryLong-term knowledge store the chat can query.Vector databases, RAG backends.AddMemoryProviderService
ImageGenGenerates images from text.Stable Diffusion (local), DALL·E, Imagen.AddImageGenService
AnimationsDrives avatar / character motion from text.Currently the HY-Motion experimental motion service.AddAnimationGenService
VideoGenGenerates video from a prompt.Text-to-video and image-to-video backends.AddVideoGenService
PortraitAnimationAnimates a character's portrait from the voice it speaks with.Talking-portrait models. See Portrait Animation.AddPortraitAnimationService
ImageProcessingTransforms an existing image.Background removal on generated portraits.AddImageProcessingService
AudioTranscriptionTranscribes audio files, as opposed to a live mic.Voice cloning, where a reference clip needs its transcript. Most providers have one or the other: OpenAI transcribes a file and can't stream, Vosk streams and has no file endpoint.AddAudioTranscriptionService
TurnDetectionDecides when the user has finished speaking.Semantic end-of-utterance models, past what a silence timer can tell.AddTurnDetectionService
AssistantThe LLM behind Voxta's own creator tools.Shares its settings with TextGen. Point it at something fast.None; runs on your TextGen service
CloudStorageStores resource backups in a cloud account.Google Drive. See Cloud storage.AddCloudStorageService
ContentHubBrowses and installs characters, scenarios and collections from a hub.Online content catalogs. See Content Hubs.AddContentHubService
ChatBridgeRuns chats of its own, answering somewhere outside Voxta.The Discord bot. Not a service a chat selects.None; register your IChatBridge in builder.Services

Choosing the right type

A few rules of thumb:

  • Wrapping an external AI API? Pick the corresponding category — TextGen for an LLM, TextToSpeech for a voice, etc. The framework gives you a ServiceBase to inherit from and an interface to implement.
  • Adding context to chats? ChatAugmentations. This is the most flexible category — you get callbacks for events, can inject text into the prompt, and can register semantic actions the LLM calls.
  • Capturing the screen or webcam? VisionCapture. Pair it with a ComputerVision service to describe what was captured.
  • Bridging a game or app? ChatAugmentations is usually the right home, since the integration's job is to feed game state into chat and translate chat output into game actions.

Registering the implementation

After declaring Supports, you tell the builder which class implements each type:

builder.AddTextGenService<MyTextGen>("my-service");
builder.AddChatAugmentationsService<MyAugmentations>("my-service");
builder.AddTextToSpeechService<MyTts>("my-service");
// etc.

The string is your ServiceName from the ServiceDefinition. Most service types register with Add{ServiceType}Service<TImpl>(string serviceName); the catalog lists the exact method for each. Your class inherits the appropriate base (ServiceBase) and implements the matching interface (ITextGenService, IChatAugmentationsService, ITextToSpeechService, …).

Examples in the wild

  • ChatAugmentations — voxta-module-elite-dangerous (open-source reference example)
  • ComputerVision — built-in Cloud / OpenAI vision modules
  • TextGen — built-in LLM providers (OpenAI, Anthropic, llama.cpp, Ollama, ...)

Not modules — these are external integrations that connect to Voxta over the WebSocket API rather than running in-process:

See Developers for when a module fits and when an integration does.

On this page