Client API
Write a Voxta client: connect to the SignalR WebSocket hub, authenticate, start a chat and follow what the character says and does.
A client is a separate process that talks to the Voxta server over its SignalR / WebSocket hub. Voxta Talk, the VAM plugin, the Minecraft companion and Voxy are all clients of the same API.
Write a client when your code has to run where Voxta can't — a game's scripting engine, a Node bot, a browser, an embedded device. Write a module instead when it can live as .NET inside the server.
Connecting
The hub is at /hub on the server (http://127.0.0.1:5384 by default), over the WebSockets transport:
const connection = new signalR.HubConnectionBuilder()
.withUrl('http://127.0.0.1:5384/hub', {
skipNegotiation: true,
transport: signalR.HttpTransportType.WebSockets,
})
.withAutomaticReconnect()
.build();Everything then flows through two hub methods:
| Direction | Method |
|---|---|
| Client → server | Invoke SendMessage with one message object. |
| Server → client | Handle ReceiveMessage. |
Every message is discriminated by a $type field.
With an API key
An app identifies itself with an API key. Hand it to SignalR through accessTokenFactory:
.withUrl('http://127.0.0.1:5384/hub', {
skipNegotiation: true,
transport: signalR.HttpTransportType.WebSockets,
accessTokenFactory: () => apiKey,
})Outside a browser, SignalR sends it as an Authorization: Bearer header. A browser cannot set headers on a WebSocket, so there it goes in the URL as access_token, and URLs end up in logs and history. In a browser, exchange the key for a ticket instead:
async function getTicket() {
const response = await fetch('http://127.0.0.1:5384/api/auth/websocket-ticket', {
method: 'POST',
headers: { Authorization: `Bearer ${apiKey}` },
});
return (await response.json()).ticket;
}
// accessTokenFactory: getTicket,A ticket opens one connection and expires 30 seconds after it is issued. It carries the key's scopes and nothing more. SignalR calls accessTokenFactory on every connect and reconnect, so each connection gets a fresh one.
The audio input socket, /ws/audio/input/stream, takes the same access_token parameter: a key, or better, a ticket.
Getting a key with a code
An app can ask for a key instead of having one pasted. Request a code with the scopes you need:
const post = (path, body) => fetch(`http://127.0.0.1:5384${path}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(body),
});
const { user_code, verification_url, device_code } = await (await post('/api/device/code', {
client_id: 'voxta',
scope: 'chat tools.write',
label: 'MyGame',
})).json();
// Show user_code and verification_url, then poll.
let response;
do {
await new Promise(r => setTimeout(r, 1000));
response = await post('/api/device/poll', { device_code });
} while (response.status === 204); // not approved yet
if (!response.ok) throw new Error('Code expired');
const { token, scope } = await response.json();The key is returned once; store it. scope is what was granted. The user can untick scopes, so it may be less than you asked for.
Without scope, you get role:app, which has no tool scope. To answer tool prompts, ask for tools.write, or tools.dangerous for destructive tools and the shell. tools.dangerous needs an administrator account with a password, and only works from the server's machine.
In .NET, with Voxta.Client:
var code = await deviceAuthorization.GetDeviceCodeAsync("MyGame", ct,
[ApiKeyScopes.Chat, ApiKeyScopes.ToolsWrite]);
var grant = await deviceAuthorization.WaitForGrantAsync(code, ct);Authenticating
authenticate must be the first message, and again after a reconnect:
connection.send('SendMessage', {
$type: 'authenticate',
client: 'MyGame',
clientVersion: '1.0.0',
scope: ['role:app'],
capabilities: {},
});scope is the role the connection asks for. role:app is still what an app sends in 1.11: it starts and reads chats, sends messages, edits resources and hears broadcasts. role:provider (messages, context and broadcasts, but no starting chats) and role:inspector (read-only, audio included) are narrower. Connecting with an API key, you get only what both the role and the key's scopes allow, and the connection is refused when that is nothing.
capabilities tells the server what your client can do with audio, so it knows whether to generate speech for you and in what form. Declaring nothing gets URL-based audio output and no microphone.
Starting a chat
Once authenticated, startChat (or resumeChat for an existing one) opens a session. From there the server streams what the character is doing — replyChunk as the text and audio are produced, speechPlaybackStart / speechPlaybackComplete to be told and to report back, action and appTrigger for what the scenario wants your app to do.
A full message catalogue is not published yet. The shipped integrations below are open source and are the working reference in the meantime.