Voxta docs

ExLlamaV3

Fast local LLM inference on NVIDIA GPUs, for EXL3 quantized models. The ExLlama backend Voxta ships.

ExLlamaV3 is the successor to ExLlamaV2 — a fast inference library for running local LLMs on NVIDIA GPUs. Voxta installs the runtime automatically on first use.

Admin-install only. It replaces ExLlamaV2, which is being removed after Voxta 1.11.

Setup

Add the service

Manage Modules → Add Modules → ExLlamaV3 → Install. Voxta installs the Python runtime and dependencies on first use (watch the Terminal for progress).

Pick a model

In the ExLlamaV3 config:

  • Model — full path to a model file, or HuggingFace identifier. Voxta downloads HuggingFace models automatically.
  • Models Directory — where downloaded models live. Default Data/HuggingFace.

Tune presets

Presets cover reply, action inference and summarization. Defaults are reasonable; only tune if you know what you're changing.

Decisions

When ExLlamaV3 runs Action Inference, it can also answer Decisions with the same loaded model, with no extra settings.

Coming from ExLlamaV2

ExLlamaV2 is being removed after Voxta 1.11. Add ExLlamaV3, point it at an EXL3 version of your model, and switch your characters over. A server that still has ExLlamaV2 configured shows it as unknown after the update, rather than failing to start.

On this page