TTS Models¶
Tryll synthesizes speech through two engines:
SherpaOnnxhosts Supertonic 3 and Pocket TTS.LlamaCpphosts Orpheus and Maya1, plus a shared SNAC 24 kHz decoder.
This page is the authoritative list of which voices the Tryll server ships, which families it can load, and why the rest are out of reach.
Short version: Tryll ships an eSpeak-free Sherpa-ONNX build so the server
can be redistributed inside a proprietary game. That choice makes two Sherpa
families available — Supertonic 3 and Pocket TTS — and puts the rest
of the Sherpa zoo out of reach. Separately, two llama.cpp SNAC voices —
Orpheus 3B FT (Q4_K_M) and Maya1 3B (i1-Q4_K_M) — are public catalog
entries. Identity and delivery use tts_voice_description /
tts_delivery_instruction; [emotion] is not supported syntax.
Heavier SNAC quants (Orpheus 3B FT (Q8), Maya1 3B (Q8),
Maya1 3B (Q4_K_M)) stay audience: "internal" and are not a public
default.
For the catalog file format these models are declared in, see Model Management. For the node parameters that select a voice at runtime, see TTS and Voice Output.
Shipped voices¶
The four entries below are in the default models.json, are downloaded on
demand, and are selected with the tts_model_name parameter on
GenerateAndSpeak or Speak.
The engine is taken from the catalog — you do not pick Sherpa vs. llama.cpp
in Project Settings. See
Choose where models run.
| Catalog name | Engine | Family | Voices | Languages | Voice cloning | Weights license |
|---|---|---|---|---|---|---|
Supertonic 3 (int8) |
sherpa-onnx |
supertonic |
10 built-in styles (speaker_id 0–9) |
31, via tts_lang |
No | OpenRAIL-M (attribution + use restrictions) |
Pocket TTS (int8) |
sherpa-onnx |
pocket |
Single-speaker (speaker_id must stay 0) |
English | Yes — zero-shot from a reference WAV (tts_voice) |
CC-BY-4.0 (attribution) |
Orpheus 3B FT (Q4_K_M) |
llama-cpp |
orpheus |
8 named presets (speaker_id 0–7: Tara, Leah, Jess, Leo, Dan, Mia, Zac, Zoe) |
English | No | Llama 3 Community License |
Maya1 3B (i1-Q4_K_M) |
llama-cpp |
maya1 |
2 catalog presets (speaker_id 0–1) or a natural-language description |
English | No — describe the voice instead | Apache-2.0 |
Both SNAC voices pull a shared SNAC 24 kHz decoder (~50 MB) the first time either parent is downloaded. Deleting one parent does not delete the decoder.
Pick Supertonic 3 when you want multiple distinct voices, non-English
speech, or the lowest latency. Pick Pocket TTS when a character needs a
specific voice you can supply as an audio sample — see
Clone a Voice. Pick Orpheus for named
English speakers plus inline tags such as <laugh>. Pick Maya1 when you
want to brief a voice in natural language (tts_voice_description) and steer
delivery per call (tts_delivery_instruction).
Attribution is required
None of these weights are public domain.
- Supertonic 3 is OpenRAIL-M (attribution plus use-based restrictions — no impersonation or harmful use).
- Pocket TTS is CC-BY-4.0 (attribution).
- Orpheus weights are a Llama 3 derivative and inherit the Llama 3 Community License; the Orpheus inference code is Apache-2.0.
- Maya1 weights are Apache-2.0.
- The shared SNAC decoder is MIT (
hubertsiuzdak/snac).
If you ship a voice in a game, include the model attribution in your credits or third-party notices. The engine code itself (Sherpa-ONNX, ONNX Runtime, llama.cpp) is permissively licensed and imposes no such obligation.
Voice controls by family¶
These are four different knobs. Do not encode identity or delivery as a
spoken [emotion] prefix — that syntax is not supported. Model-native
<tag> tokens that a catalog lists may reach the synthesizer; unknown
tags are stripped from the audio copy only and do not rewrite history.
| Control | Supertonic 3 | Pocket TTS | Orpheus | Maya1 |
|---|---|---|---|---|
speaker_id |
0–9 (style index) |
Must stay 0 |
0–7 (named preset) |
0–1 (catalog preset) |
tts_voice |
Ignored | Reference WAV | Ignored | Ignored |
tts_voice_description |
Unsupported | Unsupported | Unsupported | Stable identity (overrides speaker_id when set) |
tts_delivery_instruction |
Unsupported | Unsupported | Unsupported | Sustained delivery for this call |
tts_lang |
31 languages | Ignored (English) | Ignored (English) | Ignored (English) |
speed |
0.1–3.0 |
0.1–3.0 |
1.0 only |
1.0 only |
Inline <tag> tokens |
— | — | <laugh>, <chuckle>, <sigh>, <cough>, <sniffle>, <groan>, <yawn>, <gasp> |
<laugh>, <giggle>, <sigh>, <gasp>, <angry>, <whisper>, <cry>, <scream> |
A control the loaded model does not advertise fails rather than being
silently ignored. Orpheus and Maya1 reject a speed other than 1.0.
Maya1 catalog presets are starting points, not the only voices:
0— Male, 30s, American1— Female, 30s, American
Set tts_voice_description to something like "Realistic female voice in
the 30s with an American accent. Normal pitch, warm timbre, conversational
pacing." and the description replaces the preset. Pair it with
tts_delivery_instruction (for example whisper or angry) for
this-call delivery; leave it empty for neutral.
Family support matrix¶
Sherpa-ONNX v1.13.2 exposes seven TTS families. Tryll accepts two of them. llama.cpp SNAC families are a separate path.
| Family | Engine | Upstream examples | Status in Tryll | Why |
|---|---|---|---|---|
| Supertonic | Sherpa-ONNX | Supertonic 3 | ✅ Supported | Own tokenizer (unicode_indexer.bin); no eSpeak |
| Sherpa-ONNX | Pocket TTS (Kyutai) | ✅ Supported | SentencePiece vocabulary; no eSpeak | |
| Orpheus | llama.cpp | Orpheus 3B FT | ✅ Supported | SNAC tokens via llama.cpp; shared ONNX decoder |
| Maya1 | llama.cpp | Maya1 3B | ✅ Supported | Same SNAC engine as Orpheus; description + delivery controls |
| Kokoro | Sherpa-ONNX | Kokoro v0.19, v1.0 | ❌ Unavailable | Requires eSpeak NG — see below |
| Kitten | Sherpa-ONNX | KittenTTS nano/micro/mini | ❌ Unavailable | Requires eSpeak NG — see below |
| VITS (incl. Piper) | Sherpa-ONNX | Piper voices, MMS, LJSpeech | ❌ Unavailable | Piper-style bundles require eSpeak NG. Lexicon-based VITS bundles do not, but the family is not implemented in Tryll. |
| Matcha | Sherpa-ONNX | Matcha-Icefall (en, zh) | ❌ Unavailable | English bundles require eSpeak NG. The Chinese lexicon path does not, but the family is not implemented in Tryll. |
| ZipVoice | Sherpa-ONNX | ZipVoice, ZipVoice-Distill | ❌ Unavailable | No eSpeak requirement, but the family is not implemented in Tryll. |
Note the two distinct reasons for the Sherpa gaps. Kokoro and Kitten are blocked by licensing — they cannot be enabled without changing what Tryll redistributes. VITS, Matcha, and ZipVoice are simply not implemented; nothing in principle prevents their eSpeak-free bundles from being added later.
Why eSpeak-dependent families are unavailable¶
Sherpa-ONNX performs grapheme-to-phoneme conversion for several families with
eSpeak NG, which is licensed
GPLv3-or-later. Linking it into tryll_server.exe would make the server a
combined work under GPLv3 — and because the server payload is bundled inside
your game build, that obligation would follow you, your studio, and your
publisher.
Tryll therefore builds Sherpa-ONNX with eSpeak NG compiled out entirely.
The shipped sherpa-onnx-c-api.dll contains no eSpeak code and the payload
contains no espeak-ng-data. The upside is that nothing in the Tryll server
imposes a copyleft obligation on your game. The cost is that any voice whose
front-end calls eSpeak cannot run, regardless of how the model's own weights
are licensed.
This constraint applies to Sherpa-ONNX families only. llama.cpp SNAC voices do not use eSpeak.
A permissive model license is not enough
KittenTTS is the classic trap: its weights are Apache-2.0, which looks
unrestricted, but the model is trained on eSpeak phoneme IDs and
Sherpa-ONNX's Kitten loader requires a populated espeak-ng-data
directory. The permissive headline license does not make the runtime
dependency go away. The same reasoning excludes Kokoro, whose loader also
hard-requires that directory.
What happens if you try an unsupported family¶
A Sherpa-ONNX models.json variant whose tts_family is anything other than
supertonic or pocket fails when the model is loaded — that is, at
CreateAgent time for a graph referencing it, not silently at startup.
llama.cpp SNAC families (orpheus, maya1) load when their catalog entries
and the shared decoder asset are present.
tts_family value |
Result |
|---|---|
supertonic, pocket |
Loads normally on Sherpa-ONNX |
orpheus, maya1 |
Loads normally on llama.cpp + the shared SNAC decoder |
vits, piper, kokoro, or empty |
Rejected with an explicit eSpeak/GPLv3 message |
Any other Sherpa-ONNX value (kitten, matcha, zipvoice, typos) |
Rejected as an unknown Sherpa-ONNX family |
Those Sherpa-ONNX cases surface as agent-creation failures; see
Error Codes. An empty Sherpa-ONNX tts_family is rejected
rather than defaulting, because the historical default was VITS.
Adding your own voice¶
You can add TTS entries to models.json — either a HuggingFace repo or a
local path — using the schema in
Model Management. For
Sherpa-ONNX the constraint is the family, not the source: the bundle must be
a Supertonic or Pocket bundle, with the tts_files key set that family
expects. llama.cpp SNAC voices use tts_family orpheus or maya1 plus a
shared decoder asset (dependencies.snac_decoder).
Sherpa bundles for both supported families are published under
csukuangfj on HuggingFace alongside the
rest of the Sherpa-ONNX model zoo. A Sherpa bundle from any other family will
be rejected at load time even though the catalog entry itself parses.
Related¶
- TTS and Voice Output — the streaming pipeline, audio format, and every voice parameter.
- Add Voice Output to an Agent — task walkthrough.
- Clone a Voice from an Audio Sample — Pocket TTS reference-voice workflow.
- Model Management —
models.jsonschema, download and load lifecycle. GenerateAndSpeakandSpeak— the nodes that produce audio.