Skip to content

v0.7.0

This release makes an agent's dialog persistable and its language, placement, and voice something you choose. A full-fidelity dialog snapshot drops into the save file you already own, one session language drives speech, recognition, content files, and prompt directives together, and every model can be told whether its weights belong in RAM or VRAM. A pause can now time out instead of soft-locking an agent, two expressive SNAC voices (Orpheus and Maya1) join the public catalog with natural-language voice direction, and tool calling gains measured model rankings, prompt recipes, and in-editor testing. Memory stops being a guess — a Memory window in both editors reports the server's live RAM and VRAM with a confidence tier on every figure — and LLM intent classification graduates: described labels instead of a labelled corpus, a visible Intent block in the turn inspector, and Lab checks that assert the label a turn resolved.

The four engine dropdowns are gone, components now require a Workflow Asset, and the schema changed — see Breaking changes and behavior changes before upgrading.

Highlights

  • Save and load agent dialog. Export a full-fidelity snapshot while an agent is idle, store the canonical TDLG bytes inside your own save container, and import it into a fresh session with Strict compatibility. Tryll still writes no save files. See Save and load agent state.
  • One session language for the whole stack. Set a BCP-47 locale and speech output, speech-recognition hints, per-language content files, and the {{language}} prompt keys all follow it — changeable mid-session without tearing the session down. See Ship an agent in another language.
  • Choose where each model runs. Per-model RAM / VRAM placement, authored in the Model Manager or passed when you load a model yourself. A hint by design: a machine that cannot honour it loads the model the other way and reports what it actually did. See Choose where models run.
  • Engines come from the catalog. The four engine dropdowns are gone from Project Settings in Unity and Unreal. The catalog records which engine runs each model — which also fixes model listing showing no STT, TTS, or embedding models in a project that never touched those settings.
  • Pauses can time out. pause_timeout_ms gives a Pause a deadline and timeout_exit lets the character answer anyway, instead of an unanswered pause leaving the agent busy forever. The server also warns when a pause outlives its stall threshold.
  • Two expressive voices, plus voice direction. Orpheus 3B FT (Q4_K_M) and Maya1 3B (i1-Q4_K_M) are public catalog voices, and tts_voice_description / tts_delivery_instruction let you brief a voice's identity and its delivery in natural language. See TTS Models.
  • Fit a whole cast on one GPU. A guide to the two levers that actually decide whether five NPC agents fit in limited VRAM — sizing every LLM node's context_size, then evicting the agents nobody is talking to. See Run many agents on a small GPU.
  • Tool calling, measured. Exact-accuracy tables for the catalog on game-shaped suites and a BFCL slice, schema and prompt recipes that survived those sweeps, and a way to exercise a ToolCall workflow in the editor without Play/PIE. See Compare tool-call models and Design a tool-call prompt.
  • See what the server is actually using. A Memory window in Unity and Unreal reports live RAM and VRAM for the whole server process — resident models, per-agent KV, and the part nobody can attribute — with a confidence tier on every figure, so an estimate never reads as a measurement. See Inspect server memory.
  • Intent classification without a corpus. Describe each intent in a sentence and ClassifyIntentLLM routes the line off its own first-token logprobs — no embedding model, no labelled knowledge base, nothing to ship — with a measurement-backed prompt guide, an Intent block in the turn inspector, and Dialog Lab checks that assert the accepted label. See Classify intent without a knowledge base and Design an intent-classifier prompt.
  • Classifier and tool-call nodes size their own window. ClassifyIntentLLM and ToolCall now default to 2048 tokens instead of inheriting a generator-sized 8192 — the cheapest VRAM saving in the release, and it needs no authoring.

Save and load

  • ExportDialog / ImportDialog on every client. Both are strict idle-only — no turn running, not paused, no KV-cache operation in flight. Export returns a DialogSnapshot; ToFlatBufferBytes / FromFlatBufferBytes (encode_flatbuffer / decode_flatbuffer in Python) produce the canonical TDLG bytes. Import replaces the whole dialog and is atomic: a failed import leaves the existing dialog untouched. Unreal exposes this in C++ only — the snapshot is not Blueprint-exposed. See Export and import dialog.
  • The snapshot is dialog, and only dialog. Not agent variables, not UI transcripts, not KV-cache bytes, not agent IDs. The save-slot pattern — quiesce every agent, snapshot each one alongside variables and game state, recreate the agents in a fresh session, import, then warm the ones you actually need — is written up end to end in Save and load agent state. Import never prefills; call PrefillKvCache yourself during the loading screen.
  • Strict by default. Import checks the snapshot against the agent's create-time graph and refuses a mismatch; AllowGraphMismatch opts out. Snapshots that would overflow the 1 MiB frame cap fail rather than being truncated.
  • Unity coroutine wrappers. TryllAgentCoroutines wraps export, import, variable sync/flush, and KV prefill: it retries AgentBusy at a polite interval and bounds every wait with a wall-clock deadline, so a dropped connection fails your save UI with Timeout instead of hanging it.

One language for the session

  • locale on CreateSession, and SetSessionLocale after it. A BCP-47 tag (de, pt-BR, zh-Hans-CN) that becomes the default tts_lang on every GenerateAndSpeak / Speak, the language hint handed to multilingual STT models, the per-language file lookup, and the locale / language / language_code Mustache keys. Empty leaves every one of those at its previous behavior; a malformed tag is rejected with 2004 InvalidLocale. Unity and Unreal add a Project Settings → Tryll Client → Language picker that lists only what your registered models can serve.
  • Translations live beside the file you link. A loc/<language>/ folder next to a linked storage file supplies the translated version — deathonset/loc/pt-BR/lore.json, then loc/pt/, then the file you linked. You always link the default file; a folder with no loc/ is language-neutral; and fallback is per file, so a partially translated game works. This covers knowledge bases and their sidecars, canned responses, guardrail patterns, intent maps, hotwords, and reference-voice WAVs.
  • {{language}}, {{locale}}, {{language_code}} in templates. Available wherever Mustache rendering is — Generate, GenerateAndSpeak, ToolCall, Transform. The section collapses to nothing when no language is set, so one template works localized and unlocalized.
  • A helper for the offerable set. ListModels reports a languages list per model, and TryllLanguages (Unity) / UTryllLanguages (Unreal, Blueprint-callable) intersect them across your pipeline. HasLanguageInfo tells "no overlap" apart from "the catalog didn't say" before you show a player an empty menu.
  • A mid-session switch is not retroactive. Speech output, the template keys, and the next file load follow it from each agent's next turn; a knowledge base, canned lines, or classifier prompt already bound to a node does not — those are structural. Send the change between turns, not mid-sentence.
  • tts_lang resolution is now layered: the node's own value, then the session locale, then the model's catalog default. Set the node param only to override one character — a French witness in a German playthrough.

Models: placement, engines, and GPU backend

  • memory_placement per model. Auto / RAM / VRAM, authored on a registered model in the Model Manager or passed when you load a model yourself. The reason to reach for it: your renderer and Tryll's language model compete for the same VRAM, and moving a large model to RAM trades generation speed for graphics headroom. It is a hint on purpose — you author it on your machine and ship it to someone else's.
  • load_model() reports what happened. The Python client returns a result carrying placement and backend instead of None, so a VRAM request that could not be honoured is visible rather than guessed at. Purely additive — existing calls that ignore the return value are unaffected. Unity and Unreal log a warning naming both placements and the backend; their completion events keep their (model name, success) signatures, because a model cannot be moved after it loads.
  • The session declares engines as a hint, not a gate. CreateSession takes an optional list of the (role, engine) pairs you expect to touch, replacing the four engine arguments. It may be empty, incomplete, or wrong, and loading a model it never named works normally. What it buys is timing: engine start-up is paid at session creation instead of on first use — worth asking for when a backend is slow to initialise. See Choose where models run.
  • Pick the llama.cpp GPU API in Project Settings. Tryll Client → Llama.cpp backend (Vulkan default, or Auto), passed to an auto-launched server as --llama-backend. A Vulkan player build skips the CUDA DLLs; Auto copies them and prefers CUDA on NVIDIA. Standalone / C++ / Python set engines.llama_cpp.backend in server-config.json (committed default "vulkan") or pass the flag. Players need neither the CUDA Toolkit nor those DLLs on PATH, and speech and embedding stay on the CPU package. See Auto-launch the server and Server configuration.
  • Shared model assets. The SNAC 24 kHz decoder is the first: it has its own download record, is pulled once for whichever parent needs it, and survives deleting a parent model. ListModels reports ModelStatus.Incomplete when a parent is waiting on a shared asset, with dependency details. The catalog also gained dependencies and structured tts_capabilities (voice presets, inline tags). See Model Management.
  • Fixed: telemetry reported the wrong GPU backend. gpu_backend_active was read from build-time flags that are never set in our builds, so every install reported cpu — including machines where Vulkan was serving every token. It now reads the live device registry. If you have been comparing performance data across machines, the backend field in older data is not trustworthy.

Fitting a cast on one GPU

  • Run many agents on a small GPU — the two levers, in the order worth pulling. A KV cache is allocated at its full context_size the moment the context is created, so the cost is deterministic, short replies save nothing, and a sizing change is measurable in seconds rather than in playtests. Then evict the agents nobody is talking to. Worked through with numbers from a five-agent, nine-context demo on a 16 GB card.
  • ClassifyIntentLLM and ToolCall now default to a 2048-token window. Left at 0, they no longer inherit the model variant's context_size or the server's default_n_ctx (commonly 8192) — the two nodes whose prompts are small got the small default. Generate and GenerateAndSpeak are unchanged, so budget per node: a classifier at 2048 is a small slice of an 8192 generator. See Estimate memory footprint.
  • A per-node context table, for when the window itself is the problem. Generate, GenerateAndSpeak, ClassifyIntentLLM and ToolCall take an optional structural context table: kv_cache_type per context (overriding the catalog's model-load default, with Inherit keeping it), a veto-only offload_kqv, and escape-hatch batch sizes. Leave it unset unless you are tuning a large window — at a right-sized 2048 a q4_0 KV override is usually not a measurable VRAM win, because the dtype delta only shows up around 8192. See Models and inference engines.
  • The sizing trap is silent. When a prompt exceeds its window, the token-budget projection trims the oldest turns and keeps working — so an under-sized classifier quietly loses the history it needs to resolve "And that one?". Leave headroom, and re-check that intents still fire after tightening one.
  • An immersion guard costs two extra contexts per agent. Build an immersion guard now says so up front and points at the sizing guide before you add one to a project already close to its VRAM budget.
  • Memory estimates cover SNAC voices. llama.cpp SNAC voices are GGUF backbones that follow the same RAM/VRAM placement as a language model, plus a ~50 MB CPU decoder — unlike Sherpa STT/TTS, which stay RAM-only. See Estimate memory footprint.

Memory inspection

  • A Memory window in both editors. Window → Tryll → Memory in Unity, Window → Tryll → Tryll Memory in Unreal: resident models, per-agent KV, and the unattributed remainder, on demand. Refresh only — no poll, no graph. It reuses the shared editor session and auto-launches the bundled server the way Model Manager does, needs no agent, and does not wait for registered models to download. Closing it releases the lease without killing the server. See Inspect server memory.
  • The snapshot is server-wide, not per-connection. Every session on that process is in the numbers — Chat, Model Manager, Play-in-Editor, a Python client, your QA harness. A row you do not recognise is not a bug in the window. That is the opposite of the Agent Log, which sees only your own connection.
  • Every figure carries a confidence tier. Measured (the engine or the OS said so), Derived (computed from measured inputs), Estimated (a calibrated band, rendered with a leading ~ and an estimate basis on hover). The column is always visible on purpose: an estimate must never look like a measurement.
  • What the tables actually answer. Devices split into ours / other / free, and other is the editor and the compositor rather than a Tryll leak. KV is listed per agent, sorted background-first, with evicted and never run called out separately — background KV is the reclaim signal, because nobody is waiting on it. Context used is headroom to the limit, not memory growth: the buffer is allocated at the full context_size the moment the context is created. And memory-mapped model files are resident pages Tryll never allocated, so a VRAM-placed model often surfaces under unattributed instead of its own RAM column — the window says so when that happens.
  • GetMemoryConsumption on every client. TryllClient::GetMemoryConsumption (C++), client.get_memory_consumption() (Python), RequestGetMemoryConsumptionAsync (Unity), and UTryllSubsystem::RequestGetMemoryConsumption (Unreal C++, with a Blueprint completion event to bind) all return the same snapshot. Session-scoped — it must follow CreateSession — but the numbers cover the whole process, and it never fails with AgentBusy.
  • The server monitor gained the matching panel. Same collector, same numbers, over GET /api/memory; the browser view keeps its sparkline and stacked bars, which the editor window does not have. monitor.memory_tick_interval_ms sets the live sampling interval — 0 disables periodic sampling and the panel still loads on demand. See Server configuration.

Voice output

  • Orpheus and Maya1 are public catalog voices. Orpheus 3B FT (Q4_K_M) gives eight named English speakers (speaker_id 0–7) plus inline tags such as <laugh>; Maya1 3B (i1-Q4_K_M) takes a natural-language voice description, with two catalog presets as a starting point. Both download through the normal Model Manager, including the shared SNAC decoder. Heavier quants stay internal.
  • tts_voice_description and tts_delivery_instruction. Two new mutable params on GenerateAndSpeak and Speak: a stable speaker identity (age, accent, pitch, timbre) and a sustained delivery direction for this synthesis call. They are independent of tts_voice (a reference WAV) and speaker_id (a preset index) — four distinct knobs, and square-bracket [emotion] text is not a supported control. Model-native <tag> tokens that a catalog lists may reach the synthesizer; unknown tags are stripped from the audio copy only and never rewrite history.
  • A control the model does not advertise now fails instead of being silently ignored — Orpheus and Maya1 reject a speed other than 1.0. The full per-family control matrix, and the attribution each set of weights requires, are in TTS Models.
  • Unity: ask the speaker whether it is still talking. TryllSpeaker gains HasPendingAudio (true while decoded audio is still queued) and StopPlayback(). Combined with TurnComplete, HasPendingAudio is the signal a hands-free voice loop needs before re-opening the microphone, so the mic does not transcribe the tail of the character's own answer. AudioSource.isPlaying is not that signal — the streaming clip loops silence between turns. See TryllSpeaker.

Turn control

  • pause_timeout_ms on Pause. 0 still waits forever. A finite budget makes the failure recoverable: with timeout_exit authored the graph continues from that exit and the turn ends normally; without it the turn ends with TurnStatus.PauseTimedOut, keeping the partial interaction so the conversation stays intact. Prefer authoring a timeout_exit for anything player-facing, so the character says something. Both params are structural.
  • The deadline is server wall-clock. A minimised or alt-tabbed game stops pumping its main thread, so the deadline can elapse while your game is suspended. Make the timeout branch harmless — "shrug and answer without the tool" — rather than something that only makes sense the instant it fires.
  • pause_timeout_ms on ToolCall covers pauses that node triggers, inheriting the server's workflow.tool_call_pause_timeout_ms when left at 0. There is no timeout exit on this node — expiry ends the turn.
  • An unresumed pause is loud before it is fatal. The server logs a warning naming the parked node once a pause outlives workflow.pause_stall_warn_ms (15 s by default), and your client logs one immediately when nothing it holds can resume the pause at all.

Tool calling

  • Compare tool-call models — exact accuracy for the catalog across seven game-shaped suites and a BFCL slice, at greedy decoding. The two corpora do not rank models the same way, so the page is built for choosing by job (commands, log queries, parallel calls) rather than for crowning a winner. Models whose tool_call_support is unsupported are rejected at agent creation and are not in the tables.
  • Design a tool-call prompt — measurement-backed schema and prompt recipes. The headline finding: small models fail at restraint, not extraction. Give "the player did not narrow this down" a real enum value, list it first, and mark the parameter required — the value a pressured model reaches for becomes the right answer. Claims are marked Tested / Hint / no effect, so you can skip what did not pay off.
  • Test tool calls in Chat and the Dialog Lab — register a handler on the scene agent component in edit mode, talk to it in Chat, then script Lab runs with canned results or against the live handler. No Play/PIE.
  • Dialog Lab live tools are opt-in. A Lab step with no canned ToolResults entry still means the model must not call that tool, even when the variant's scene agent has a RegisterTool handler. Tick Use live tools on the variant only when the run should exercise it; canned results still win. See Compare dialog variants in the Lab.
  • Diagnostics answer "why didn't it call my tool". parameters.tools[] echoes the signatures the model actually saw, output.tool_calls[] carries each detected call with its raw arguments, output.parse_fallback flags a call the native template parse missed, and engine.chat_turn.has_grammar shows whether require_call actually constrained the sampler. See Tool Call.

Intent classification

  • notify_client is now a three-state enum. Disabled / OnFound / Always on ClassifyIntent and ClassifyIntentLLM, replacing the bool. Always also fires on every not_found path, carrying not_found_reason and the rejected top-1 — which is how you see a deflection that didn't happen.
  • Typed callbacks, and a separate LLM event. ClassifyIntentLLM now emits its own intent_llm_classified event with the full probability vector, subscribed via SetOnIntentLlmClassified / IntentLlmClassified / OnIntentLlmClassified / set_on_intent_llm_classified. Both classification callbacks now take a structured event object instead of a positional argument list. A typed subscriber swallows the generic OnNodeEvent fallback.
  • Richer classification diagnostics. The embedding classifier records ranked topk[] candidates, the winner, and a runner_up with its margin — the number to watch when a classifier is unstable, because a margin near zero means small rewordings will flip the route. The LLM classifier records the label ids in letter order alongside the probability vector.
  • ClassifyIntentLLM without a knowledge base, end to end. Classify intent without a knowledge base rebuilds the gate-guard NPC with one language model and no corpus: a sentence per intent, no embedding pass, nothing on disk. It opens with a table for choosing between the two classifiers — many intents and pre-labelled data still favour embedding ClassifyIntent; a label set you are still designing favours this one, because writing a sentence is a far shorter loop than authoring twenty phrases and rebuilding an index.
  • system_prompt is a Mustache template, and there is no default. It renders against exactly two variables — {{intents_block}} (the letter-prefixed label list) and {{letters}} — and none of the Generate variables (user_message, slot.<name>, var.<name>, instructions, knowledge) resolve there. Leave it empty and the node sends an empty system message: the model is asked for a letter having never been shown what the letters mean, and nothing fails loudly. A template that fails to parse is sent verbatim, literal {{intents_block}} included. Verify the substitution in input.prompt[] after any edit. See Classify Intent (LLM).
  • Design an intent-classifier prompt — claims marked Tested / Hint / no effect, like the NPC and tool-call prompt guides. The mechanism decides what matters: the node reads the logits at the end of the prompt and never generates, so "think step by step", "explain your reasoning" and "answer in JSON" are paid for in prefill and then discarded unread. What pays off is mutually exclusive descriptions — overlap is the top cause of a collapsed margin — an explicit catch-all placed last, because models carry strong, model-specific positional bias over the answer letter and that bias has to be re-checked when you change classifier models, and tuning threshold only after the overlap is fixed.
  • intents_ids and intents_prompt, spelled out. intents_prompt is the text the model reads, rendered as A. <first>, B. <second>, …; intents_ids is never shown to it and only supplies the label attached on a win. They pair positionally — entry 0 is letter A — must have equal counts and 2–26 unique ids, and each letter must be a single token for that model. A comma inside a description silently becomes an extra label, so use dashes or semicolons. All of it is validated at agent creation.
  • An Intent block in the turn inspector. One subsection per classifier execution, with the decision in the header — ask_wing · found, or not_found · below_threshold — and not_found coloured as a warning rather than an error, because it is usually correct behaviour. Embedding classifiers get the funnel line, a winner card or the rejected pick, and a top-k table that dims rows beyond the threshold; LLM classifiers get a probability bar per candidate and the predicted_letter / top_prob / second_prob / margin_result strip, with the winning or rejected label shown in words so you never map a letter by hand. notify_client does not gate the block — a classifier left at Disabled still appears, it just produces no Agent Log event. See Turn Inspector.
  • Dialog Lab can assert the label. Intent is passes when an addressed classifier on that turn accepted the exact id; Intent not found passes only when every addressed classifier took not_found. Leave the node name empty to address every classifier that ran, or name one to scope the check. Labels compare exactly and case-sensitively — they are identifiers, not prose. A missing or mis-typed classifier fails with intent unavailable rather than skipping, so a green suite cannot come from a classifier that never ran, and seeded steps refuse these checks before Run: a seeded turn is injected without running the graph, so there is no classification to assert. See Compare dialog variants in the Lab.
  • A rejected pick is now recorded. Both classifiers add output.classification.rejected_intent — the label the scorer or the search chose when routing rejected it — and the embedding node adds rejected_distance beside it, which is what turns a bare not_found into "it wanted ask_wing, at 0.31, against a 0.25 threshold". The classified text is now input.query on both kinds (output.classification.query is kept for one release), and the LLM probs[] vector is documented as positionally aligned with parameters.intents_ids[], equal length on found.

Diagnostics

  • Turn Diagnostics JSON — a new reference for the whole TurnComplete.debug_info payload: the envelope, the per-node parameters / input / output / engine sections, the turn-level tool-call and interaction snapshots, and the rules for why a given key is absent. Every node reference page now carries a Diagnostics section listing the keys it records.
  • The turn inspector's Raw JSON is verbatim. The editors add no keys, rename none, and drop none — so the reference above describes exactly what you see in the panel and exactly what arrives on the wire. Raw JSON is also the one block the Route breadcrumb's node scope does not narrow; it always shows the whole turn. See Turn Inspector.
  • Two server flags trim the payload. include_engine_diagnostics drops the engine sections and include_interaction_in_diagnostics drops the end-of-turn interaction snapshot — which is why the inspector's tool-call block can show calls with no results.

Server operations

  • Log flushing is configurable. log_flush_on (default "warn") flushes the log file as soon as a message at that level or higher is written, and log_flush_every_ms (default 5000) flushes on a timer so a quiet server does not leave the file minutes behind. A running server's log is now readable without waiting for process exit. See Server configuration.
  • Startup states standalone vs managed idle-exit. The server logs Standalone mode: server will not self-exit or Managed mode: self-exit after <N>s of idle right after logging starts, instead of announcing the mode only when --idle-shutdown-timeout was on the command line.
  • Fixed: a managed server could finish its idle exit and never actually leave. An open monitor event stream kept the HTTP listener from shutting down, so the process stayed alive with both the game port and the monitor port refusing connections — Dialog Lab then failed with "the target machine actively refused it" and would not relaunch. Subscribers are now closed before the listener is joined, and the Unity and Unreal launchers treat an owned process that is alive but no longer accepting as dead and start a fresh one.

Editor fixes

  • Chat honors a cleared Agent Component. Clearing the Chat window's Agent Component (or Voice Component) to None now sticks. Start and Compare in Lab refuse with the usual assign warning instead of silently restoring the previous scene object.
  • Dialog Lab shows the seed it actually uses. Typing 0 into Repeats or base seed still raises the value to 1 (seed 0 means random on the wire). The field now displays 1 instead of keeping the typed 0.

Breaking changes and behavior changes

  • The schema changed — rebuild every client. CreateSession, the TTS nodes, the classification nodes, the pause params, and the LLM nodes' new context table all changed shape this release, and a server-wide memory snapshot call was added. Rebuild your integration against the new schema and upgrade both ends together; mismatched peers are rejected at connection time with 5004 ProtocolVersionMismatch. See Wire Protocol.
  • ClassifyIntentLLM and ToolCall default to a 2048-token context. With context_size left at 0 they previously fell back to the model variant's window, then the server's default_n_ctx — commonly 8192. Most projects simply get the VRAM back; a node whose classifier prompt or tool-schema block genuinely needs more than 2048 tokens must now set context_size explicitly. Generate and GenerateAndSpeak are unchanged, and context_size is structural, so this is a graph edit rather than a runtime one. See Run many agents on a small GPU.
  • ClassifyIntentLLM has no built-in system_prompt, and an empty one fails quietly. Every client leaves the field empty by default and the node passes it through verbatim, so an unauthored classifier sends an empty system message and asks the model for a letter whose meaning it was never shown. Author a template containing {{intents_block}}. See Design an intent-classifier prompt.
  • A component now requires a Workflow Asset. TryllAgentComponent.InlineGraphDescription / InlineVariables and UTryllAgentComponent::InlineGraphDescription / InlineVariables are removed — deprecated in v0.6.0, gone now. Author the graph as a Workflow Asset and assign it; declare variables on the asset. Building a TryllGraphDescription in code and creating an agent from it directly — TryllClient.RequestCreateAgentAsync / UTryllSubsystem::RequestCreateAgent, with no component — remains fully supported. If you still want a component around a code-built graph, put the graph on a transient Workflow Asset (ScriptableObject.CreateInstance<TryllWorkflowAsset>() / NewObject<UTryllWorkflowAsset>()) and assign that. See TryllAgentComponent and Unreal C++ API.
  • The four engine dropdowns are gone, and nothing replaces them. Unity's InferenceEngine / SttEngine / TtsEngine / EmbeddingEngine runtime settings and Unreal's Engine / SttEngine / TtsEngine / EmbeddingEngine project settings are removed; the session's engine set is derived from the models you registered. If you build sessions by hand, the four engine arguments are replaced by one optional engines list. This also fixes a real bug: three of the four defaulted to Mock and ListModels filtered by them, so a project that never touched them saw no STT, TTS, or embedding models in the Model Manager at all. Model listing is now engine-agnostic.
  • notify_client is an enum, not a bool. ClassifyIntent and ClassifyIntentLLM take Disabled / OnFound / Always; the old true becomes OnFound. The IntentClassified callback signature changed to a structured event object on every client, and ClassifyIntentLLM now reports through the separate intent_llm_classified event — code that relied on intent_classified firing for an LLM classifier must subscribe to the new one.
  • Unity RemoveInteractionsFromEnd reports the count. Its completion callback is now Action<uint, TryllError> (was Action<TryllError>), matching AppendInteractions.
  • enable_diagnostics is structural. It is fixed for the agent's lifetime and cannot be turned on for an agent that already exists. The editor tooling handles this for you — Chat enables it on the agent it creates, and the Agent Log holds a lease so agents created while it is open request it too. See Agent Parameters.
  • device_preference is removed from models.json. A catalog that still declares it fails to load, with a message naming the replacement (memory_placement). Every shipped entry said "auto", so it decided nothing; "directml" was accepted there and never implemented, and is gone too.
  • engines.llama_cpp.n_gpu_layers is deprecated and ignored. It applied one value to every model in the process, which cannot be right once different models want different memory. Offload is now decided per model from its placement. A server that still finds the key warns once at startup — remove it.
  • Parakeet TDT 0.6B v2 is replaced by v3. Update model_name in any voice-input config; the v2 entry is no longer in the catalog.
  • Dialog Lab live tools are opt-in. A variant whose scene agent has a RegisterTool handler no longer exercises it implicitly — a step with no canned ToolResults entry asserts that the model must not call that tool. Tick Use live tools on the variant to restore handler execution. See Test tool calls in Chat and the Dialog Lab.
  • loc is a reserved folder name inside storage folders. Do not use loc/ for your own content; it is the per-language lookup root. If your build pipeline filters what gets staged, make sure it keeps **/loc/** — a missing translation folder degrades silently to the default language rather than failing the build. See Ship storage folders for builds.

New error codes

  • 2004 InvalidLocale — CreateSessionRequest.locale (or a later SetSessionLocale) is not a well-formed BCP-47 tag. Expected language[-script][-region]; the message names what was wrong.
  • 3017 ModelLanguageUnsupported — a node would rely on the session language but its model does not support it. Pick a model that covers the language, set the node's own tts_lang, or change the session locale. Not raised when the model declares no languages at all.
  • 3018 DialogSnapshotInvalid — the snapshot is malformed, missing required fields, or has inconsistent tool call / result data. Import leaves the existing dialog unchanged.
  • 3019 DialogFormatUnsupported — the snapshot's format_version is newer than this server, or the file identifier is not TDLG.
  • 3020 DialogGraphMismatch — Strict import: the snapshot's graph signature does not match the agent's create-time graph.
  • 3021 DialogSnapshotTooLarge — the encoded snapshot would overflow the 1 MiB frame cap. Tryll never truncates.
  • 3022 DialogUnsupportedComponent — an unknown component type on export, or an unknown union arm on import.

See Error Codes for the full list.