v0.7.0¶
This release makes an agent's dialog persistable and its language, placement, and voice something you choose. A full-fidelity dialog snapshot drops into the save file you already own, one session language drives speech, recognition, content files, and prompt directives together, and every model can be told whether its weights belong in RAM or VRAM. A pause can now time out instead of soft-locking an agent, two expressive SNAC voices (Orpheus and Maya1) join the public catalog with natural-language voice direction, and tool calling gains measured model rankings, prompt recipes, and in-editor testing. Memory stops being a guess — a Memory window in both editors reports the server's live RAM and VRAM with a confidence tier on every figure — and LLM intent classification graduates: described labels instead of a labelled corpus, a visible Intent block in the turn inspector, and Lab checks that assert the label a turn resolved.
The four engine dropdowns are gone, components now require a Workflow Asset, and the schema changed — see Breaking changes and behavior changes before upgrading.
Highlights¶
- Save and load agent dialog. Export a full-fidelity snapshot while an agent is idle, store the canonical
TDLGbytes inside your own save container, and import it into a fresh session withStrictcompatibility. Tryll still writes no save files. See Save and load agent state. - One session language for the whole stack. Set a BCP-47 locale and speech output, speech-recognition hints, per-language content files, and the
{{language}}prompt keys all follow it — changeable mid-session without tearing the session down. See Ship an agent in another language. - Choose where each model runs. Per-model RAM / VRAM placement, authored in the Model Manager or passed when you load a model yourself. A hint by design: a machine that cannot honour it loads the model the other way and reports what it actually did. See Choose where models run.
- Engines come from the catalog. The four engine dropdowns are gone from Project Settings in Unity and Unreal. The catalog records which engine runs each model — which also fixes model listing showing no STT, TTS, or embedding models in a project that never touched those settings.
- Pauses can time out.
pause_timeout_msgives aPausea deadline andtimeout_exitlets the character answer anyway, instead of an unanswered pause leaving the agent busy forever. The server also warns when a pause outlives its stall threshold. - Two expressive voices, plus voice direction.
Orpheus 3B FT (Q4_K_M)andMaya1 3B (i1-Q4_K_M)are public catalog voices, andtts_voice_description/tts_delivery_instructionlet you brief a voice's identity and its delivery in natural language. See TTS Models. - Fit a whole cast on one GPU. A guide to the two levers that actually decide whether five NPC agents fit in limited VRAM — sizing every LLM node's
context_size, then evicting the agents nobody is talking to. See Run many agents on a small GPU. - Tool calling, measured. Exact-accuracy tables for the catalog on game-shaped suites and a BFCL slice, schema and prompt recipes that survived those sweeps, and a way to exercise a
ToolCallworkflow in the editor without Play/PIE. See Compare tool-call models and Design a tool-call prompt. - See what the server is actually using. A Memory window in Unity and Unreal reports live RAM and VRAM for the whole server process — resident models, per-agent KV, and the part nobody can attribute — with a confidence tier on every figure, so an estimate never reads as a measurement. See Inspect server memory.
- Intent classification without a corpus. Describe each intent in a sentence and
ClassifyIntentLLMroutes the line off its own first-token logprobs — no embedding model, no labelled knowledge base, nothing to ship — with a measurement-backed prompt guide, an Intent block in the turn inspector, and Dialog Lab checks that assert the accepted label. See Classify intent without a knowledge base and Design an intent-classifier prompt. - Classifier and tool-call nodes size their own window.
ClassifyIntentLLMandToolCallnow default to 2048 tokens instead of inheriting a generator-sized 8192 — the cheapest VRAM saving in the release, and it needs no authoring.
Save and load¶
ExportDialog/ImportDialogon every client. Both are strict idle-only — no turn running, not paused, no KV-cache operation in flight. Export returns aDialogSnapshot;ToFlatBufferBytes/FromFlatBufferBytes(encode_flatbuffer/decode_flatbufferin Python) produce the canonicalTDLGbytes. Import replaces the whole dialog and is atomic: a failed import leaves the existing dialog untouched. Unreal exposes this in C++ only — the snapshot is not Blueprint-exposed. See Export and import dialog.- The snapshot is dialog, and only dialog. Not agent variables, not UI transcripts, not KV-cache bytes, not agent IDs. The save-slot pattern — quiesce every agent, snapshot each one alongside variables and game state, recreate the agents in a fresh session, import, then warm the ones you actually need — is written up end to end in Save and load agent state. Import never prefills; call
PrefillKvCacheyourself during the loading screen. Strictby default. Import checks the snapshot against the agent's create-time graph and refuses a mismatch;AllowGraphMismatchopts out. Snapshots that would overflow the 1 MiB frame cap fail rather than being truncated.- Unity coroutine wrappers.
TryllAgentCoroutineswraps export, import, variable sync/flush, and KV prefill: it retriesAgentBusyat a polite interval and bounds every wait with a wall-clock deadline, so a dropped connection fails your save UI withTimeoutinstead of hanging it.
One language for the session¶
localeonCreateSession, andSetSessionLocaleafter it. A BCP-47 tag (de,pt-BR,zh-Hans-CN) that becomes the defaulttts_langon everyGenerateAndSpeak/Speak, the language hint handed to multilingual STT models, the per-language file lookup, and thelocale/language/language_codeMustache keys. Empty leaves every one of those at its previous behavior; a malformed tag is rejected with2004 InvalidLocale. Unity and Unreal add a Project Settings → Tryll Client → Language picker that lists only what your registered models can serve.- Translations live beside the file you link. A
loc/<language>/folder next to a linked storage file supplies the translated version —deathonset/loc/pt-BR/lore.json, thenloc/pt/, then the file you linked. You always link the default file; a folder with noloc/is language-neutral; and fallback is per file, so a partially translated game works. This covers knowledge bases and their sidecars, canned responses, guardrail patterns, intent maps, hotwords, and reference-voice WAVs. {{language}},{{locale}},{{language_code}}in templates. Available wherever Mustache rendering is —Generate,GenerateAndSpeak,ToolCall,Transform. The section collapses to nothing when no language is set, so one template works localized and unlocalized.- A helper for the offerable set.
ListModelsreports alanguageslist per model, andTryllLanguages(Unity) /UTryllLanguages(Unreal, Blueprint-callable) intersect them across your pipeline.HasLanguageInfotells "no overlap" apart from "the catalog didn't say" before you show a player an empty menu. - A mid-session switch is not retroactive. Speech output, the template keys, and the next file load follow it from each agent's next turn; a knowledge base, canned lines, or classifier prompt already bound to a node does not — those are structural. Send the change between turns, not mid-sentence.
tts_langresolution is now layered: the node's own value, then the session locale, then the model's catalog default. Set the node param only to override one character — a French witness in a German playthrough.
Models: placement, engines, and GPU backend¶
memory_placementper model.Auto/RAM/VRAM, authored on a registered model in the Model Manager or passed when you load a model yourself. The reason to reach for it: your renderer and Tryll's language model compete for the same VRAM, and moving a large model to RAM trades generation speed for graphics headroom. It is a hint on purpose — you author it on your machine and ship it to someone else's.load_model()reports what happened. The Python client returns a result carryingplacementandbackendinstead ofNone, so aVRAMrequest that could not be honoured is visible rather than guessed at. Purely additive — existing calls that ignore the return value are unaffected. Unity and Unreal log a warning naming both placements and the backend; their completion events keep their(model name, success)signatures, because a model cannot be moved after it loads.- The session declares engines as a hint, not a gate.
CreateSessiontakes an optional list of the(role, engine)pairs you expect to touch, replacing the four engine arguments. It may be empty, incomplete, or wrong, and loading a model it never named works normally. What it buys is timing: engine start-up is paid at session creation instead of on first use — worth asking for when a backend is slow to initialise. See Choose where models run. - Pick the llama.cpp GPU API in Project Settings. Tryll Client → Llama.cpp backend (
Vulkandefault, orAuto), passed to an auto-launched server as--llama-backend. AVulkanplayer build skips the CUDA DLLs;Autocopies them and prefers CUDA on NVIDIA. Standalone / C++ / Python setengines.llama_cpp.backendinserver-config.json(committed default"vulkan") or pass the flag. Players need neither the CUDA Toolkit nor those DLLs onPATH, and speech and embedding stay on the CPU package. See Auto-launch the server and Server configuration. - Shared model assets. The SNAC 24 kHz decoder is the first: it has its own download record, is pulled once for whichever parent needs it, and survives deleting a parent model.
ListModelsreportsModelStatus.Incompletewhen a parent is waiting on a shared asset, with dependency details. The catalog also gaineddependenciesand structuredtts_capabilities(voice presets, inline tags). See Model Management. - Fixed: telemetry reported the wrong GPU backend.
gpu_backend_activewas read from build-time flags that are never set in our builds, so every install reportedcpu— including machines where Vulkan was serving every token. It now reads the live device registry. If you have been comparing performance data across machines, the backend field in older data is not trustworthy.
Fitting a cast on one GPU¶
- Run many agents on a small GPU — the two levers, in the order worth pulling. A KV cache is allocated at its full
context_sizethe moment the context is created, so the cost is deterministic, short replies save nothing, and a sizing change is measurable in seconds rather than in playtests. Then evict the agents nobody is talking to. Worked through with numbers from a five-agent, nine-context demo on a 16 GB card. ClassifyIntentLLMandToolCallnow default to a 2048-token window. Left at0, they no longer inherit the model variant'scontext_sizeor the server'sdefault_n_ctx(commonly 8192) — the two nodes whose prompts are small got the small default.GenerateandGenerateAndSpeakare unchanged, so budget per node: a classifier at 2048 is a small slice of an 8192 generator. See Estimate memory footprint.- A per-node
contexttable, for when the window itself is the problem.Generate,GenerateAndSpeak,ClassifyIntentLLMandToolCalltake an optional structuralcontexttable:kv_cache_typeper context (overriding the catalog's model-load default, withInheritkeeping it), a veto-onlyoffload_kqv, and escape-hatch batch sizes. Leave it unset unless you are tuning a large window — at a right-sized 2048 aq4_0KV override is usually not a measurable VRAM win, because the dtype delta only shows up around 8192. See Models and inference engines. - The sizing trap is silent. When a prompt exceeds its window, the token-budget projection trims the oldest turns and keeps working — so an under-sized classifier quietly loses the history it needs to resolve "And that one?". Leave headroom, and re-check that intents still fire after tightening one.
- An immersion guard costs two extra contexts per agent. Build an immersion guard now says so up front and points at the sizing guide before you add one to a project already close to its VRAM budget.
- Memory estimates cover SNAC voices. llama.cpp SNAC voices are GGUF backbones that follow the same RAM/VRAM placement as a language model, plus a ~50 MB CPU decoder — unlike Sherpa STT/TTS, which stay RAM-only. See Estimate memory footprint.
Memory inspection¶
- A Memory window in both editors. Window → Tryll → Memory in Unity, Window → Tryll → Tryll Memory in Unreal: resident models, per-agent KV, and the unattributed remainder, on demand. Refresh only — no poll, no graph. It reuses the shared editor session and auto-launches the bundled server the way Model Manager does, needs no agent, and does not wait for registered models to download. Closing it releases the lease without killing the server. See Inspect server memory.
- The snapshot is server-wide, not per-connection. Every session on that process is in the numbers — Chat, Model Manager, Play-in-Editor, a Python client, your QA harness. A row you do not recognise is not a bug in the window. That is the opposite of the Agent Log, which sees only your own connection.
- Every figure carries a confidence tier.
Measured(the engine or the OS said so),Derived(computed from measured inputs),Estimated(a calibrated band, rendered with a leading~and an estimate basis on hover). The column is always visible on purpose: an estimate must never look like a measurement. - What the tables actually answer. Devices split into ours / other / free, and
otheris the editor and the compositor rather than a Tryll leak. KV is listed per agent, sorted background-first, withevictedand never run called out separately — background KV is the reclaim signal, because nobody is waiting on it. Context used is headroom to the limit, not memory growth: the buffer is allocated at the fullcontext_sizethe moment the context is created. And memory-mapped model files are resident pages Tryll never allocated, so a VRAM-placed model often surfaces under unattributed instead of its own RAM column — the window says so when that happens. GetMemoryConsumptionon every client.TryllClient::GetMemoryConsumption(C++),client.get_memory_consumption()(Python),RequestGetMemoryConsumptionAsync(Unity), andUTryllSubsystem::RequestGetMemoryConsumption(Unreal C++, with a Blueprint completion event to bind) all return the same snapshot. Session-scoped — it must followCreateSession— but the numbers cover the whole process, and it never fails withAgentBusy.- The server monitor gained the matching panel. Same collector, same numbers, over
GET /api/memory; the browser view keeps its sparkline and stacked bars, which the editor window does not have.monitor.memory_tick_interval_mssets the live sampling interval —0disables periodic sampling and the panel still loads on demand. See Server configuration.
Voice output¶
- Orpheus and Maya1 are public catalog voices.
Orpheus 3B FT (Q4_K_M)gives eight named English speakers (speaker_id0–7) plus inline tags such as<laugh>;Maya1 3B (i1-Q4_K_M)takes a natural-language voice description, with two catalog presets as a starting point. Both download through the normal Model Manager, including the shared SNAC decoder. Heavier quants stay internal. tts_voice_descriptionandtts_delivery_instruction. Two new mutable params onGenerateAndSpeakandSpeak: a stable speaker identity (age, accent, pitch, timbre) and a sustained delivery direction for this synthesis call. They are independent oftts_voice(a reference WAV) andspeaker_id(a preset index) — four distinct knobs, and square-bracket[emotion]text is not a supported control. Model-native<tag>tokens that a catalog lists may reach the synthesizer; unknown tags are stripped from the audio copy only and never rewrite history.- A control the model does not advertise now fails instead of being silently ignored — Orpheus and Maya1 reject a
speedother than1.0. The full per-family control matrix, and the attribution each set of weights requires, are in TTS Models. - Unity: ask the speaker whether it is still talking.
TryllSpeakergainsHasPendingAudio(true while decoded audio is still queued) andStopPlayback(). Combined withTurnComplete,HasPendingAudiois the signal a hands-free voice loop needs before re-opening the microphone, so the mic does not transcribe the tail of the character's own answer.AudioSource.isPlayingis not that signal — the streaming clip loops silence between turns. See TryllSpeaker.
Turn control¶
pause_timeout_msonPause.0still waits forever. A finite budget makes the failure recoverable: withtimeout_exitauthored the graph continues from that exit and the turn ends normally; without it the turn ends withTurnStatus.PauseTimedOut, keeping the partial interaction so the conversation stays intact. Prefer authoring atimeout_exitfor anything player-facing, so the character says something. Both params are structural.- The deadline is server wall-clock. A minimised or alt-tabbed game stops pumping its main thread, so the deadline can elapse while your game is suspended. Make the timeout branch harmless — "shrug and answer without the tool" — rather than something that only makes sense the instant it fires.
pause_timeout_msonToolCallcovers pauses that node triggers, inheriting the server'sworkflow.tool_call_pause_timeout_mswhen left at0. There is no timeout exit on this node — expiry ends the turn.- An unresumed pause is loud before it is fatal. The server logs a warning naming the parked node once a pause outlives
workflow.pause_stall_warn_ms(15 s by default), and your client logs one immediately when nothing it holds can resume the pause at all.
Tool calling¶
- Compare tool-call models — exact accuracy for the catalog across seven game-shaped suites and a BFCL slice, at greedy decoding. The two corpora do not rank models the same way, so the page is built for choosing by job (commands, log queries, parallel calls) rather than for crowning a winner. Models whose
tool_call_supportisunsupportedare rejected at agent creation and are not in the tables. - Design a tool-call prompt — measurement-backed schema and prompt recipes. The headline finding: small models fail at restraint, not extraction. Give "the player did not narrow this down" a real enum value, list it first, and mark the parameter
required— the value a pressured model reaches for becomes the right answer. Claims are marked Tested / Hint / no effect, so you can skip what did not pay off. - Test tool calls in Chat and the Dialog Lab — register a handler on the scene agent component in edit mode, talk to it in Chat, then script Lab runs with canned results or against the live handler. No Play/PIE.
- Dialog Lab live tools are opt-in. A Lab step with no canned
ToolResultsentry still means the model must not call that tool, even when the variant's scene agent has aRegisterToolhandler. Tick Use live tools on the variant only when the run should exercise it; canned results still win. See Compare dialog variants in the Lab. - Diagnostics answer "why didn't it call my tool".
parameters.tools[]echoes the signatures the model actually saw,output.tool_calls[]carries each detected call with its raw arguments,output.parse_fallbackflags a call the native template parse missed, andengine.chat_turn.has_grammarshows whetherrequire_callactually constrained the sampler. See Tool Call.
Intent classification¶
notify_clientis now a three-state enum.Disabled/OnFound/AlwaysonClassifyIntentandClassifyIntentLLM, replacing the bool.Alwaysalso fires on everynot_foundpath, carryingnot_found_reasonand the rejected top-1 — which is how you see a deflection that didn't happen.- Typed callbacks, and a separate LLM event.
ClassifyIntentLLMnow emits its ownintent_llm_classifiedevent with the full probability vector, subscribed viaSetOnIntentLlmClassified/IntentLlmClassified/OnIntentLlmClassified/set_on_intent_llm_classified. Both classification callbacks now take a structured event object instead of a positional argument list. A typed subscriber swallows the genericOnNodeEventfallback. - Richer classification diagnostics. The embedding classifier records ranked
topk[]candidates, the winner, and arunner_upwith itsmargin— the number to watch when a classifier is unstable, because a margin near zero means small rewordings will flip the route. The LLM classifier records the label ids in letter order alongside the probability vector. ClassifyIntentLLMwithout a knowledge base, end to end. Classify intent without a knowledge base rebuilds the gate-guard NPC with one language model and no corpus: a sentence per intent, no embedding pass, nothing on disk. It opens with a table for choosing between the two classifiers — many intents and pre-labelled data still favour embeddingClassifyIntent; a label set you are still designing favours this one, because writing a sentence is a far shorter loop than authoring twenty phrases and rebuilding an index.system_promptis a Mustache template, and there is no default. It renders against exactly two variables —{{intents_block}}(the letter-prefixed label list) and{{letters}}— and none of theGeneratevariables (user_message,slot.<name>,var.<name>,instructions,knowledge) resolve there. Leave it empty and the node sends an empty system message: the model is asked for a letter having never been shown what the letters mean, and nothing fails loudly. A template that fails to parse is sent verbatim, literal{{intents_block}}included. Verify the substitution ininput.prompt[]after any edit. See Classify Intent (LLM).- Design an intent-classifier prompt — claims marked Tested / Hint / no effect, like the NPC and tool-call prompt guides. The mechanism decides what matters: the node reads the logits at the end of the prompt and never generates, so "think step by step", "explain your reasoning" and "answer in JSON" are paid for in prefill and then discarded unread. What pays off is mutually exclusive descriptions — overlap is the top cause of a collapsed
margin— an explicit catch-all placed last, because models carry strong, model-specific positional bias over the answer letter and that bias has to be re-checked when you change classifier models, and tuningthresholdonly after the overlap is fixed. intents_idsandintents_prompt, spelled out.intents_promptis the text the model reads, rendered asA. <first>,B. <second>, …;intents_idsis never shown to it and only supplies the label attached on a win. They pair positionally — entry 0 is letterA— must have equal counts and 2–26 unique ids, and each letter must be a single token for that model. A comma inside a description silently becomes an extra label, so use dashes or semicolons. All of it is validated at agent creation.- An
Intentblock in the turn inspector. One subsection per classifier execution, with the decision in the header —ask_wing · found, ornot_found · below_threshold— andnot_foundcoloured as a warning rather than an error, because it is usually correct behaviour. Embedding classifiers get the funnel line, a winner card or the rejected pick, and a top-k table that dims rows beyond the threshold; LLM classifiers get a probability bar per candidate and thepredicted_letter/top_prob/second_prob/margin_resultstrip, with the winning or rejected label shown in words so you never map a letter by hand.notify_clientdoes not gate the block — a classifier left atDisabledstill appears, it just produces no Agent Log event. See Turn Inspector. - Dialog Lab can assert the label. Intent is passes when an addressed classifier on that turn accepted the exact id; Intent not found passes only when every addressed classifier took
not_found. Leave the node name empty to address every classifier that ran, or name one to scope the check. Labels compare exactly and case-sensitively — they are identifiers, not prose. A missing or mis-typed classifier fails withintent unavailablerather than skipping, so a green suite cannot come from a classifier that never ran, and seeded steps refuse these checks before Run: a seeded turn is injected without running the graph, so there is no classification to assert. See Compare dialog variants in the Lab. - A rejected pick is now recorded. Both classifiers add
output.classification.rejected_intent— the label the scorer or the search chose when routing rejected it — and the embedding node addsrejected_distancebeside it, which is what turns a barenot_foundinto "it wantedask_wing, at 0.31, against a 0.25 threshold". The classified text is nowinput.queryon both kinds (output.classification.queryis kept for one release), and the LLMprobs[]vector is documented as positionally aligned withparameters.intents_ids[], equal length onfound.
Diagnostics¶
- Turn Diagnostics JSON — a new reference for the whole
TurnComplete.debug_infopayload: the envelope, the per-nodeparameters/input/output/enginesections, the turn-level tool-call and interaction snapshots, and the rules for why a given key is absent. Every node reference page now carries a Diagnostics section listing the keys it records. - The turn inspector's Raw JSON is verbatim. The editors add no keys, rename none, and drop none — so the reference above describes exactly what you see in the panel and exactly what arrives on the wire. Raw JSON is also the one block the Route breadcrumb's node scope does not narrow; it always shows the whole turn. See Turn Inspector.
- Two server flags trim the payload.
include_engine_diagnosticsdrops theenginesections andinclude_interaction_in_diagnosticsdrops the end-of-turn interaction snapshot — which is why the inspector's tool-call block can show calls with no results.
Server operations¶
- Log flushing is configurable.
log_flush_on(default"warn") flushes the log file as soon as a message at that level or higher is written, andlog_flush_every_ms(default5000) flushes on a timer so a quiet server does not leave the file minutes behind. A running server's log is now readable without waiting for process exit. See Server configuration. - Startup states standalone vs managed idle-exit. The server logs
Standalone mode: server will not self-exitorManaged mode: self-exit after <N>s of idleright after logging starts, instead of announcing the mode only when--idle-shutdown-timeoutwas on the command line. - Fixed: a managed server could finish its idle exit and never actually leave. An open monitor event stream kept the HTTP listener from shutting down, so the process stayed alive with both the game port and the monitor port refusing connections — Dialog Lab then failed with "the target machine actively refused it" and would not relaunch. Subscribers are now closed before the listener is joined, and the Unity and Unreal launchers treat an owned process that is alive but no longer accepting as dead and start a fresh one.
Editor fixes¶
- Chat honors a cleared Agent Component. Clearing the Chat window's Agent Component (or Voice Component) to None now sticks. Start and Compare in Lab refuse with the usual assign warning instead of silently restoring the previous scene object.
- Dialog Lab shows the seed it actually uses. Typing
0into Repeats or base seed still raises the value to1(seed0means random on the wire). The field now displays1instead of keeping the typed0.
Breaking changes and behavior changes¶
- The schema changed — rebuild every client.
CreateSession, the TTS nodes, the classification nodes, the pause params, and the LLM nodes' newcontexttable all changed shape this release, and a server-wide memory snapshot call was added. Rebuild your integration against the new schema and upgrade both ends together; mismatched peers are rejected at connection time with5004 ProtocolVersionMismatch. See Wire Protocol. ClassifyIntentLLMandToolCalldefault to a 2048-token context. Withcontext_sizeleft at0they previously fell back to the model variant's window, then the server'sdefault_n_ctx— commonly 8192. Most projects simply get the VRAM back; a node whose classifier prompt or tool-schema block genuinely needs more than 2048 tokens must now setcontext_sizeexplicitly.GenerateandGenerateAndSpeakare unchanged, andcontext_sizeis structural, so this is a graph edit rather than a runtime one. See Run many agents on a small GPU.ClassifyIntentLLMhas no built-insystem_prompt, and an empty one fails quietly. Every client leaves the field empty by default and the node passes it through verbatim, so an unauthored classifier sends an empty system message and asks the model for a letter whose meaning it was never shown. Author a template containing{{intents_block}}. See Design an intent-classifier prompt.- A component now requires a Workflow Asset.
TryllAgentComponent.InlineGraphDescription/InlineVariablesandUTryllAgentComponent::InlineGraphDescription/InlineVariablesare removed — deprecated in v0.6.0, gone now. Author the graph as a Workflow Asset and assign it; declare variables on the asset. Building aTryllGraphDescriptionin code and creating an agent from it directly —TryllClient.RequestCreateAgentAsync/UTryllSubsystem::RequestCreateAgent, with no component — remains fully supported. If you still want a component around a code-built graph, put the graph on a transient Workflow Asset (ScriptableObject.CreateInstance<TryllWorkflowAsset>()/NewObject<UTryllWorkflowAsset>()) and assign that. See TryllAgentComponent and Unreal C++ API. - The four engine dropdowns are gone, and nothing replaces them. Unity's
InferenceEngine/SttEngine/TtsEngine/EmbeddingEngineruntime settings and Unreal'sEngine/SttEngine/TtsEngine/EmbeddingEngineproject settings are removed; the session's engine set is derived from the models you registered. If you build sessions by hand, the four engine arguments are replaced by one optional engines list. This also fixes a real bug: three of the four defaulted toMockandListModelsfiltered by them, so a project that never touched them saw no STT, TTS, or embedding models in the Model Manager at all. Model listing is now engine-agnostic. notify_clientis an enum, not a bool.ClassifyIntentandClassifyIntentLLMtakeDisabled/OnFound/Always; the oldtruebecomesOnFound. TheIntentClassifiedcallback signature changed to a structured event object on every client, andClassifyIntentLLMnow reports through the separateintent_llm_classifiedevent — code that relied onintent_classifiedfiring for an LLM classifier must subscribe to the new one.- Unity
RemoveInteractionsFromEndreports the count. Its completion callback is nowAction<uint, TryllError>(wasAction<TryllError>), matchingAppendInteractions. enable_diagnosticsis structural. It is fixed for the agent's lifetime and cannot be turned on for an agent that already exists. The editor tooling handles this for you — Chat enables it on the agent it creates, and the Agent Log holds a lease so agents created while it is open request it too. See Agent Parameters.device_preferenceis removed frommodels.json. A catalog that still declares it fails to load, with a message naming the replacement (memory_placement). Every shipped entry said"auto", so it decided nothing;"directml"was accepted there and never implemented, and is gone too.engines.llama_cpp.n_gpu_layersis deprecated and ignored. It applied one value to every model in the process, which cannot be right once different models want different memory. Offload is now decided per model from its placement. A server that still finds the key warns once at startup — remove it.- Parakeet TDT 0.6B v2 is replaced by v3. Update
model_namein any voice-input config; the v2 entry is no longer in the catalog. - Dialog Lab live tools are opt-in. A variant whose scene agent has a
RegisterToolhandler no longer exercises it implicitly — a step with no cannedToolResultsentry asserts that the model must not call that tool. Tick Use live tools on the variant to restore handler execution. See Test tool calls in Chat and the Dialog Lab. locis a reserved folder name inside storage folders. Do not useloc/for your own content; it is the per-language lookup root. If your build pipeline filters what gets staged, make sure it keeps**/loc/**— a missing translation folder degrades silently to the default language rather than failing the build. See Ship storage folders for builds.
New error codes¶
2004 InvalidLocale—CreateSessionRequest.locale(or a laterSetSessionLocale) is not a well-formed BCP-47 tag. Expectedlanguage[-script][-region]; the message names what was wrong.3017 ModelLanguageUnsupported— a node would rely on the session language but its model does not support it. Pick a model that covers the language, set the node's owntts_lang, or change the session locale. Not raised when the model declares no languages at all.3018 DialogSnapshotInvalid— the snapshot is malformed, missing required fields, or has inconsistent tool call / result data. Import leaves the existing dialog unchanged.3019 DialogFormatUnsupported— the snapshot'sformat_versionis newer than this server, or the file identifier is notTDLG.3020 DialogGraphMismatch—Strictimport: the snapshot's graph signature does not match the agent's create-time graph.3021 DialogSnapshotTooLarge— the encoded snapshot would overflow the 1 MiB frame cap. Tryll never truncates.3022 DialogUnsupportedComponent— an unknown component type on export, or an unknown union arm on import.
See Error Codes for the full list.