Skip to content

Struct FTryllContextOverrides

ClassList > FTryllContextOverrides

More...

  • #include <TryllContextOverrides.h>

Public Attributes

Type Name
ETryllKvCacheType KvCacheType = ETryllKvCacheType::Inherit
int32 NBatch = 0
int32 NOutputsMax = 0
int32 NUbatch = 0
bool OffloadKqv = 0
bool bOverrideNBatch = false
bool bOverrideNOutputsMax = false
bool bOverrideNUbatch = false
bool bOverrideOffloadKqv = false

Detailed Description

Backend context construction. Null fields inherit ModelVariant then server config, except n_outputs_max which CreateContext defaults to 1. Structural at the field: these are baked into llama_context_params at context creation, so a change means a new context and a lost KV cache.

Public Attributes Documentation

variable KvCacheType

ETryllKvCacheType FTryllContextOverrides::KvCacheType;

KV cache dtype. Inherit = ModelVariant, then q8_0.


variable NBatch

int32 FTryllContextOverrides::NBatch;

variable NOutputsMax

int32 FTryllContextOverrides::NOutputsMax;

variable NUbatch

int32 FTryllContextOverrides::NUbatch;

variable OffloadKqv

bool FTryllContextOverrides::OffloadKqv;

variable bOverrideNBatch

bool FTryllContextOverrides::bOverrideNBatch;

Logical llama_decode submission ceiling.


variable bOverrideNOutputsMax

bool FTryllContextOverrides::bOverrideNOutputsMax;

Logits buffer positions. Engine default is 1 for every current LLM node; this field is an escape hatch if a future path needs more.


variable bOverrideNUbatch

bool FTryllContextOverrides::bOverrideNUbatch;

Physical micro-batch size that sizes activation buffers.


variable bOverrideOffloadKqv

bool FTryllContextOverrides::bOverrideOffloadKqv;

Place the KV cache on GPU. Veto only — a node can force false but cannot grant offload to a RAM-placed model. Null = inherit server/placement.



The documentation for this class was generated from the following file C:/_tryll/_monorepo3/tryll/clients/unreal/Source/TryllClient/Public/Generated/Nodes/TryllContextOverrides.h