Definitie
The total number of tokens allocated for a model request, encompassing both input (prompt + context) and output. Managing token budgets is central to controlling inference cost in production LLM applications, especially when processing long documents or maintaining conversational history.
Gerelateerde Termen
Hulp Nodig bij het Begrijpen van AI?
Boek een Physical AI-kennismaking om te bespreken hoe deze AI-concepten op uw branche en uitdagingen van toepassing zijn.