定义
The total number of tokens allocated for a model request, encompassing both input (prompt + context) and output. Managing token budgets is central to controlling inference cost in production LLM applications, especially when processing long documents or maintaining conversational history.
相关术语
了解术语只是第一步,将其落地应用才是第二步。
预约一次 Physical AI 适配性沟通,探讨这些 AI 概念如何转化到您所在的具体行业与业务挑战中。