The total number of tokens allocated for a model request, encompassing both input (prompt + context) and output. Managing token budgets is central to controlling inference cost in production LLM applications, especially when processing long documents or maintaining conversational history.
Réservez un appel de cadrage Physical AI pour discuter de l'application de ces concepts IA à votre secteur et vos défis métier.