Définition
A technique that caches the key-value (KV) attention states of a repeated prompt prefix, so subsequent requests reuse the pre-computed computation rather than re-running it. Prompt caching reduces latency and cost significantly for applications with long system prompts.
Termes Connexes
Besoin d'Aide pour Comprendre l'IA?
Réservez un appel de cadrage Physical AI pour discuter de l'application de ces concepts IA à votre secteur et vos défis métier.