انتقل إلى المحتوى

التعريف

A technique that caches the key-value (KV) attention states of a repeated prompt prefix, so subsequent requests reuse the pre-computed computation rather than re-running it. Prompt caching reduces latency and cost significantly for applications with long system prompts.

تحتاج مساعدة في فهم الذكاء الاصطناعي؟

احجز مكالمة تقييم ملاءمة Physical AI لمناقشة كيفية تطبيق مفاهيم الذكاء الاصطناعي هذه على قطاعك وتحدياتك.

Prompt Cache | مسرد مصطلحات AI — LLM وRAG والضبط الدقيق والمفاهيم الرئيسية