Token Cache Hit
Sokudo
LLM Systems · Explainer
Faster inference starts before generation

Understanding Token Cache Hits
in Large Language Models

What gets cached, how a request reuses prior computation, and how to design prompts that actually benefit.

Prompt prefixCache hit
78% of input tokens reused