Sokudo
LLM Systems · Explainer
Faster inference starts before generation
Understanding Token Cache Hits
in Large Language Models
What gets cached, how a request reuses prior computation, and how to design prompts that actually benefit.
Prompt prefix→Cache hit
78% of input tokens reused