Sokudo
Impact
Optimize the expensive input path
The larger and more repetitive the prefix, the more valuable caching becomes.
Latency
↓ TTFT
Less prefill work can reduce time to first token, especially for long contexts.
Cost
↓ Input
Some providers price cached input tokens below uncached input tokens.
Capacity
↑ Throughput
Reusing prefix computation frees accelerator work for more requests.
Important: output-token generation is still computed. Caching does not make long answers free.