Why It Matters
Sokudo
Impact

Optimize the expensive input path

The larger and more repetitive the prefix, the more valuable caching becomes.

Latency
↓ TTFT

Less prefill work can reduce time to first token, especially for long contexts.

Cost
↓ Input

Some providers price cached input tokens below uncached input tokens.

Capacity
↑ Throughput

Reusing prefix computation frees accelerator work for more requests.

Important: output-token generation is still computed. Caching does not make long answers free.