Prefix caching cuts LLM latency by 70% GKE Inference Gateway uses prefix caching to cut time-to-first-token latency by over 70%, eliminating redundant computation in AI pipelines. Jun 11, 2026 Marcus Chen 12 min read