Skip to content
Intelligence RadarNews & Launches1 MIN READ

OpenAI Announces Enhanced Prompt Caching for GPT-6

OpenAI has introduced improved prompt caching for GPT-6, offering higher cache hit rates and new tools to help persistent agents run faster and cost less.

Fathom Intelligence
Fathom IntelligenceFathom Layer Expert

OpenAI announced enhancements to the prompt caching system for GPT-6, aiming to deliver higher cache hit rates and introduce new tools to help persistent agents run more efficiently and cost-effectively. The improvements include cache discounts for eligible shared prefixes reused within a 30-minute window, reducing response times and lowering costs by up to 90% on cached input tokens. Developers can now monitor cache performance and diagnose cache misses, optimizing their applications to maximize cache hit rates. Additionally, explicit cache breakpoints allow developers to choose which prompt prefixes to reuse, and changes to reasoning effort between responses can be made without breaking cache. These enhancements have led to significant improvements in cache hit rates, with some teams seeing increases from 83% to 91%, resulting in reduced cache writes and lower inference costs.


Source: openai

END OF SIGNAL

Don't Miss the Next Signal

Get high-impact hardware and AI launches distilled into your inbox. No noise, just the changes that matter.