AI Industry Published on July 6, 2026
OpenAI releases gpt-realtime-2.1 with 25% lower P95 latency
OpenAI added gpt-realtime-2.1 and gpt-realtime-2.1-mini to the Realtime API.
The update improves caching and reduces P95 latency by at least 25%.
The full model gains better alphanumeric recognition, noise handling, interruption behavior, and configurable reasoning effort; the mini tier now includes reasoning and tool use for the first time. Pricing is unchanged from the previous Realtime models.
Source: OpenAI / Releasebot