Show notes
OpenAI has introduced a technical preview of its new "Ultrafast" processing tier for the GPT-5.6 Sol model, demonstrating performance up to 14 times faster than standard processing. This performance boost is powered by specialized hardware from Cerebras, enabling the model to generate up to 750 output tokens per second. While the GPT-5.6 family, which includes the balanced Terra and speed-optimized Luna models, became broadly available in July, the Ultrafast tier is a new specialized service for the flagship Sol model. It is designed specifically for latency-sensitive applications such as real-time voice interaction, developer agents, and immediate security response. OpenAI notes that its internal developers have already used the tier to compress lengthy research cycles into just a few hours. Currently, access is restricted to a select group of customers via a waitlist, where businesses must provide specific details regarding their workloads and latency requirements to be considered for the preview.



