OpenAI is previewing a new high-speed API tier called Ultrafast for its GPT-5.6 Sol model. This service runs up to 14 times faster than the standard version.

The system delivers up to 750 output tokens per second to reduce latency for time-sensitive applications. Performance gains occur without compromising the intelligence of the model.

Cerebras hardware powers the new tier using its Wafer-Scale Engine architecture. This specialized hardware minimizes data movement bottlenecks that slow down large models on conventional hardware.

The Ultrafast tier is currently available in a limited preview to select customers. OpenAI plans to expand access to the service over time.

This partnership highlights a trend of co-designing AI models with specialized hardware.