OpenAI positions Jalapeño chip within its integrated compute strategy

OpenAI has released the first measured performance results for Jalapeño, its first custom inference chip, and positioned the processor as one component of a broader compute strategy. In the public InferenceX benchmark using GPT-OSS 120B, OpenAI said Jalapeño achieved higher peak throughput per kilowatt and lower token latency than the commercial systems included in the comparison.
The company also said Jalapeño performed strongly on DeepSeek R1 and Kimi K2. The results are intended to show that the chip’s advantages extend across model families, rather than being limited to one OpenAI workload.
A chip designed as part of the serving stack
OpenAI describes its compute approach as an integrated system spanning data centres, chips, frontier models, a developer platform, consumer and enterprise products, and AI-native devices. Its argument is that changes at one layer can improve the others: software can make hardware more productive, while hardware built around target workloads can improve speed and efficiency.
Jalapeño gives OpenAI more direct control over how its models run and over the economics of serving them. The company says it is co-developing the model, serving software, chip, memory and network so that throughput, latency, energy efficiency and cost can be improved together. Future Jalapeño generations are already under development.
OpenAI does not present first-party silicon as a replacement for outside accelerators. Instead, it calls Jalapeño a first-party option alongside partner hardware, enabling workloads to be assigned to systems with the strongest economics and performance characteristics.
A portfolio for different AI workloads
The company distinguishes among frontier training, high-volume inference and always-on agents, noting that each places different demands on chips, software, networks, power and latency. Its stated goal is to remain on the Pareto frontier, balancing capability, speed, reliability, efficiency and cost for each workload.
Microsoft compute and NVIDIA chips remain foundational to OpenAI’s growth. The company’s current portfolio also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank, spanning cloud infrastructure, accelerated computing, low-latency inference, data-centre development and energy delivery.
OpenAI says credible choice across providers, hardware and deployment models helps it direct demand toward stronger performance per dollar and maintain pricing discipline as conditions change. It will use premium systems where capability is most important and optimise for efficiency where scale and cost take priority.
Efficiency as an economic measure
OpenAI frames the value of its infrastructure in terms of useful intelligence from every unit of compute. It points to better models requiring fewer attempts, smarter routing and context management reducing wasted work, and purpose-built hardware and software improving speed and energy efficiency.
On the Artificial Analysis Coding Agent Index, OpenAI said GPT-5.6 Sol with max reasoning set a new high while using 54% fewer output tokens than another leading model. The company links such gains to faster results, fewer retries, longer agent workflows and a lower total cost for successful work.
For businesses deploying AI, the practical implication is to measure systems against completed, dependable work rather than model output alone, while considering latency, token use, routing and infrastructure efficiency as part of the total cost.

