OpenAI revealed benchmark details for Jalapeño, its custom Application-Specific Integrated Circuit (ASIC) built in partnership with Broadcom for running AI inference workloads. According to a company blog post and briefing by hardware VP Richard Ho, the chip delivered 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency than NVIDIA’s GB200 and GB300 systems on the InferenceX benchmark across models like GPT-OSS 120B and DeepSeek R1.

Originally disclosed in June, Jalapeño is designed to lower latency and increase throughput for running trained models and deploying autonomous agents. OpenAI plans to deploy the chip in small volumes by the end of this year before ramping up production into 2027, while continuing development on second and third-generation designs.

Despite the custom silicon gains, OpenAI stated it does not intend to fully replace third-party hardware, confirming that strategic partnerships with vendors like NVIDIA remain central to its overall compute strategy.

Why it matters

  • Demonstrates frontier AI labs actively diversifying away from pure reliance on NVIDIA for high-density inference workloads.

  • Improved work-per-watt efficiency helps lower operational unit economics for running complex, high-latency reasoning agents.

  • Signals rising custom silicon competition that could put margin pressure on incumbent GPU hardware vendors long-term.

Source: theverge.com