The Spectrum Dispatch News

technology

OpenAI's Jalapeño Inference Chip Outperforms Nvidia Blackwell on Efficiency

OpenAI announced Jalapeño, a custom AI inference chip developed with Broadcom, that achieves better performance-per-watt than competing chips including Nvidia's Blackwell and Vera

OpenAI's Jalapeño Inference Chip Outperforms Nvidia Blackwell on Efficiency

OpenAI has unveiled Jalapeño, an inference chip jointly developed with Broadcom and announced at Hot Chips. According to Semianalysis, the design process began in mid-2024 and reached manufacturing tape-out in approximately 16 months, an unusually fast development cycle for custom chip design.

OpenAI’s Jalapeño Inference Chip Outperforms Nvidia Blackwell on Efficiency

Unlike typical first-generation chips, Jalapeño performs competitively against established competitors. According to Semianalysis’ testing, the chip “beats every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models.” The chip uses HBM4 memory, comparable to flagship GPUs from Nvidia and AMD.

The key performance metric OpenAI prioritizes is tokens per megawatt (tok/MW), reflecting datacenter power constraints. According to Semianalysis, Jalapeño achieves superior perf/W across nearly all scenarios without Multi Token Prediction (MTP) optimization. On single-token prediction, the chip demonstrates strong performance, hitting over 700 tokens per second per user at low concurrency on the DeepSeek R1 model, and approximately 1,400 tok/sec/user on other models including Kimi-K2.5 and GPT-OSS.

Designed as a generalized inference chip rather than one tailored exclusively to OpenAI models, Jalapeño successfully ran third-party workloads including Semianalysis’ InferenceX benchmark suite and even a port of the video game Doom.

However, Semianalysis notes important caveats. All performance numbers were provided by OpenAI, though the publication verified the InferenceX runs in-person at OpenAI’s lab. The publication did not run Semianalysis’ full benchmark suite or observe AgentX results, which the publication states better reflects realistic production workloads with long context and multi-turn characteristics. Additionally, Semianalysis argues that comparing Jalapeño to Blackwell is “somewhat incomplete,” suggesting Vera Rubin—which also uses HBM4—is a more appropriate comparison point. Vera Rubin systems are already shipping to customers, while Jalapeño remains at the engineering sample stage.

Regarding model capability, the models tested on Jalapeño are not among the latest frontier models. Semianalysis notes that Nvidia and AMD have published results on larger, newer models like DeepSeek V4 Pro and Kimi K3 using AgentX benchmarks.

Key facts

  • Jalapeño was developed by OpenAI with Broadcom, from design initiation to manufacturing in ~16 months
  • The chip achieves superior performance-per-watt compared to Nvidia Blackwell, AMD, and Google chips according to Semianalysis testing
  • Jalapeño uses HBM4 memory and is designed as a generalized inference chip, not exclusively for OpenAI models
  • Performance testing was conducted in-person at OpenAI labs using Semianalysis’ InferenceX benchmark
  • Vera Rubin is a more comparable competitor than Blackwell as both use HBM4; Vera Rubin is already shipping while Jalapeño remains in engineering samples

Sources

← All posts