Coding Horizon

OpenAI Just Put Nvidias AI Chip Monopoly On Notice

Sources for every figure, date and benchmark this video puts on screen. Checked 26 August 2026.

The chip

Jalapeno is OpenAI’s first custom accelerator, designed with Broadcom, and it is for inference rather than training. Broadcom supplies the silicon implementation plus the networking and connectivity; Celestica does board, rack and system integration. OpenAI describes it as generation one of a multi generation platform.

Deployment begins inside OpenAI’s own infrastructure by the end of 2026, in limited volume, with a larger rollout into 2027. OpenAI says it will keep deploying Nvidia and other accelerators alongside it.

Gen 2 is described as deep in development and Gen 3 as taking shape. Gen 2 targets performance per watt; Gen 3 targets economical low latency serving through aggregate HBM bandwidth rather than raw bandwidth alone.

The published part specifications

Presented by OpenAI at Hot Chips 2026 and reported from the session.

Figure on screen Value Where
Package power 700 W (never above 550 W on any tested workload) Hot Chips 2026, Tom’s Hardware
Memory 216 GiB HBM4 Hot Chips 2026
Memory bandwidth 15.4 TB/s per chip Hot Chips 2026
Peak matrix compute 13.4 PFLOP/s (mxfp4) Hot Chips 2026
Local domain 128 chips at 600 GB/s Hot Chips 2026
Global domain 2,048 chips at 200 GB/s Hot Chips 2026
Full system 27 EFLOP/s, 432 TiB Hot Chips 2026
Comparison power GB200 1,200 W, GB300 1,400 W Tom’s Hardware

The KV cache is kept local on Jalapeno rather than moved between specialised systems, and idle silicon blocks are powered down per phase. That is the design point the video’s serving system beats draw.

The benchmark

InferenceX is a public inference benchmark suite from SemiAnalysis. SemiAnalysis ran it with OpenAI engineers in OpenAI’s own lab. The published comparison covers three open models: GPT OSS 120B, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5 at 1 trillion parameters.

Across the suite: 1.5x to 1.9x more throughput per kilowatt at peak, and 1.7x to 3.6x lower end to end latency, against Nvidia GB200 and GB300 rack systems. Minimum time between tokens improved 2.7x to 4.1x across the three models.

The DeepSeek R1 numbers, which the video states largest

Every figure the shots render for DeepSeek R1 against Nvidia GB300:

On screen Value Ratio as drawn
Peak mixed tokens per second per kilowatt 19,641 against 11,781 1.7x, divided from the pair
End to end latency 1.65 s against 5.99 s 3.6x, divided from the pair
Minimum time between tokens top of the published range 4.1x

The shots compute every multiple from the two values beside it rather than printing a typed number, so the chart and the caption cannot disagree.

The other custom silicon named on screen

Google TPU, Amazon Trainium and Inferentia, Microsoft Maia and Meta MTIA are all real, shipping programmes, and all of them are captive to their builders rather than rentable as parts.

What the comparison is conditioned on

Stated plainly because the video’s own caveat chapter draws it:

Not chased to a primary source

These appear in the narration and are illustrative rather than measured. The shots that carry them draw a mechanism rather than printing a figure as fact.