Technology News • • 5-8 minutes

Jalapeño Enters Deployment: OpenAI's Inference Chip Runs on AMD EPYC Turin Hosts

Diego Cortés
Diego Cortés
Full Stack Developer & SEO Specialist
Share:
Jalapeño Enters Deployment: OpenAI's Inference Chip Runs on AMD EPYC Turin Hosts
Image generated with AI

OpenAI is now deploying Jalapeño, its inference chip, alongside AMD EPYC Turin CPUs with about 1.5 TB of memory per host. The story from October 2, 2026 isn't the chip itself, it's how the system around it is put together.

What Was Confirmed on October 2

Jalapeño Enters Internal Deployment and the News Is the Host

On Friday, October 2, Tom's Hardware reported that OpenAI's Jalapeño ASICs are already deployed alongside AMD EPYC "Turin" processors acting as hosts. The deployment is internal: there is no public offering attached, no announced price and no unit count. SemiAnalysis described the rack-scale configuration, and the CPU choice was confirmed in an interview with Richard Ho, OpenAI's VP and Head of Hardware.

It helps to separate three layers that headlines tend to blend: what an analyst reports, what the company says in an interview and what the press interprets. The easy headline is "OpenAI stops depending on Nvidia" and the lazy one is "OpenAI has its own chip". Both are wrong.

AMD EPYC Turin Instead of Nvidia Vera: the Three Reasons Richard Ho Gave

The obvious question is why a company with a deep relationship with Nvidia builds its first in-house chip on AMD CPUs. Ho's answer is less epic than the headline and far more useful: platform maturity, the partners' prior experience with Turin, and the will to move fast without taking on unnecessary design risk.

On Nvidia's Vera CPU, Ho placed it "a little bit behind" on that maturity level as a standalone option. That is a statement about platform maturity in the host role, not a verdict on the GPU and not a break with Nvidia: the same deployment coexists with the compute contracts OpenAI keeps with the vendor.

Inside the System

According to SemiAnalysis, each host carries two Turin-class CPUs, about 1.5 TB of DRAM and 400G front-end networking. Each accelerator tray groups several Jalapeño chips into a chassis, and a full rack adds up to a high number of XPUs. The exact tray and rack figures should always be attributed to their source: they are the analyst's description, not an OpenAI statement.

In one sentence: inside an inference rack, four pieces decide the cost, the accelerator, the host CPU, the memory and the network. The chip is only the first.

What Jalapeño Is and What It Is Not

An Inference ASIC, Not a Training Chip

Jalapeño is an ASIC, a chip designed for one job. Here, serving already-trained models as efficiently as possible. It does not train them. Giving up the flexibility of training is exactly what buys work per watt and latency, the two metrics you pay for once a model is in production.

Broadcom, TSMC N3P and a 16-Month Tape-Out

The chip was shown publicly at Hot Chips 2026 on August 25, not on October 2. It is designed by OpenAI and built with Broadcom on TSMC's N3P process, with a 16-month tape-out. One detail that surprises many people: it is not tuned for OpenAI's own models, it is meant to serve open models. That is part of the plan, not an accidental limitation.

The Numbers Circulating Come From a Third Party

The most repeated performance figures are 13.4 PFLOPs in MXFP4 at 700 W and HBM4 at 15.4 TB/s, with attention and MoE kernels 1.5x to 1.8x faster than hand-written implementations. All of them come from SemiAnalysis and InferenceX, measured on engineering samples and open models, and the analysis itself admits the comparison against Blackwell is incomplete. Read them always as a third-party measurement, never as an official OpenAI figure.

Efficiency sits in the same bucket: up to 1.9x more work per watt against GB200 and GB300. That is a measurement on open models, not a production promise.

Why the Host Matters More Than It Looks

What the CPU Actually Does in an Inference System

The host CPU does not generate tokens. It prepares input batches, tokenizes, moves data between memory and accelerators, runs the code that isn't a matrix operation, and handles networking and orchestration. If the host falls short, accelerators wait; if it is oversized, you pay energy for work that produces no tokens.

Host Memory and Accelerator Idle Time

The 1.5 TB per host is not decoration. Models, context caches and working data live in memory, and the closer they sit to the accelerator, the fewer milliseconds are lost moving them. It is the same bottleneck logic we already covered when writing about the memory shortage running to 2028.

Choosing a Host Is Not Choosing a GPU

Running Jalapeño on Turin does not mean OpenAI is leaving Nvidia. The company keeps buying and contracting compute from the vendor, as shown by the $180 billion contracted with Anthropic. What changed is which CPU accompanies one specific accelerator in one specific deployment.

What Changes for People Serving Models or Paying the Bill

Cost Per Token vs Cost Per Watt

When picking an inference provider, look at several things at once:

  • The real cost per token, not just the model's list price.
  • Cost per watt and consumption per megawatt, which decide the bill at scale.
  • Latency and the ability to sustain long context without degrading.
  • What fraction of the theoretical peak turns into real tokens.

The Software Layer Decides How Much of the Peak You Get

OpenAI writes its own kernels and its own language, Gluon, instead of relying only on Nvidia's. It is not a new chip or a headline: it is the layer that turns the silicon's theoretical peak into real work. A fast accelerator with weak software performs below a modest one with tuned kernels, a lesson that holds for any inference stack.

A Market With More Than One Silicon

The AI chip map no longer has a single protagonist: Nvidia, AMD and Broadcom are all in, plus the labs' own chips, from Meta's MTIA to Huawei's Ascend 960DT or Alibaba's Zhenwu V900. For developers, the decision splits in two: which model you use and which infrastructure it runs on. They are different choices with different costs.

The software layer that ships with this hardware is moving fast too, as seen at OpenAI's DevDay 2026.

What Is Still Unknown

There is no public price, no unit count and no confirmed savings. It is also unclear how much of the measured performance holds up in large-scale production with frontier models. And the August measurements should not be read as October results: they are different moments, and mixing them up is a mistake.

Conclusion

Jalapeño is a real step, but the most instructive part of the story is the plumbing. The cost of serving a model is not decided by one piece: it is decided by the combination of accelerator, host, memory, network and kernels. Comparing bare chips is like comparing cars by engine power while ignoring weight and fuel consumption. If you want to know how AI is paid for from the inside, keep reading the blog: we take the infrastructure apart piece by piece.

Categories