For most of its history, OpenAI pushed intelligence forward while renting the physical limits beneath it. Its models, products and serving software could improve rapidly; the accelerators powering them still came from somebody else. Jalapeño changes that equation.
OpenAI now has working first-generation inference silicon, measured results and a multigenerational platform on its roadmap. That is more than a component announcement. The company behind ChatGPT, Codex and one of the world’s most consequential model platforms has crossed from buying compute to helping design it. The frontier-model race has reached the physical layer.
Field signal/ 01The model race has reached silicon.
OpenAI has crossed a strategic boundary
Jalapeño is an inference ASIC: an application-specific integrated circuit built primarily to run trained language models. OpenAI says it architected the chip around modern LLM inference, informed by its models, kernels, serving systems and product needs. Broadcom handled silicon implementation and contributed networking technology; Celestica contributes boards, racks and systems.
OpenAI does not suddenly own fabrication, the supply chain or a global data-centre estate. Nvidia accelerators and infrastructure partners remain essential. What it now owns is architectural intent: the ability to encode assumptions about interactive answers, long-running agents, token generation and serving volume into the machine itself, then repeat that process across generations. Product demand is no longer merely consuming infrastructure. It is beginning to shape the physics beneath the product.
Google’s full-stack advantage is no longer unique
GPU
Discovery engine- Frontier training
- Changing architectures
- Experimentation
- Broad workloads
Programmability · Optionality · Ecosystem
ASIC
Production engine- Repeated inference
- Serving at scale
- Stable workloads
- Product-specific demand
Latency · Performance per watt · Predictability
Google deployed its first TPU internally in 2015. A decade later, Ironwood was its seventh generation and its first TPU designed specifically for inference. Google could observe demand in Search, Gemini, Workspace and Cloud; change models and serving software; then feed what it learned into the next hardware generation. That loop made Google the clearest proof that AI advantage could compound across an entire stack.
Jalapeño does not erase Google’s decade-long lead. It changes the asymmetry. OpenAI can now connect signals from ChatGPT, Codex and the API to decisions in models, kernels, serving software, memory movement, networking, racks and silicon. Google still has the more mature hardware programme. It no longer has the full-stack arena to itself.
Field signal/ 02Google still has the hardware lead. It no longer has the field to itself.
The model contest is becoming a hardware contest
For years, AI competition was narrated through model launches and benchmark scores. But inference happens every time somebody uses the product. For an assistant, an API or an agent taking many sequential steps, latency compounds, power limits capacity and memory movement becomes a product constraint. A model can improve in the lab while the experience around it remains slow, scarce or expensive.
At frontier scale, the winning unit is no longer the model in isolation. It is the whole request path: product, model, kernels, scheduler, memory, network, rack and accelerator. The company that can co-design those layers can trade what it learns in one layer for an advantage in the next. A product decision can influence a model. A model pattern can influence a kernel. A serving bottleneck can influence silicon.
That makes hardware part of model strategy. The next frontier system will not win only because its weights are better. It will win because the organisation can turn intelligence into a responsive, reliable and economically sustainable product at enormous scale.
ASIC adoption now looks inevitable at frontier scale
AI ASICs are not new. Google has TPUs. Amazon has Inferentia and Trainium. Meta has MTIA. What changes with Jalapeño is the signal. When another major owner of frontier models and global inference demand commits to a multigenerational custom-compute platform, specialised silicon stops looking like one company’s unusual advantage. It starts looking like the destination of every sufficiently large AI platform.
The economics push in that direction. GPUs are programmable, powerful and supported by a mature software ecosystem. That flexibility remains indispensable for training, experimentation and workloads that change faster than hardware can. ASICs win a different game. When a workload becomes stable, repeated and enormous, a narrower architecture can trade flexibility for lower latency, better performance per watt and more predictable unit economics.
The GPU era is not ending. It is splitting. GPUs remain the flexible engine for discovery and breadth. Custom accelerators become the production engine for industrialised inference. CPUs, memory, networking and rack systems bind them together. The future is heterogeneous—but at the frontier, an inference ASIC is starting to look less like an optional optimisation and more like a rite of passage.
What Jalapeño has actually proved
On 25 August, OpenAI published results from SemiAnalysis’s InferenceX fixed-sequence workload. The profile was nominally 8,000 input tokens and 1,000 output tokens using single-token prediction. It tested GPT-OSS 120B against an Nvidia GB200 baseline, then DeepSeek R1 670B and Kimi K2.5 1T against GB300 systems. The comparison used each package’s published power rating: 700 watts for Jalapeño, 1,200 for GB200 and 1,400 for GB300.
Across those three model-and-system configurations, OpenAI reported 1.5 to 1.9 times higher peak mixed-token throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency. Those are meaningful system results. They suggest that OpenAI has working first-run silicon capable of moving the latency-efficiency frontier across several large models, including models developed outside OpenAI.
The boundaries matter. OpenAI published the results, on A0 silicon, with model-specific configurations and selected Nvidia baselines. GPT-OSS was not shown against GB300 in the headline comparison. The efficiency ratios divide by rated package power, not a complete data-centre energy bill. Independent reporting notes that GB300 still leads absolute throughput per package by roughly 20 to 25 percent in the disclosed tests. Performance per watt answers a vital capacity question; it does not mean one package does the most total work.
Field signal/ 03Efficiency tells you how much useful work fits inside a constraint. It is not the same claim as absolute speed or capacity.
What remains unproven
Working silicon and selected benchmark results are not the same as fleet-scale success. OpenAI still has to qualify, manufacture, deploy, programme and operate Jalapeño at scale. The initial deployment target, a more efficient later stepping and the second- and third-generation roadmap remain plans until the hardware is serving substantial production traffic.
Nor do better chip economics automatically mean cheaper ChatGPT or API pricing. Greater efficiency may fund more capacity, faster agents, higher margins, new product modes or lower prices. OpenAI has announced no customer price reduction tied to Jalapeño. The allocation is a business decision, not a benchmark output.
The lesson beyond OpenAI
Most companies should not build a chip. Jalapeño is the visible extreme of a more general principle: own a layer when its constraint is large, repeated, measurable and central to the product promise. The strategic question is not simply build or buy. It is how narrowly you can change your ownership position without adding more risk than advantage.
There are four practical positions.
- Buy when the layer is a commodity, suppliers are healthy and switching remains credible.
- Adapt around it when the constraint can be absorbed through product design, caching, workflow or software.
- Partner when specialised capability matters but ownership would add risk faster than advantage.
- Own when the constraint is enduring, differentiating, measurable and important enough to justify the fixed cost.
Before moving toward ownership, ask four uncomfortable questions. Does this bottleneck recur at enough volume to compound? Can you measure it from the user experience down to the responsible layer? Will controlling it create a durable advantage rather than a temporary optimisation? Do you have the operating capability to maintain the layer after the launch story ends? If any answer is vague, ownership is probably premature. Specialisation always trades optionality for fit.
A practical first move
Start with a constraint ledger. Name the customer promise, the repeated workload beneath it, the limiting resource, the evidence, the cost of the limit and the alternatives. Then run the smallest experiment that distinguishes a supplier problem from an architecture problem or a product-design problem. Sometimes the strategic move is a chip. More often it is a better cache, a different workflow, a tighter model, a partnership or a clearer boundary.
Our earlier field note AI Without the Hype makes the same argument from the product side: technology should earn its place against a real job and explicit constraints. Jalapeño extends that discipline all the way down the stack. The signal is not that every company should build a chip. It is that the frontier AI race is no longer only a model race. The organisations at the front are learning how to turn what a product learns into changes in software, infrastructure and eventually matter.
Sources and boundaries
This field note was sparked by Evolving AI’s video, then independently checked against OpenAI’s June announcement and August results, Broadcom’s corresponding release, Google’s TPU history and Ironwood announcement, AWS Inferentia documentation, Meta’s MTIA programme, the public InferenceX methodology, and independent technical reporting from TechRadar and Tom’s Hardware. The benchmark figures remain OpenAI-reported; links are provided so the workload, baselines and caveats can be examined directly.
The model race has reached silicon. OpenAI is now building on both sides of the boundary.
Dhaval ShahFounder & Principal Architect

