OpenAI unveils Jalapeño, its first custom AI inference chip built with Broadcom
The ChatGPT maker is moving down the silicon stack with a purpose-built ASIC it says beats current state-of-the-art on performance per watt, with gigawatt-scale deployment planned by end of 2026.
What matters
- OpenAI announced Jalapeño, its first custom AI inference processor, developed in partnership with Broadcom and Celestica.
- The ASIC is purpose-built for LLM inference, with early testing claiming performance per watt substantially better than current state-of-the-art.
- The chip went from design to production in nine months, a timeline OpenAI says was accelerated by its own models.
- Gigawatt-scale deployment with data center partners including Microsoft is planned to begin by the end of 2026.
- Broadcom CEO Hock Tan described the collaboration as the start of a multi-generation roadmap for scaling AI physical infrastructure.
What happened
On June 24, 2026, OpenAI and Broadcom jointly unveiled Jalapeño, OpenAI's first custom "Intelligence Processor" — an application-specific integrated circuit (ASIC) designed from the ground up for large language model inference. The chip was co-developed with Broadcom handling silicon implementation and high-performance networking, and Celestica contributing board and rack system integration.
OpenAI says Jalapeño went from design to production in just nine months, a timeline the company claims was accelerated by its own models. Early lab testing shows the first-generation accelerator running ML workloads at production target frequency and power, with performance per watt "substantially better than current state-of-the-art." No independent benchmarks have been published yet.
The chip was formally delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom CEO Hock Tan and President Charlie Kawwas. OpenAI describes Jalapeño as the first in a multi-generation compute platform the two companies are building together.
Deployment is planned at gigawatt scale with data center partners — including Microsoft — beginning by the end of 2026, according to a Broadcom press release.
Why it matters
This is OpenAI's most significant step yet toward vertical integration of its entire AI stack — from consumer products like ChatGPT, to frontier models, to the silicon that runs them. Until now, OpenAI has relied heavily on Nvidia GPUs for inference. A purpose-built inference ASIC could meaningfully reduce operating costs, improve latency, and give OpenAI more control over capacity planning at a time when GPU supply remains constrained.
The move mirrors a strategy already pursued by Google (TPU) and, to a lesser extent, Amazon (Trainium/Inferentia) and Meta (MTIA). What's notable here is the speed: nine months from design to tape-out is aggressive by semiconductor standards, and OpenAI credits its own models with accelerating the process — a claim that, if substantiated, points to AI-assisted chip design becoming a competitive lever.
Broadcom CEO Hock Tan framed the partnership as a long-term commitment: "Our collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure required for the next decade of AI. This is just the beginning of a multi-generation roadmap."
For Nvidia, the competitive signal is nuanced. Training the next generation of frontier models still requires general-purpose GPUs at massive scale, and OpenAI is unlikely to abandon Nvidia for training anytime soon. But inference — the repetitive, high-volume work of serving model responses to users — is exactly where custom silicon tends to pay off, and OpenAI is one of the world's largest inference consumers.
What to watch
- Deployment timing: Whether gigawatt-scale deployment with Microsoft and other partners actually begins by end of 2026, and what workloads Jalapeño handles first.
- API-level signals: Watch OpenAI's API changelog for latency improvements, throughput increases, or pricing changes that could indicate Jalapeño is serving production traffic.
- Benchmark transparency: Whether OpenAI or Broadcom publish independent or third-party benchmarks, or whether performance-per-watt claims remain internal.
- Nvidia's response: Any competitive positioning around inference-specific GPUs or software optimizations.
- Multi-generation roadmap: Details on Jalapeño's successors and whether the platform expands beyond inference.
What to do next
Developers
Monitor OpenAI's API changelog and developer documentation for any inference performance, latency, or pricing updates that may signal Jalapeño deployment.
Custom inference silicon typically brings throughput or efficiency gains that surface as API-level changes — faster response times, higher rate limits, or lower costs — before they are publicly benchmarked.
Founders
Reassess build-versus-buy assumptions around inference infrastructure, especially if your product depends heavily on OpenAI API costs.
If Jalapeño meaningfully reduces OpenAI's inference costs, those savings may eventually be passed to API customers, shifting unit economics for AI-native startups that rely on OpenAI's platform.
PMs
Track whether OpenAI introduces tiered inference options or performance guarantees that could map to Jalapeño-powered endpoints.
Custom silicon may enable differentiated service tiers — lower-latency or higher-throughput plans — that affect product roadmaps depending on OpenAI's API.
Investors
Evaluate the competitive implications for Nvidia and other GPU-dependent inference providers as frontier labs vertically integrate into silicon.
OpenAI joining Google and others in custom inference ASICs signals a structural trend that could compress demand growth for general-purpose AI GPUs in inference workloads, though training-side demand remains distinct.
Operators
Review current GPU-based inference capacity plans and model whether a shift toward ASIC-backed inference at major providers could alter cloud pricing or instance availability.
If hyperscalers and frontier labs increasingly run inference on custom silicon, the secondary market and cloud availability of GPU instances may shift, affecting capacity planning and cost projections.
Testing notes
Caveats
- Jalapeño is an internal infrastructure component, not a publicly available product or API endpoint.
- No benchmarks, SDKs, or developer-facing tools have been announced, so direct testing is not possible at this time.
- Any observable impact would come indirectly through changes in OpenAI API performance, pricing, or availability.
- OpenAI's performance-per-watt claim is based on early testing and has not been independently verified.