OpenAI’s custom inference chip, called Jalapeño, emerged in greater technical detail around August 25, 2026, after a presentation at Hot Chips and an assessment by semiconductor research outlet SemiAnalysis. The processor is an application-specific integrated circuit built for large-language-model inference rather than training, and was developed with Broadcom.
According to SemiAnalysis, design work began in mid-2024 and reached manufacturing tape-out in roughly 16 months. The outlet characterized Jalapeño as a general inference processor rather than hardware restricted to OpenAI’s own models. It said it visited an OpenAI laboratory and observed the chip running its InferenceX benchmark, as well as open models including DeepSeek R1, Kimi-K2.5 and GPT-OSS.
The early results focus heavily on output per unit of power. SemiAnalysis reported that Jalapeño led the other tested systems in token throughput per megawatt across most measured conditions, including both low-latency and high-throughput settings. It said DeepSeek R1 exceeded 700 tokens per second for one user at concurrency one, while Kimi-K2.5 and GPT-OSS ran at about 1,400 tokens per second per user. The tested configuration used single-token prediction rather than multi-token prediction or speculative decoding. The report also said GSM8k evaluation results were comparable with Nvidia hardware.
Those numbers come with significant qualifications. SemiAnalysis said OpenAI supplied all performance figures, even though its analysts witnessed some InferenceX runs in person. The outlet did not execute its full InferenceX suite and had not seen results from AgentX, its preferred benchmark for longer-context, multi-turn production workloads. Such workloads can expose constraints involving routing, prefix caches, memory management and offloading that shorter tests may not capture.
The comparison with Nvidia’s Blackwell generation is also imperfect. Jalapeño uses HBM4 memory and remains at the engineering-sample stage, while SemiAnalysis said Nvidia’s newer Vera Rubin systems were beginning to reach customers. The outlet argued that Rubin, which also uses HBM4, is the more relevant competitor. It further noted that other vendors have published tests on newer and larger models than those demonstrated on Jalapeño.
The power focus reflects practical deployment limits. Utility connections, cooling systems and backup generation can constrain a data centre’s available megawatts long after additional accelerator orders become feasible. In that setting, tokens produced per joule can matter directly to the amount of inference a fixed facility supports.
OpenAI’s design priority appears to be performance per watt, reflecting limits on data-centre power rather than only chip acquisition costs or floor space. Higher token output within a fixed power envelope could increase serving capacity without waiting for additional grid connections and cooling infrastructure. Still, the evidence available at this stage is an observed but incomplete benchmark exercise, not an independent, comprehensive review. Production availability, performance across broader models and direct comparisons with Rubin remain unresolved.



