OpenAI's First Custom Chip Reportedly Posts 104x Efficiency Edge on Inference Benchmark
OpenAI's 'Jalapeño' silicon, built with Broadcom, allegedly hit up to 104x more tokens per watt than rivals on SemiAnalysis's InferenceX test — with small-volume shipments expected by year end.
OpenAI's long-rumored bid to escape its dependence on Nvidia moved from roadmap to benchmark this week. According to @DreyXAI, the company's first custom AI accelerator — codenamed 'Jalapeño' and co-designed with Broadcom — reportedly hit up to 104x more tokens per watt than competing hardware on SemiAnalysis's InferenceX efficiency test. The same report indicates the chip will ship in small volumes before the end of the year, which reads less like a commercial launch than a controlled pilot to validate the design in production.
The number itself deserves scrutiny. A 104x figure is not the kind of margin one silicon generation typically opens over another on general-purpose workloads, and it almost certainly reflects a narrow, inference-specific benchmark tuned to the chip's strengths rather than a broad measure of capability. Tokens-per-watt is an efficiency metric, not a throughput or latency one — it says how cheaply you can serve, not necessarily how fast. Read carefully, the claim is that OpenAI has built something extremely good at one thing: serving its own models at scale, cheaply.
Get our free daily newsletter
Get this article free — plus the lead story every day — delivered to your inbox.
Want every article and the full archive? Upgrade anytime.
No spam. Unsubscribe anytime.