OpenAI and Samsung announce deepened chip collaboration and large ChatGPT Enterprise deployment
OpenAI Unveils Jalapeño, Its First Custom Inference Chip
Building its own inference chip gives OpenAI a potential lever on compute cost and throughput independent of third-party suppliers. The efficiency and usage disclosures show the scale at which these hardware and software choices operate.
The full picture
OpenAI announced Jalapeño, its first in-house custom inference chip, which it says delivered 1.5 to 1.9 times as much peak token throughput per watt and 1.7 to 3.6 times lower end-to-end latency than commercial systems in InferenceX tests across three public models. Deployment is planned by year-end alongside accelerators from NVIDIA, AMD, and other partners. OpenAI also disclosed that GPT-5.6 Sol reduced its end-to-end serving costs by 20% and increased token-generation efficiency by more than 15%. On the scale side, OpenAI's products reach more than one billion weekly active users and 2.5 million businesses, and its research organization uses 3.1 agent-workdays of effort for every workday of human labor. Separately, OpenAI and Samsung announced a deepened collaboration covering next-generation chip research and production, alongside one of Samsung's largest ChatGPT Enterprise deployments.
How it developed
OpenAI CFO Sarah Friar discloses usage tiers: Pro subscribers average 11x daily activity of free users
OpenAI publishes case study on GPT-5.6 Sol running quantum computing experiments at MIT
Sources
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free