Cerebras Systems filed its S-1 registration statement with the SEC on April 17, targeting a Nasdaq listing under ticker CBRS at a $22-26 billion valuation. The filing revealed $510 million in 2025 revenue (up 76% YoY), a non-GAAP net income of $237.8 million, and a revenue backlog of $24.6 billion. The centerpiece is a multi-year compute agreement with OpenAI valued at over $10 billion, under which OpenAI will deploy 750 megawatts of Cerebras WSE-3 chips. OpenAI also loaned Cerebras $1 billion secured by warrants allowing OpenAI to buy over 33 million shares.
The WSE-3 is the third generation of Cerebras's wafer-scale engine — a single chip 57x larger than Nvidia's H100, with 4 trillion transistors, 900,000 AI-optimized cores, and 44GB of on-chip SRAM delivering 21 PB/s of bandwidth. The chip targets inference workloads, where the AI hardware market has been rotating since 2025. 'Obviously Nvidia didn't want to lose the fast inference business at OpenAI, and we took that from them,' CEO Andrew Feldman told the WSJ. Morgan Stanley, Citigroup, Barclays, and UBS are lead underwriters.
The GPU Technology Conference this quarter marked a turning point for NVIDIA: for the first time, Jensen Huang devoted the majority of his keynote to inference computing rather than training. The shift reflects a structural change in the AI market — as models move from development to production, the compute demand profile flips from training (dominated by NVIDIA) to inference (where custom chips are competitive). NVIDIA's $20 billion technology licensing deal with Groq and the loss of OpenAI's inference business to Cerebras underscore the competitive pressure.
NVIDIA's response is the Vera Rubin platform, already in full production, featuring the Vera CPU and Rubin GPU with six breakthrough chips co-designed for inference workloads. The platform includes 6th-gen NVLink switches, ConnectX-9 networking, and the world's first 200Gb Ethernet co-packaged optics. Despite the challenges, NVIDIA's data center revenue continues to grow, but the margin structure is shifting as inference-optimized chips command different pricing than training GPUs.
OpenAI and Broadcom unveiled the Jalapeño inference chip on June 24 (development details emerged in April), a custom ASIC optimized specifically for transformer-based LLM inference. The chip is designed to reduce inference costs by up to 75% compared to NVIDIA H100s for ChatGPT workloads, with a focus on low-latency token generation rather than raw training throughput. OpenAI has been quietly building a silicon team led by former Google TPU and Apple A-series engineers since 2024.
The Jalapeño chip represents a strategic hedge: while OpenAI's $122 billion funding round includes $30 billion from NVIDIA and commitments to purchase 5GW of NVIDIA compute, the company simultaneously invests in reducing its dependency. The chip will initially be deployed for internal ChatGPT inference before being offered to API customers. OpenAI joins Google (TPU), Amazon (Trainium/Inferentia), Microsoft (MAI), and Meta (MTIA) in building custom silicon optimized for its specific workload profile.