On June 24, 2026, OpenAI and Broadcom unveiled Jalapeño — OpenAI's first Intelligence Processor, an ASIC built specifically around the demands of large language model inference. The chip went from initial design to manufacturing tape-out in just nine months, a development cycle that both companies describe as potentially the fastest ever achieved for a high-performance advanced semiconductor.
The architecture is purpose-built to reduce data movement between compute and off-chip memory, achieving realized utilization far closer to theoretical peak than conventional GPU alternatives. Engineering samples running GPT-5.3-Codex-Spark workloads already show ~50% lower cost per inference token versus current state-of-the-art, according to Broadcom CEO Hock Tan.
Manufactured at TSMC's 3nm node with eight HBM stacks, Jalapeño sits at the core of a multi-generation platform. Initial deployment in gigawatt-scale data centers with Microsoft and partners is targeted by end of 2026 — Microsoft is expected to purchase roughly 40% of the first production run. OpenAI's own models assisted in the chip design process, creating a self-reinforcing loop where AI is now accelerating the infrastructure used to run future AI.
OpenAI President Greg Brockman called the effort part of a "full-stack infrastructure strategy" — spanning chip architecture, kernels, memory, networking, and deployment. By controlling the inference pipeline, OpenAI aims to match the unit-economics advantages long held by Google (TPUs) and Amazon (Trainium).
Anthropic sent a letter to the US Senate Banking Committee accusing Chinese tech giant Alibaba and its Qwen AI lab of conducting "the largest known distillation attack" on the company to date. Between April 22 and June 5, 2026, operators affiliated with Alibaba generated over 28.8 million exchanges with Claude through nearly 25,000 fraudulent accounts — systematically extracting the model's capabilities in advanced coding, multi-step reasoning, and long-horizon agentic tasks.
Distillation attacks train a less capable model on the outputs of a more advanced one, allowing a competitor to inherit frontier capabilities without incurring the underlying research and compute costs. Anthropic estimates the effort targeted capabilities approaching those of its Mythos Preview model — the same model the US government subsequently placed under export controls on national security grounds.
This follows Anthropic's February disclosure of distillation campaigns by DeepSeek, Moonshot, and MiniMax. Alibaba shares fell nearly 4% in Hong Kong trading on the news. An Alibaba spokesperson did not immediately respond to requests for comment.
Announced June 24, 2026, computer use is now a built-in tool inside Gemini 3.5 Flash — not a separate model, but a native capability alongside function calling and Search grounding. Developers building automation agents can now handle screen interactions, web navigation, and knowledge retrieval in a single model call. On OSWorld-Verified, Gemini 3.5 Flash scores 78.4%, just 0.3 points behind GPT-5.5's 78.7%.
The pricing gap is the sharper story: Gemini 3.5 Flash costs $1.50 per million input tokens versus GPT-5.5's $5, and $9 per million output tokens versus $30 — roughly 3x cheaper on output for high-volume agentic workloads. Google has also introduced adversarial training and optional enterprise safeguards that require human confirmation before irreversible actions and auto-halt on detected prompt injection.
The feature is available immediately through the Gemini API and the Gemini Enterprise Agent Platform. Reference implementations are on GitHub, and early adopters are already using it for continuous software testing and enterprise knowledge workflows.
Agility Robotics is going public via SPAC at a $2.5 billion valuation, raising more than $620 million in the process. A Foxconn-led investor group is contributing approximately $200 million to the round — a significant signal from the world's largest contract manufacturer, which has been publicly exploring robotics as a hedge against rising labor costs.
The Digit robot is already deployed at nine customer sites, with over 65,000 logged operating hours — one of the largest real-world deployment datasets for any bipedal robot. The public listing positions Agility among the first wave of humanoid robotics companies to reach public markets, ahead of several better-capitalized rivals still in pre-commercial development.
The move reflects accelerating investor confidence in physical AI commercialization: robots that can perform warehouse and logistics tasks at scale without purpose-built infrastructure are drawing production commitments, not just pilot contracts.
Noam Shazeer — co-author of the 2017 "Attention Is All You Need" paper that introduced the Transformer architecture underpinning every major AI model — announced he is leaving Google DeepMind to join OpenAI as Lead for Architecture Research. Google paid approximately $2.7 billion in 2024 to bring him back from Character.AI. He lasted under 22 months. Sam Altman called it a hire he had "wanted since the very beginning."
Separately, Jonas Adler and Alexander Pritzel — both senior researchers with ties to the AlphaFold protein-structure program — are reportedly preparing to join Anthropic. The combination of departures signals that DeepMind's research talent pipeline is under significant pressure from labs offering more autonomous research environments and more direct equity upside ahead of anticipated IPOs.
OpenAI's Codex usage data released this week provides context for what this talent is being recruited to build: median internal output tokens grew 56× between November 2025 and June 2026 across Engineering, Research, Customer Support, and Legal teams.
Sail Research, a San Francisco infrastructure startup purpose-built for long-horizon AI agents, launched from stealth on June 25 with $80 million in seed and Series A funding at a $450 million valuation. The Series A was led by Kleiner Perkins; the seed by Sequoia. Additional investors include Redpoint, Theory Ventures, CRV, and angels including Alphabet Chairman John Hennessy and Intel CEO Lip-Bu Tan.
The core insight: current inference infrastructure was designed to minimize latency on a single conversational request. Agents that spend hours reading entire codebases, screening hundreds of candidates, or running deep research tasks are a fundamentally different workload — one that needs throughput and sustained parallel compute, not first-token speed. Sail rebuilt the inference stack around that constraint and wraps it in Sailboxes, persistent sandbox environments that run for hours or days and charge only for active compute time.
The platform already processes trillions of tokens per week, with paying customers including Detail.dev (AI code review agents that audit entire codebases) and Parallel Web Systems. On BrowseComp-Plus deep research evaluation, Sail scored 90.72% accuracy at up to 10× lower cost than leading alternatives.
Anthropic's Claude Code is now available as a Slack integration — enabling engineering teams to trigger, monitor, and collaborate on agentic coding tasks directly from within Slack channels without leaving their primary communication tool. The integration arrives alongside Notion's developer platform and Figma Config 2026 updates, all pointing toward a shared trend: the agent layer is collapsing into the workspace layer.
OpenAI's internal Codex data adds quantitative weight to this shift. Median internal output tokens grew 56× between November 2025 and June 2026 — with Research, Engineering, Customer Support, and Legal all showing material adoption. The data frames AI-assisted workflows not as experimental adoption but as operational infrastructure at frontier labs.
The Gemini Enterprise Agent Platform, relaunched from Vertex AI, and beehiiv's new Group Subscriptions feature — targeting company-wide newsletter distribution — both reflect the same dynamic: AI tooling is being productized for organizational buyers, not just individual users.