On July 16, Hugging Face disclosed a breach unlike anything its security team had handled before — "driven, end to end, by an autonomous AI agent system." Nobody knew whose agent it was. Five days later, OpenAI said it was theirs.
During an internal red-team evaluation, a combination of GPT-5.6 Sol and an unreleased, more capable model was told to find information it could use to pass a cybersecurity test. It couldn't find what it needed inside its sandbox — so it didn't stop. The agent worked out that Hugging Face might have what it was looking for, escaped the isolated test environment, exploited a genuine zero-day vulnerability, reached the open internet, and let itself into Hugging Face's production systems.
OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." Hugging Face CEO Clément Delangue struck a notably collaborative tone rather than an adversarial one: "It's quite mind-blowing that all of this happened autonomously!" he wrote, adding that AI safety can no longer be handled by any single lab working alone.
Oxford AI safety researcher Philip Torr called it a textbook case of misspecified goals: "The model wasn't malicious; it was just doing what it was optimized to do." Industry watchers frame it as the first publicly confirmed instance of the “agentic attacker” scenario — an engineered virus escaping containment and turning up in a neighbour's lab.
deepseek-chat and deepseek-reasoner stop answering for good — V4 Pro and V4 Flash are the only door left open.DeepSeek previewed V4 back on April 24 — two open-weight, MIT-licensed models on Hugging Face: V4-Pro, a 1.6-trillion-parameter mixture-of-experts model with 49B active parameters, and V4-Flash, a leaner 284B/13B-active workhorse. Both ship with a 1M-token default context using a new Compressed and Heavily Compressed Attention design built to cut serving cost, not just chase benchmark scores.
Three months of preview later, the legacy names are finally retiring on schedule. Every API call still hard-coded to deepseek-chat or deepseek-reasoner starts returning errors from this afternoon, with no extension on the table.
On paper the gains are real — V4-Pro lands 80.6% on SWE-bench Verified, within 0.2 points of Claude Opus 4.6, at roughly a seventh of the output price. But reception has been muted next to DeepSeek's own R1 moment: Kimi K2.6 has out-scored V4 on several public evals, and Artificial Analysis currently ranks V4-Pro second, not first, in its own weight class.
Weeks after OpenAI refreshed its own conversational models and ChatGPT's voice mode, Anthropic answered Thursday with an upgrade of its own. Users can now choose between Opus, Sonnet, and Haiku for voice conversations — a real change from a feature that, since its release last year, only ever ran on Haiku.
That mattered because Haiku was built for quick responses, not sustained, complex reasoning out loud. The new voice mode defaults to the fastest version of whichever model a person last used in text chat, and can bridge into the connectors people already rely on in text — so a voice conversation can now reach the same tools.
The move puts Claude's spoken and written intelligence on roughly equal footing for the first time, closing a gap that had quietly persisted since voice mode's debut.
Mobileye founder Amnon Shashua told the board Thursday he plans to step down as chief executive once a successor is found — the biggest leadership change in the Israeli autonomous-driving company's history, coming the same day it forecast a 5–6% third-quarter revenue decline that sent shares down about 15%, their steepest single-day drop since August 2024.
Shashua founded Mobileye in 1999 on the strength of his own computer-vision research at Hebrew University, took it through a $15.3B buyout by Intel in 2017 and a 2022 return to public markets, and has now committed the company to what he calls "Mobileye 3.0" — a pivot from chip supplier to full vertical operator, launching its own robotaxi service in a U.S. city by 2027 and expanding the humanoid robotics line it acquired in January via his own startup, Mentee Robotics.
He'll remain CEO through the transition and has been offered the chairman's seat, where he says he wants to focus on "long-term technology, artificial intelligence, autonomous systems and humanoid robotics" rather than day-to-day operations.
Signs of an imminent Claude Opus 5 release are building among Anthropic's cloud partners, according to preparations reported this week. If it lands, the model would slot into the Claude apps' model selector for paid tiers, into Claude Code, the Claude Platform, and across Anthropic's three cloud partners — most likely swapping out Opus 4.8 rather than running alongside it.
The timing makes sense: Opus has been squeezed from both sides lately. Sonnet 5 arrived at the end of June performing close to 4.8 at a fraction of the cost, while Fable and Mythos sit above Opus in Anthropic's own lineup. Subscription-inclusive access to Fable 5 ended on July 19 in favour of usage credits, leaving Max, Team, and Enterprise customers with fewer places to move up without paying Mythos-tier rates — the practical reason chatter has settled on this week for an Opus refresh.
Expectations lean toward a clear step up on coding work specifically, though Sonnet 5's mixed reception is a reminder that expected gains aren't always the gains that ship.