Anthropic's newest mid-tier model can plan multi-step jobs, operate a browser or terminal, and run autonomously at a level that, until recently, required a much larger model. Early testers said the model finishes complex tasks other Sonnets would abandon halfway, and checks its own output without being asked to.
On agentic coding benchmarks Sonnet 5 posts 63.2% against Opus 4.8's 69.2%, a meaningful narrowing of the gap at a fraction of the price — $2 per million input tokens through August 31, rising to $3 afterward. It ships as the default model for Free and Pro users and is now live in Claude Code, GitHub Copilot, and AWS Bedrock.
Anthropic frames the safety profile as an asset too: Sonnet 5 shows lower rates of undesirable behavior than its predecessor and carries far weaker cybersecurity capability than Opus-tier or Mythos-tier models, keeping its risk surface deliberately narrow even as its task range widens.
Most enterprise AI agents run inside a static, hand-built harness — the prompts, tools, and retry logic that connect a model to its environment. Engineers have to manually patch that scaffolding every time a task or a document format changes, and the fixes rarely generalize.
Self-Harness flips this: an agent proposes narrow, failure-specific edits to its own harness, validates each one against regression tests, and only keeps changes that improve results without breaking anything else. The lead researcher, Hangfan Zhang, told VentureBeat that a skilled engineer can still out-propose an LLM today — the real gain comes from replacing ad hoc debugging with a repeatable, evidence-based loop.
A related effort from Xiaomi, HarnessX, tested the same idea across Claude Sonnet 4.6, GPT-5.4, and an open-weight Qwen model, and found the smallest models gained the most — evidence that better scaffolding, not just a bigger model, may be the cheaper lever for enterprises to pull next.
The dura mater is the tough membrane that shields the brain — more than ten times thicker than Neuralink's electrode threads, which are themselves finer than a human hair. Until now, reaching the cortex meant a durectomy: cutting and removing a coin-sized patch of that membrane, one of the most delicate steps in the whole procedure.
In May, a team at Toronto Western Hospital carried out Neuralink's first transdural implant, guiding the threads through the intact dura using dye-based imaging and a redesigned needle instead of an open surgical window. The trial participant was moving a cursor with their thoughts within an hour, with recovery proceeding normally.
Neuralink's own framing is blunt: "the best step is no step at all." Removing the durectomy cuts one of the most manual, error-prone parts of the operation, which the company says points toward safer, more repeatable surgeries as it works to scale beyond its current 21 trial participants.
A University of Minnesota team led by Kate Adamala and Aaron Engelhart built SpudCell from a lipid membrane, 36 purified enzymes, and a 90,000-base-pair genome spread across nine DNA molecules — no unknown or biologically scavenged parts, everything chemically defined from the ground up.
Over five generations, the system grew, replicated its genome, divided without relying on a cytoskeleton, and even showed natural selection: a genetic tweak that sped up growth let the modified line outcompete the original, especially once nutrients ran short.
Adamala compares it to the Wright Flyer, not the Dreamliner — a first proof that the fundamental behaviors of life don't require "a mysterious magical spark," years away from anything resembling a self-sustaining organism. To speed outside verification, she and Stanford's Drew Endy launched Biotic, a public-benefit institute releasing SpudCell's full data and protocols in the open.
Glen Givens of Poway, California, told FOX 5 he was contacted at work by a man claiming to be a police officer, who said his daughter had caused a serious crash injuring a pregnant woman. The caller knew a private detail about a childhood medical condition — enough that Givens says it "checked out."
Then a woman came on the line. Givens believed it was his daughter. It wasn't. He handed over $18,000 in cash, nearly all of it, before realizing the entire call had been engineered — voice, urgency, and personal detail — to bypass his usual caution.
Cases like this are multiplying: consumer research cited by industry trackers finds roughly one in three people who engage with an AI-cloned voice call end up losing money, since scrapers need only a few seconds of public audio — a video, a voicemail, a reel — to produce a convincing clone.