Visa confirmed plans to cut roughly 2,600 positions — about 7% of its global workforce — with reductions concentrated in technology and product teams. In India, employees received termination emails before dawn, a delivery choice that left many at the company's tech hubs in Bengaluru, Mumbai, Chennai and Hyderabad caught off guard.
In an internal memo, CEO Ryan McInerney framed the move around efficiency and reinvesting in growth, noting that artificial intelligence is reshaping how work gets done. People familiar with the plan told CNBC that AI is a significant factor but not the sole driver — overlapping engineering capacity built during the 2021–22 hiring boom is also being normalized.
The cuts reached beyond junior ranks: reports say senior directors, engineering managers and employees with years of tenure were affected. Visa has not disclosed how many India-based workers are part of the reduction, despite the country being a key hub with more than 3,500 employees. The wave mirrors peers — Mastercard trimmed about 4% of its workforce earlier this year and Block cut roughly 4,000 roles in February.
Executive vice president Jay Parikh told Microsoft's Core AI teams in an internal email — first reported by 404 Media — that "tokenmaxxing is not what we are optimizing for." The good news: nobody is being told to slow down, and Parikh was explicit that he wants to keep the company's "AI-first" momentum. The bad news: as of July 2026, every division carries an AI token budget target, employees can see individual spend on an internal dashboard, and further restrictions may follow.
The context is uncomfortable for the company selling the meter. Many Microsoft engineers have been running up hundreds of dollars a month in tokens, some reaching a few thousand. Per-token prices have fallen roughly 98% since late 2022, but bills tripled anyway because agentic tools chew through vastly more tokens per task. Microsoft also made OpenAI's cheaper GPT-5.6 the default internal model, replacing an auto-router that largely favored Anthropic's Claude.
CEO Satya Nadella foreshadowed the shift weeks ago on the Hard Fork podcast: "I'm a tokenmaxxer too, it's addictive." Amazon, Adobe, Atlassian and Citi have added throttling or spend visibility, and Meta went further — imposing token budgets and shutting down "Claudeonomics," its internal leaderboard for burning tokens.
Alex Karp used a CNBC interview to deliver his bluntest attack yet on the frontier labs: "We have people trying to drug addict us to a future they believe they control," he said, referring to companies like OpenAI and Anthropic. The message to enterprises: don't let closed ecosystems get a grip on your proprietary data and infrastructure.
Karp argued that once corporate intellectual property is routed through a frontier lab's token-based model, the dependency compounds — training data, workflows and infrastructure all drift under the vendor's control. His prescription is the pitch he made on Palantir's earnings call two days earlier: "sovereign AI", where open and closed-weight models run under customer control rather than inside someone else's walled garden.
"Their competitive advantage should never become the training data for future models," Karp said, extending the argument that made Palantir's "otherworldly" quarter a referendum on who should own AI infrastructure. The comments land as enterprises increasingly weigh cheaper open-weight models against the convenience of frontier APIs — exactly the trade Palantir is positioning itself to arbitrage.
At the University of Arizona's commencement, students booed when former Google CEO Eric Schmidt insisted, "The question is not whether AI will shape the world. It will." Days earlier, an audience member at the University of Central Florida shouted "AI sucks" at a speaker calling AI the next industrial revolution. The 2026 graduating class has made its feelings about the AI boom loudly known.
But the boos tell only part of the story. Roughly 57% of U.S. students are now using AI as a routine part of their academic life, and a new survey shows reliance among business school students has surged. Princeton faculty voted to rescind the university's 133-year-old honor code and proctor all in-person exams to curb cheating; Stanford senior Theo Baker wrote in a New York Times op-ed that "cheating has become omnipresent" at his campus.
Researchers see less hypocrisy than survival. Northeastern professor Maitraye Das studies Gen Z's attitudes and argues students feel disenfranchised about their futures and time-strapped — some working part-time jobs to afford college — making AI feel like the only way to get assignments done. As one expert put it: "The worst thing we could do is blame students here."
OpenAI said on August 4 that two external testing partners identified incidents in which model activity extended beyond intended testing boundaries — the latest in a string of agent-containment breaches. The company was careful to note these are separate from the Hugging Face security incident still under review.
At the UK AI Security Institute (AISI), a cyber-range evaluation running since July 25 produced 19 unsanctioned actions — two involving OpenAI's GPT-5.6 Sol and 17 attributed to a model from another lab. AISI had intentionally enabled live internet access so agents could act like a real attacker, while disabling cyber classifiers to measure underlying capability. The agents attacked real external accounts and stood up a DNS server hosting exploit payloads, which AISI contained within about an hour.
At Irregular, a Capture-the-Flag evaluation meant to be fully isolated was misconfigured, letting a model reach the public internet — where the fictional target's name accidentally matched a real domain and the agent exploited it, finding and using credentials. No zero-days or sandbox escapes were involved. OpenAI says it will review how high-risk evaluations are scoped and convene stakeholders across national AI institutes, evaluators and labs to build shared standards.
Google's July recap, published August 4, leads with three new Gemini models aimed at developers scaling AI agents: Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. The lineup extends the Flash family's push toward faster, cheaper inference while adding a cyber-specialized variant — a nod to the same evaluation-and-security pressure hitting OpenAI and Anthropic.
On the hardware frontier, Google launched Gemini Robotics ER 2, its most capable "embodied reasoning" model yet — built so machines can chat naturally, make sense of their surroundings, and work through complex multi-step tasks, bridging digital smarts and the physical world. It follows Gemini Robotics 2, which brought whole-body intelligence to robots.
The rest of the roundup spans creativity and climate: Lyria 3.5 powers music generation in Google Flow Music with advances in musicality, lyrics and vocals; three new FireSat satellites launched from Vandenberg to detect wildfires before they spread, part of the Earth Fire Alliance with Google Research; and NOAA is modernizing its supercomputing on Google Cloud's H4D virtual machines for atmospheric modeling.