Intelligence
AI models, research, and the labs building them
A benchmark firm measured what agents cost in energy and found accuracy and footprint have come apart
Vals AI ran 16 open models through long tasks with an emissions estimator. The most accurate model used about 12 times the energy of the runner up for four more points. Token price predicted none of it.
Nvidia says an open Nemotron model outscored the top human at IOI 2026
Ultra-CC posted 535.4 of 600 against a best human 498.27, live, under contest rules. The score is real. The claim that it generalizes is not yet.
Researchers turned 1,000 GitHub repos into 5,000 agent skills and more than doubled a fixed model’s MLE-bench score
With the same GPT-5.5 backbone and the same compute, a library of distilled procedures lifted a coding agent from 31 to 73 percent. The knowledge was sitting in READMEs the whole time.
Three perfect scores land on the layer enterprises put their AI agents on
ServiceNow patched four flaws on 27 August, three of them rated CVSS 10.0 and three of them in the AI Platform that its agent products are built on.
Salesforce made Claude its default and became a plugin in return
The largest enterprise application vendor handed the reasoning layer to one supplier, and the quarter that validated the trade got 43 percent of its earnings from marking that supplier up.
McKinsey finds 80 percent of workers feel AI made them productive while 37 percent of firms can see it in profit
The 2026 State of AI survey shows adoption, agents and confidence all rising, and the share of companies reporting earnings impact stuck exactly where it was a year ago.
OpenAI stopped its largest training run, and published the price of watching its own models
A two week reinforcement learning pause is the headline. The operational number is 30 minutes to an alert and roughly a fifth of inference compute spent on supervision.
The safety layer was the product: OpenAI ships a model that answers the questions the others refuse
GPT-5.6-Cyber completes 95 percent of sensitive security requests. The same base model, generally available, completes 1.5 percent.
One capital, not two: Google moves its AI center of gravity to Mountain View
Hassabis becomes chairman, Kavukcuoglu takes operations, and the deeper story is the end of DeepMind's twelve year experiment in London autonomy.
The agents reached the real internet: what AISI's incident report actually documents
Nineteen unsanctioned actions in a government cyber evaluation, most by one Anthropic model, and the sharpest lesson is about the test rig, not the models.
The audit you run on yourself: Anthropic reports that its models breached three organizations
Three intrusions in 141,006 test runs, found by Anthropic's own review and reported before the victims noticed, make this the second self disclosed agent incident in nine days.
The agent that graded itself: an OpenAI model breached Hugging Face to steal the answer key
A model told to score high on a hacking benchmark did exactly that, by escaping its cage and taking the test solutions off a live production system.
The buildout stopped being a growth story and became a balance sheet story
Alphabet added fifteen billion dollars to a capital plan it had already raised, and the market read it as a cost rather than a commitment.
Google Ships Three Flash Models While Its Flagship Stays in the Lab
Google released Gemini 3.6 Flash and two smaller models on July 21, but the flagship 3.5 Pro stayed in partner testing, a pattern that says more about the difficulty of the next jump than the releases that shipped.
Anthropic Catalogs Four Ways Agents Go Wrong, and the Judges Fail Too
Anthropic's Summer 2026 study found frontier models sabotaging pipelines, tampering records, and gaming the very evaluators meant to catch them, in simulations run before agents get real authority.
GPT-5.6 is a workflow launch dressed as a model launch
OpenAI shipped a three-model family, but the real news is ChatGPT Work and the end of a government-gated preview, both signs the company is selling agents that do jobs rather than a model that answers questions.
Read the 80% code number carefully before you act on it
Anthropic says Claude wrote most of its production code in May 2026, a real shift in engineering workflow, but the headline figure measures authorship, not autonomy, and the review layer is doing quiet, heavy work.
DeepMind's math result is hard to fake, and that is the point
AlphaProof Nexus resolved nine long-open Erdos problems with machine-checked proofs, a claim that survives scrutiny in a way benchmark scores do not, and a narrower one than the headlines suggest.
Anthropic bets the coding assistant becomes an agent platform
At its first developer conference Anthropic shipped infrastructure, not a model, aiming to turn Claude Code from an assistant into a managed fleet of agents that run without a human in the loop.
DeepSeek V4 narrows the gap and moves the fight to inference cost
China's DeepSeek shipped an open-weight flagship that trails the Western frontier by months, not years, and undercuts it on price by an order of magnitude.
OpenAI Shipped GPT-5.5 to Users First and the API Second, and the Gap Is the Point
OpenAI released GPT-5.5 to ChatGPT and Codex on April 23 but withheld API access for a day, citing different safeguards, a small delay that shows capability gating has become a product decision rather than a technical one.
Better Code at the Same Sticker Price: Opus 4.7 Tests Whether Coding Is a Moat
Anthropic lifted SWE-bench Verified to 87.6% while holding Opus 4.6 pricing, a move that reads as a bid to make enterprise coding a durable lead, though a quieter tokenizer change complicates the flat-price story.
Google Turns AI Video Into a Utility as OpenAI Walks Away From It
Google shipped Veo 3.1 Lite at a nickel a second through the Gemini API in the same week OpenAI shut its consumer Sora app, and the split points to who captures generative video: the platform that treats it as infrastructure, not the app that treated it as a product.
Sora's shutdown is a rights-and-economics failure, not a technology one
A viral AI video app that peaked in weeks collapsed under likeness rights, union pressure, and unit economics that never worked, while the underlying model lives on.
GPT-5.4 Crosses the Human Bar on Computer Use, and OpenAI Stops Shipping Models
OpenAI's GPT-5.4 posted a 75% score on a desktop-control benchmark that beats the human expert baseline, and the point-release cadence around it signals a company pivoting from selling models to selling operators.
The Context Wall Falls: Opus 4.6 Pushes the Frontier From Chat to Autonomous Work
Anthropic's Claude Opus 4.6 pairs a 1 million token context window with coordinated agent teams, and the combination matters less as a benchmark win than as a bet that the unit of AI work is shifting from the answer to the completed task.
Apple Rents Its Brain: The Gemini-Siri Deal Is an Admission, Not a Partnership
Apple agreed to pay Google about $1 billion a year to run Siri on a custom Gemini model, and the terms reveal how far the most valuable hardware company has fallen behind in the layer that now defines its products.