Vals AI, an independent benchmarking firm, published on 3 September an estimate of the energy, carbon and water footprint of 16 open weight models running the long, multi step tasks in its Vals Index, using its own measured token counts and the open source EcoLogits estimator.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source The headline, as Bloomberg reported it, is that a complex task such as building a web app can carry 10,000 times the environmental impact of a simple query, and that building one app used energy equivalent to powering a home for two and a half hours.2 Bloomberg via The Star 2026-09-04 10,000 times figure; web app equals two and a half hours of home electricity; 14 Chinese models plus Thinking Machines and Mistral; Qwen3.8 Max highest impact, Ling 3.0 Flash 2607 lowest; Almatov quote. Open source The more useful finding is a comparison: Kimi K3 led the index at 57 percent accuracy against 53 percent for DeepSeek V4 Flash, and used roughly 12 times the energy per task to get there.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source Our assessment, with high confidence, is that agentic workloads have decoupled accuracy from efficiency in a way that per token pricing hides; with moderate confidence, that this report will be the reference point for the next round of AI energy policy despite covering no American frontier model; and with low confidence on the absolute numbers, which are modeled from assumptions rather than metered.
What was measured and how
The method is transparent about its limits. Vals took the token usage from its own benchmark runs and fed it to EcoLogits, assuming Nvidia H100 GPUs, 16 bit weights and vLLM serving, to produce carbon, water and electricity estimates.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source The 16 models were all open weight: 14 from Chinese companies, one from Thinking Machines Lab and one from Mistral.2 Bloomberg via The Star 2026-09-04 10,000 times figure; web app equals two and a half hours of home electricity; 14 Chinese models plus Thinking Machines and Mistral; Qwen3.8 Max highest impact, Ling 3.0 Flash 2607 lowest; Almatov quote. Open source OpenAI and Anthropic models were excluded because closed weights make the architecture, and therefore the compute, unknowable from outside.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source The estimates do not account for quantization, real data center efficiency, cooling systems or local externalities such as heat and noise.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source
Within those limits the results are stark. A full run of the 2,157 task index on Kimi K3 needs about 1.7 million watt hours, which Vals compares to the emissions of a flight from San Francisco to Miami, the water in a fire truck's tank, and a month of electricity for an average US home.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source A ten exchange conversation costs 169.5 watt hours on Kimi K3 against 10.5 on DeepSeek V4 Flash.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source Alibaba's Qwen3.8 Max had the highest impact of the 16 and Ant Group's Ling 3.0 Flash 2607 the lowest.2 Bloomberg via The Star 2026-09-04 10,000 times figure; web app equals two and a half hours of home electricity; 14 Chinese models plus Thinking Machines and Mistral; Qwen3.8 Max highest impact, Ling 3.0 Flash 2607 lowest; Almatov quote. Open source Long reasoning requests raised consumption by more than an order of magnitude over direct answers.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source
Why price does not tell you
The finding with commercial teeth is that token pricing is an imperfect proxy for footprint. DeepSeek V4 Pro costs 15 times less per token than Kimi K3 yet has a similar environmental impact, because price reflects the vendor's margin and subsidy strategy while energy reflects the model's architecture and how many tokens it burns thinking.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source A buyer choosing a model on cost per token is not choosing on energy, and a buyer choosing on accuracy is, per this data, choosing to spend an order of magnitude more electricity for a few points.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source
Omar Almatov, a founding engineer at Vals, put the mechanism plainly: footprints scale because these are no longer single shot questions.2 Bloomberg via The Star 2026-09-04 10,000 times figure; web app equals two and a half hours of home electricity; 14 Chinese models plus Thinking Machines and Mistral; Qwen3.8 Max highest impact, Ling 3.0 Flash 2607 lowest; Almatov quote. Open source An agent that plans, calls tools, reads results and retries multiplies the tokens of a chat answer by the number of steps, and reasoning models multiply again inside each step.
Who gains and who loses
Efficient model builders gain a metric they can win on: DeepSeek V4 Flash and Ling 3.0 Flash 2607 come out of this report as the models a sustainability conscious buyer should shortlist.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source 2 Bloomberg via The Star 2026-09-04 10,000 times figure; web app equals two and a half hours of home electricity; 14 Chinese models plus Thinking Machines and Mistral; Qwen3.8 Max highest impact, Ling 3.0 Flash 2607 lowest; Almatov quote. Open source Vals gains a position as the reference evaluator for a question regulators are starting to ask, and its chief executive Rayan Krishnan framed the report exactly that way, as evidence for policy conversations.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source Data center operators with clean power gain an argument, because the same tokens on a hydro grid produce a fraction of the carbon.
Kimi K3 and Qwen3.8 Max lose a reputational round, fairly or not, since they are the named high impact models in a study that could not test their American competitors.2 Bloomberg via The Star 2026-09-04 10,000 times figure; web app equals two and a half hours of home electricity; 14 Chinese models plus Thinking Machines and Mistral; Qwen3.8 Max highest impact, Ling 3.0 Flash 2607 lowest; Almatov quote. Open source OpenAI and Anthropic lose in a subtler way: their absence from the data means the public conversation about AI energy will be conducted with Chinese models as the examples, and the assumption that closed frontier models are at least as hungry will go unrebutted until they disclose.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source Enterprise buyers of agents lose the comfort of treating compute cost as the only cost.
The counter case
The absolute numbers should be held loosely. EcoLogits estimates from assumed hardware and serving; a provider running quantized weights on newer accelerators with efficient cooling could be several times better than the model predicts, and Vals says so.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source The 10,000 times figure compares the extremes of task length, which is arithmetic more than discovery: a task with 10,000 times the tokens will use about 10,000 times the energy. The sample is also skewed by exclusion, and the ranking might look different if GPT-6 Astra or Opus 5 could be measured.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source Finally, the framing assumes energy per task is the right unit; if an agent that costs 170 watt hours replaces an hour of a human at a computer, the comparison is not obviously unfavorable. The report measures the AI and not the alternative.
What to watch
- A closed lab publishes per task energy. If OpenAI, Anthropic or Google disclose energy per task or per token for a frontier model within a year, the comparison becomes complete; continued silence leaves the Vals numbers as the public benchmark.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source
- Efficiency as a benchmark axis. If a major leaderboard adds energy or emissions alongside accuracy by early 2027, the decoupling Vals found becomes a standing metric.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source
- Regulatory citation. A US state or EU document citing this report or its method in an AI energy disclosure proposal within a year would confirm the policy reference point.3 Bloomberg via The Spokesman-Review 2026-09-03 Long reasoning raises consumption by more than an order of magnitude; OpenAI and Anthropic models excluded as closed weight; infrastructure disclosure limits; Krishnan quote. Open source
- The accuracy gap closes or the energy gap does. If the next DeepSeek Flash generation reaches Kimi K3 accuracy at similar efficiency, the trade off dissolves; if Kimi's successor is more accurate and hungrier, it hardens.1 Vals AI 2026-09-03 16 open weight models; EcoLogits with H100, 16 bit, vLLM assumptions; Kimi K3 57 percent vs DeepSeek V4 Flash 53 percent at about 12 times the energy; 1.7 million Wh for 2,157 tasks; 169.5 vs 10.5 Wh per ten exchange conversation; token price a poor proxy; closed models excluded; efficiency and cooling not modeled. Open source
Four points of accuracy cost twelve times the energy in this data. Whether buyers care will depend on who is paying the electricity bill, and on who is counting.