Gimlet Labs has raised a 300 million dollar Series B at a 3 billion dollar valuation, led by Andreessen Horowitz with Sapphire Ventures, prior investors Menlo Ventures and Factory, and two strategic newcomers whose presence is the story: Arm and Microsoft's venture arm M12.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source SiliconANGLE adds Samsung Ventures and more than a dozen others to the list.2 SiliconANGLE 2026-09-04 Samsung Ventures also in; prefill and decode separation, module level disaggregation, drafter plus frontier workflows; agent driven design search and a custom compiler; several hundred megawatts planned; edge inference servers. Open source Total funding is now 392 million dollars, which means the company had raised 92 million before this round, a figure this outlet derives from the disclosed total.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source Gimlet's product is an inference cloud that breaks a model into pieces and runs each piece on whichever chip suits it, across Nvidia, AMD, Intel, Arm, Cerebras and d-Matrix silicon.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source Our assessment, with moderate confidence, is that the valuation prices a thesis rather than a business: that inference, unlike training, will fragment across vendors, and that the software which orchestrates the fragments captures margin. With moderate confidence, we assess Arm and Microsoft as buying an option on Nvidia independence rather than a return.
What Gimlet does
Inference has two phases with different hardware appetites. Prefill, processing the prompt, is compute bound; decode, generating tokens, is memory bandwidth bound. Gimlet separates them across different chips, and goes further: it decomposes a model into modules and assigns each to the silicon it runs best on, supports drafter model plus frontier model workflows where a small model proposes and a large one verifies, and uses AI agents to search the design space with a custom compiler applying generic and chip specific optimizations.2 SiliconANGLE 2026-09-04 Samsung Ventures also in; prefill and decode separation, module level disaggregation, drafter plus frontier workflows; agent driven design search and a custom compiler; several hundred megawatts planned; edge inference servers. Open source The company claims up to 10 times throughput and interactivity gains and sells the result as serverless or managed capacity.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source 2 SiliconANGLE 2026-09-04 Samsung Ventures also in; prefill and decode separation, module level disaggregation, drafter plus frontier workflows; agent driven design search and a custom compiler; several hundred megawatts planned; edge inference servers. Open source
The customers named are unnamed: one of the top three frontier labs and one of the top three hyperscalers, with the customer base tripled by March and billions of dollars in contracted revenue.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source FourWeekMBA notes that the coverage it reviewed disclosed no revenue or customer counts.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source A 3 billion dollar valuation on undisclosed revenue and anonymous marquee customers is normal for this cycle, and it is a sourcing caveat all the same.
Why the backers matter more than the amount
Arm and Microsoft do not need a venture return from a 300 million dollar round. Arm's interest is architectural: an inference layer that treats chips as interchangeable is a layer that makes Arm based processors a first class target for AI workloads.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source Microsoft's interest is Azure: it has spent two years building Maia accelerators to reduce its dependence on Nvidia, and software that lets a workload move between Maia, AMD and Nvidia without rewriting is the missing piece.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source Cerebras and d-Matrix appear on Gimlet's supported list for the same reason; every alternative silicon vendor wants a neutral orchestration layer to exist.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source
Nvidia's lock is CUDA, and CUDA's grip is strongest in training, where the software stack is deep and the runs are long.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source Inference is high volume, cost sensitive and latency bound, and a buyer serving billions of tokens a day will move a workload for a 20 percent cost saving in a way a lab mid training run will not. Gimlet is a bet that the inference market behaves like a commodity market once someone builds the exchange.
Who gains and who loses
Gimlet's founders and early investors gain a roughly seven fold markup in six months, from about 400 million dollars to 3 billion, on the reported prior valuation.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source Arm, Microsoft, Samsung and the alternative accelerator makers gain a well funded neutral party whose commercial interest aligns with theirs.2 SiliconANGLE 2026-09-04 Samsung Ventures also in; prefill and decode separation, module level disaggregation, drafter plus frontier workflows; agent driven design search and a custom compiler; several hundred megawatts planned; edge inference servers. Open source Hyperscaler customers gain leverage over Nvidia at the negotiating table even if they never move a workload.
Nvidia loses nothing today and something structural tomorrow if the thesis holds: an inference layer that abstracts the chip turns Nvidia's software moat into a hardware price comparison.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source The neoclouds that built on Nvidia alone lose differentiation if capacity becomes fungible. And the labs that built their own serving stacks face a build versus buy question they thought they had answered.
The counter case
The thesis fails if Nvidia's inference software closes the gap first. Nvidia ships its own disaggregated serving frameworks, and a customer already on Nvidia hardware has little reason to add a middleman unless the alternatives are materially cheaper per token, which depends on chips Gimlet does not make. The thesis also fails if the model decomposition proves brittle: a modular split that works for one architecture may not for the next, and frontier labs change architectures every few months. The 10 times claim is Gimlet's own and unverified.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source Finally, the strategic backers are also the exit risk: a company whose value lies in neutrality is worth less the day Arm or Microsoft owns it, and a 3 billion dollar valuation leaves few acquirers who are not also chip vendors.
What to watch
- A named customer and a number. If Gimlet or a customer discloses production traffic on non Nvidia silicon with a cost per token figure by the first quarter of 2027, the thesis has evidence; anonymous marquee names through 2027 would suggest it does not.1 Gimlet Labs via GlobeNewswire 2026-09-04 300 million at 3 billion; a16z lead, Sapphire, M12, Arm, Menlo, Factory; 392 million total; disaggregated inference across NVIDIA, AMD, Intel, Arm, Cerebras, d-Matrix; a top three lab and a top three hyperscaler as customers; billions in contracted revenue; Zain Asgar. Open source
- Maia workloads on Gimlet. Any public Azure deployment routing inference to Maia through Gimlet within a year would confirm the Microsoft logic.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source
- The several hundred megawatts. Announced capacity of that size under contract by mid 2027 would show Gimlet is a cloud, not a compiler; a smaller footprint would suggest the software is the product.2 SiliconANGLE 2026-09-04 Samsung Ventures also in; prefill and decode separation, module level disaggregation, drafter plus frontier workflows; agent driven design search and a custom compiler; several hundred megawatts planned; edge inference servers. Open source
- Nvidia's response. A material Nvidia move on disaggregated serving pricing or an inference software bundle within two quarters would signal it sees the threat.3 FourWeekMBA 2026-09-04 Strategic logic of the backers: Arm gains from interchangeability, Microsoft from reducing Azure dependence on NVIDIA via Maia; CUDA lock strongest in training, inference cost and latency driven. Open source
Training belongs to CUDA. Gimlet's investors are betting that inference belongs to whoever writes the routing table, and they include two companies that would very much like that to be true.