Nvidia disclosed NVHBM on 26 August, a custom high bandwidth memory design that takes the memory controller off the accelerator die and puts it inside the HBM base die, claiming up to 30 percent more stack bandwidth, 15 percent lower HBM power and up to 25 percent more free area on the compute die against standard HBM4E.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source The first partner is not an Nvidia GPU. It is Amazon's Annapurna Labs, whose next generation Trainium4 will support NVLink Fusion and use NVHBM, so that Amazon designed accelerators and Nvidia GPUs sit inside one rack scale architecture.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source We assess with moderate confidence that the point of NVHBM is not the bandwidth number but the specification: Nvidia is claiming authorship of the interface between logic and memory, a layer that until now belonged to JEDEC and to the three memory makers, and it is establishing that claim through a customer that builds a competing chip.4 The Data Center Engineer 2026-08-28 NVLink Fusion expanded with NVHBM for semi custom AI infrastructure; traditional HBM keeps the memory controller on the XPU die where it consumes silicon area; Annapurna Labs the first collaborator with NVLink Fusion support beginning at Trainium4; NVIDIA to establish a standardized NVHBM implementation validated by multiple memory partners. Open source
What actually changed inside the package
The mechanical description is short. In a conventional stack the memory controller and its physical layer live on the accelerator die, where they consume silicon that could otherwise be matrix or vector units.4 The Data Center Engineer 2026-08-28 NVLink Fusion expanded with NVHBM for semi custom AI infrastructure; traditional HBM keeps the memory controller on the XPU die where it consumes silicon area; Annapurna Labs the first collaborator with NVLink Fusion support beginning at Trainium4; NVIDIA to establish a standardized NVHBM implementation validated by multiple memory partners. Open source NVHBM moves that block into the base die at the bottom of the 3D memory stack. Nvidia's own framing of the payoff is the three figures above, quoted against HBM4E rather than against a shipping product, and they are vendor figures unverified by any third party benchmark.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source StorageReview, working from the same disclosure, puts a sharper number on the area effect: a 67 percent reduction in physical layer and peripheral support area, which Nvidia says frees enough package real estate to grow the central compute die by up to 30 percent.3 StorageReview 2026-08-27 67 percent reduction in physical layer and peripheral support area, freeing package area to grow central compute dies by up to 30 percent; reduced compute starvation in memory bound inference; a hypothetical 1 gigawatt site running 2,000 watt accelerators could support up to 15,000 additional XPUs from the 15 percent memory power saving; Annapurna Labs selected NVHBM for Trainium4; SK hynix, Samsung and Micron positioned to qualify compatible stacks. Open source
Two of those three claims matter far more than the headline bandwidth. The area claim is a direct transfer of value: every square millimeter reclaimed from memory plumbing is a square millimeter of compute sold at compute margins on a wafer supply that is the binding constraint of the entire buildout. The power claim compounds at facility scale, where the constraint is not dollars but megawatts. StorageReview's arithmetic on a hypothetical 1 gigawatt site running 2,000 watt accelerators puts the 15 percent memory power saving at up to 15,000 additional XPUs inside the same electrical envelope.3 StorageReview 2026-08-27 67 percent reduction in physical layer and peripheral support area, freeing package area to grow central compute dies by up to 30 percent; reduced compute starvation in memory bound inference; a hypothetical 1 gigawatt site running 2,000 watt accelerators could support up to 15,000 additional XPUs from the 15 percent memory power saving; Annapurna Labs selected NVHBM for Trainium4; SK hynix, Samsung and Micron positioned to qualify compatible stacks. Open source That is an illustrative calculation rather than a measured result, and it should be read as a sense of scale, not a promise. But the direction is what counts: in a market where energized capacity is scarcer than capital, memory power is now a lever on how many chips a site can hold.
The customer choice is the strategy
Nvidia could have shipped NVHBM quietly on Rubin and said nothing about anyone else. Instead the announcement leads with Annapurna Labs, and the AWS side of it is explicit that NVLink Fusion is being extended with custom Nvidia high bandwidth memory as part of a broader deployment.2 AWS and NVIDIA press release 2026-08-26 AWS to deploy 2 million additional NVIDIA GPUs in 2027 and 2028, on top of more than 1 million announced at GTC 2026; Blackwell Ultra, Rubin and Rubin Ultra plus RTX PRO 4500 for EC2 G7; Vera CPU based systems coming to AWS; NVLink Fusion extended with custom NVIDIA high bandwidth memory; 100,000 GPUs for US Government secure workloads; quotes from Matt Garman and Jensen Huang. Open source Nafea Bshara, vice president at Annapurna Labs, is quoted describing NVHBM as a new architectural approach to high bandwidth memory performance.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source Read structurally, this is Nvidia offering the most successful in house accelerator program in the industry a better memory subsystem than it can buy on the open market, on condition that it adopt Nvidia's interconnect and Nvidia's memory specification.
The commercial envelope around that offer is large. The same day, AWS and Nvidia said AWS would deploy 2 million additional Nvidia GPUs across 2027 and 2028, on top of the more than 1 million announced at GTC 2026, spanning Blackwell Ultra, Rubin and Rubin Ultra, with Vera CPU based systems coming to AWS and 100,000 GPUs earmarked for US Government secure workloads.2 AWS and NVIDIA press release 2026-08-26 AWS to deploy 2 million additional NVIDIA GPUs in 2027 and 2028, on top of more than 1 million announced at GTC 2026; Blackwell Ultra, Rubin and Rubin Ultra plus RTX PRO 4500 for EC2 G7; Vera CPU based systems coming to AWS; NVLink Fusion extended with custom NVIDIA high bandwidth memory; 100,000 GPUs for US Government secure workloads; quotes from Matt Garman and Jensen Huang. Open source Also the same day, Nvidia reported second quarter fiscal 2027 revenue of $96.2 billion with Data Center at $89.0 billion, gross margin at 75.0 percent, and guidance of $108.0 billion for the next quarter that excludes Data Center compute revenue from China.5 NVIDIA Newsroom 2026-08-26 Second quarter fiscal 2027 revenue of $96.2 billion, up 106 percent year over year; Data Center revenue of $89.0 billion, up 117 percent; GAAP and non GAAP gross margins of 75.0 percent; GAAP diluted EPS of $2.46; $26.0 billion returned to shareholders; third quarter guidance of $108.0 billion plus or minus 2 percent at about 74 percent gross margin, excluding Data Center compute revenue from China. Open source The custom silicon threat has been the standing bear case against those margins for three years. NVHBM is the answer to it: if the alternative accelerator runs Nvidia's memory interface inside Nvidia's rack, the substitution is partial by construction.
Who gains and who loses
Nvidia gains a position it did not previously hold. It becomes the specifier of a memory interface that competitors, not just its own GPUs, are asked to adopt, and it does so while selling the interconnect and the rack around it.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source 4 The Data Center Engineer 2026-08-28 NVLink Fusion expanded with NVHBM for semi custom AI infrastructure; traditional HBM keeps the memory controller on the XPU die where it consumes silicon area; Annapurna Labs the first collaborator with NVLink Fusion support beginning at Trainium4; NVIDIA to establish a standardized NVHBM implementation validated by multiple memory partners. Open source AWS gains real engineering: better memory economics on Trainium4 and the ability to place its own silicon and Nvidia GPUs in one architecture, which is the practical form of the customer choice argument Matt Garman makes in the release.2 AWS and NVIDIA press release 2026-08-26 AWS to deploy 2 million additional NVIDIA GPUs in 2027 and 2028, on top of more than 1 million announced at GTC 2026; Blackwell Ultra, Rubin and Rubin Ultra plus RTX PRO 4500 for EC2 G7; Vera CPU based systems coming to AWS; NVLink Fusion extended with custom NVIDIA high bandwidth memory; 100,000 GPUs for US Government secure workloads; quotes from Matt Garman and Jensen Huang. Open source
The memory makers are the ambiguous party. Nvidia says a standard NVHBM implementation will be available from multiple memory providers, which positions SK hynix, Samsung and Micron to qualify compatible stacks.3 StorageReview 2026-08-27 67 percent reduction in physical layer and peripheral support area, freeing package area to grow central compute dies by up to 30 percent; reduced compute starvation in memory bound inference; a hypothetical 1 gigawatt site running 2,000 watt accelerators could support up to 15,000 additional XPUs from the 15 percent memory power saving; Annapurna Labs selected NVHBM for Trainium4; SK hynix, Samsung and Micron positioned to qualify compatible stacks. Open source Near term that is volume, and volume is what SK hynix is building for: in late August it broke ground on an advanced packaging plant in West Lafayette, Indiana, more than $4 billion, cleanroom by October 2028, mass production of next generation HBM in the second half of 2029.6 SK hynix Newsroom 2026-08-28 Groundbreaking for an advanced packaging plant for AI memory in West Lafayette, Indiana; investment of more than $4 billion; cleanroom by October 2028 and mass production of next generation HBM in the second half of 2029; about 7,000 direct and indirect jobs and more than 100 supplier companies; wafers made in South Korea, packaged and tested in Indiana; quote from chief executive Kwak Noh-Jung. Open source Longer term the risk is category, not volume. A base die authored by the customer moves the differentiating logic out of the memory vendor's hands and pushes the stack toward a commodity built to someone else's drawing. We assess with moderate confidence that the vendor which secures the base die design work rather than only the DRAM and the packaging keeps its margin through this transition, and the others do not.
The clearer loser is anyone building a custom accelerator without Nvidia's cooperation. If NVHBM class memory reaches the market first and predominantly through Nvidia's specification, then a merchant XPU program starts a generation behind on bandwidth, power and usable die area at once, and the gap is structural rather than a matter of design skill.3 StorageReview 2026-08-27 67 percent reduction in physical layer and peripheral support area, freeing package area to grow central compute dies by up to 30 percent; reduced compute starvation in memory bound inference; a hypothetical 1 gigawatt site running 2,000 watt accelerators could support up to 15,000 additional XPUs from the 15 percent memory power saving; Annapurna Labs selected NVHBM for Trainium4; SK hynix, Samsung and Micron positioned to qualify compatible stacks. Open source
The counter-case
The strongest argument against this reading is that custom base dies are already ordinary. The industry has been shipping customer tailored HBM base dies under the label cHBM, and NVHBM is on its face Nvidia's instance of an existing practice rather than a new layer of control. If Google, Meta and Broadcom simply commission their own base dies, Nvidia has bought an optimization, not a chokepoint. There is also no independent verification of any of the performance claims: the 30, 15 and 25 percent figures come from Nvidia, and the derived numbers in the trade press are calculated from them rather than measured.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source 3 StorageReview 2026-08-27 67 percent reduction in physical layer and peripheral support area, freeing package area to grow central compute dies by up to 30 percent; reduced compute starvation in memory bound inference; a hypothetical 1 gigawatt site running 2,000 watt accelerators could support up to 15,000 additional XPUs from the 15 percent memory power saving; Annapurna Labs selected NVHBM for Trainium4; SK hynix, Samsung and Micron positioned to qualify compatible stacks. Open source Nothing has shipped. Trainium4 is a future part, and the AWS deployment it belongs to runs through 2027 and 2028.2 AWS and NVIDIA press release 2026-08-26 AWS to deploy 2 million additional NVIDIA GPUs in 2027 and 2028, on top of more than 1 million announced at GTC 2026; Blackwell Ultra, Rubin and Rubin Ultra plus RTX PRO 4500 for EC2 G7; Vera CPU based systems coming to AWS; NVLink Fusion extended with custom NVIDIA high bandwidth memory; 100,000 GPUs for US Government secure workloads; quotes from Matt Garman and Jensen Huang. Open source For the thesis here to fail, two things have to be true: rival base die programs reach parity on roughly the same schedule, and the memory makers keep enough of the base die design to defend their margins. Both are plausible. Neither is yet visible.
What to watch
- A second NVHBM adopter that is not Nvidia or AWS. If another custom silicon program announces NVHBM before the end of 2027, the specification is becoming an industry interface rather than a bilateral deal. Silence through 2027 means it stayed a two party arrangement.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source
- Which memory vendors publicly qualify. Nvidia has promised multiple providers.3 StorageReview 2026-08-27 67 percent reduction in physical layer and peripheral support area, freeing package area to grow central compute dies by up to 30 percent; reduced compute starvation in memory bound inference; a hypothetical 1 gigawatt site running 2,000 watt accelerators could support up to 15,000 additional XPUs from the 15 percent memory power saving; Annapurna Labs selected NVHBM for Trainium4; SK hynix, Samsung and Micron positioned to qualify compatible stacks. Open source Named qualifications from SK hynix, Samsung and Micron within twelve months confirm the multi source claim; a single sourced launch would show the leverage sits with one supplier, not with the spec.
- Whether the SK hynix Indiana schedule holds. Cleanroom by October 2028 and mass production in the second half of 2029 is the published plan.6 SK hynix Newsroom 2026-08-28 Groundbreaking for an advanced packaging plant for AI memory in West Lafayette, Indiana; investment of more than $4 billion; cleanroom by October 2028 and mass production of next generation HBM in the second half of 2029; about 7,000 direct and indirect jobs and more than 100 supplier companies; wafers made in South Korea, packaged and tested in Indiana; quote from chief executive Kwak Noh-Jung. Open source Slippage past 2029 would put US packaged HBM behind the accelerator generation it was built to serve.
- Independent bandwidth and power measurements. Until a third party measures a shipping NVHBM stack, treat 30 percent and 15 percent as vendor claims.1 NVIDIA Blog 2026-08-26 NVHBM moves NVIDIA's memory controller into the HBM base die; up to 30 percent greater bandwidth, 15 percent lower HBM power and up to 25 percent more freed XPU compute die area versus standard HBM4E; a standard NVHBM implementation to be validated and offered by multiple memory providers; Amazon's Annapurna Labs the first collaborator with NVLink Fusion support from Trainium4; quote from Annapurna Labs vice president Nafea Bshara. Open source Published measurements during 2027 that land materially below those figures would remove the area and power argument that makes the deal attractive to a competitor.
- Whether gross margin holds while custom silicon scales. Nvidia guided to $108.0 billion at roughly 74 percent gross margin.5 NVIDIA Newsroom 2026-08-26 Second quarter fiscal 2027 revenue of $96.2 billion, up 106 percent year over year; Data Center revenue of $89.0 billion, up 117 percent; GAAP and non GAAP gross margins of 75.0 percent; GAAP diluted EPS of $2.46; $26.0 billion returned to shareholders; third quarter guidance of $108.0 billion plus or minus 2 percent at about 74 percent gross margin, excluding Data Center compute revenue from China. Open source If margin holds through 2027 while AWS deploys 2 million GPUs and ships Trainium4, the co-option worked.
The competitive question in AI hardware has been whether hyperscalers can build their own accelerators. NVHBM changes the question to what those accelerators are built out of, and Nvidia has just proposed itself as the answer.