More than 15% in many cases
This is not yet an official Nvidia pricing announcement. Bloomberg reports, however, that some of the company's largest customers have been informed of increases exceeding 15% in many cases for servers containing Nvidia artificial-intelligence chips.
Reuters, which relayed the report, says the new pricing is expected to affect systems shipped from early 2027. The size of the increase would vary depending on the chip generation and, importantly, the memory configuration.
Existing Grace Blackwell platforms would be affected alongside the upcoming Vera Rubin generation. Server manufacturers supplying major data-center operators including Microsoft, Google and Oracle have reportedly already passed the new pricing information on to their customers.
Nvidia had not publicly commented on the reported increases when the story was published. The 15% figure should therefore be treated as information reported by Bloomberg rather than an officially announced Nvidia price increase.
An AI rack now carries terabytes of memory
The scale of Nvidia's current systems helps explain how memory can create such a large difference in price. A GB200 NVL72 combines 72 Blackwell GPUs and 36 Grace CPUs inside a single liquid-cooled rack.
Nvidia specifies 13.4 TB of HBM3E GPU memory for that configuration. The Grace processors add roughly 17 TB of LPDDR5X. A single rack therefore contains an amount of memory far beyond what would be found in a conventional server.
HBM is especially important. Its extremely high bandwidth is required to feed GPUs quickly enough while training or running large models. It is also difficult to manufacture, relies on advanced stacked-chip designs and consumes production capacity that cannot be expanded overnight.
Vera Rubin will not make the problem smaller
The upcoming Rubin architecture pushes the same principle further. Nvidia specifies 288 GB of HBM4 and as much as 22 TB/s of memory bandwidth for a single Rubin GPU. At Vera Rubin NVL72 rack scale, those design choices make memory even more central to the machine rather than a secondary component attached to the GPU.
That is exactly why higher memory prices can hurt so much. Nvidia can improve compute performance or accelerator efficiency, but memory capacity cannot simply be removed without affecting the workloads these machines are designed to handle.
Configuration therefore becomes a major pricing variable. Two systems based on the same accelerator generation can see different cost increases depending on how much memory they contain and which memory technologies they use.
AI is beginning to feel the pressure created by its own demand
There is a circular element to the current situation. The surge in AI investment created enormous demand for accelerators, but also for HBM, DRAM, networking hardware, storage and power infrastructure.
That demand is putting pressure on production capacity throughout the semiconductor supply chain. Samsung, for example, has recently increased prices for some advanced chipmaking services by as much as 15% amid strong demand, particularly for AI-related silicon.
The cost of an AI server can therefore no longer be reduced to the GPU at its center. Memory, advanced packaging, networking, cooling and power delivery account for an increasingly important share of a machine now designed at rack scale.
A 15% rack-level increase is not a small correction
On a consumer product, a 15% price increase is immediately noticeable. In a data center, the impact can be even larger because a project rarely involves a single machine.
Hyperscalers and cloud providers deploy infrastructure made up of dozens, hundreds or thousands of systems. A double-digit increase applied at that scale can add substantial amounts to the budget required for new compute capacity.
That may be the most significant signal behind the report. For several years, the AI race was largely about securing enough GPUs. The next constraint may look different: the GPUs can be obtained, but everything required around them is becoming expensive enough to change the economics of the entire data center.