44 million units shortfall! General-purpose DRAM is being sold out until 2028
Relevant institutions estimate that the crowding effect of HBM capacity squeezing out general-purpose storage capacity will continue until 2028. HBM's share of overall DRAM chip will rise from 19% in 2026 to 28% in 2028. With general-purpose DRAM supply continuously compressed and NAND capacity resources constrained, the supply-demand structure of the two major storage categories will be unlikely to shift to a more relaxed one in 2028. Estimates show that in 2027, the demand for HBM for GPUs and ASICs will be about 196 million units. If all are stacked with 12 layers, supply will be only 152 million, leaving a gap of 44 million chips. The 8-layer solution can produce more HBM chips on the same wafer, and after supply eases, it is still possible to switch back to the 12-layer layer. The rumored 4-layer solution is more of a bargaining tactic, making mass production difficult and significantly weakening product positioning. Wafer delivery data shows that monthly HBM wafer shipments from 2026 to 2028 will be 414,000, 579,000, and 745,000 units respectively. Even if new wafer fab capacity is launched in 2028, excluding HBM capacity, GM DRAM wafer production growth will only be in the mid single digits. HBM prices continue to rise, with HBM3e and HBM4 nearly doubling in price from 2026 to 2027, and HBM4e offering about a 20% premium over HBM4. Due to higher DRAM wafer output value, Korean manufacturers will prioritize capacity, limiting the room for NAND expansion. Industry risk points lie in GPU shipments falling short of expectations; If AI agents surges and drives DDR5 demand, manufacturers' capacity reallocation will further intensify wafer competition between DDR5 and HBM.


The European PC market has "crashed," with a maximum drop of another 30%.
The latest forecast from market research firm Context shows that, driven by rising prices of core components like memory and driving up overall device prices, demand in the European PC market continues to shrink, with shipments expected to drop sharply in the second half of the year. By category, third-quarter shipments of laptops and desktops fell 6.4% and 20% year-on-year respectively, with the decline widening to 20% and 30% in the fourth quarter, indicating a significant cooling in substitution demand in mature markets. Currently, the industry is forming a clear negative cycle: rising storage chip prices drive up PC terminal prices, enterprise users are proactively extending device replacement cycles, and only purchasing equipment for essential needs. Analysts point out that while PC replacement demand has not disappeared, high costs have completely changed the market procurement rhythm. Although companies have temporarily delayed upgrades, short-term budget compression has been restricted, but frequent failures of old equipment and loss of official support have brought significant security risks. It is worth noting that the market shows a divergence pattern of "volume falls, profits rise," with price increases fully offsetting the sales decline. By focusing on high-end models, Lenovo has successfully offset the industry impact caused by rising memory prices. Meanwhile, upstream hardware manufacturers are adjusting prices, with AMD and NVIDIA successively raising graphics card prices. NVIDIA plans to raise AI server prices by more than 15% next year, passing on memory costs across the board. The industry as a whole has entered a phase of price increases and volume reduction, so manufacturers do not need to cut prices to boost sales.

US PC average price breaks 1000 for the first time! Annual shipments expected to drop 10.7%
Omdia data shows that in Q2 2026, US PC shipments will reach 18.8 million units, a slight year-on-year increase of 1.0%, ending the 7.0% decline in Q1, but this recovery is only a short-term pulse. The market average price surpassed $1,000 for the first time, up 12% year-on-year, driven by tight DRAM and NAND supply pressures driving up component costs. Channel stocking is expected to hedge against price hike risks. Market structure clearly tilts toward the high-end segment: entry-level models under $699 dropped 6.5% in shipments, while high-end PCs above $1,500 surged by 36.8%. AI PC penetration continued to rise, accounting for 48.3% in Q2 and expected to exceed 50% in Q3, further pushing up the average device price. By sector, the consumer market performed well in Q2, with shipments up 2.8%; The government sector was dragged down by budget pressure, with shipments plunging 24.4%, while the education market saw a modest 9.9% increase. The long-term outlook is not optimistic. Omdia expects US PC shipments to decline 10.7% in 2026, widen to 18.4% in the second half, and drop another 4.9% in 2027. Before 2028, both consumer and commercial markets are unlikely to return to growth. Besides limited storage supply, enterprises have concentrated upgrades before Windows 10 stoppages, overdrawing subsequent update demand. Despite high prices supporting revenue, price increases cannot offset shrinking sales; institutions estimate overall US PC market revenue will decline by 6% in 2027.

Apple has "bowed its head"! No longer cutting prices, only aiming to lock in volume
On September 8, according to South Korea's Economic Forum, Apple is negotiating a long-term NAND supply agreement with Kioxia, shifting its procurement approach from the previous "price reduction + multi-source" approach to a "lock-in volume + long-term contract" model.
According to a report by South Korea's "Link: Economic Forum," Apple has signed a long-term NAND flash supply agreement. Although Apple has not officially confirmed the agreement, industry insiders believe it breaks Apple's usual practice: Apple has traditionally relied on its scale advantage to negotiate flexible short-term contracts. This time, reports indicate that Apple signed a three- to five-year agreement with no price cap.

AI computing power competition extends from GPUs to high-speed flash memory
Recently, Kioxia announced its latest technology roadmap for ultra-high-performance SSDs. The originally planned target of 100 million random read IOPS in 2027 has been postponed to 2028. The product will feature a PCIe 7.0 interface and third-generation XL-Flash memory as a supplement to HBM memory, enabling high-speed direct GPU connectivity. Compared to the second generation, the third-generation XL-Flash offers three times the read performance, a 150% boost in write performance, and 40% and 80% higher energy efficiency respectively. Through optimized flash memory architecture, the die plane is increased and the routing is shortened, read latency is reduced to under 5 microseconds, and the erase/write cycle can reach 150,000 to 250,000 cycles. Currently, Kioxia has not disclosed details of the chip stacking or packaging. Currently, the transitional GP1 SSD uses PCIe 6.0 and second-generation XL-Flash, achieving up to 10.3 million IOPS, and is available in two versions: 800GB SLC and 1600GB MLC. The SLC version emphasizes high stability and long lifespan. This delay aligns with the pace of the PCIe 7.0 industry ecosystem. In AI inference scenarios, memory bottlenecks become prominent; ultra-high IOPS storage can offload HBM pressure and build multi-layered memory systems. The iteration of the GP series route also reflects that AI computing power competition is extending from GPU and HBM to the low-latency, high-speed flash memory track.

288GB becomes 192GB! 12-layer HBM becomes 8-layer
Reports indicate that NVIDIA plans to change the next-generation Rubin Ultra AI accelerator's HBM solution from 12-layer stacking to 8 layers. In an environment of rising storage costs, NVIDIA is shifting its focus to cost per unit bandwidth and overall device economy. SemiAnalysis shows that industry constraints have shifted from cost per unit capacity to cost per unit bandwidth. After adopting 8-layer HBM4, Rubin Ultra's single SIM memory is about 192GB, a one-third reduction compared to the original 12-layer solution of 288GB. HBM bandwidth is determined by the number of stacks, interfaces, and transfer rates, not by vertical stacking layers. With interfaces and stack counts unchanged, the 8-layer solution can retain or even slightly increase total bandwidth, increasing bandwidth per unit capacity by about 50%. This trade-off is suited to AI inference scenarios. When generating tokens, model weights and KV caches are repeatedly read, making memory bandwidth a performance bottleneck. Large memory capacity is beneficial for supporting large models and long contexts, but simply increasing the number of stacked layers cannot improve bandwidth. Currently, Rubin GPUs use 12-layer HBM4, with 288GB of memory per card and 22TB/s bandwidth. The changes to Rubin Ultra indicate that NVIDIA prioritizes bandwidth with better cost-effectiveness. Supply chain adjustments have accordingly, with Samsung and SK hynix planning to increase the supply share of 8-layer HBM4 in the second half of 2026. Nvidia has pushed suppliers to shift some HBM4 capacity from 12-layer to 8-layer. This marks a shift in AI chip design thinking: no longer blindly pursuing memory capacity, and balancing bandwidth and system cost has become a new priority.