Stock Markets August 29, 2026 07:16 PM

AI's Growing Memory Challenge: How Workloads Are Reshaping the Memory Stack

Analysts say AI is driving demand beyond high-bandwidth memory into DRAM, NAND and long-term storage, forcing new tiers and technical trade-offs

By Jordan Park
Share
Twitter Reddit Facebook LinkedIn
SNDK NVDA MU WDC STX

Bernstein analysts warn that artificial intelligence is creating an expanding memory bottleneck as different AI tasks place distinct demands across the memory hierarchy. Training, inference, retrieval-augmented generation and agentic AI all stress combinations of HBM, system DRAM, local SSDs and networked storage in different ways, prompting vendors and hyperscalers to explore new memory tiers and architectures.

AI's Growing Memory Challenge: How Workloads Are Reshaping the Memory Stack
SNDK NVDA MU WDC STX
Summarize with
ChatGPT Perplexity Claude Grok Gemini

Key Points

  • AI workloads impose different memory demands - training is bandwidth and compute-heavy while inference splits into compute-bound prefill and memory-bound decode stages.
  • The KV cache used during decoding can become a dominant memory consumer, potentially exceeding model weights in large deployments and limiting concurrent user capacity.
  • New memory tiers and architectures - including CXL, Nvidia's Storage Next, CMX and proposed high-bandwidth flash - are emerging to balance bandwidth, capacity and cost across HBM, DRAM and SSDs.

Artificial intelligence is amplifying pressure on the entire memory ecosystem, not just the high-bandwidth memory (HBM) that has been central to recent compute advances, according to Bernstein analysts. As AI workloads diversify, demand is spreading into conventional DRAM, NAND flash and even networked storage, reshaping where capacity and bandwidth are required.

The analysts argue that investment opportunities across the memory stack will increasingly hinge on the type of AI workload being run. Training, inference, retrieval-augmented generation - commonly abbreviated RAG - and the emerging class of agentic AI each impose distinct technical constraints and cost trade-offs that favor different parts of the memory hierarchy.

Training remains intensely compute-bound, with memory bandwidth limiting both the scale of models and the pace of computation. HBM retains a central role in large-scale model training because of its bandwidth advantage, but Bernstein emphasises that training pipelines also rely on other memory resources: system DRAM, local solid-state drives and networked storage are necessary for holding datasets, supporting caches and writing checkpoints during long training runs.

Inference divides into two phases with separate bottlenecks. The initial "prefill" stage - when a model ingests a prompt and produces the first token - is primarily compute-limited, which elevates the importance of GPUs and HBM. The later "decode" phase, however, is bounded by memory. During decoding, models retain previously generated tokens in a key-value cache - the KV cache - and the size of that cache grows with context length and with the number of simultaneous users served.

Bernstein notes that in substantial deployments the KV cache may demand more memory than the model weights themselves, creating a practical ceiling on how many users a given AI service can support before memory becomes a limiting resource.

These pressures are motivating a wave of new memory tiers and approaches alongside the established HBM, conventional DRAM and SSD options. The analysts point to technologies such as CXL memory, Nvidia's "Storage Next" initiative and CMX context storage as examples of efforts to balance performance, capacity and cost when larger memory pools are needed.

Retrieval-augmented generation could further expand storage requirements. Bernstein explains that constructing RAG databases requires considerable capacity on SSDs or HDDs, with system DRAM playing a larger role when those databases are searched frequently.

Agentic AI - systems that hold intermediate results, call external tools and pass state between autonomous agents - can accelerate KV-cache growth and place additional demands on CPUs and conventional server memory. This pattern further highlights how non-HBM memory and traditional server resources become important as AI systems gain new capabilities.

Memory vendors are also pursuing hybrid solutions such as high-bandwidth flash, which aims to marry HBM-like bandwidth with the capacity and cost profile of NAND. Bernstein cautions, however, that the technical challenges for such approaches remain substantial.

On the investment front, Bernstein's ratings assign Outperform to Samsung Electronics, SK hynix, Micron, SanDisk, Seagate and Western Digital, while Kioxia is rated Underperform.


Companies and tickers referenced in market data: SNDK, NVDA, MU, WDC, STX, 000660, 005930, 285A.

Risks

  • KV-cache growth could cap the number of users an AI service can support, affecting cloud and AI service providers that rely on conventional server memory and cache architectures.
  • Technical hurdles for hybrid solutions like high-bandwidth flash remain high, creating uncertainty for memory vendors pursuing such technologies.
  • Shifting and divergent memory requirements across AI workloads may complicate capital allocation and product roadmaps for semiconductor and storage companies.

More from Stock Markets

Glencore Sets Aside $480 Million for Radiant World Exposure Aug 29, 2026 Truist and Fifth Third Halt Sales of Products Linked to Delaware Life Amid Federal Probe Aug 28, 2026 Colombian equities retreat as COLCAP slips 1.28% at Friday close Aug 28, 2026 BAE Systems Awarded $167.5M Modification for MK 41 VLS Mechanical Design Support Aug 28, 2026 U.S. Defense Officials in Talks to Secure Long-Term Interest in Venezuelan Oil Fields Aug 28, 2026