Artificial intelligence is amplifying pressure on the entire memory ecosystem, not just the high-bandwidth memory (HBM) that has been central to recent compute advances, according to Bernstein analysts. As AI workloads diversify, demand is spreading into conventional DRAM, NAND flash and even networked storage, reshaping where capacity and bandwidth are required.
The analysts argue that investment opportunities across the memory stack will increasingly hinge on the type of AI workload being run. Training, inference, retrieval-augmented generation - commonly abbreviated RAG - and the emerging class of agentic AI each impose distinct technical constraints and cost trade-offs that favor different parts of the memory hierarchy.
Training remains intensely compute-bound, with memory bandwidth limiting both the scale of models and the pace of computation. HBM retains a central role in large-scale model training because of its bandwidth advantage, but Bernstein emphasises that training pipelines also rely on other memory resources: system DRAM, local solid-state drives and networked storage are necessary for holding datasets, supporting caches and writing checkpoints during long training runs.
Inference divides into two phases with separate bottlenecks. The initial "prefill" stage - when a model ingests a prompt and produces the first token - is primarily compute-limited, which elevates the importance of GPUs and HBM. The later "decode" phase, however, is bounded by memory. During decoding, models retain previously generated tokens in a key-value cache - the KV cache - and the size of that cache grows with context length and with the number of simultaneous users served.
Bernstein notes that in substantial deployments the KV cache may demand more memory than the model weights themselves, creating a practical ceiling on how many users a given AI service can support before memory becomes a limiting resource.
These pressures are motivating a wave of new memory tiers and approaches alongside the established HBM, conventional DRAM and SSD options. The analysts point to technologies such as CXL memory, Nvidia's "Storage Next" initiative and CMX context storage as examples of efforts to balance performance, capacity and cost when larger memory pools are needed.
Retrieval-augmented generation could further expand storage requirements. Bernstein explains that constructing RAG databases requires considerable capacity on SSDs or HDDs, with system DRAM playing a larger role when those databases are searched frequently.
Agentic AI - systems that hold intermediate results, call external tools and pass state between autonomous agents - can accelerate KV-cache growth and place additional demands on CPUs and conventional server memory. This pattern further highlights how non-HBM memory and traditional server resources become important as AI systems gain new capabilities.
Memory vendors are also pursuing hybrid solutions such as high-bandwidth flash, which aims to marry HBM-like bandwidth with the capacity and cost profile of NAND. Bernstein cautions, however, that the technical challenges for such approaches remain substantial.
On the investment front, Bernstein's ratings assign Outperform to Samsung Electronics, SK hynix, Micron, SanDisk, Seagate and Western Digital, while Kioxia is rated Underperform.
Companies and tickers referenced in market data: SNDK, NVDA, MU, WDC, STX, 000660, 005930, 285A.