The chart does not lie, but it does not tell the truth either. SanDisk’s HBF—High Bandwidth Flash—isn’t a memory chip. It’s a confession. A confession that the AI memory race, so far dominated by HBM’s blistering bandwidth, has ignored the silent majority of workloads: inference. The market is drunk on speed. But speed is a tax, and taxes are for the rich. HBF is the poor man’s HBM, and in a market where margins are being squeezed by hyperscalers, poverty might be the only sustainable strategy.
### Context: The Architecture of Desperation SanDisk, freshly split from Western Digital, needed a narrative. The old story—NAND flash for SSDs—was tired. The new story is AI memory. But SanDisk doesn’t own HBM. It doesn’t produce DRAM. It doesn’t have advanced packaging lines like CoWoS. So it did what every underdog does: it redefined the battlefield. HBF stacks NAND dies vertically, using TSV interconnects similar to HBM, but with one critical difference—the storage medium is flash, not DRAM. That means latency in microseconds, not nanoseconds. It means bandwidth in the tens of GB/s, not hundreds. But it also means cost per GB that is a fraction of HBM’s.
This is not a technology born from innovation alone. It is a technology born from necessity. During my 2022 winter solitude in the Mekong Delta, I studied Zero-Knowledge Proofs, but I also watched the collapse of Terra and the liquidity crisis that followed. I learned that when the market denies you the top tier, you carve a new tier. SanDisk is doing exactly that. HBF targets the AI inference market—where model parameters must be loaded into memory, inference run, and result returned. Here, latency is secondary to capacity and cost. A 70B parameter model requires ~140GB of memory. At HBM prices, that’s $2,000+ per GPU. At NAND prices, it’s $200. The math is seductive.
### Core: The Order Flow of NAND vs. DRAM Let me be precise. The core insight is not about bandwidth benchmarks. It’s about the order flow of silicon dollars. HBM3e delivers 1.6 TB/s bandwidth per stack. HBF, based on the architecture, might deliver 50-100 GB/s. That’s a 10x gap. But inference workloads—especially batch inference with large context windows—are not bandwidth-bound. They are capacity-bound. The bottleneck is the number of parameters that can be resident in memory. A single H100 GPU has 80GB HBM3e. That limits you to a 40B parameter model at full precision. With HBF, you could have 1TB of flash memory on a single accelerator card, running a 500B parameter model—albeit at slower speeds. The trade-off is real, but for many production inference tasks (e.g., retrieval-augmented generation, batch summarization), the throughput per dollar is what matters.
I’ve audited enough smart contracts to know that code is never neutral. Similarly, architecture is never neutral. HBF, by choosing NAND, implicitly accepts that it will never compete in training. That’s fine. Training is a $50B market, but inference is projected to be $200B by 2028. The real question is whether HBF can survive the inevitable counterattack from HBM vendors. SK Hynix and Samsung have already started developing “HBM Lite” versions—lower bandwidth, lower cost. If they can bring HBM to a price point that competes with NAND, HBF dies. But the physics of DRAM limits how low cost can go. DRAM requires a capacitor per bit, a more complex process, and higher defect density. NAND, especially 3D NAND with 200+ layers, is a mature, high-yield process. The cost advantage is structural, not temporary.
### Contrarian: The Retail Blind Spot Retail traders see HBF as a “cheap HBM” and think it will disrupt the AI memory market. That’s naive. The contrarian angle is that HBF is not a threat to HBM—it’s a threat to the entire concept of memory hierarchy. If HBF succeeds, it will blur the line between storage and memory. That means enterprise SSDs, which have been the high-margin product for Samsung and Micron, will face a new competitor. The real losers are not the HBM kings, but the traditional SSD players. SanDisk is cannibalizing its own SSD market to create a new one. That’s a bold move, but it also signals desperation.
Another blind spot is the geopolitical angle. The US export controls on HBM to China are tightening. HBF, being NAND-based, does not fall under current restrictions. That opens a massive market: China’s AI inference infrastructure. The Chinese hyperscalers (Alibaba, Tencent, Baidu) are starved for HBM. They will welcome a cheaper alternative, even if it’s slower. SanDisk’s HBF could become the de facto memory for China’s AI cloud, bypassing the US chip ban. But this is a double-edged sword. If HBF is seen as a vehicle for smuggling AI capability to China, the US Commerce Department will quickly restrict it. The risk is high, but the reward is huge.
### Takeaway: Between the Block and the Breath I am not bullish on HBF. I am not bearish. I am watching. The signal to track is not the press release—it’s the customer engagement. If a hyperscaler like Microsoft or Google announces a proof-of-concept using HBF in their inference servers, that’s a gamma squeeze. If not, HBF will be a footnote in the history of AI memory. The ledger remembers what the market forgets. But the market has a short memory. I will remember this moment: when SanDisk, a company I once dismissed as a storage relic, dared to bet that speed is not the only virtue.
Liquidity is a mirror, not a floor. HBF reflects the market’s growing realization that AI’s scaling law is not just about compute—it’s about memory. And memory, like water, finds its own level. We traded souls for pixels, now we seek the ghost. The ghost of HBF is the ghost of inference: cheaper, slower, but vast. Will it be enough? The algorithm does not care about your conviction. It only cares about the cost per inference.
Disclaimer: This is not financial advice. I am a crypto trader, not a hardware analyst. But I know a liquidity trap when I see one. HBF is a liquidity trap for the AI memory market, but it might also be the escape route.