This AI SSD tech makes 8 RTX 5090s perform like 46 GPUs in inference
GenStorAIGE AI90 shifts AI memory beyond traditional GPU HBM limitations using SSDs PT200Z SSD supports constant cache updates during demanding inference workloads efficiently Eight RTX 5090 GPUs gain dramatically larger effective inference memory capacity GenStorAIGE has introduced its AI90 inference acceleration platform at WAIC 2026, taking a storage-centric approach to expanding effective AI memory capacity. Rather than depending solely on GPU high-bandwidth memory, the platform incorporates PCIe Gen5 solid-state drives directly into the memory hierarchy itself. This allows portions of the Key-Value Cache used by large language models to sit outside GPU memory entirely. A three-tier memory architecture built around SSD offloading AI90 combines HBM, system DRAM, and SSD into a unified three-tier memory structure for handling inference workloads. By transparently offloading KV Cache data onto...