MemVerge and Micron Boost NVIDIA GPU Utilization with CXL Memory

TheLostSwede · Mar 18, 2024

MemVerge, a leader in AI-first Big Memory Software, has joined forces with Micron to unveil a groundbreaking solution that leverages intelligent tiering of CXL memory, boosting the performance of large language models (LLMs) by offloading from GPU HBM to CXL memory. This innovative collaboration is being showcased in Micron booth #1030 at GTC, where attendees can witness firsthand the transformative impact of tiered memory on AI workloads.

Charles Fan, CEO and Co-founder of MemVerge, emphasized the critical importance of overcoming the bottleneck of HBM capacity. "Scaling LLM performance cost-effectively means keeping the GPUs fed with data," stated Fan. "Our demo at GTC demonstrates that pools of tiered memory not only drive performance higher but also maximize the utilization of precious GPU resources."

The demonstration, conducted by engineers from MemVerge and Micron featured a FlexGen high-throughput generation engine and OPT-66B large language model running on a Supermicro Petascale Server equipped with an AMD Genoa CPU, Nvidia A10 GPU, Micron DDR5-4800 DIMMs, CZ120 CXL memory modules, and MemVerge Memory Machine X intelligent tiering software.

The results of the demonstration were impressive. The FlexGen benchmark, utilizing tiered memory, completed tasks in less than half the time compared to conventional NVMe storage methods. Simultaneously, GPU utilization soared from 51.8% to 91.8%, thanks to the transparent management of data tiering across DIMMs and CXL modules facilitated by MemVerge Memory Machine X software.

This collaboration between MemVerge, Micron, and Supermicro marks a significant milestone in advancing the capabilities of AI workloads, enabling organizations to achieve unprecedented levels of performance, efficiency, and time-to-insight. By harnessing the power of CXL memory and intelligent tiering, businesses can unlock new opportunities for innovation and accelerate their journey towards AI-driven success.

"Through our collaboration with MemVerge, Micron is able to demonstrate the substantial benefits of CXL memory modules to improve effective GPU throughput for AI applications resulting in faster time to insights for customers. Micron's innovations across the memory portfolio provide compute with the necessary memory capacity and bandwidth to scale AI use cases from cloud to the edge," said Raj Narasimhan, senior vice president and general manager of Micron's Compute and Networking Business Unit.

View at TechPowerUp Main Site | Source

docnorth · Mar 18, 2024

@TheLostSwede I think CXL memory needs PCIe-5 protocol? Thanks.

Wirko · Mar 18, 2024

docnorth said:
@TheLostSwede I think CXL memory needs PCIe-5 protocol? Thanks.

That's correct. In this case, CXL devices connect to Epyc's PCIe bus. They don't connect directly to the GPU (if that's what confused you).

docnorth · Mar 19, 2024

Wirko said:
That's correct. In this case, CXL devices connect to Epyc's PCIe bus. They don't connect directly to the GPU (if that's what confused you).

No, I meant the CPU, that was the discussion when Alder Lake came out. Thanks.

System Name	Overlord Mk MLI
Processor	AMD Ryzen 7 7800X3D
Motherboard	Gigabyte X670E Aorus Master
Cooling	Noctua NH-D15 SE with offsets
Memory	32GB Team T-Create Expert DDR5 6000 MHz @ CL30-34-34-68
Video Card(s)	Gainward GeForce RTX 4080 Phantom GS
Storage	1TB Solidigm P44 Pro, 2 TB Corsair MP600 Pro, 2TB Kingston KC3000
Display(s)	Acer XV272K LVbmiipruzx 4K@160Hz
Case	Fractal Design Torrent Compact
Audio Device(s)	Corsair Virtuoso SE
Power Supply	be quiet! Pure Power 12 M 850 W
Mouse	Logitech G502 Lightspeed
Keyboard	Corsair K70 Max
Software	Windows 10 Pro
Benchmark Scores	https://valid.x86.fr/yfsd9w

System Name	Office / HP Prodesk 490 G3 MT (ex-office)
Processor	Intel 13700 (90° limit) / Intel i7-6700
Motherboard	Asus TUF Gaming H770 Pro / HP 805F H170
Cooling	Noctua NH-U14S / Stock
Memory	G. Skill Trident XMP 2x16gb DDR5 6400MHz cl32 / Samsung 2x8gb 2133MHz DDR4
Video Card(s)	Asus RTX 3060 Ti Dual OC GDDR6X / Zotac GTX 1650 GDDR6 OC
Storage	Samsung 2tb 980 PRO MZ / Samsung SSD 1TB 860 EVO + WD blue HDD 1TB (WD10EZEX)
Display(s)	Eizo FlexScan EV2455 - 1920x1200 / Panasonic TX-32LS490E 32'' LED 1920x1080
Case	Nanoxia Deep Silence 8 Pro / HP microtower
Audio Device(s)	On board
Power Supply	Seasonic Prime PX750 / OEM 300W bronze
Mouse	MS cheap wired / Logitech cheap wired m90
Keyboard	MS cheap wired / HP cheap wired
Software	W11 / W7 Pro ->10 Pro

Processor	i5-6600K
Motherboard	Asus Z170A
Cooling	some cheap Cooler Master Hyper 103 or similar
Memory	16GB DDR4-2400
Video Card(s)	IGP
Storage	Samsung 850 EVO 250GB
Display(s)	2x Oldell 24" 1920x1200
Case	Bitfenix Nova white windowless non-mesh
Audio Device(s)	E-mu 1212m PCI
Power Supply	Seasonic G-360
Mouse	Logitech Marble trackball, never had a mouse
Keyboard	Key Tronic KT2000, no Win key because 1994
Software	Oldwin

System Name	Office / HP Prodesk 490 G3 MT (ex-office)
Processor	Intel 13700 (90° limit) / Intel i7-6700
Motherboard	Asus TUF Gaming H770 Pro / HP 805F H170
Cooling	Noctua NH-U14S / Stock
Memory	G. Skill Trident XMP 2x16gb DDR5 6400MHz cl32 / Samsung 2x8gb 2133MHz DDR4
Video Card(s)	Asus RTX 3060 Ti Dual OC GDDR6X / Zotac GTX 1650 GDDR6 OC
Storage	Samsung 2tb 980 PRO MZ / Samsung SSD 1TB 860 EVO + WD blue HDD 1TB (WD10EZEX)
Display(s)	Eizo FlexScan EV2455 - 1920x1200 / Panasonic TX-32LS490E 32'' LED 1920x1080
Case	Nanoxia Deep Silence 8 Pro / HP microtower
Audio Device(s)	On board
Power Supply	Seasonic Prime PX750 / OEM 300W bronze
Mouse	MS cheap wired / Logitech cheap wired m90
Keyboard	MS cheap wired / HP cheap wired
Software	W11 / W7 Pro ->10 Pro

MemVerge and Micron Boost NVIDIA GPU Utilization with CXL Memory

TheLostSwede

News Editor

docnorth

Wirko

docnorth