New Attack Method Exposes Vulnerabilities in Nvidia GPUs
Researchers from the University of Toronto have introduced a memory bit flipping technique named GPUThor, which poses significant security risks to enterprise Nvidia GPUs. This method successfully bypasses the error-correcting codes (ECC) traditionally employed to safeguard these systems, potentially leading to root access.
GPUThor is categorized as a Rowhammer attack, which exploits the compact structure of modern RAM chips. The original Rowhammer vulnerability was identified back in 2015, demonstrating that closely packed memory cells could leak electrical charges, inadvertently flipping bits from 0 to 1, and vice versa. This attack involves a process termed row hammering, which uses repeated read operations on a memory row to provoke bit flips, resulting in security risks such as privilege escalation and manipulation of artificial intelligence models.
Over the years, various adaptations of Rowhammer attacks have emerged, targeting different memory types, including attempts on DDR3, DDR4, and even DDR5. Initially perceived as safe, newer memory technologies have not escaped scrutiny, prompting ongoing investigations into their physical security.
The University of Toronto research team asserted, “GPUThor is the first Rowhammer attack on Nvidia GPUs to breach ECC, which has been Nvidia's primary defense.” They referenced their prior studies, GPUHammer and GPUBreach, which established the feasibility of bit flips within the GDDR6 memory of Nvidia's graphics cards. Their earlier findings suggested enabling ECC as a preventative measure—yet with GPUThor, this countermeasure proves ineffective.
What Sets GPUThor Apart?
Prior approaches, namely GPUHammer and GPUBreach, executed uniform memory row hammering techniques. This uniformity allowed the Target Row Refresh (TRR) mechanism in DDR5 and later RAM generations to detect single-bit flips, with ECC subsequently correcting them. In contrast, GPUThor pioneers practical non-uniform row hammering on GPU DRAM, leading to multiple simultaneous bit flips, which ECC mechanisms aren’t designed to correct. GPUThor attacks its targets with intensity, being 6.6 times more aggressive than previous methods and achieving between 500 to 23,500 times more bit flips overall.
This advancement drastically reduces the time needed to find exploitable bit flips, as not all flips occur in memory locations tied to operations requiring OS permission. Specifically, GPUHammer needed around 21.9 hours to exploit on an Nvidia RTX A6000 without ECC, while GPUThor can accomplish the same in just 1.1 minutes.
“We confirmed bit flips across four Nvidia Ampere GPUs: the RTX A4000, A4500, A5000, and A6000,” the researchers noted. “These GPUs are prevalent in workstations and cloud instances, and the attack methodology has broad applicability.” Nevertheless, the team found that the attack didn’t trigger bit flips in Nvidia's A100, H100, or newer architectures like the RTX 5090 and RTX 6000, possibly due to their implementation of newer memory types such as HBM and GDDR6X. Future investigations into these architectures are planned.
Implications of GPUThor
With the growing reliance on enterprise GPUs for AI model training and inference, the findings of the GPUThor research carry significant weight. These graphics cards often operate in data centers, managing sensitive workloads across multiple virtual machines. During their experiments, the research team caused GPUs to fail so often that the internal fault detection flagged these instances as defects, necessitating replacement.
Beyond mere GPU crashes, the research revealed the potential for privilege escalation through memory corruption, allowing less privileged programs to elevate themselves to root access. “When GPUs are shared among users, an attacker on the same GPU can manipulate bits in another user’s data,” the researchers explained. “Even in isolated scenarios, any untrusted code could lead to root-level access, creating avenues for malware infiltration.”
Addressing the Threat
Combatting such design vulnerabilities will require enhanced hardware defenses in future GPU iterations. However, until more secure hardware becomes available, users should exercise caution with untrusted software running on their GPUs. Monitoring error correction counters from Nvidia can also provide valuable insight; a sudden uptick in these counters may signal an ongoing attack.
Nvidia was informed of GPUThor in April, and a security advisory was released in response, detailing the attack and offering remediation strategies such as enabling host IOMMU/DMA isolation and employing tools like nvidia-smi for monitoring ECC telemetry in affected GPUs.