核心结论
通过 NVLink 进行 GPU-to-GPU peer memory 访问时,本地 GPU 的 L2 Cache 不缓存从远端 GPU 读取的数据。
这与 LTC Fabric(die 内跨分区访问)的行为形成鲜明对比:
| 访问类型 | 本地 L2 是否缓存 | 数据缓存位置 |
|---|---|---|
| LTC Fabric(die 内跨 L2 分区) | ✅ 是 --- 在本地 L2 复制 ^23^ | 本地 L2(Shared 状态) |
| NVLink(GPU 间 peer memory) | ❌ 否 --- 直接 bypass 本地 L2 ^181^^191^ | 仅远端 GPU 的 L2 |
证据来源
1. Stanford Hazy Research Blog(最明确)
^181^^182^ Stanford Hazy Research 关于多 GPU kernel 设计的系列文章明确写道:
"Remote L2 caching behavior: Because remote HBM access bypasses the local L2 cache, data is cached only on the remote peer's L2. This design simplifies inter-GPU memory consistency but makes every remote access bottlenecked by NVLink bandwidth."
^182^ 进一步解释了数据路径:
"The data path always flows through the source device's L2 cache, then its crossbar, across NVLink, through the destination device's crossbar, and finally to the SMs. This means peer data is never cached in the local L2, and peer memory access is always bottlenecked by NVLink bandwidth."
数据路径图:
远程 GPU (Source) 本地 GPU (Destination)
┌─────────────────┐ ┌─────────────────┐
│ HBM │ │ │
│ ↑ │ │ │
│ L2 Cache │ │ ❌ L2 Cache │
│ (命中后缓存) │ │ (不缓存peer │
│ ↓ │ │ 数据) │
│ Crossbar │ │ Crossbar │
│ ↓ │ │ ↓ │
│ NVLink 接口 │ ─── NVLink ───→ │ NVLink 接口 │
└─────────────────┘ │ ↓ │
│ SM (L1) │
└─────────────────┘
2. "Spy in the GPU-box" 论文(反向工程实证)
^191^ Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems 通过微基准测试直接测量了缓存行为:
"Our experiment shows that this data accessed on the remote GPU is cached on the remote GPU, rather on the local L2 GPU . Of course, caching the data locally, would introduce cache coherence issues since copies of the same data could exist in multiple L2 caches."
"In summary, our reverse engineering results demonstrate that an access to the memory of a remote GPU through NVLink is cached on the L2 cache of remote GPU, but not L2 cache of local GPU."
他们测量了四种访问延迟:
- 本地 L2 命中 (~250 cycles)
- 本地 HBM 访问
- 远程 L2 命中(数据在远程 GPU 的 L2 中!)
- 远程 HBM 访问
这说明:访问远程 GPU 的数据时,如果该数据在远程 GPU 的 L2 中,你仍然可以"命中"远程 L2(延迟比远程 HBM 低),但绝不会在本地 L2 中留下副本。
3. NVIDIA Developer Forums(用户实测)
^192^ NVIDIA 官方论坛上的讨论:
用户 cudaMancpy 问:"Is there a way to see if it's being cached? (I mean, I also want to check the hit ratio for the data that was accessed from peer.)"
经过实验后,用户自己得出结论:
"It seems no caching totally. (Remote GPU memory's total data amount is 256MB. Received User Bytes value is for about 268MB.)"
"Peer Memory is enabled. In this state, when the local GPU reads the data from the remote GPU memory, I thought caching occurred by default with the local GPU L2 cache. However, it does not appear to be the result of the experiment. It seems to be a one-time read only data without caching with direct access."
NVIDIA 员工 Robert_Crovella 也确认了测试方法:如果数据被缓存,反复访问性能应该提升;如果每次都走 NVLink,性能维持在 NVLink 带宽水平。
4. "Processing Large Data on GPUs with Fast Interconnects" 论文
^193^ 这篇 2020 年的论文明确写道:
"We observe that, in contrast to GPU memory, the small hash table is not cached in the GPU's L2 cache for NVLink 2.0 . The L2 cache is memory-side, and cannot cache remote data."
这是关键洞察:L2 Cache 是"memory-side"的------它只缓存"属于"本 GPU HBM 地址范围的数据。远程 GPU 的 HBM 数据不属于本地 GPU 的 memory space,所以本地 L2 不缓存。
5. Multi-GPU Programming Blog
^206^ 经典的多 GPU 编程教程也证实了这一点:
"The difference in L2 cache is quite interesting. The data is actually cached in L2. The only difference is that it's L2 of a different GPU."
为什么 NVIDIA 这样设计?
| 如果本地 L2 缓存远端数据 | 结果 |
|---|---|
| 需要跨 GPU 缓存一致性协议 | 复杂度高,延迟大 |
| 一个 cache line 可能在多个 GPU L2 中存在 | 需要 Directory + Invalidate 广播 |
| 写操作需要失效所有副本 | 严重影响性能 |
| GPU 间一致性粒度 | 需精确到 cache line |
| 当前设计(远端数据不缓存在本地 L2) | 结果 |
|---|---|
| 简化的弱一致性模型 | 易于实现,硬件成本低 |
| 每次远程访问都走 NVLink | 可预测,但性能受限 |
软件层面通过同步原语(如 nvshmem_quiet)保证一致性 |
程序员显式控制 |
| GPU 间一致性粒度 | 显式同步点,而非隐式 cache line |
^204^ 也提到:
"GPU caches not globally coherent across GPUs, only the CPU--GPU NVLink-C2C path is cache coherent."
这说明:GPU 之间的缓存一致性不是硬件自动维护的(至少不是 cache-line granularity 的),而是通过软件层(NVSHMEM/NCCL)的显式同步来保证。
一个细节:什么是 "memory-side" L2?
^193^ 提出的 "memory-side" 概念非常重要:
地址空间归属模型:
GPU-A 的 L2 Cache 只缓存 "GPU-A HBM 地址空间" 的数据
GPU-B 的 L2 Cache 只缓存 "GPU-B HBM 地址空间" 的数据
即使 GPU-A 通过 NVLink 读取 GPU-B 的数据,
该数据的地址仍属于 GPU-B 的地址空间,
所以 GPU-A 的 L2 不会缓存它。
但 GPU-B 的 L2 会缓存它(因为这是它的 memory space)!
这与论文 ^23^ 中发现的 LTC Fabric 行为截然不同:
- LTC Fabric(die 内):通过地址哈希,同一虚拟地址可能映射到不同物理 L2 分区。读取远程分区的数据时,在本地 L2 建立一个 Shared 副本(因为两个分区共享同一个虚拟地址空间,都是"本地"的)。
- NVLink(GPU 间):两个 GPU 有各自独立的 HBM 地址空间(即使通过 UVA/P2P 映射,底层仍区分归属)。本地 L2 严格按地址归属缓存,不缓存"属于其他 GPU"的数据。
结论
"NVLink 本地 GPU 不缓存远端数据"这个说法是正确的。
这是 NVIDIA 在 GPU-to-GPU NVLink 上的有意设计选择,目的是简化多 GPU 缓存一致性。代价是每次远程访问都必须经过 NVLink,无法通过本地 L2 的后续命中来加速。
这与 die 内部的 LTC Fabric 行为(跨分区访问在本地 L2 复制数据)形成鲜明对比,也解释了为什么跨 GPU 访问比 die 内跨分区访问慢得多。
参考资料:Stanford Hazy Research Blog ^181^^182^、Spy in the GPU-box 论文 ^191^、NVIDIA Developer Forums ^192^、Fast Interconnects 论文 ^193^、Multi-GPU Programming Blog ^206^、AI Infrastructure 文档 ^204^。
原文