- 现象:exits with return code -7
- 原因 :Setting the shm-size to a large number instead of default 64MB when creating docker container solves the problem in my case. It appears that multi-gpu training relies on the shared memory. ref
- 排查是否是shm-size过小: 在服务器上执行
ds_report,查看最后一行的是不是shared memory (/dev/shm) size .... 64.00 MB
- 排查是否是shm-size过小: 在服务器上执行
- 解决方案:增加docker的shm
deepseed 单机多卡程序报错:exits with return code -7
遇到好事了2024-01-16 13:30
相关推荐
云和数据.ChenGuang6 小时前
fastapi的参数剖析goodlook01237 小时前
LLaMa factory 大模型高效微调(二)Chasing__Dreams8 小时前
大模型应用开发--6--Transformer架构介绍147API8 小时前
蒸馏模型版本升级怎么做,权重、评测器和服务配置一起管XLYcmy9 小时前
京东 算法实习一面 上chen_zn9514 小时前
《VLA 系列》π0.5 + KI | 知识隔离 | 离散与连续动作联合训练 | 论文与源码解析EasyGBS15 小时前
告别“通用AI”的三大痛点:国标GB28181公网平台EasyGBS+DLTM让企业拥有专属AI安防大脑就是一顿骚操作16 小时前
VGG:用小卷积块把 CNN 做深的经典解读高洁0117 小时前
工信部教考中心证书DogDaoDao18 小时前
NNVC-17.1 深度解析:神经网络视频编码的最新进展与性能全景