技术栈

moe model

缘友一世
2 小时前
模型部署·模型推理·deepseek·moe model
在8×A800上把 DeepSeek-V4-Flash 榨到极限:从5倍提速到TP+EP双实例的完整实测单机 8 卡,双实例,一路从 50 tok/s 到 479 tok/s——我把每一步的决策、测量和踩坑都写在这里。
我是有底线的