技术栈
v4.1-flash
deepseek23
13 天前
agent
·
kv缓存
·
deepseek
·
稀疏注意力
·
moe架构
·
v4.1-flash
DeepSeek-V4.1-Flash 拆解:8B/16B 非对称激活,KV 缓存压到 890 字节/Token 才是 1M Agent 经济账
DeepSeek-V4.1-Flash 预填激活 8B,解码激活 16B。预填的 8B 比上一代 V4-Flash 的 13B 更低。
Together_CZ
20 天前
缓存
·
llm
·
compression
·
kv cache
·
deepseek
·
v4.1-flash
·
推动 kv 缓存压缩的极限
DeepSeek-V4.1-Flash:Pushing the Limits of KV Cache Compression——推动 KV 缓存压缩的极限
长时程智能体(long-horizon agents)的广泛应用,使模型工作负载越来越“输入密集”。虽然稀疏注意力等先前工作已经大幅降低了长上下文计算成本,但真正的瓶颈转移到了:
我是有底线的