技术栈

v4.1-flash

deepseek23
13 天前
agent·kv缓存·deepseek·稀疏注意力·moe架构·v4.1-flash
DeepSeek-V4.1-Flash 拆解:8B/16B 非对称激活,KV 缓存压到 890 字节/Token 才是 1M Agent 经济账DeepSeek-V4.1-Flash 预填激活 8B,解码激活 16B。预填的 8B 比上一代 V4-Flash 的 13B 更低。
Together_CZ
20 天前
缓存·llm·compression·kv cache·deepseek·v4.1-flash·推动 kv 缓存压缩的极限
DeepSeek-V4.1-Flash:Pushing the Limits of KV Cache Compression——推动 KV 缓存压缩的极限长时程智能体(long-horizon agents)的广泛应用,使模型工作负载越来越“输入密集”。虽然稀疏注意力等先前工作已经大幅降低了长上下文计算成本,但真正的瓶颈转移到了:
我是有底线的