- 2
- 0
- 约18.97万字
- 约 113页
- 2026-04-29 发布于湖南
- 举报
DeepSeek-V4:
TowardsHighlyEfficientMillion-TokenContextIntelligence
Abstract
WepresentapreviewversionofDeepSeek-V4series,includingtwostrongMixture-of-Experts(MoE)languagemodels—DeepSeek-V4-Prowith1.6Tparameters(49Bactivated)andDeepSeek-V4-Flashwith284Bparameters(13Bactivated)—bothsupportingacontextlengthofonemilliontokens.DeepSeek-V4seriesincorporateseveralkeyupgradesinarchitectureandop-timization:(1)ahybridattentionarchitecturethatcombinesCompressedSparseAttention(CSA)andHeavilyCompressedAttention(HCA)toimprovelong-co
原创力文档

文档评论(0)