Open sourceSep 28, 2026Entry № 780
Naive AI open-sources Naive-N0.5-Flash, a SWA–DSA hybrid-attention model built through AI for AI
Naive-N0.5-Flash, newly open-sourced by Naive AI, pairs five sliding-window attention (SWA) layers with one dynamic sparse attention (DSA) layer. In DSA, MLA is replaced by grouped-query attention with four KV groups and a lightweight indexer with 16 query heads; context length stays at 1M tokens. Training used 3.25T tokens in three stages: 50B indexer warmup, 3T sparse-attention training, and 200B learning-rate-decay SFT. The company, founded by Jifeng Dai, says training-system optimization and low-level bug fixes were also done by AI. License not specified.
Original sources (Chinese)
中训练、后训练持续升温,模型快速迭代,成为 AI for AI 最佳试炼场