Research result · Oct 9, 2026 · as partner
ByteDance Seed identifies 'Phase Sensitivity' in DeepSeek's chunked KV Cache compression, causing up to 40.2-point swings in long-context retrieval accuracy
ByteDance Seed reports that DeepSeek-V4-series models can change answers when irrelevant tokens are prepended, with performance oscillating with compression stride in a 4-token period. In 128K-context "needle-in-a-haystack" retrieval, accuracy differences across positions reached 40.2 percentage points for DeepSeek-V4-Flash-Base and 34.8 points for DeepSeek-V4-Pro-Base.
The team attributes this "phase sensitivity" to chunked KV cache compression and observes internal "phase-specialized" attention heads. Control Qwen3-0.6B models trained from scratch showed periodic swings that shifted with compression strides of 4/6/8, while full-attention baselines did not. ByteDance Seed recommends testing such models across different compression phases.
Original sources (Chinese)
字节找到了DeepSeek时强时弱的原因