← Back to the latest

Research resultOct 9, 2026Entry № 904

ByteDance Seed identifies 'Phase Sensitivity' in DeepSeek's chunked KV Cache compression, causing up to 40.2-point swings in long-context retrieval accuracy


ByteDance Seed reports that DeepSeek-V4-series models can change answers when irrelevant tokens are prepended, with performance oscillating with compression stride in a 4-token period. In 128K-context "needle-in-a-haystack" retrieval, accuracy differences across positions reached 40.2 percentage points for DeepSeek-V4-Flash-Base and 34.8 points for DeepSeek-V4-Pro-Base.

The team attributes this "phase sensitivity" to chunked KV cache compression and observes internal "phase-specialized" attention heads. Control Qwen3-0.6B models trained from scratch showed periodic swings that shifted with compression strides of 4/6/8, while full-attention baselines did not. ByteDance Seed recommends testing such models across different compression phases.

Original sources (Chinese)

字节找到了DeepSeek时强时弱的原因qbitai