← Back to the directory

Company · Backed by Tencent, CATL

DeepSeek

深度求索


Coverage6

Research result · Oct 9, 2026 · as partner

ByteDance Seed identifies 'Phase Sensitivity' in DeepSeek's chunked KV Cache compression, causing up to 40.2-point swings in long-context retrieval accuracy

ByteDance Seed reports that DeepSeek-V4-series models can change answers when irrelevant tokens are prepended, with performance oscillating with compression stride in a 4-token period. In 128K-context "needle-in-a-haystack" retrieval, accuracy differences across positions reached 40.2 percentage points for DeepSeek-V4-Flash-Base and 34.8 points for DeepSeek-V4-Pro-Base.

The team attributes this "phase sensitivity" to chunked KV cache compression and observes internal "phase-specialized" attention heads. Control Qwen3-0.6B models trained from scratch showed periodic swings that shifted with compression strides of 4/6/8, while full-attention baselines did not. ByteDance Seed recommends testing such models across different compression phases.

Original sources (Chinese)

字节找到了DeepSeek时强时弱的原因qbitai

Research result · Oct 9, 2026 · as partner

Lenovo Tianxi's self-developed code agent TianxiCode tops SWE-bench-Live with 71% resolution rate

Lenovo Tianxi AI's self-developed code agent framework TianxiCode, paired with DeepSeek-v4.1-Flash, achieved a 71% issue resolution rate on the SWE-bench-Live Lite sub-benchmark, ranking first globally and passing the official Verified review mechanism.

The company describes three engineering capabilities: cross-file multi-hop retrieval and context management, self-driven planning and multi-round tool invocation, and closed-loop patch generation with test-driven self-healing. It plans to deploy the framework in Lenovo AI hardware products and developer toolchains.

Original sources (Chinese)

联想天禧自研代码智能体TianxiCode斩获SWE-bench-Live全球第一qbitai

Research result · Oct 3, 2026

DeepSeek publishes DSec elastic-computing sandbox infrastructure tech report for large-scale agent training

DeepSeek has published a technical report on DSec, its elastic-computing sandbox infrastructure for large-scale agent training. A single extended shard comprises about 160 servers, 30,000 CPU cores and 250TB of memory; with multiple shards in production, DSec serves around 3 million sandboxes daily, with peak concurrent use above 380,000 and creation rates over 5,000 sandboxes per second. The company says DSec supports all training, evaluation and data preprocessing for DeepSeek-V4, achieving over 50x resource oversubscription. It uses three-layer image splitting, on-demand loading from 3FS, virtio-pmem/DAX memory sharing and tiered CPU scheduling, and offers FnCall, Container, MicroVM and Full VM execution backends with pack_diff trajectory forking. DeepSeek is also hiring engineers for the team.

Original sources (Chinese)

DeepSeek扩招!弹性计算团队大量HC,尤其需要资深工程师qbitai

Open source · Sep 30, 2026

DeepSeek open-sources Ascend infrastructure components in collaboration with Huawei Ascend

The Ascend counterpart to DeepSeek's GPU infrastructure stack has been open-sourced, including the TileLang high-level compiler, DeepGEMM, FlashMLA, TileKernel and DeepSelect kernel libraries, and the DeepEP distributed communication library. Huawei contributed the jointly defined SuperPoD Flex and UBL128 fabric with 128-card 3.2 Tbps single-layer scale-up and 256,000-card two-layer scale-out. Reported DeepEP throughput is 375 GB/s for dispatch and 347 GB/s for combine; on Ascend 950, DeepSeek-V4.1-Flash offline inference runs at 2,469 tokens/s per card at TPOT=5ms. Licence terms were not stated.

Original sources (Chinese)

DeepSeek官方开源昇腾基础组件,与昇腾共建高效易用的AI芯片软件生态qbitaiDeepSeek 开源面向华为昇腾算力平台的基础设施组件,与英伟达平台一一对应ithome

Release · Sep 25, 2026

DeepSeek releases V4.1 Flash multimodal model; OpenCode makes its $60 usage quota permanent and fully integrates it

A new V4.1 Flash model from DeepSeek uses a 552B-parameter mixture-of-experts architecture that activates 8B parameters on input and 16B on output, with a causal encoder-decoder design and native multimodal visual understanding. DeepSeek says its KV cache lowers HBM demand to a quarter of the previous generation and SSD demand to an eighth.

OpenCode began phase two of “Operation Cheepseek” on September 25, making the model's previously time-limited $60 usage quota permanent. OpenCode, an official DeepSeek partner, has fully integrated V4.1 Flash; as of September 25, the model accounted for about 13% of the roughly 2 million usage events OpenCode observed, ranking first.

Original sources (Chinese)

接连发生 AI 失控,OpenAI 暂停最强模型训练;腾讯推出云端小龙虾:已接入微信 QQ;王兴兴回应造 390 万元变形机甲geekpark

Funding · Sep 24, 2026

DeepSeek reportedly finalizing new RMB 50B round at RMB 500B target valuation, annualized revenue reaches $1B

A reported RMB 50B ($7.1B) round would value DeepSeek at RMB 500B ($71B), with closing targeted by end of October. Founder Liang Wenfeng told investors annualized revenue has reached $1B, more than double from under $500M a few months ago. The new investors were not named; in June the company raised about RMB 50B at a post-money valuation near RMB 400B ($57B), with Liang contributing RMB 20B, Tencent RMB 10B, and CATL RMB 5B. Shanghai STAR Market IPO preparations are also progressing.

Original sources (Chinese)

DeepSeek被曝冲刺5000亿元估值!年化营收达67亿zhidx