Pyyan / News / 10 September 2026

Open weightsDeepSeek

Two days after being named by US agencies, DeepSeek released a bigger open model

552Bparameters, MIT licence, 1M context

On 8 September three US agencies named DeepSeek for copying American models. On 10 September DeepSeek gave away a 552B parameter model under the MIT licence.

V4.1 Flash is a mixture of experts model with 552B parameters in total, 8B active while reading a prompt and 16B while generating, a 1M token context and native image input. The weights are on Hugging Face under MIT. DeepSeek's model card reports 90.6% on Terminal-Bench 2.1 and 74.2% on DeepSWE v1.1, and DeepSeek says it beats its own V4-Pro, which it is retiring: V4-Flash and V4-Flash-Vision-Exp requests already route to V4.1 Flash, and V4-Pro requests follow on 14 September.

Why this one is different

Bigger models usually cost more to run. This one is twice the size of the Flash it replaces and needs less memory to serve, because its KV cache, the working memory a model holds during a conversation, is about a quarter the size. In a week when Chinese AI chip prices rose up to half on a shortage of exactly that memory, a model engineered to need less of it is a model built for its supply chain.

Twice the parameters, a quarter of the memory.

How we got here

  1. Dec 2024The US tightens export controls on high bandwidth memory to China.
  2. 4 Sep 2026DeepSeek orders at least 160,000 Huawei accelerators for a gigawatt site.
  3. 8 Sep 2026The NSA, CISA and FBI name DeepSeek among six labs accused of industrial scale distillation.
  4. 10 Sep 2026V4.1 Flash ships under MIT, retiring V4-Pro.
  5. 10 Sep 2026Reuters reports Huawei's Ascend card price up 20% to 50% on a memory shortage.

What it does and does not mean

Benchmarks on a model card are the maker's own figures, and no independent evaluation had been published at release. The advisory naming DeepSeek alleges distillation but does not say which models, so this release neither confirms nor answers it. A 552B download is also open in licence rather than in practice for anyone without the hardware to run it. What it does show is that being named by US security agencies did not slow DeepSeek's release schedule by a day, and that the model it shipped is shaped by the memory shortage as much as by the benchmark table.

DeepSeekHugging Facefrom the source itself

Related

← All the news, newest first