Pyyan / Compare / Inkling vs DeepSeek V4.1 Flash

Inkling vs DeepSeek V4.1 Flash

2 of 5

Open-Weight Models · verified 13 Sept 2026

TInklingThinking Machines Labcurrent
DeepSeek V4.1 FlashDeepSeekcurrent
3 slots left
SpecificationInklingDeepSeek V4.1 Flash
Summary975B under Apache 2.0, shipped as a starting point rather than a finished product.Twice the parameters of the Flash it replaces, and cheaper to run.
Context1M1M
Input ($/Mtok)Weights only$0.3
Output ($/Mtok)Weights only$1.2
LicenceApache 2.0MIT
Max outputNot published256K
Parameters975B total, 41B active552B total, 8B / 16B active
Released15 Jul 202610 Sep 2026
Model IDNot publisheddeepseek-flash
CategoryOpen-Weight ModelsOpen-Weight Models
OfficialThinking Machines LabDeepSeek

Highlighted rows are where these differ.

Inkling

  • Pretrained on 45 trillion tokens of text, images, audio and video, with an encoder free architecture
  • Apache 2.0, weights on Hugging Face, same day fine tuning through the lab's Tinker platform
  • The lab earns from the compute customers use to adapt it, not from per token access

Best for fine tuning on your own data.

Full spec sheet →

DeepSeek V4.1 Flash

  • 552B total, 8B active while reading the prompt and 16B while generating, with a 1M token context
  • DeepSeek says it beats V4-Pro; V4-Pro requests route to it from 14 September 2026
  • MIT licence, native image input, and a KV cache about a quarter the size of V4-Flash's

Best for cheap agentic coding with images.

Full spec sheet →