Pyyan / Compare / DeepSeek V4.1 Flash vs Inkling

DeepSeek V4.1 Flash vs Inkling

2 of 5

Open-Weight Models · verified 13 Sept 2026

DeepSeek V4.1 FlashDeepSeekcurrent
TInklingThinking Machines Labcurrent
3 slots left
SpecificationDeepSeek V4.1 FlashInkling
SummaryTwice the parameters of the Flash it replaces, and cheaper to run.975B under Apache 2.0, shipped as a starting point rather than a finished product.
Context1M1M
Input ($/Mtok)$0.3Weights only
Output ($/Mtok)$1.2Weights only
LicenceMITApache 2.0
Max output256KNot published
Parameters552B total, 8B / 16B active975B total, 41B active
Released10 Sep 202615 Jul 2026
Model IDdeepseek-flashNot published
CategoryOpen-Weight ModelsOpen-Weight Models
OfficialDeepSeekThinking Machines Lab

Highlighted rows are where these differ.

DeepSeek V4.1 Flash

  • 552B total, 8B active while reading the prompt and 16B while generating, with a 1M token context
  • DeepSeek says it beats V4-Pro; V4-Pro requests route to it from 14 September 2026
  • MIT licence, native image input, and a KV cache about a quarter the size of V4-Flash's

Best for cheap agentic coding with images.

Full spec sheet →

Inkling

  • Pretrained on 45 trillion tokens of text, images, audio and video, with an encoder free architecture
  • Apache 2.0, weights on Hugging Face, same day fine tuning through the lab's Tinker platform
  • The lab earns from the compute customers use to adapt it, not from per token access

Best for fine tuning on your own data.

Full spec sheet →