Pyyan / Compare / Qwen3.8-Flash-Next vs Inkling

Qwen3.8-Flash-Next vs Inkling

2 of 5

Open-Weight Models · verified 6 Sept 2026

Qwen3.8-Flash-NextAlibabacurrent
TInklingThinking Machines Labcurrent
3 slots left
SpecificationQwen3.8-Flash-NextInkling
SummaryThe most downloaded thing on Hugging Face right now, and a stated Qwen4 preview.975B under Apache 2.0, shipped as a starting point rather than a finished product.
Context262K native, ~1M extended1M
Input ($/Mtok)$0.15Weights only
Output ($/Mtok)$0.47Weights only
LicenceOpen weightsApache 2.0
Max output32KNot published
Parameters125B total, 6B active975B total, 41B active
Released26 Aug 202615 Jul 2026
Model IDQwen/Qwen3.8-Flash-NextNot published
CategoryOpen-Weight ModelsOpen-Weight Models
OfficialThinking Machines Lab

Highlighted rows are where these differ.

Qwen3.8-Flash-Next

  • 125B total parameters with only 6B active per token
  • Alibaba describes it as a preview of the Qwen4 architecture
  • Top of the Hugging Face trending list, with a GGUF build close behind it

Best for open weights at speed.

Full spec sheet →

Inkling

  • Pretrained on 45 trillion tokens of text, images, audio and video, with an encoder free architecture
  • Apache 2.0, weights on Hugging Face, same day fine tuning through the lab's Tinker platform
  • The lab earns from the compute customers use to adapt it, not from per token access

Best for fine tuning on your own data.

Full spec sheet →