Pyyan / Compare / Inkling vs Qwen3.8-Flash-Next

Inkling vs Qwen3.8-Flash-Next

2 of 5

Open-Weight Models · verified 6 Sept 2026

TInklingThinking Machines Labcurrent
Qwen3.8-Flash-NextAlibabacurrent
3 slots left
SpecificationInklingQwen3.8-Flash-Next
Summary975B under Apache 2.0, shipped as a starting point rather than a finished product.The most downloaded thing on Hugging Face right now, and a stated Qwen4 preview.
Context1M262K native, ~1M extended
Input ($/Mtok)Weights only$0.15
Output ($/Mtok)Weights only$0.47
LicenceApache 2.0Open weights
Max outputNot published32K
Parameters975B total, 41B active125B total, 6B active
Released15 Jul 202626 Aug 2026
Model IDNot publishedQwen/Qwen3.8-Flash-Next
CategoryOpen-Weight ModelsOpen-Weight Models
OfficialThinking Machines Lab

Highlighted rows are where these differ.

Inkling

  • Pretrained on 45 trillion tokens of text, images, audio and video, with an encoder free architecture
  • Apache 2.0, weights on Hugging Face, same day fine tuning through the lab's Tinker platform
  • The lab earns from the compute customers use to adapt it, not from per token access

Best for fine tuning on your own data.

Full spec sheet →

Qwen3.8-Flash-Next

  • 125B total parameters with only 6B active per token
  • Alibaba describes it as a preview of the Qwen4 architecture
  • Top of the Hugging Face trending list, with a GGUF build close behind it

Best for open weights at speed.

Full spec sheet →