DeepSeek released V4.1 Flash on Hugging Face
DeepSeek Releases V4.1 Flash, a 552B MoE With Native Vision
The asymmetric parameter activation and smaller KV cache reduce inference cost relative to similarly scaled models. V4.1 Flash is described as the smallest in DeepSeek's new architecture family, indicating larger models at the same architecture are planned.
The full picture
DeepSeek released V4.1 Flash on September 10, 2026, describing it as the smallest model in a new architecture family with native visual understanding. The model is a 552-billion-parameter Mixture-of-Experts system built on a Causal-Encoder-Decoder architecture, with an asymmetric design that activates only 8 billion parameters on input and 16 billion on output. DeepSeek said it is designed for greater capability, faster inference, higher throughput, and scaling to larger models. The model incorporates a new pre-training method and larger-scale reinforcement learning post-training, and is available on Hugging Face. Analyst Chris Porter said the model's KV cache is 437 times smaller than competing models and that this helps it beat GPT-5.6 Sol on several agentic benchmarks.
How it developed
Sources
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free