Key Notes
- DeepSeek released V4.1 Flash, a 552B-parameter open-weight model (MIT license) using a new "asymmetric" Causal Encoder-Decoder architecture, with only 8B active parameters for input and 16B for output, a 1M-token context, and native image support.
- On DeepSeek's own benchmarks it beats its own larger V4 Pro and matches or edges Claude Opus 5 on some agentic/coding tests, but trails significantly on harder reasoning and on Terminal-Bench 3.0. Pricing runs $0.003-0.006 per million cached input tokens up to $0.60-1.20 per million output, varying by time of day. .
DeepSeek released V4.1 Flash on September 10, its first model built on a new architecture the company plans to scale up to larger systems later.
The model has 552 billion total parameters, but only a fraction activate at once. Roughly 8 billion parameters process input, and 16 billion generate output.
DeepSeek calls this an “asymmetric” Causal Encoder-Decoder design. The split is meant to keep inference cheap without sacrificing much capability.
The model supports a 1-million-token context window, native image and text input, and an explicit reasoning mode. Weights are published on Hugging Face under the permissive MIT license, meaning anyone can download, modify and run it.
DeepSeek’s own benchmark table shows V4.1 Flash beating its larger sibling, DeepSeek V4 Pro, on several tests. On some agentic and coding benchmarks, the company’s numbers also place it ahead of or close to Claude Opus 5 and GPT-5.6 Sol, two considerably more expensive models to run.
On CyberGym, a cybersecurity benchmark, V4.1 Flash scored 88.1, ahead of GPT-5.6 Sol’s 84.5. On DeepSWE v1.1, a software engineering test, it scored 74.2, edging out Opus 5’s 74.0.
But the picture is far more mixed on harder reasoning tasks. On Humanity’s Last Exam, an academic difficulty benchmark, V4.1 Flash scored 36.8 against Opus 5’s 56.3, a wide gap. On Terminal-Bench 3.0, Opus 5 led comfortably at 43.3 to V4.1 Flash’s 30.0.
All of these figures come from DeepSeek’s own published tables and have not been independently verified.
What the Numbers Actually Show
The pattern that emerges is a model that’s genuinely strong on specific agentic and coding benchmarks relative to its price, but not a broad match for frontier reasoning models across the board.
Against its closest Chinese rival, Moonshot’s Kimi K3, DeepSeek’s own table shows V4.1 Flash ahead on every agent and coding test where both have scores, though Kimi K3 still leads on GPQA Diamond and Humanity’s Last Exam.
Pricing reflects the efficiency pitch. In the API, cached input tokens run $0.003 to $0.006 per million, cache misses $0.15 to $0.30, and output $0.60 to $1.20 per million tokens, with rates varying by time of day.
DeepSeek has also made an unusual operational move. Starting September 14, every API request sent to V4 Pro will be automatically routed to V4.1 Flash instead, billed at the cheaper Flash rates, and will continue until a V4.1 Pro model arrives, for which DeepSeek gave no timeline.
The Decoder’s Jonathan Kemper, who reviewed DeepSeek’s technical report, flagged a less flattering detail: for long agent sessions, the model’s memory cache requires 890 bytes per token, a real resource cost for the kind of extended, file-heavy agentic work DeepSeek says the architecture was optimized for.
Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.