AI-Models

DeepSeek: New Flash Model Scores Ten Points in AI Test

3 min read
A bar chart with a sharp upward spike on a monitor next to the DeepSeek logo in a data center symbolizes the score jump of V4 Flash 0731. Image generated with GPT Image 2
A bar chart with a sharp upward spike on a monitor next to the DeepSeek logo in a data center symbolizes the score jump of V4 Flash 0731.

TL;DR Too Long; Didn’t read

DeepSeek has updated V4 Flash: In the Intelligence Index from Artificial Analysis, the score rises from 40 to 50 points, placing second among 162 measured models. The cause is primarily fewer hallucinations, not more accuracy. Price and architecture remain unchanged at $0.14 and $0.28 per million tokens, respectively. Open weights are expected to follow in the coming weeks.

Key takeaways

  • V4 Flash 0731 achieves 50 out of 100 points in the Intelligence Index, placing it second among 162 models.
  • According to Artificial Analysis, the progress is mainly based on clean answers, not more expertise.
  • In agent tasks in the GDPval-AA v2 test, the Elo rating increased from 1189 to 1559.
  • Price and architecture remain unchanged: 284 billion total and 13 billion active parameters, a context window of one million tokens.
  • The update currently only affects the Flash API; app, web interface, and pro variant remain unchanged.
  • Open model weights for local use are expected to follow in the coming weeks.

DeepSeek has revised its language model V4 Flash and increased its value in the Intelligence Index of the analysis firm Artificial Analysis from 40 to 50 points. The new V4 Flash 0731 is thus only one point behind OpenAI’s top model GPT-5.6 Luna and costs about 60 percent less per solved task, according to the provider. The architecture and price remained unchanged compared to the previous version.

Ten Points More in the Intelligence Index of Artificial Analysis

The Intelligence Index of Artificial Analysis summarizes results from tasks related to logic, knowledge, mathematics, and programming into a point value. On this scale, V4 Flash 0731 now achieves 50 out of a possible 100 points, placing it second among 162 measured models, with a median of 17 points. In practical agent tasks in the GDPval-AA v2 test, the Elo rating increased from 1189 to 1559, while the hit rate in Terminal-Bench 2.1 climbed by 17 points to 79 percent.

In direct comparison, the model is one point behind GPT-5.6 Luna (51) and practically on par with Z.AI’s GLM-5.2 and Google’s Gemini 3.6 Flash (both 50). Moonshots Kimi K3 (57) and Anthropic’s Claude Opus 5 (61) remain ahead. Once open weights are available, V4 Flash 0731 would be the second strongest freely available model in the index after Kimi K3. The model also improved in individual tests such as the physics task collection CritPt (17 percent), the programming benchmark SciCode (50 percent), and the knowledge test GPQA Diamond (91 percent) compared to the previous version.

Fewer Hallucinations Instead of Higher Accuracy Drive the Leap

According to Artificial Analysis, the progress is primarily due to cleaner answers, not more knowledge: The hallucination rate in the AA Omniscience test decreased by 12 points, while pure accuracy remained largely constant. The architecture and number of parameters did not change – 284 billion total and 13 billion active parameters continue to be used in the Mixture-of-Experts method, supplemented by a context window of one million tokens. This is sufficient for several thousand pages of text in a single request.

The price remains at $0.14 per million input tokens and $0.28 per million output tokens, while already cached requests cost only $0.0028 per million tokens – a discount of 98 percent compared to the full price. At the same time, the model consumed about 12 percent fewer tokens in the tests by Artificial Analysis to solve the same tasks than the previous version – an effect that DeepSeek attributes, among other things, to a still integrated module for speculative decoding that predicts multiple tokens per computation step.

DeepSeek Positions Itself in the Race for Affordable AI Models

DeepSeek confirmed the step in its official API change log and clarified that the update currently only affects the Flash API; the app, web interface, and the Pro variant remained unchanged, with the company announcing an official version of V4 Pro as “ASAP.” Open model weights for local use are expected to follow in the coming weeks, according to Artificial Analysis.

The leap is part of a series of Chinese models that have recently caught up with Western providers – for example, DeepSeek’s own V4 Pro and Z.AI’s GLM-5.2 in security-related capabilities. DeepSeek itself is also in a new billion-dollar funding round and is advancing the development of a proprietary inference chip to reduce dependence on Nvidia hardware.

It remains to be seen whether DeepSeek can replicate the progress in the officially announced version of V4 Pro later and whether the actual accuracy will also increase instead of just the number of hallucinations decreasing. Crucial for companies will be whether the low price is also confirmed in more complex, multi-stage agent tasks in everyday use.

Frequently asked questions

Is DeepSeek V4 Flash 0731 free to use?

No, usage via the official API costs $0.14 per million input and $0.28 per million output tokens; cached requests are significantly cheaper.

Will the model weights be publicly available?

Yes, DeepSeek plans to release open weights for local use according to Artificial Analysis in the coming weeks; the company has not yet provided a specific date.

How does the model compare to Claude Opus 5?

Claude Opus 5 continues to lead the Intelligence Index with 61 points; V4 Flash 0731 is eleven points behind but significantly cheaper in price per task.

Will anything change for app and web users of DeepSeek?

No, the update only affects the Flash API according to DeepSeek; the app and web interface will continue to run with the previous model version for now.

When will the official version of DeepSeek V4 Pro be released?

DeepSeek announced the official version of V4 Pro as 'coming soon' but has not yet provided a specific date.


← Back to the blog