At a glance
- What changed
- NIST’s CAISI says its evaluation of DeepSeek V4 Pro finds the model lags the frontier by about eight months, based on benchmarks spanning cyber, coding, science, reasoning, and math.
- Why it matters
- Independent evaluations can reduce hype and make cross-model comparisons more reliable. They also help policymakers and buyers understand what “open-weight” systems can and cannot do in sensitive areas like cyber and coding.
- Who is affected
- researchers, technical leaders, AI-watchers
- What to do next
- Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.
What changed
On May 1, 2026, NIST’s CAISI published results from its evaluation of the open-weight model DeepSeek V4 Pro, reporting that it lags the frontier by about eight months across a multi-domain benchmark suite.
Why it matters
Independent evaluations can reduce hype and make cross-model comparisons more reliable. They also help policymakers and buyers understand what “open-weight” systems can and cannot do in sensitive areas like cyber and coding.
In plain English
CAISI ran a set of tests across different skills and summarized where DeepSeek V4 Pro sits compared with other leading models and earlier generations.
What this means for you
Who is affected: researchers, technical leaders, AI-watchers
Next move: Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.
- CAISI calls DeepSeek V4 Pro the most capable PRC model it has evaluated so far.
- Reported capability lag is based on benchmarks across five domains, including cyber and software engineering.
- The report contrasts CAISI results with the developer’s self-reported evaluations.
What remains uncertain
Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.