AI for LearningAI Research Source checked

NIST says DeepSeek V4 Pro trails the frontier by about eight months

NIST’s CAISI says its evaluation of DeepSeek V4 Pro finds the model lags the frontier by about eight months, based on benchmarks spanning cyber, coding, science, reasoning, and math.

Original source ↗
In this briefing

At a glance

What changed
NIST’s CAISI says its evaluation of DeepSeek V4 Pro finds the model lags the frontier by about eight months, based on benchmarks spanning cyber, coding, science, reasoning, and math.
Why it matters
Independent evaluations can reduce hype and make cross-model comparisons more reliable. They also help policymakers and buyers understand what “open-weight” systems can and cannot do in sensitive areas like cyber and coding.
Who is affected
researchers, technical leaders, AI-watchers
What to do next
Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.
01

What changed

On May 1, 2026, NIST’s CAISI published results from its evaluation of the open-weight model DeepSeek V4 Pro, reporting that it lags the frontier by about eight months across a multi-domain benchmark suite.

02

Why it matters

Independent evaluations can reduce hype and make cross-model comparisons more reliable. They also help policymakers and buyers understand what “open-weight” systems can and cannot do in sensitive areas like cyber and coding.

03

In plain English

CAISI ran a set of tests across different skills and summarized where DeepSeek V4 Pro sits compared with other leading models and earlier generations.

Tap a word for its meaning
04

What this means for you

Who is affected: researchers, technical leaders, AI-watchers

Next move: Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.

  • CAISI calls DeepSeek V4 Pro the most capable PRC model it has evaluated so far.
  • Reported capability lag is based on benchmarks across five domains, including cyber and software engineering.
  • The report contrasts CAISI results with the developer’s self-reported evaluations.
What remains uncertain

Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.