The gap between open and closed models has been narrowing.
True if an open-weights model scores within 2% of the leading closed model on MMLU, HumanEval, and GPQA.
DeepSeek-V3 matches o4-mini on DataGovBench (arXiv) — closest to independent benchmark validation in forecast history
DeepSeek, Qwen, Zhipu within single digits of closed frontier on coding at 10-30x lower cost per token — economic parity argument now dominant
GLM-5.2 MIT license with 1M context window confirmed by multiple sources including leaderboard data
196 days remaining provides runway for additional qualifying performance demonstrations
Open-weight models economically dominant even if slightly behind on raw benchmark — cost-adjusted parity argument strengthens
Claude Fable 5.1 and GPT-5.6 still outperform on hardest reasoning tasks — frontier edge remains closed-model territory; resolution criterion may require top-tier task parity
GLM-5.2 math reasoning scores are early comparisons without official published benchmarks — independent third-party standardized validation still pending
DataGovBench is a specialized domain — frontier parity on hardest general reasoning tasks not demonstrated
Cumulative +59pt drift from 28% initial is extreme — calibration ceiling discipline limits move to 2pts
Copyright litigation targeting Chinese open-weight labs could disrupt GLM and DeepSeek ecosystems
Cloudflare blocking AI training crawlers by default (Story 11) may slow open-weight model training data access