textak
← BACK TO FEED
87% 2 ptsby Q1 2027
moderate

Open-source model matches closed frontier performance

The gap between open and closed models has been narrowing.

RESOLUTION CRITERIA

True if an open-weights model scores within 2% of the leading closed model on MMLU, HumanEval, and GPQA.

▲ FOR

DeepSeek-V3 matches o4-mini on DataGovBench (arXiv) — closest to independent benchmark validation in forecast history

DeepSeek, Qwen, Zhipu within single digits of closed frontier on coding at 10-30x lower cost per token — economic parity argument now dominant

GLM-5.2 MIT license with 1M context window confirmed by multiple sources including leaderboard data

196 days remaining provides runway for additional qualifying performance demonstrations

Open-weight models economically dominant even if slightly behind on raw benchmark — cost-adjusted parity argument strengthens

▼ AGAINST

Claude Fable 5.1 and GPT-5.6 still outperform on hardest reasoning tasks — frontier edge remains closed-model territory; resolution criterion may require top-tier task parity

GLM-5.2 math reasoning scores are early comparisons without official published benchmarks — independent third-party standardized validation still pending

DataGovBench is a specialized domain — frontier parity on hardest general reasoning tasks not demonstrated

Cumulative +59pt drift from 28% initial is extreme — calibration ceiling discipline limits move to 2pts

Copyright litigation targeting Chinese open-weight labs could disrupt GLM and DeepSeek ecosystems

Cloudflare blocking AI training crawlers by default (Story 11) may slow open-weight model training data access

RECENT SIGNALS (7)
OpenAI launches GPT-5.6 Sol with new Sol/Terra/Luna tiering, unveils GPT-6 Astra multimodal model
OpenAI
Temporal raises $550M Series E at $12.55B valuation as agentic AI demands reliable infrastructure
Temporal
DeepSeek-V3 Matches Closed-Source Frontier on Data Governance Workflows While Outperforming GPT-5
arXiv (DataGovBench)
Open-Source AI Models Close Gap on Frontier Performance; Zhipu GLM-5.2 Matches Closed Models on Math
ETF Trends
Mistral Raises €3 Billion; Positions Open-Weight Infrastructure as Sovereign AI Alternative
HIPTHER
Temporal Raises $550 Million to Power AI Agents with Failure Recovery; $12.55B Valuation
LLM Stats
Abacus.AI launches open-weight enterprise LLMs with 100x cost reduction for agents
AI Agents News Brief