The gap between open and closed models has been narrowing.
True if an open-weights model scores within 2% of the leading closed model on MMLU, HumanEval, and GPQA.
GLM-5.2 91.2% GPQA Diamond — highest open-source score with weights available
Kimi K3 2.8T full weights publicly available, largest open-weight release in history
APEC Chengdu Statement: US and China jointly endorsing open-source governance
Chinese open-weight models 41% of Hugging Face downloads — structural market shift confirmed
Nvidia, Microsoft, SpaceX institutional coalition endorsing open-weight access for security use cases
Open Secure AI Alliance (Story 6) formation to develop open-source cybersecurity tools validates strategic open-weight importance
GPT-5.6 Sol leads reasoning benchmarks at 58.1; Claude Opus 5 tops Artificial Analysis Intelligence Index at 61 — closed frontier demonstrably ahead on current top benchmarks
Kimi K3 MXFP4 full weights require 1.4TB memory — practical self-hosting limited to enterprise-scale infrastructure
GPT-6 August release window signals closed frontier actively extending lead — open-weight models catching up to prior generation, not current frontier
Benchmark gaming history means full-weight independent evaluation still required
Resolution criteria imprecision: 'matches closed frontier performance' lacks benchmark specificity — leadership fragmented across domains
Cumulative +61pt drift from 28% initial — ceiling effects at 89% are real; correction warranted on DRIFT ALERT
MAI-Cyber-1-Flash (Story 5) is a closed Microsoft model outperforming open alternatives on CyberGym — closed model advantage persists in specialized domains