textak
← BACK TO FEED
88% 1 ptsby Q1 2027
moderate

Open-source model matches closed frontier performance

The gap between open and closed models has been narrowing.

RESOLUTION CRITERIA

True if an open-weights model scores within 2% of the leading closed model on MMLU, HumanEval, and GPQA.

▲ FOR

GLM-5.2 91.2% GPQA Diamond — highest open-source score with weights available

Kimi K3 2.8T full weights publicly available, largest open-weight release in history

APEC Chengdu Statement: US and China jointly endorsing open-source governance

Chinese open-weight models 41% of Hugging Face downloads — structural market shift confirmed

Nvidia, Microsoft, SpaceX institutional coalition endorsing open-weight access for security use cases

Open Secure AI Alliance (Story 6) formation to develop open-source cybersecurity tools validates strategic open-weight importance

▼ AGAINST

GPT-5.6 Sol leads reasoning benchmarks at 58.1; Claude Opus 5 tops Artificial Analysis Intelligence Index at 61 — closed frontier demonstrably ahead on current top benchmarks

Kimi K3 MXFP4 full weights require 1.4TB memory — practical self-hosting limited to enterprise-scale infrastructure

GPT-6 August release window signals closed frontier actively extending lead — open-weight models catching up to prior generation, not current frontier

Benchmark gaming history means full-weight independent evaluation still required

Resolution criteria imprecision: 'matches closed frontier performance' lacks benchmark specificity — leadership fragmented across domains

Cumulative +61pt drift from 28% initial — ceiling effects at 89% are real; correction warranted on DRIFT ALERT

MAI-Cyber-1-Flash (Story 5) is a closed Microsoft model outperforming open alternatives on CyberGym — closed model advantage persists in specialized domains

RECENT SIGNALS (6)
OpenAI cuts GPT-5.6 budget tier pricing by 80%, Luna model now $0.20 per million input tokens
Build Fast with AI
LG AI Research releases K-EXAONE 2.0 with 262K-token context, 37B active parameters under Apache 2.0
AI Weekly
OpenAI Cuts GPT-5.6 Prices 80% on Budget Model, Undercuts Open-Source Pricing
buildfastwithai.com
Meituan Releases LongCat-2.0 Open-Source Model with 1.6T Parameters for Agentic Coding
aiagentstore.ai
OpenAI Cuts GPT-5.6 Luna Pricing by 80 Percent, Targeting High-Volume Workloads
Build Fast with AI
Open-source models now exceed 45% of OpenRouter traffic, Chinese models command majority of volume
State of Open Source AI