AI developer tools · What shipped
For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.
Open coding models arrive; Fable 5 holds Epoch lead
1 min read
Fable 5 benchmark lead
Claude Fable 5 extends its margin on Epoch's index.
Following previous issue, Fable 5 now scores 161 on the Epoch Capabilities Index, holding a one-point lead over GPT-5.5 Pro [Quelle: Epoch AI]. The index itself expanded this month to include seven fresh evaluations covering agentic work, cybersecurity, algorithm engineering, forecasting, and physics—domains where general benchmarks blind-spot. Fable 5 leads across all seven new evals.
The real signal is evaluation velocity, not point margin.
Open-weight coding models ship
Alibaba, Mistral, and Moonshot each released production coding models this month.
Alibaba's Qwen3-Coder and Qwen3-Coder-Next (80B MoE, 3B active per pass) target local agentic coding with lower cost per repo-level workflow [Quelle: Turing Post]. Mistral's Devstral 2 ships in dense (123B) and local-friendly (24B) flavors, both with 256K context and multi-file codebase manipulation. Moonshot's Kimi K2.7 Code is a 1-trillion parameter MoE with 256K context and autonomous multi-file execution.
Open weights are becoming competitive on speed and context window.
Proprietary coding models narrow
Grok 4.5 and Claude Sonnet 5 now cluster with Opus on production benchmarks.
Grok 4.5 ships with 80 TPS speeds and a 500k context window, benchmarking at Opus and GPT-5.5 tier [Quelle: Developers Digest]. Claude Sonnet 5 lands near Opus 4.8 on most tasks but at lower cost, though its new tokenizer runs approximately 30 percent more tokens per request. Pricing pressure from open weights is forcing velocity over margin.
Cost-per-correct-task replaces raw score as the real differentiator.
Data on AI Capabilities and Benchmarking - Epoch AI4 hours ago ... Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated ...epoch.ai
Claude Fable 5 achieved a new high score of 161 on the Epoch Capabilities Index, surpassing GPT-5.5 Pro by 1 point and marking the first time Anthropic has led the index in over a year. Epoch AI recently expanded its benchmarking hub by tracking 13 new evaluations, with 7 incorporated into the Capabilities Index, and added nine external benchmarks spanning agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics.
Best AI Coding Tools in 2026: Assistants, Agents, IDEs & Open Models22 hours ago ... The category includes coding assistants, AI-native IDEs, terminal agents, repo-level agents, and open-source coding models that can run locally or inside ...turingpost.com

Qwen3-Coder and Qwen3-Coder-Next have been released as open-weight coding models for agentic software work. Qwen3-Coder-Next is an 80B MoE model that activates 3B parameters per forward pass, designed for local coding agents and lower-cost repo-level workflows, with a research paper available. Kimi K2.7 Code is a 1-trillion parameter open-weight MoE model by Moonshot AI featuring 256K context window and autonomous multi-file execution capabilities. Devstral 2, released by Mistral AI, is an open-weight coding model family with dense (123B) and local-friendly (24B) variants, both with 256K context windows, purpose-built for multi-file codebase manipulation and tool use in autonomous software engineering agents.
Developers Digest on AI Models11 hours ago ... Meta's first paid API model arrives with $1.25/M input tokens, 1M context window, and strong tool-use benchmarks. HN debates what it means for the open-weights ...developersdigest.tech
Grok 4.5 ships with 80 TPS speeds, a 500k context window, and benchmark results positioning it at Opus and GPT 5.5 tier, priced at $2/$6 per million tokens. Claude Sonnet 5 lands near Opus 4.8 on some tasks at a lower price point but runs approximately 30 percent more tokens due to a new tokenizer. GLM 5.2, an open-weight model, is positioned as a rival to GPT-5.5 with published benchmarks and pricing details.