Signing you in...

Please wait while we verify your authentication

Article · Wednesday, July 8, 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

By Marius BongartsTech22 editions
← See today's latest
Editions
19 / 22
Generated by AI overnight from public sources, refreshed daily.
AI developer tools · What shipped
Wednesday, July 8, 2026
AI developer tools · What shipped

Claude Fable 5 extends lead, JetBrains unifies team AI, small models gain traction

1 min read

Claude Fable 5 benchmark lead

Fable 5 now leads by a wider margin than ever.

Following previous issue, Epoch's latest Capabilities Index shows Claude Fable 5 at 161 points, still ahead of GPT-5.5 Pro at 160 [Quelle: Epoch AI]. The index absorbed seven new evaluations this week covering agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics—and Fable 5 leads across the refresh. Separately, Anthropic shipped Claude Science, a workbench for researchers with customizable tools, auditable artifacts, and flexible compute access [Quelle: Anthropic].

Watch whether benchmarks or real deployment data drift next.

JetBrains AI for Teams

JetBrains just bet on vendor-agnostic team AI.

The company is rolling out JetBrains AI for Teams and Organizations starting July 2026, adding team automations, cloud agents for long-running tasks, and JetBrains Context—repository intelligence for faster code understanding [Quelle: JetBrains]. Organizations get centralized governance through JetBrains Central, with cost attribution across tools. The shift from per-seat licenses to on-demand AI credits valid for twelve months lets teams scale adoption without per-user lock-in.

This open-protocol approach—MCP and ACP support—signals IDE makers are choosing connectors over exclusive deals.

Small models gain coding speed

A 9.7B model just scored 45% on industry benchmarks.

little-coder, a coding agent harness tuned for smaller LLMs, published fresh results on consumer hardware: Qwen 9.7B hits 45.56% on Aider Polyglot (v0.0.2), while the 35B variant reaches 78.67% [Quelle: GitHub]. On Terminal-Bench 2.0, both models rank on the official leaderboard without cloud inference. The v0.1.0+ release migrates to pi, an agent framework with TUI, multi-provider support, and extension architecture.

Local coding agents just became deployable on a dev machine.

Math benchmarks show reasoning spread

Reasoning competition is tightening at the top.

On MATH-500, a benchmark for competition math across algebra, geometry, number theory, and calculus, GPT-5 leads at 99.4%, followed closely by o3 and Grok 3 Mini both at 99.2% [Quelle: Price Per Token]. Claude Sonnet 4 Thinking and Grok 4 sit at 99.1% and 99.0% respectively. The benchmark evaluates 111 models on multi-step reasoning, and the leaderboard now includes pricing data for performance-per-dollar comparisons.

The spread has collapsed—differentiation is shifting to cost and latency.

Sources
Newsroom - Anthropic
Newsroom - Anthropic
22 hours ago ... Sonnet 5 delivers frontier performance across coding, agents, and professional work at scale. Announcements Jun 30, 2026. Claude Science, an AI workbench for ...
anthropic.com
AI Summary

Claude Science, an AI workbench for scientists, launched and is now available with customizable features integrating tools and packages researchers commonly use, along with auditable artifacts and flexible access to computing resources. Separately, Fable 5 returns globally on July 1, and Anthropic is proposing an industry-wide framework for scoring jailbreak severity in collaboration with Amazon, Microsoft, Google, and other partners.

Visit source
Data on AI Capabilities and Benchmarking - Epoch AI
Data on AI Capabilities and Benchmarking - Epoch AI
10 hours ago ... Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated ...
epoch.ai
AI Summary

Claude Fable 5 achieved a new high score of 161 on the Epoch Capabilities Index (ECI), surpassing GPT-5.5 Pro by 1 point and marking the first time Anthropic has led the benchmark in over a year. Epoch AI recently began tracking 13 new evaluations and added nine external benchmarks spanning agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics, with 7 of the new evals incorporated into the ECI.

Visit source
itayinbarr/little-coder: A harness optimized to smaller LLMs - GitHub
itayinbarr/little-coder: A harness optimized to smaller LLMs - GitHub
17 hours ago ... Write/Edit confirmations are pi's responsibility; little-coder doesn't intercept those. Paper / benchmark results. Release, Model, Benchmark, Result. v0.
github.com
AI Summary

little-coder is a coding agent tuned for small local models, achieving significant performance gains through architectural adaptation. The project recently published benchmark results across multiple frameworks: on Aider Polyglot, a 9.7B Qwen model achieved 45.56% (v0.0.2), with the larger Qwen3.6-35B-A3B variant reaching 78.67% (v0.0.5). On Terminal-Bench 2.0, Qwen3.6-35B-A3B scored 24.6% ± 3.2 (rank 120) and the smaller Qwen3.5-9B variant scored 9.2% ± 2.4 (rank 142), both accepted to the official leaderboard. On GAIA's validation set, Qwen3.6-35B-A3B achieved 40.0% (66/165 tasks) with per-level breakdowns of L1 60.4%, L2 37.2%, L3 7.7%. The v0.1.0+ release migrates from a Python substrate to pi, an agent framework providing the core loop, multi-provider support, TUI, and extension architecture. All benchmarks ran on consumer hardware (i9-14900HX, 32 GB RAM, 8 GB VRAM RTX 5070) without cloud inference. The detailed research methodology is documented in the accompanying Substack article "Honey, I Shrunk the Coding Agent."

Visit source
JetBrains AI for Teams and Organizations: From Fragmented AI ...
JetBrains AI for Teams and Organizations: From Fragmented AI ...
17 hours ago ... JetBrains AI for Teams and Organizations: From Fragmented AI Usage to Coordinated Software Development ... As usage of various agentic development tools ...
blog.jetbrains.com
AI Summary

JetBrains is rolling out AI for Teams and Organizations starting July 2026, introducing vendor-agnostic capabilities including team automations and cloud agents for long-running tasks, JetBrains Context to provide repository intelligence for more efficient code understanding, and JetBrains Central for organization-wide governance and cost management across multiple AI tools. The company is transitioning from AI licenses to flexible on-demand AI credits valid for twelve months, allowing organizations to manage AI adoption at scale while developers continue using their preferred tools through open integrations via MCP and ACP protocols.

Visit source
MATH-500 Leaderboard 2026 - Compare AI Model Scores
MATH-500 Leaderboard 2026 - Compare AI Model Scores
19 hours ago ... Model Releases · AI Coding · Open Source · Benchmarks · Hardware · Chips · Regulation · Funding · Image Generation · Video Generation · MCP. Follow us:.
pricepertoken.com
AI Summary

As of July 8, 2026, GPT-5 leads the MATH-500 benchmark for competition mathematics problems with a 99.4% score, followed by o3 and Grok 3 Mini both at 99.2%. The benchmark evaluates 111 models on multi-step reasoning across algebra, geometry, number theory, and calculus. Other top performers include Claude Sonnet 4 Thinking at 99.1%, Grok 4 at 99.0%, and o4 Mini at 98.9%, with pricing data available to compare performance against cost.

Visit source
Compiled overnight by MorningMail.aiDelivered at 05:10 AM