Signing you in...

Please wait while we verify your authentication

Article · Wednesday, August 5, 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

By Marius BongartsTech22 editions
← See today's latest
Editions
5 / 22
Generated by AI overnight from public sources, refreshed daily.
AI developer tools · What shipped
Wednesday, August 5, 2026
AI developer tools · What shipped

Claude Code ships v2.1.222, vision-language benchmarks climb

1 min read

Claude Code v2.1.222

Claude Code hardens its isolation layer in the latest release.

Version 2.1.222 landed with fixes across worktree isolation, permissions, and connectivity—the kind of plumbing that matters quietly [Source: GitHub]. MCP server attribution, pull request linking, and sandbox credential masking on Linux/WSL all got addressed. Streaming idle timeouts on custom gateways and org-restricted model handling round out the reliability push.

This is the fifth release in three weeks.

Vision-language code benchmark

Multimodal coding just hit 98.8% on the Flame-VLM-Code benchmark.

Claude Opus 4.6 leads the vision-language track, which measures code generation from visual and multimodal inputs [Source: BenchLM]. GLM-5V-Turbo trails at 93.8% and Kimi K2.5 at 88.8%. The gap between top and third is tightening faster than single-modal benchmarks, suggesting the multimodal frontier is solidifying.

Watch whether full-project benchmarks follow the same trajectory.

Repo-level code generation benchmark

VIBE-Pro measures what matters: can a model ship a whole product?

The benchmark spans web, mobile, and simulation tasks end-to-end, not snippet-by-snippet [Source: BenchLM]. MiniMax M2.7 leads the August snapshot at 55.6%, with 24 model releases in the last month chasing the leaderboard. The 55% ceiling hints at how much harder project delivery is than isolated code generation.

The gap between VIBE-Pro and Flame-VLM-Code tells you how far the real work still is.

Sinch Agent Tools for AI assistants

Sinch is tying its API directly into Claude Code and other agents.

Agent Tools ship with VS Code, JetBrains, and Open VSX extensions, plus an MCP server that gives AI assistants live access to Sinch's API definitions and Skills documentation [Source: PRNewswire]. Simulator Mode lets developers test integrations without account setup or live calls. It's the communications-platform play on agent-aware tooling.

Expect API vendors to rush similar IDE integrations.

Sources
VIBE-Pro Leaderboard & Scores — August 2026 | BenchLM.ai
VIBE-Pro Leaderboard & Scores — August 2026 | BenchLM.ai
23 hours ago ... Benchmark profile. VIBE-Pro. A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.
benchlm.ai
AI Summary

VIBE-Pro is a repo-level code generation and full-project delivery benchmark that spans web, mobile, and simulation-style implementation tasks. As of August 4, 2026, MiniMax M2.7 leads the public snapshot with a score of 55.6%, with 24 confirmed model releases in the last 30 days. The benchmark measures whether models can complete substantial product requirements across end-to-end software delivery scenarios rather than single-file code snippets.

Visit source
Flame-VLM-Code Leaderboard & Scores — August 2026
Flame-VLM-Code Leaderboard & Scores — August 2026
20 hours ago ... Flame-VLM-Code tracked score snapshot across 3 AI models. Display only on BenchLM and excluded from overall rankings. A vision-language coding benchmark for ...
benchlm.ai
AI Summary

Flame-VLM-Code is a vision-language coding benchmark for generating correct code from visual and multimodal inputs, with results tracked as of August 4, 2026. Claude Opus 4.6 leads the benchmark at 98.8%, followed by GLM-5V-Turbo at 93.8% and Kimi K2.5 at 88.8%. The benchmark currently covers 3 evaluated models and falls within the multimodal and grounded category, though BenchLM displays it as a reference-only benchmark pending full verification of source attachments.

Visit source
Releases · anthropics/claude-code - GitHub
Releases · anthropics/claude-code - GitHub
6 hours ago ... Enterprise platformAI-powered developer platform. AVAILABLE ADD-ONS. GitHub ... Fixed tool errors not being displayed for tools no longer available locally, for ...
github.com
AI Summary

Claude Code v2.1.222 was released with multiple bug fixes and reliability improvements focused on isolation, permissions, and connectivity. Key changes include fixes for worktree-isolated sessions preventing destructive git commands, PreToolUse auto-allow hooks bypassing tool restrictions, usage credit request blocking on Team/Enterprise, HTTPS proxy connectivity issues, and connection mid-response errors. The release also addressed MCP server usage attribution, pull request linking after branch pushes, org-restricted model handling, stream idle timeouts on custom gateways, and issues with connector authorization, tool error display, long summaries, and file watcher crashes. Additional improvements cover permission classifier safety for cross-session messaging, sandbox credentials masking on Linux/WSL, and various UI fixes for screen readers and diff views. Previous v2.1.221 added Focus view for chat activity summaries, sandbox credential file masking, and MCP server connection improvements in print mode, alongside fixes for Bash permission check bypasses, PowerShell path handling, thinking toggle persistence, and team spend-limit messaging. v2.1.219 introduced Claude Opus 5 as the default model with 1M context and fast-mode pricing, plus nested subagent forwarding and new safety features. v2.1.212 changed /fork to create background sessions and added /subtask for in-session subagents, session-wide WebSearch limits, automatic background promotion for long-running MCP calls, and session resume picker functionality.

Visit source
Sinch launches Agent Tools for developers and AI coding assistants
14 hours ago ... Agent Tools gives those developers a set of integrations optimized for AI agents and modern development environments, reducing the need to move between coding ...
prnewswire.com
AI Summary

Sinch announced Agent Tools, a new developer toolkit for building applications with Sinch's communication platform. The suite includes extensions for Visual Studio Code, JetBrains IDEs, and Open VSX-compatible editors, with built-in support for AI coding assistants including Claude Code, Cursor, GitHub Copilot, and ChatGPT Desktop. Key features include Simulator Mode for testing integrations without creating an account or making live API calls, MCP server integration providing AI assistants access to live Sinch API definitions, and structured product knowledge through Sinch Skills covering configuration patterns across services including Conversation API, Voice, Verification, Numbers, and others.

Visit source
Compiled overnight by MorningMail.aiDelivered at 05:10 AM