Releases, benchmarks, open-source drops, and industry shifts — refreshed regularly so you can see what’s newest at a glance.
Updated July 2026 · refreshed regularly
Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel…
The Decoder →Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more…
The Decoder →Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds…
The Decoder →Introducing Claude Opus 5 I've been offline kayaking with sea otters for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its…
Simon Willison →AI Mania Is Eviscerating Global Decision-Making Here's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he…
Simon Willison →Firefox in WebAssembly This is absurdly cool: Puter compiled Firefox to WebAssembly such that the whole browser runs in another browser. Here's my blog, running in…
Simon Willison →OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human red teamers manage just 13 percent. The…
The Decoder →A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 Sol Pro to disprove a central open conjecture about the Benjamini-Hochberg method in roughly 90…
The Decoder →PrismML compressed a 27B reasoning model to under 4 GB, small enough for phone-class local inference while retaining most benchmark performance.
The Decoder →Inkling supports text, image and audio inputs, a 1M context window, controllable thinking effort, and day-0 deployment support in Transformers, SGLang and llama.cpp.
Hugging Face →Soofi S 30B-A3B is an open Mamba-Transformer MoE model activating 3.2B of 31.6B parameters for German and English The post Soofi Consortium Releases Soofi S 30B-A3B: An…
MarkTechPost →xAI open-sourced grok-build after criticism of its CLI behavior, giving developers a closer look at how the coding tool handles local project context.
Simon Willison →A Claude web_fetch data-exfiltration test shows how tool design and prompt boundaries matter when LLMs browse private or sensitive content.
Simon Willison →See how Vibe, Claude Code, Cursor, and Codex compare on cost, open weights, self-hosting, and async agent surfaces. The post Mistral Vibe for Code vs Claude Code vs…
MarkTechPost →Final Token Preference Optimization targets the token that starts repetitive loops; Liquid reports LFM2.5-2.6B loop rates falling from 10.2% to 1.4%.
Liquid AI →Anthropic is rolling out its AI agent Claude Cowork to mobile and web. Until now, the feature was limited to the desktop app. The agent keeps working in the background…
The Decoder →Cohere has released Transcribe Arabic, an open-source model for Arabic speech recognition that the company says outperforms Whisper and OmniASR on dialects,…
The Decoder →Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it "J-Space" and can now read it using a new analysis…
The Decoder →OpenAI added two Realtime models to its API. GPT-Realtime-2.1-mini is a mini reasoning model for voice, priced like the earlier gpt-realtime-mini. OpenAI also cut p95…
MarkTechPost →We build an end-to-end GRPO training workflow that teaches Gemma-3 to reason through GSM8K math problems. We prepare the environment, authenticate with Hugging Face,…
MarkTechPost →Synthetic Sciences has released OpenScience, an Apache-2.0 AI workbench for scientific research. It works with any frontier or open-weight model, using your own API…
MarkTechPost →Sakana Translate adds Japanese-English-Chinese translation, proofreading and follow-up Q&A modes to Sakana Chat, powered by the Namazu model series.
MarkTechPost →Better Models: Worse Tools Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool…
Simon Willison →Fable 5 returns globally on Claude.ai, Claude Code, Claude Cowork and the Claude Platform with additional cybersecurity safeguards after the June suspension.
Anthropic →China’s LLM scene moves on its own beat — the latest on Chinese models (DeepSeek, Qwen, GLM, Kimi, MiniMax and more), in English, pulled from Chinese and international coverage.
The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on…
The Decoder →Exploring a new path for native multimodal integration.
QbitAI →Give it a picture, and it returns your entire world.
QbitAI →Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been…
Simon Willison →Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite…
Simon Willison →Claude make Fable 5 permanent An update from the @claudeai account on Twitter: Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at…
Simon Willison →STEPX Neo is positioned as a large-model-native intelligent-agent phone, pushing StepFun beyond chatbots into device-level AI interaction.
QbitAI →The real-time speech model upgrades intelligence, agent tool invocation, empathetic dialogue and duplex interaction fluency for voice-first applications.
QbitAI →DeepSeek is already raising again. The Chinese AI lab just closed its first funding round and needs capital for its own data centers and chips to keep its aggressive…
The Decoder →Hy3 activates 21B parameters per token, ships under Apache 2.0, and targets reasoning, coding agents and long-context workflows with vLLM/SGLang deployment recipes.
MarkTechPost →The openJiuwen team introduces a multimodal skill pattern that pairs text instructions with visual references and reusable experience libraries for agent workflows.
QbitAI →Chinese startup Deepseek is building its own AI chip, Reuters reports. The article Deepseek is designing its own AI chip appeared first on The Decoder.
The Decoder →LongCat-2.0 targets agentic coding with a native 1M-token context window, about 48B active parameters per token, and MIT-licensed release plans.
MarkTechPost →AI startup Lindy ditched Claude entirely for Deepseek after AI costs exceeded personnel costs. CEO Flo Crivello calls it "a matter of survival for the business." The…
The Decoder →Chinese AI lab Zhipu AI releases GLM-5.2 with a stable 1-million-token context under the MIT license. On FrontierSWE, a benchmark for hours-long coding tasks, the…
The Decoder →MiniMax released MSA, a sparse attention built on Grouped Query Attention. A lightweight Index Branch selects Top-k key-value blocks per query and GQA group; the Main…
MarkTechPost →GLM-5.2: Built for Long-Horizon Tasks
Hugging Face →Microsoft is weighing a fine-tuned version of Deepseek V4 as a cheaper model option for Copilot Cowork. The company is also switching to usage-based billing, since…
The Decoder →Chinese AI startup DeepSeek has raised more than 50 billion yuan - about $7.4 billion - in its first external funding round. The article DeepSeek takes outside money for…
The Decoder →The Qwen team's newest agentic coding model lands alongside a wave of MiniMax 'Highspeed' variants (M2.5/M2.7), keeping the open-weight release pace relentless.
LLM-Stats →Agentic coding model (~1T params, 256K context) under a Modified MIT license — the update cuts reasoning token usage ~30% versus K2.6 and boosts MCP tool-calling.
LLM-Stats →DeepSeek-V4: a million-token context that agents can actually use
Hugging Face →The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+
Hugging Face →The artificial intelligence coding revolution comes with a catch: it's expensive.Claude Code, Anthropic's terminal-based AI agent that can write, debug, and deploy code…
VentureBeat →Compiled from public reporting; Chinese-source items are machine-translated. Confirm details with each vendor.