Releases, benchmarks, open-source drops, and industry shifts — refreshed regularly so you can see what’s newest at a glance.
Updated August 2026 · refreshed regularly
One paragraph of "Lord of the Rings" in, 5,500 lines of code out. Andrej Karpathy had Claude Opus 5 turn Tolkien's opening into a 3D browser scene. The article Unicorn,…
The Decoder →Two research teams independently solved the same open quantum cryptography problem using OpenAI's GPT-5.6 Sol Ultra, submitting their papers just three hours apart. "If…
The Decoder →AI agents, MCP servers, and LLM apps break the core AppSec assumption that applications do what their code says. This guide walks through a practical see-fix-protect…
MarkTechPost →Cogent AI team released Cogent VR-1, a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general…
MarkTechPost →Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a…
MarkTechPost →OpenAI's refutation of the Unit Distance Conjecture has sparked a wave of AI-assisted advances in mathematics. Fields Medal winner Timothy Gowers says GPT 5.6 Pro solved…
The Decoder →Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview,…
Simon Willison →smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this…
Simon Willison →Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop.…
Simon Willison →Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel…
The Decoder →Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more…
The Decoder →Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds…
The Decoder →Introducing Claude Opus 5 I've been offline kayaking with sea otters for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its…
Simon Willison →AI Mania Is Eviscerating Global Decision-Making Here's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he…
Simon Willison →Firefox in WebAssembly This is absurdly cool: Puter compiled Firefox to WebAssembly such that the whole browser runs in another browser. Here's my blog, running in…
Simon Willison →OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human red teamers manage just 13 percent. The…
The Decoder →A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 Sol Pro to disprove a central open conjecture about the Benjamini-Hochberg method in roughly 90…
The Decoder →PrismML compressed a 27B reasoning model to under 4 GB, small enough for phone-class local inference while retaining most benchmark performance.
The Decoder →Inkling supports text, image and audio inputs, a 1M context window, controllable thinking effort, and day-0 deployment support in Transformers, SGLang and llama.cpp.
Hugging Face →Soofi S 30B-A3B is an open Mamba-Transformer MoE model activating 3.2B of 31.6B parameters for German and English The post Soofi Consortium Releases Soofi S 30B-A3B: An…
MarkTechPost →xAI open-sourced grok-build after criticism of its CLI behavior, giving developers a closer look at how the coding tool handles local project context.
Simon Willison →A Claude web_fetch data-exfiltration test shows how tool design and prompt boundaries matter when LLMs browse private or sensitive content.
Simon Willison →See how Vibe, Claude Code, Cursor, and Codex compare on cost, open weights, self-hosting, and async agent surfaces. The post Mistral Vibe for Code vs Claude Code vs…
MarkTechPost →Final Token Preference Optimization targets the token that starts repetitive loops; Liquid reports LFM2.5-2.6B loop rates falling from 10.2% to 1.4%.
Liquid AI →China’s LLM scene moves on its own beat — the latest on Chinese models (DeepSeek, Qwen, GLM, Kimi, MiniMax and more), in English, pulled from Chinese and international coverage.
Alibaba is marketing its new AI model Qwen 3.8 with a video that shows the AI working while a person enjoys their hobbies. It's a deliberate contrast to the job loss…
The Decoder →MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. The article China's MiniMax H3 is the first open model…
The Decoder →Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips…
The Decoder →Alibaba's Qwen team moved Qwen3.8-Max from preview to general availability, with published per-token pricing and open weights due next week. The 2.4T parameter MoE model…
MarkTechPost →阿里巴巴正式发布新一代基座大模型Qwen3.8,整体性能处于全球大模型第一梯队。Qwen3.8-Max预计下周开源,同时还将开源 Qwen3.8-27B。
QbitAI →阿里巴巴旗下“千问办公”(QwenWork)开启公测,个人和企业用户均可体验。用户可在“千问办公”体验阿里最新旗舰模型Qwen3.8。
QbitAI →Anthropic's Claude Opus 5 generates complete 3D games from single prompts, including a first-person shooter, a kart racer, and a Minecraft clone, all without a single…
The Decoder →deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on…
Simon Willison →Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild…
Simon Willison →moonshotai/Kimi-K3 As promised earlier this month, Moonshot have released the weights for their excellent 2.8 trillion parameter Kimi K3. They're a hefty 1.56TB on…
Simon Willison →The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on…
The Decoder →Exploring a new path for native multimodal integration.
QbitAI →Give it a picture, and it returns your entire world.
QbitAI →Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been…
Simon Willison →Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite…
Simon Willison →Claude make Fable 5 permanent An update from the @claudeai account on Twitter: Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at…
Simon Willison →STEPX Neo is positioned as a large-model-native intelligent-agent phone, pushing StepFun beyond chatbots into device-level AI interaction.
QbitAI →The real-time speech model upgrades intelligence, agent tool invocation, empathetic dialogue and duplex interaction fluency for voice-first applications.
QbitAI →DeepSeek is already raising again. The Chinese AI lab just closed its first funding round and needs capital for its own data centers and chips to keep its aggressive…
The Decoder →Hy3 activates 21B parameters per token, ships under Apache 2.0, and targets reasoning, coding agents and long-context workflows with vLLM/SGLang deployment recipes.
MarkTechPost →The openJiuwen team introduces a multimodal skill pattern that pairs text instructions with visual references and reusable experience libraries for agent workflows.
QbitAI →Chinese startup Deepseek is building its own AI chip, Reuters reports. The article Deepseek is designing its own AI chip appeared first on The Decoder.
The Decoder →LongCat-2.0 targets agentic coding with a native 1M-token context window, about 48B active parameters per token, and MIT-licensed release plans.
MarkTechPost →Compiled from public reporting; Chinese-source items are machine-translated. Confirm details with each vendor.