Semifly
Semifly · LLMs

Latest in large language models

Releases, benchmarks, open-source drops, and industry shifts — refreshed regularly so you can see what’s newest at a glance.

Global · English

What’s new worldwide

Updated August 2026 · refreshed regularly

Aug 3, 2026Release

Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test

One paragraph of "Lord of the Rings" in, 5,500 lines of code out. Andrej Karpathy had Claude Opus 5 turn Tolkien's opening into a 3D browser scene. The article Unicorn,…

The Decoder →
Aug 3, 2026Release

Two teams solved the same quantum crypto problem using GPT-5.6 just three hours apart

Two research teams independently solved the same open quantum cryptography problem using OpenAI's GPT-5.6 Sol Ultra, submitting their papers just three hours apart. "If…

The Decoder →
Aug 3, 2026Release

How to Secure AI Agents, MCP Servers, and LLM Apps in Production

AI agents, MCP servers, and LLM apps break the core AppSec assumption that applications do what their code says. This guide walks through a practical see-fix-protect…

MarkTechPost →
Aug 3, 2026Release

Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths

Cogent AI team released Cogent VR-1, a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general…

MarkTechPost →
Aug 3, 2026Release

Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce…

Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a…

MarkTechPost →
Aug 1, 2026Release

AI keeps cracking unsolved math problems, and mathematicians have mixed feelings

OpenAI's refutation of the Unit Distance Conjecture has sparked a wave of AI-assisted advances in mathematics. Fields Medal winner Timothy Gowers says GPT 5.6 Pro solved…

The Decoder →
Aug 1, 2026Release

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview,…

Simon Willison →
Jul 31, 2026Release

smevals - a small eval suite for evaluating models, prompts, and harnesses

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this…

Simon Willison →
Jul 30, 2026Release

Advancing the price-performance frontier with GPT‑5.6

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop.…

Simon Willison →
Jul 24, 2026Model rankings

Anthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price

Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel…

The Decoder →
Jul 24, 2026Open source

Microsoft's open-weight AI push is so obviously an Azure play it hurts

Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more…

The Decoder →
Jul 24, 2026Model rankings

Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool

Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds…

The Decoder →
Jul 24, 2026Release

Introducing Claude Opus 5

Introducing Claude Opus 5 I've been offline kayaking with sea otters for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its…

Simon Willison →
Jul 19, 2026Release

AI Mania Is Eviscerating Global Decision-Making

AI Mania Is Eviscerating Global Decision-Making Here's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he…

Simon Willison →
Jul 16, 2026Release

Firefox in WebAssembly

Firefox in WebAssembly This is absurdly cool: Puter compiled Firefox to WebAssembly such that the whole browser runs in another browser. Here's my blog, running in…

Simon Willison →
Jul 15, 2026Release

OpenAI is now using AI to attack its own AI, and it's working better than humans ever did

OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human red teamers manage just 13 percent. The…

The Decoder →
Jul 15, 2026Release

GPT-5.6 Sol reportedly disproves a 30-year-old statistics conjecture in 90 minutes after humans couldn't crack it

A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 Sol Pro to disprove a central open conjecture about the Benjamini-Hochberg method in roughly 90…

The Decoder →
Jul 15, 2026Open source

Bonsai 27B is a full open reasoning model that fits on an iPhone

PrismML compressed a 27B reasoning model to under 4 GB, small enough for phone-class local inference while retaining most benchmark performance.

The Decoder →
Jul 15, 2026Open source

Thinking Machines releases Inkling, a 975B open-weight multimodal MoE

Inkling supports text, image and audio inputs, a 1M context window, controllable thinking effort, and day-0 deployment support in Transformers, SGLang and llama.cpp.

Hugging Face →
Jul 15, 2026Architecture

Soofi Consortium Releases Soofi S 30B-A3B: An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English

Soofi S 30B-A3B is an open Mamba-Transformer MoE model activating 3.2B of 31.6B parameters for German and English The post Soofi Consortium Releases Soofi S 30B-A3B: An…

MarkTechPost →
Jul 15, 2026Open source

xai-org/grok-build, now open source

xAI open-sourced grok-build after criticism of its CLI behavior, giving developers a closer look at how the coding tool handles local project context.

Simon Willison →
Jul 15, 2026Architecture

How I tricked Claude into leaking your deepest, darkest secrets

A Claude web_fetch data-exfiltration test shows how tool design and prompt boundaries matter when LLMs browse private or sensitive content.

Simon Willison →
Jul 14, 2026Open source

Mistral Vibe for Code vs Claude Code vs Cursor vs Codex: Four Agents Scored on One Scaffold-to-PR Task

See how Vibe, Claude Code, Cursor, and Codex compare on cost, open weights, self-hosting, and async agent surfaces. The post Mistral Vibe for Code vs Claude Code vs…

MarkTechPost →
Jul 7, 2026Open source

Liquid AI open-sources Antidoom to reduce reasoning-model doom loops

Final Token Preference Optimization targets the token that starts repetitive loops; Liquid reports LFM2.5-2.6B loop rates falling from 10.2% to 1.4%.

Liquid AI →
China · Watch

China watch

China’s LLM scene moves on its own beat — the latest on Chinese models (DeepSeek, Qwen, GLM, Kimi, MiniMax and more), in English, pulled from Chinese and international coverage.

Aug 3, 2026Release

Alibaba's new Qwen model is also taking your job, but this time it's great

Alibaba is marketing its new AI model Qwen 3.8 with a video that shows the AI working while a person enjoys their hobbies. It's a deliberate contrast to the job loss…

The Decoder →
Aug 3, 2026Open source

China's MiniMax H3 is the first open model to top an AI video ranking

MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. The article China's MiniMax H3 is the first open model…

The Decoder →
Aug 3, 2026Open source

Alibaba’s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters

Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips…

The Decoder →
Aug 3, 2026Open source

Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to…

Alibaba's Qwen team moved Qwen3.8-Max from preview to general availability, with published per-token pricing and open weights due next week. The 2.4T parameter MoE model…

MarkTechPost →
Aug 3, 2026Release

刚刚,阿里Qwen3.8-Max来了!冲进全球第一梯队,模型表现直逼Claude

编程、专业工作、长程任务、多模态分析统统梭哈

QbitAI →
Aug 3, 2026Release

阿里Qwen3.8正式发布,编程与办公再进化,推理更快更稳定

阿里巴巴正式发布新一代基座大模型Qwen3.8,整体性能处于全球大模型第一梯队。Qwen3.8-Max预计下周开源,同时还将开源 Qwen3.8-27B。

QbitAI →
Aug 3, 2026Release

阿里“千问办公”开启公测

阿里巴巴旗下“千问办公”(QwenWork)开启公测,个人和企业用户均可体验。用户可在“千问办公”体验阿里最新旗舰模型Qwen3.8。

QbitAI →
Aug 2, 2026Release

Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music

Anthropic's Claude Opus 5 generates complete 3D games from single prompts, including a first-person shooter, a kart racer, and a Minecraft clone, all without a single…

The Decoder →
Jul 31, 2026Release

deepseek-ai/DeepSeek-V4-Flash-0731

deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on…

Simon Willison →
Jul 31, 2026Open source

Oxide and Friends: The Open Weight Revolution with Simon Willison

Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild…

Simon Willison →
Jul 27, 2026Open source

moonshotai/Kimi-K3

moonshotai/Kimi-K3 As promised earlier this month, Moonshot have released the weights for their excellent 2.8 trillion parameter Kimi K3. They're a hefty 1.56TB on…

Simon Willison →
Jul 24, 2026Release

Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on…

The Decoder →
Jul 24, 2026Release

Hefei hits another AI unicorn: In the multimodal sector, it raised 2.1 billion in just three months.

Exploring a new path for native multimodal integration.

QbitAI →
Jul 24, 2026Model rankings

The domestic world model has topped the leaderboard by Fei-Fei Li's team! It is compatible with domestic Ascend…

Give it a picture, and it returns your entire world.

QbitAI →
Jul 22, 2026Release

Are AI labs pelicanmaxxing?

Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been…

Simon Willison →
Jul 20, 2026Release

Who’s Afraid of Chinese Models?

Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite…

Simon Willison →
Jul 18, 2026Release

Claude make Fable 5 permanent

Claude make Fable 5 permanent An update from the @claudeai account on Twitter: Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at…

Simon Willison →
Jul 15, 2026Release

StepFun showcases STEPX Neo, an LLM-native agent phone, at WAIC 2026

STEPX Neo is positioned as a large-model-native intelligent-agent phone, pushing StepFun beyond chatbots into device-level AI interaction.

QbitAI →
Jul 15, 2026Release

Alibaba releases Qwen-Audio-3.0-Realtime for live voice agents

The real-time speech model upgrades intelligence, agent tool invocation, empathetic dialogue and duplex interaction fluency for voice-first applications.

QbitAI →
Jul 14, 2026Industry

DeepSeek needs more cash just weeks after closing its first $7 billion round

DeepSeek is already raising again. The Chinese AI lab just closed its first funding round and needs capital for its own data centers and chips to keep its aggressive…

The Decoder →
Jul 7, 2026Open source

Tencent releases Hy3, a 295B open MoE model with 256K context

Hy3 activates 21B parameters per token, ships under Apache 2.0, and targets reasoning, coding agents and long-context workflows with vLLM/SGLang deployment recipes.

MarkTechPost →
Jul 7, 2026Release

openJiuwen debuts Skill-Omni for multimodal agent skills

The openJiuwen team introduces a multimodal skill pattern that pairs text instructions with visual references and reusable experience libraries for agent workflows.

QbitAI →
Jul 7, 2026Release

Deepseek is designing its own AI chip

Chinese startup Deepseek is building its own AI chip, Reuters reports. The article Deepseek is designing its own AI chip appeared first on The Decoder.

The Decoder →
Jul 5, 2026Open source

Meituan releases LongCat-2.0, a 1.6T open MoE coding model

LongCat-2.0 targets agentic coding with a native 1M-token context window, about 48B active parameters per token, and MIT-licensed release plans.

MarkTechPost →

Compiled from public reporting; Chinese-source items are machine-translated. Confirm details with each vendor.

Run any of these on Semifly

Tokens & API

Access hosted models through a simple, metered token API.

Get API access →

GPU servers

Buy or lease Supermicro GPU systems to self-host open-weight models.

Browse GPU servers →

AI Foundry

Managed compute for training, fine-tuning, and inference.

Explore AI Foundry →