Semifly
Semifly · Tokens & LLM Models

Large Language Models, made easy to compare, download, and run

Track the latest models, compare frontier and open-source LLMs, and download open-weight models — then run any of them on Semifly with tokens, GPU servers, and AI Foundry.

Genealogy

The whole family tree

Every major lab grouped by camp — US vs China, proprietary vs open — with each family’s key versions and years. The China open-weight branch is by far the busiest.

US proprietaryUS/EU openChina openChina proprietary

Key versions per lab, hand-curated as of 2026-06. Refreshed as new flagships ship.

Quality × Price

The whole field on one chart

Every model placed by quality (Artificial Analysis Intelligence Index) and input price — top-left is the value sweet spot. As of 2026-06-18.

Proprietary (API)Open-weight (self-host)Bubble size = context window
Capability over time

How the field got here — and how open caught up

LMArena Elo of the #1 proprietary model (blue) vs the #1 open-weight model (orange), 2023–2026. Open-weight nearly drew level in early 2025; since then the open #1 has been almost entirely Chinese — DeepSeek, Qwen, GLM, Kimi.

Proprietary #1Open-weight #1 (diamond = China)

Source: LMArena (Arena) Elo via BenchLM leaderboard history. Blue tracks the frontier milestones; orange is the open-weight #1 over time.

Adoption

Who is actually being used — and who gets paid

OpenRouter routes about 25T tokens a week across 8M+ developers. By real token volume, Chinese open-weight models dominate the usage charts — yet premium US models still capture most of the dollars.

China vendorOverseas

Top models · weekly tokens

By vendor · top-10 aggregate

The dollar–token split: China-origin models take 45%+ of tokens, while Anthropic holds ~12% of tokens but ~46% of dollar spend through premium pricing. Source: OpenRouter rankings (live-scraped) + market reporting. Per-language and per-use-case breakdowns are chart-only on OpenRouter and not yet scrapable.

Latest in LLMs

What’s new in large models

Updated August 2026 · refreshed regularly

Aug 3, 2026Release

Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test

One paragraph of "Lord of the Rings" in, 5,500 lines of code out. Andrej Karpathy had Claude Opus 5 turn Tolkien's opening into a 3D browser scene. The article Unicorn,…

The Decoder →
Aug 3, 2026Release

Two teams solved the same quantum crypto problem using GPT-5.6 just three hours apart

Two research teams independently solved the same open quantum cryptography problem using OpenAI's GPT-5.6 Sol Ultra, submitting their papers just three hours apart. "If…

The Decoder →
Aug 3, 2026Release

How to Secure AI Agents, MCP Servers, and LLM Apps in Production

AI agents, MCP servers, and LLM apps break the core AppSec assumption that applications do what their code says. This guide walks through a practical see-fix-protect…

MarkTechPost →
Aug 3, 2026Release

Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths

Cogent AI team released Cogent VR-1, a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general…

MarkTechPost →
Aug 3, 2026Release

Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce…

Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a…

MarkTechPost →
Aug 1, 2026Release

AI keeps cracking unsolved math problems, and mathematicians have mixed feelings

OpenAI's refutation of the Unit Distance Conjecture has sparked a wave of AI-assisted advances in mathematics. Fields Medal winner Timothy Gowers says GPT 5.6 Pro solved…

The Decoder →
Aug 1, 2026Release

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview,…

Simon Willison →
Jul 31, 2026Release

smevals - a small eval suite for evaluating models, prompts, and harnesses

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this…

Simon Willison →
View all news & China watch →

Run any of these on Semifly

Tokens & API

Access hosted models through a simple, metered token API.

Get API access →

GPU servers

Buy or lease Supermicro GPU systems to self-host open-weight models.

Browse GPU servers →

AI Foundry

Managed compute for training, fine-tuning, and inference.

Explore AI Foundry →