title: "Top 10 Open Source LLMs (August 2026): Kimi K3 & More" slug: "top-10-open-source-llms-the-deepseek-revolution-2026" excerpt: "Kimi K3 hits 57 and beats last generation's closed champion. The verified August 2026 open-weight LLM ranking: benchmarks, prices, licenses, routing." author: "O-mega Team" date: "2026-08-05" tags: ["Open Source AI", "Large Language Models", "Kimi K3", "DeepSeek", "AI Rankings"] category: "Artificial Intelligence"
The August 2026 ranking of open-weight AI models, rebuilt after the most violent month in open-model history, with every number verified against a primary source this week.
An open-weight model now outscores the closed champion of one generation ago. Kimi K3 scores 57 on the Artificial Analysis Intelligence Index v4.1, ahead of where Claude Opus 4.8, the strongest closed model in the world until July 24, stood in our previous edition at roughly 55.7 - Artificial Analysis. The new closed leader, Claude Opus 5, sits at 61, which puts the open-versus-closed gap at 4 points, about 6.6% relative. That crossing is not a footnote. It is the strongest evidence yet for the thesis this guide has argued since the original DeepSeek moment: open weights trail the closed frontier by months, not by a class.
Full disclosure before anything else: the July 8 edition of this guide aged badly, fast. We published it as a complete rebuild, and within three weeks its central claims were false. GLM-5.2 was no longer the open peak, Opus 4.8 was no longer the closed peak, Kimi K2.6 was no longer Moonshot's flagship, and a deprecation we described in the future tense had already happened. July 2026 compressed roughly a year of normal model-market churn into 25 days: nine dated releases and retirements that section 1 walks through as a change log, because owning exactly what broke is more useful to you than pretending this edition is timeless. It is not. It is accurate as of August 5, 2026, with sources you can check.
One more thing distinguishes this guide from the ranking pages currently filling search results, several of which carry fresh "August 2026" labels above lists that still lead with Kimi K2.5 and Llama 4. We operate O-mega, a platform that runs production AI agent workforces, which means our routing table had to absorb every one of July's releases with real workloads attached. Section 14 shows what we actually route where, including the models we deliberately do not use and why. That is information an aggregator cannot give you, because an aggregator does not pay the output-token bill.
Contents
- What Our July Edition Got Wrong: A 25-Day Change Log
- Kimi K3: The Open Model That Passed Last Generation's Closed Champion
- DeepSeek V4 Flash 0731: The Retirement, the Re-Post-Train, and Surge Pricing
- GLM-5.2: Still Z.ai's Latest, Still the Open Coding Benchmark
- MiniMax M3: Long-Context Value and the 2.7T Report
- Tencent Hy3: The Apache 2.0 Search Specialist Almost Everyone Missed
- Inkling and the American Rearguard
- gpt-oss-120b Turns One: The Most-Run Open Model Has No Successor
- MiMo-V2.5-Pro: The Quiet Trillion
- The Half-Exits: Qwen Goes API-First, Meta Stays Closed
- The Frontier Gap: The Lag Thesis Just Got Its Proof
- Licenses in August 2026: The Fine Print Is Coming Back
- Pricing Economics: Cache Hits, Surge Hours, and the $15 Output Token
- What an Agent Platform Actually Routes Where
- Conclusion: Choosing Your Model in August 2026
The August 2026 Scorecard
Before the deep dives, the master ranking: exactly ten models, scored on four criteria weighted by what determines success when you deploy an open model in production. Raw intelligence and agentic capability (35%), cost efficiency (25%), deployability covering license, context window, and hardware reality (20%), and ecosystem adoption measured by downloads, provider availability, and tooling (20%). Scores run 0-10 and every cell shows the evidence behind the number. Every benchmark score, price, and download count in this table was pulled live this week from the linked primary sources, not carried over from the July edition.
The ranking rewards the model you would actually build on, which is why the intelligence leader does not take the top slot. Kimi K3 is the smartest open model ever released and it ranks third, because its output pricing and license terms are genuinely different from the MIT-and-pennies norm the open frontier established. That is the honest trade-off aggregator lists flatten, and it is the single most important change in open-model economics this year.
| # | Model | What It Does | Intelligence (35%) | Cost (25%) | Deployability (20%) | Ecosystem (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0731 | Re-post-trained agentic workhorse, launch pricing held | 9 - AA 50, Terminal-Bench 2.1 82.7 | 10 - $0.14/M in, $0.0028/M cache hit, $0.28/M out | 9 - MIT, 1M context, ~110GB at 3-bit | 9 - 433K HF downloads/mo, every major provider | 9.3 |
| 2 | GLM-5.2 | Open agentic-coding leader, unchanged since June 13 | 9 - AA 51, SWE-bench Verified 84.2 | 8 - $0.60/M in, $1.50/M out on OpenRouter | 8 - MIT, 1M context | 9 - 2.23M HF downloads/mo | 8.6 |
| 3 | Kimi K3 | 2.8T MoE, first open model past the old closed champ | 10 - AA 57, DeepSWE v1.1 67.3 | 6 - $15/M output; $0.30/M cache-hit input | 7 - 1M context, native vision, but 2.8T weights + license gates | 8 - 1.13M HF downloads in week one | 8.0 |
| 4 | MiniMax M3 | 1M-context multimodal workhorse (text+image+video in) | 8 - AA 44, ties old V4 Pro preview class | 9 - $0.24/M in, $0.96/M out (60% promo) | 8 - 1M context, 428B total | 6 - multi-provider, thinner tooling than DeepSeek | 7.9 |
| 5 | Tencent Hy3 | Agentic search and long-context retrieval specialist | 8 - BrowseComp 84.2, AA-LCR 73.4 (vendor-reported) | 8 - FP8 fits under 300GB, single-node serving | 9 - Apache 2.0 with zero carve-outs, 256K context | 6 - 23K HF downloads/mo, day-one OpenRouter + Cline | 7.8 |
| 6 | gpt-oss-120b | Single-GPU American workhorse, one year old this week | 4 - AA 24, two generations off the pace | 9 - one 80GB GPU via MXFP4, zero API spend | 9 - Apache 2.0, 117B total / 5.1B active | 10 - 4.22M HF downloads/mo, most-run open model | 7.5 |
| 7 | MiMo-V2.5-Pro | Xiaomi's MIT-licensed trillion-scale value play | 8 - AA 42, top-five open weights | 8 - 42B active of 1.02T | 8 - MIT, 1M context | 5 - 54K HF downloads/mo | 7.4 |
| 8 | Gemma 4 | Phone-to-workstation family with audio + vision input | 4 - AA 29 for the 31B | 9 - free on-device inference | 9 - E2B/E4B run offline on phones and edge boards | 8 - deepest small-model tooling | 7.1 |
| 9 | Inkling | New leading US open model, audio-native, 3 weeks old | 7 - AA 41, SWE-bench Verified 77.6 | 7 - Tinker at 50% intro discount; 975B self-host is heavy | 8 - Apache 2.0, 1M-token weights | 5 - 69K HF downloads/mo, brand-new ecosystem | 6.8 |
| 10 | Nemotron 3 Ultra 550B | US open leader until July 15, NVIDIA-stack aligned | 6 - AA 38 on v4.1 | - (no verified current list price) | 7 - 550B open weights, NVIDIA tooling | 6 - enterprise NVIDIA channel | 6.3 |
Sort check: 9.3, 8.6, 8.0, 7.9, 7.8, 7.5, 7.4, 7.1, 6.8, 6.3. Descending order confirmed, ten rows exactly. Nemotron's cost cell is a dash because we could not verify a current hosted list price this run, so its final score is the weighted average of the remaining three criteria. Intelligence scores anchor to the Artificial Analysis open-weights leaderboard on Index v4.1; download counts come from each model's Hugging Face page as of August 5; prices come from official pricing pages and OpenRouter listings. Two labs that appeared in the July top ten are gone from the table entirely: Qwen, because Alibaba's latest open release no longer places in the AA open top ten (section 10), and Mistral, whose Medium 3.5 scores 30 on the same index and now trails even the American rearguard's new leader. Now, the confession.
1. What Our July Edition Got Wrong: A 25-Day Change Log
The July 8 edition of this guide was not lazy. It rewrote every section, verified every benchmark against a source, and built a weighted scorecard. It still went stale in under three weeks, and one of its omissions was stale on the day it published: Tencent's Hy3 shipped July 6, two days before our update, under Apache 2.0 at 295B parameters, and we missed it completely - Digital Applied. We also kept describing Alibaba's Qwen line as fully open when its flagship tier had quietly gone API-only in May. Those were unforced errors. The rest was July.
Here is the dated sequence that invalidated the July edition, compressed into one diagram, because no month since the original DeepSeek release has moved the board this much.
Walking the diagram: July 6, Hy3 arrives. July 9, Meta opens a public developer API for Muse Spark 1.1, softening its closed turn without reversing it - Fortune. July 15, Thinking Machines Lab releases Inkling, instantly the best American open model - Thinking Machines. July 16, Moonshot ships Kimi K3 hosted, at once the smartest and the most expensive open-frontier model - Simon Willison. July 24 was the pivot day: Anthropic released Claude Opus 5, resetting the closed frontier - Anthropic, and DeepSeek retired deepseek-chat and deepseek-reasoner at 15:59 UTC with no grace period - Enterprise DNA. July 26, K3's open weights landed a day ahead of Moonshot's own deadline - Roo. July 31, DeepSeek shipped V4-Flash-0731, a re-post-trained Flash that jumped ten points on the intelligence index at unchanged prices - MarkTechPost. Alongside all of that, Alibaba announced Qwen3.8-Max at 2.4 trillion parameters with open weights promised but, as of this writing, not delivered - Wikipedia.
For the record, these specific July-edition claims are now corrected. GLM-5.2 at 51 is no longer "the highest mark ever recorded by open weights"; Kimi K3 at 57 is, and it did not just beat GLM-5.2, it passed the closed champion it was previously measured against. Opus 4.8 is no longer the closed peak; Opus 5 at 61 is, with Claude Fable 5 at 60 and GPT-5.6 Sol at 59 behind it - Artificial Analysis. Kimi K2.6 is no longer Moonshot's flagship. The "44-club three-way tie" we described at the open frontier is dead: the live ladder now reads 57 / 51 / 50 / 44 / 42 / 41 / 38. And our July scorecard credited Nemotron 3 Ultra with a 48 based on a June OpenRouter snapshot; on the current AA Index v4.1 it scores 38, a reminder that index versions and evaluation settings move numbers and that this guide should have anchored to one instrument from the start.
Why publish the errata instead of quietly overwriting? Because the failure mode generalizes to every model list you will read this year, including this one. A ranking of open models is a photograph of a moving object, and the honest response is not to claim your photograph is special but to timestamp it, source it, and tell readers exactly which parts move fastest. Version numbers and leaderboard order rot in weeks. Structural facts, cache-pricing economics, license mechanics, the closed-to-open lag, rot in years. This edition is organized so that the fast-rotting material is concentrated in sections you can re-verify in one click, and the durable reasoning stands apart from it.
The same forensics yield a checklist for reading anyone's ranking, ours included, and it is worth internalizing because most of the pages competing for this search query fail it. First, check whether the claims carry dates, not just the title: a page labeled for the current month that still leads with last generation's models is describing its label, not its contents. Second, check the index version behind any composite score: the same model can legitimately score 48 on one snapshot and 38 on another, as Nemotron just demonstrated across index revisions, so scores from different instruments must never be compared directly. Third, distinguish vendor-reported from independently measured numbers: Tencent's launch benchmarks for Hy3 are honest about their own weaknesses, which raises their credibility, but they remain the vendor's numbers until a third party reproduces them, and this guide labels them accordingly. Fourth, check whether prices are list or promotional: two of the prices in our own scorecard carry active discounts (GLM-5.2 and MiniMax M3 on OpenRouter), which we flag inline because a promo that expires changes the math. A reader armed with those four checks can audit any model ranking in about five minutes, and a ranking that survives the audit is worth trusting for roughly a month. None deserves longer.
2. Kimi K3: The Open Model That Passed Last Generation's Closed Champion
Moonshot AI's Kimi K3 is the headline event of this cycle and the most consequential open release since DeepSeek R1. Announced and hosted on July 16, 2026, with open weights following on July 26, K3 is a 2.8 trillion parameter Mixture-of-Experts model activating roughly 104B parameters per token, with a 1M-token context window and native vision - Kimi. Moonshot calls the architecture Stable LatentMoE, activating about 16 of 896 experts per token, and credits refinements including Kimi Delta Attention with an approximate 2.5x improvement in scaling efficiency over K2. On the Artificial Analysis Intelligence Index it scores 57, seven clear points above GLM-5.2 and, more symbolically, above the roughly 55.7 that Claude Opus 4.8 recorded as the world's best closed model in our July edition - Artificial Analysis.
The benchmark detail is worth reading precisely, because Moonshot's own framing is unusually honest. K3 posts 67.3 on DeepSWE v1.1 with the mini-SWE-agent harness and, per third-party review of the weights release, 81.2 on FrontierSWE and 91.2 on BrowseComp at maximum thinking effort, while Moonshot itself states the model still trails Claude Fable 5 and GPT-5.6 Sol on overall intelligence - Roo. In other words: the best open model is now competitive with the current closed frontier on specific agentic suites and behind it overall, exactly the shape the lag thesis predicts. Adoption was immediate. The Hugging Face repository logged 1,125,935 downloads in its first reporting month - Hugging Face. Our dedicated Kimi K3 benchmarks and cost guide covers the evaluation detail this section compresses.
Then comes the price, and it changes the story of open AI in a way benchmarks do not. K3's API costs $0.30 per million input tokens on a cache hit, $3.00 on a miss, and $15.00 per million output tokens - Kimi. Simon Willison, testing on launch day, called it the most expensive model a Chinese lab has ever shipped and measured a single SVG generation task consuming 13,241 reasoning tokens for about 25 cents, noting the model currently exposes only one reasoning effort, max - Simon Willison. For two years, "open" and "cheap" moved together so reliably that most buyers treat them as synonyms. K3 breaks the correlation: the open frontier now has a premium tier, priced within sight of closed flagships like Opus 5's $25 output rate. The 21% reduction in output tokens versus K2.6 softens the bill; it does not change the category.
The license moved in the same direction, and this is the fine print that matters more than any benchmark. K3 does not ship under K2's Modified MIT. It ships under a custom Kimi K3 license with two commercial gates: any model-as-a-service business earning over $20M per year from serving K3 needs a separate agreement with Moonshot, and any product exceeding 100M monthly active users or $20M per month in revenue must display "Kimi K3" prominently - Roo. The attribution clause is inherited from K2 and remains cosmetic for almost everyone. The MaaS clause is new, aimed squarely at inference providers, and it is the first time a leading open lab has claimed a revenue share from the serving layer. Section 12 places this in the broader license story; the short version is that Moonshot built the best open model in the world and priced both the tokens and the terms accordingly.
There is also a quieter shift hiding inside K3's headline specs: it is the clearest case yet of open weights you cannot practically run. The repository is public and the license permits self-hosting, but 2.8 trillion parameters means the raw weights occupy roughly 1.4 terabytes even at 4-bit quantization, before KV-cache for a million-token context, which puts genuine self-deployment beyond every organization that is not itself an inference provider. "Open" for K3 therefore means something different than it means for a 117B model that fits on one GPU: it means auditability, fine-tuning access for the few labs equipped to do it, provider competition on serving, and insurance against the hosted API disappearing, but not the ability to run it in your own rack. That is still substantially better than closed, and provider competition alone will discipline K3's hosted pricing over time. But buyers should notice that as open models grow into the multi-trillion class, the practical meaning of openness is migrating from "run it yourself" toward "the ecosystem can serve it, and you can leave," a distinction that matters most in exactly the compliance conversations where "open weights" gets used as a talisman.
So where does K3 actually belong in a deployment? It is the open model you route to when per-step intelligence decides the outcome: the hardest agentic coding tasks, screen-operating agents that need frontier judgment plus native vision in one loop, and evaluation harnesses where you want the strongest open baseline money can currently buy. It is emphatically not a volume workhorse, and treating it as one will produce closed-frontier bills without closed-frontier support contracts. That volume role belongs, still, to the lab that gave this article its name.
3. DeepSeek V4 Flash 0731: The Retirement, the Re-Post-Train, and Surge Pricing
DeepSeek had the most operationally eventful July of any lab, executing three moves in one week that every DeepSeek-dependent team felt directly. The first was the long-telegraphed guillotine: on July 24 at 15:59 UTC, the legacy deepseek-chat and deepseek-reasoner API names were retired with no grace period, and calls against them now simply error - Enterprise DNA. Both aliases mapped to deepseek-v4-flash, with the old chat behavior as its non-thinking mode and the old reasoner behavior as a thinking-mode request parameter. That last detail is the migration gotcha that caught real teams: reasoning is now a parameter, not a model name, so a naive find-and-replace to deepseek-v4-flash silently loses extended reasoning, while teams that assumed V4-Pro was the reasoner's natural successor discovered they had tripled their output bill, since Pro's $0.87 per million output tokens runs about 3.1x Flash's $0.28 - DeepSeek.
The second move landed on July 31: DeepSeek-V4-Flash-0731, an official release replacing the April checkpoints, which DeepSeek now frames as previews. The architecture is untouched, a 284B total / 13B active MoE with one shared and 256 routed experts per layer and a 1M-token window; the jump is post-training only, re-targeted at coding, agentic tool use, and long-horizon tasks - MarkTechPost. The results are startling for a post-train: the AA Intelligence Index score rose from the April preview's 40 to roughly 50, and Flash-0731 now beats the V4-Pro preview on every agentic benchmark DeepSeek publishes, posting 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE. Weights are MIT-licensed and ungated on Hugging Face, where the repo shows 433,284 downloads this month and self-hosting starts around 110GB at 3-bit quantization - Hugging Face. Simon Willison's assessment lands where our scorecard does: possibly the best value-per-intelligence model available from anyone - Simon Willison.
Crucially, the pricing did not move. Flash-0731 holds the launch rates that made V4 famous: $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit, $0.28 output, with Pro at $0.435 / $0.003625 / $0.87, both at 1M context and up to 384K output - DeepSeek. A ten-point intelligence jump at frozen prices is a straight capability dividend, and it is why Flash-0731 holds this guide's #1 slot: on the combined weight of intelligence, cost, MIT terms, and ecosystem, nothing else compounds as well. The same pricing page also notes the Responses API currently supports Flash only, with Pro support expected in early August, and that Flash-0731 ships adapted for Codex-style harnesses. V4-Pro remains, officially, a preview whose general availability "will follow." For the April-era context this builds on, our DeepSeek V4 launch guide still covers the family architecture in depth.
The third move is the one almost nobody in the current search results is covering properly: DeepSeek has announced that its API "will soon adopt a peak/off-peak pricing policy", with prices at 2x the regular rate during Beijing peak hours, 9:00-12:00 and 14:00-18:00 UTC+8, across all billing items, effective date to be announced - DeepSeek. This is the first surge pricing in the LLM API market, and it is worth thinking about from first principles rather than as a curiosity. Inference capacity is a fixed asset with time-varying demand, exactly the economics of electricity grids, and DeepSeek serves cheap enough that its constraint is capacity, not customers. Congestion pricing is the textbook response. For buyers the implications are concrete: Beijing's peak windows are roughly 01:00-04:00 and 06:00-10:00 UTC, which means European morning workloads overlap the second window while American working hours barely touch either. Batch and agent workloads that can time-shift will pay half what naive always-on deployments pay.
Because the retirement is the template for how open-model deprecations will run from now on, the operational lesson deserves spelling out. The migration was mechanically trivial for teams whose stack expressed "use extended reasoning" as a capability flag resolved to a model-plus-parameters combination at one configuration point: they changed one mapping and re-ran their evaluation suites. It was a genuine incident for teams that had scattered the string deepseek-reasoner across services, because the failure mode was not an error message but a silent capability downgrade: requests kept succeeding against deepseek-v4-flash in its default non-thinking mode, quality degraded on exactly the hard cases that justified the reasoner in the first place, and nothing in the logs said why. Silent degradation is worse than breakage, and it is the specific risk that model-name abstraction exists to prevent. The second-order lesson is about reading deprecation notices structurally: DeepSeek did not retire a capability, it retired a naming convention, folding two products into one model with a mode switch. Labs will keep doing this as reasoning modes proliferate, and every hardcoded model string in your codebase is a small bet that they will not.
Put the three moves together and the pattern is coherent. DeepSeek is behaving like an infrastructure operator: hard-cutting legacy endpoints, upgrading in place at fixed prices, and introducing load-based billing. That posture is precisely what makes it the default substrate for high-volume agent fleets, and precisely why we run the cost math against closed alternatives on Flash first. It also throws the intelligence crown's economics into relief: the gap between K3 at 57 and Flash-0731 at 50 is seven index points, and the gap between their output prices is 54x.
4. GLM-5.2: Still Z.ai's Latest, Still the Open Coding Benchmark
In a month when nearly every position changed, Z.ai's GLM-5.2 is notable for not changing at all, and for how well it held up anyway. Released June 13, it remains Z.ai's current model as of early August, scoring 51 on the AA Index, now second among open weights - Artificial Analysis. What July's arrivals clarified is GLM-5.2's actual specialty: agentic coding. In Tencent's own comparative benchmarking for the Hy3 launch, GLM-5.2 posted 84.2 on SWE-bench Verified and 46.2 on DeepSWE against Hy3's 78.0 and 28.0, numbers Tencent published while conceding the coding category to Z.ai - Digital Applied. When a competitor's launch material cedes a benchmark to you, that is about as credible as vendor-reported evidence gets.
The deployment picture is strong and, on one detail, corrects our own July edition: the Hugging Face repository for GLM-5.2 carries an MIT license tag, not the Apache 2.0 we previously reported for the GLM-5 line, and the model logged 2,234,662 downloads this month with the FP8 variant adding 2.7M more - Hugging Face. That download volume, second only to gpt-oss-120b among the models in this guide, reflects a real production install base rather than novelty pulls: GLM-5.2 has been the default open coding-agent engine for seven weeks, which in this market is an eternity of accumulated harness integrations and fine-tunes. Hosted access via OpenRouter currently lists $0.60 per million input and $1.50 output at a 50% promotional discount, on a 1M-token context - OpenRouter. Our GLM-5.2 practical guide and the companion benchmarks and cost breakdown go deeper on serving options.
What about GLM-5.3 or GLM-5.5? Treat everything you have read as rumor, explicitly. As of an August 4 check, Z.ai's documentation lists GLM-5.2 as current; GLM-5.3, GLM-5.5, and GLM-6 all circulate in community teases and a JPMorgan analyst projection pointing at an August window, with a "more than one trillion parameters" figure attached to none of them by Z.ai itself - Evolink. The July edition of this guide got burned stating leaderboard facts that expired in days; we are not going to compound that by publishing analyst forecasts as roadmap. If a successor ships, the model to beat is K3's 57, and Z.ai's release cadence (GLM-5 in February, 5.1 in April, 5.2 in June) makes an autumn attempt plausible. Until weights exist, GLM-5.2 is what you can build on, and it remains the best intelligence-per-dollar-per-license package in open AI's top tier: seven points below K3 on the index at a tenth of the output price, under cleaner terms.
5. MiniMax M3: Long-Context Value and the 2.7T Report
MiniMax M3 held its ground through July as the value anchor of the open frontier's second tier. Released at the end of May, M3 is a 428B total parameter MoE scoring 44 on the AA Index, with a genuine 1M-token context window and, unusually for this class, native multimodal input across text, image, and video - OpenRouter. Current hosted pricing on OpenRouter runs $0.24 per million input and $0.96 output under an active 60% promotional discount, with the platform noting that effective customer cost often lands 60-80% below list once prompt caching engages on repeated context. Even at undiscounted rates, M3 is the cheapest way to hold a million tokens of mixed-media context in front of frontier-adjacent intelligence.
The video-input capability deserves emphasis because it remains rare in open weights and it changes what counts as a "document." A model that accepts video natively can treat a recorded product demo, a screen capture of a failing workflow, or an hour of user-research footage as first-class context, without a transcription-and-caption pipeline flattening it to text first and discarding everything visual. Pipelines like that are not just lossy, they are where multimodal agent projects historically go to die, accumulating glue code, latency, and modality-mismatch bugs. Native ingestion collapses the pipeline into a prompt. Among this guide's ten models, only M3 and the K3-and-Inkling pair offer serious native multimodality, and M3 does it at a small fraction of K3's price, which is why it, not K3, is the default answer for mixed-media workloads that do not need frontier judgment.
That specific combination, long context plus video input plus low sustained cost, defines M3's production niche: research and document intelligence. An agent that ingests a full data room, three hours of recorded meetings, and every page it browses, and holds all of it in working memory rather than a retrieval index, is an M3-shaped workload. The intelligence ceiling is real (44 is a class below GLM-5.2's 51 and two below K3's 57), but research synthesis is dominated by context capacity and reading throughput, not by peak reasoning, which is why section 14's routing table sends exactly this workload here. The ecosystem is the honest weak point relative to DeepSeek: fewer providers, thinner fine-tuning tooling, and a smaller harness footprint, which is what caps its ecosystem score at 6 in our table.
The forward-looking item deserves the same rumor discipline we applied to GLM. The Information reports that MiniMax is developing an internal model called M3 Pro at 2.7 trillion parameters, with a possible open-source release as early as Q3 - Techmeme. That is a report about intent, not a release, and it belongs in your planning as an option, not a dependency. Its significance, if it lands, is scale signaling: a 2.7T open release would put MiniMax in the same weight class as Moonshot's K3 and confirm that multiple Chinese labs can now fund 3T-class training runs. For now, M3 is the pick when your constraint is tokens held per dollar, and the strongest evidence that the open frontier's middle tier is where price competition is fiercest.
6. Tencent Hy3: The Apache 2.0 Search Specialist Almost Everyone Missed
Hy3 is the model this guide owes an apology to, having shipped on July 6, 2026, two days before our July edition failed to mention it. It is also, on the fine print, one of the most enterprise-friendly releases of the year: Tencent published the 295B total / 21B active MoE under Apache 2.0 with no geographic or field-of-use restrictions, explicitly reversing the April preview's exclusion of the EU, UK, and South Korea - Digital Applied. The architecture routes top-8 of 192 experts across an 80-layer transformer, carries a separate 3.8B multi-token-prediction layer for speculative decoding, and holds a 256K native context. Deployment is the standout: FP8 quantization brings the whole model under 300GB, single-node territory, roughly 2.5x lighter than GLM-5.2's memory footprint.
The benchmark profile is sharply specialized, and Tencent's launch material is refreshingly explicit about the trade. Hy3 posts vendor-reported scores of 84.2 on BrowseComp (agentic search), 91.0 on DeepSearchQA, 79.1 on MCP-Atlas for tool orchestration, and 73.4 on AA-LCR long-context retrieval, while conceding agentic coding to GLM-5.2 outright (78.0 vs 84.2 SWE-bench Verified, 28.0 vs 46.2 DeepSWE). All of those figures are the vendor's own, with no independent verification published yet, and should be read with that discount. Independent color exists at the edges: Gigazine's coverage notes Hy3 outperformed GPT-5.5 on FrontierScience-Olympiad and completed a document-processing test using 47.4% fewer tokens than GLM-5.2 - Gigazine. Token frugality is an underrated production trait: an agent that reads and reasons in fewer tokens is cheaper at any price per token.
Adoption is early but the distribution was unusually well executed for a first open release from Tencent's Hunyuan team: day-one availability across OpenRouter, Cline, OpenCode, and Cherry Studio, with vLLM and SGLang serving support, though the Hugging Face repo shows a modest 23,303 downloads this month - Hugging Face. The strategic read: Tencent looked at an open field where DeepSeek owns volume economics and Z.ai owns coding, and picked the unclaimed specialty, agentic search and retrieval at half the deployed memory. For teams building research agents that need self-hosted weights inside a clean license, Hy3 is arguably the most interesting new option of the summer, and its WeChat-scale parent gives it staying power that most 23K-download models lack. It is the model on this list we most expect to rank higher in the next edition.
7. Inkling and the American Rearguard
The American open-weights story got its first genuinely good chapter in over a year on July 15, when Thinking Machines Lab released Inkling: a 975B total / 41B active MoE, pretrained on 45 trillion tokens of text, images, audio, and video, with weights supporting up to 1M tokens of context - Thinking Machines. It debuts at 41 on the AA Index, which Artificial Analysis itself headlined as the new leading US open-weights model, ahead of Nemotron 3 Ultra at 38, Gemma 4 31B at 29, and gpt-oss-120b at 24 - Artificial Analysis. The weights ship under Apache 2.0 on Hugging Face, where the main repository has logged 69K downloads in three weeks with quantized variants adding several hundred thousand more - Hugging Face.
The lab behind it matters to the reading. Thinking Machines chose to enter the open-weights market at the top of the American field on its first release, with a training recipe it published in unusual detail: 45T multimodal tokens, synthetic data bootstrapped from existing open models including Kimi K2.5, and a reinforcement-learning phase spanning more than 30 million rollouts - Thinking Machines. That recipe transparency is itself a competitive posture. Where most labs publish weights and a benchmark table, Inkling's documentation reads like an argument that the American open ecosystem's deficit was never talent or capital but willingness to ship weights at all, and that a well-funded newcomer can reach the front of the US pack in one release by simply deciding to.
Two things make Inkling more interesting than its index score. First, it is natively multimodal in a way no other open model matches, trained from scratch on interleaved text, image patches, and audio spectrograms rather than bolting encoders onto a text model; it posts 73.5 on MMMU Pro and 91.4 on VoiceBench alongside a solid 77.6 SWE-bench Verified. Audio-native open weights are a category of one right now, and voice agents are among the fastest-growing production workloads. Second, the lab's distribution strategy is fine-tuning-first: Inkling is served through the Tinker platform at a 50% introductory discount with 64K and 256K context options, positioning the model as a substrate you adapt rather than a commodity endpoint you call. A preview of Inkling-Small, a 276B total / 12B active sibling, shipped on July 27 with weights already public. The bet is that American enterprises want a domestic, clean-licensed, adaptable foundation more than they want three extra index points, and it is a credible bet.
The rest of the rearguard sorts into supporting roles. NVIDIA's Nemotron 3 Ultra 550B held the American open crown for barely a month of this guide's lifetime and now scores 38 on Index v4.1, a correction from the June-snapshot 48 our July edition cited; it remains a serious model with obvious staying power, since NVIDIA funds it to sell hardware, not tokens - Artificial Analysis. Gemma 4 continues to own the edge: its E2B and E4B variants run fully offline on phones, Raspberry Pi, and Jetson-class boards with audio and visual understanding, 140-language support, and native function calling, while the 12B-31B tier targets consumer GPUs - Google DeepMind. The 31B's index score of 29 is beside the point; nothing else turns a handset into a competent agent host. Mistral, meanwhile, has drifted out of the top table entirely: its best-placed current model on the open index, Medium 3.5, scores 30, and Europe's sovereignty argument increasingly rests on procurement policy rather than capability. The pattern across all of them is consistent: American and European open models now win at the edges (on-device, single-GPU, audio-native, fine-tuning substrates) while the intelligence frontier lives in Beijing, Hangzhou, and Shenzhen.
8. gpt-oss-120b Turns One: The Most-Run Open Model Has No Successor
This week marks one year since OpenAI released gpt-oss-120b in early August 2025, and the anniversary data point is remarkable: the most-downloaded large open model in the world is a year-old artifact whose maker has never updated it. The Hugging Face repo logged 4,221,708 downloads this month, nearly double GLM-5.2 and almost four times Kimi K3, for a model scoring 24 on the current index, dead last among this guide's ten - Hugging Face. No successor has shipped or been announced; the only derivatives in a year have been safety-focused variants. A model two generations off the intelligence pace is the most-run open model on Earth, and explaining that gap between leaderboard and reality is the whole value of an ecosystem criterion.
The explanation has three parts, all structural. Hardware fit: gpt-oss-120b is a 117B total / 5.1B active MoE post-trained with MXFP4 quantization specifically so it runs on a single 80GB GPU, an H100 or MI300X, with the 20B sibling fitting in 16GB; per the model card, the 120B can even be fine-tuned on a single H100 node - Hugging Face. One card, no cluster, no serving team. License and brand: Apache 2.0 from OpenAI clears compliance reviews that Beijing-lab weights still trigger in many Western enterprises, whatever the technical merits of that hesitation. Compounding inertia: twelve months of tutorials, quantizations, and deployment recipes make it the path of least resistance for every team entering local AI, a dynamic our open-source personal AI guide leans on for exactly that reason.
The chart's shape carries a lesson the ranking table cannot: deployment share and capability share are different competitions. GLM-5.2's strong second place shows a current-generation model earning production install base; K3's first-month surge shows what frontier status buys; but the year-old single-GPU model still doubles them both, because the largest population of open-model users is not choosing among frontiers, it is running whatever fits the hardware it already owns. It is worth being precise about what a download count does and does not measure, since this guide weights it at a fifth of the final score. Downloads measure pulls of the artifact, which conflates production deployments, quantization pipelines, academic experiments, and CI jobs re-fetching weights; they say nothing about tokens served, and hosted-API usage bypasses them entirely, which understates API-first models like K3 relative to self-host-friendly ones. What downloads do reliably proxy is the size of the surrounding ecosystem: every pull is a person or pipeline that may produce a quantization, a fine-tune, a bug report, or a deployment guide, and that surrounding material is what makes a model cheap to adopt eighteen months later. By that reading, the chart is not saying gpt-oss-120b is the best model. It is saying it has the deepest moat of accumulated operational knowledge, which is a real asset the index cannot see. The strategic question for OpenAI's open line is whether the anniversary passes without a successor while Hy3's sub-300GB footprint and Inkling-Small's 12B-active design close in on the same "fits on what you have" niche from above. Until one ships, gpt-oss-120b remains the correct answer to a narrow, extremely common question, and a 24-index-point answer to every other one.
9. MiMo-V2.5-Pro: The Quiet Trillion
Xiaomi's MiMo-V2.5-Pro spent July doing something rare in this market: nothing. No new checkpoint, no repricing, no license drama, and yet it holds fifth place on the open intelligence ladder at 42, ahead of Inkling and Nemotron - Artificial Analysis. The specs remain among the best value propositions in the trillion class: 1.02 trillion total / 42B active parameters, a full 1M-token context window, and a clean MIT license, with the Hugging Face repo showing 53,794 downloads this month - Hugging Face. For a company whose core business is phones and appliances, shipping and quietly maintaining a top-five open model is the whole point: vertical integration of intelligence from handset to cloud, with the open release serving as ecosystem bait.
MiMo's position in a portfolio is the value hedge. It offers most of the 44-class capability at MIT terms with a 1M window, from a lab whose incentives (selling devices, not tokens) make sudden license enclosure unlikely. The honest caveats are ecosystem thinness, two orders of magnitude fewer downloads than the leaders and a correspondingly small harness and fine-tune community, and Xiaomi's demonstrated willingness to deprecate commercial API generations quickly, which makes the downloadable weights, not the hosted endpoint, the thing you should depend on. As a self-hosted secondary model for teams already committed to running their own serving stack, it is quietly one of the best deals in open AI, and its month of stability amid July's chaos is itself a small argument in its favor.
10. The Half-Exits: Qwen Goes API-First, Meta Stays Closed
The July edition of this guide told a clean two-pole story: Meta exited open weights, Chinese labs doubled down. August requires a third category, the half-exit, and Alibaba's Qwen is its defining case. The timeline, per the consolidated release record: Qwen3.5 shipped in February 2026 as an open Apache 2.0 family alongside a proprietary Plus variant; Qwen3.6 followed in April, again open, and remains the lab's latest open-weights release; then the line split. Qwen3.7 Max arrived in May as an API-only model, Qwen3.7 Plus followed closed in June, and the newly announced Qwen3.8-Max, at a reported 2.4 trillion parameters, launched with open weights promised but not yet delivered - Wikipedia. Meanwhile Qwen has fallen out of the Artificial Analysis open top ten entirely, because the models Alibaba still opens are no longer the models Alibaba considers its best.
This matters beyond one lab's product strategy, because Qwen was the open ecosystem's workhorse substrate: the deepest fine-tune tree, the widest size ladder, the default multilingual pick. A world where Qwen's frontier is closed and its open releases trail by a generation is a world where "everything Qwen is open," the assumption baked into hundreds of deployment guides including our own July edition, is simply false. The first-principles read is that Alibaba is running the standard commoditization playbook in reverse: open weights built the developer base, and the flagship tier is now being monetized through Alibaba Cloud, with each open release repositioned as the on-ramp rather than the product. Whether the promised Qwen3.8 weights actually land will tell you which way the lab has finally decided; hold that promise to the same standard we hold GLM-5.5 rumors.
Meta, for its part, stayed exactly where it landed in April: closed, but commercially warmer. Muse Spark 1.1 shipped on July 9 with API access for developers in public preview, Zuckerberg promising "aggressive pricing" against competitors, though per-token rates were not disclosed in launch coverage - Fortune. The Llama line remains ended, making Llama 4 Scout's 10M-token context window, still the longest of any open model, a legacy artifact with no successor - Meta. Our Muse Spark guide covers the closed model itself; for this guide's purposes, the lesson Meta taught in April now has a corollary from Alibaba in May: open weights are a strategy, not an identity, and the strategy can be abandoned in halves as well as wholes. Plan on the weights you can download today, not on any lab's stated philosophy.
11. The Frontier Gap: The Lag Thesis Just Got Its Proof
Every edition of this guide has carried the same falsifiable claim: open weights trail the closed frontier by a stable 3-6 months, not by a permanent class. July 2026 delivered the cleanest test that thesis will ever get, because both frontiers moved in the same week. On July 16, Kimi K3 reached 57 on the Artificial Analysis Index; on July 24, Claude Opus 5 reset the closed peak to 61, with Claude Fable 5 at 60 and GPT-5.6 Sol at 59 - Artificial Analysis. The July edition of this guide recorded the then-champion Opus 4.8 at roughly 55.7. K3 did not merely close the gap to the previous closed peak; it passed it, within a single product generation. An open model exceeding the prior generation's best closed model is precisely what a 3-6 month lag predicts and what a widening-moat story forbids.
Read the ladder structurally rather than as a scoreboard. The closed trio occupies a four-point band at the top; K3 sits alone in the gap; then a six-point drop to the open pack at 51-50, which itself sits where the closed frontier stood a generation ago. The commercial meaning of the 4-point top gap depends entirely on workload type. For tasks at the absolute capability edge, the difference between 61 and 57 can be the difference between working and not working, and the closed premium buys real outcomes; our comparisons of Opus 5 against Opus 4.8 and GPT-5.6 against Opus 5 map that edge in detail. For the enormous majority of production work that was already within open-model capability months ago, the premium buys headroom you never touch.
Why does the lag stay stable instead of widening, when closed labs outspend most open ones? Because the ingredients of frontier capability diffuse faster than they can be hoarded. Architectural ideas circulate through papers, weights, and hiring within months; efficiency techniques born under export-control constraints (aggressive sparsity, low-precision training, attention compression) spread globally within quarters of publication; and the synthetic-data flywheel now runs in both directions, with Inkling openly bootstrapping from Kimi K2.5's outputs the way closed labs distill from their own flagships - Thinking Machines. What does not diffuse is compute and integration depth, which is why the gap exists at all and why closed labs increasingly market reliability engineering, tooling, and enterprise certification rather than raw index points. If the gap were protected by secret architecture, it should widen with every closed release. Three generations of evidence, and now a direct crossing, say it is protected only by a head start.
The strategic corollary has a new wrinkle this cycle, though: the open frontier crossing the old closed frontier does not automatically bring the old economics with it. K3 charges $15 per million output tokens, which means "open catches closed" now describes capability more than it describes price at the very top. The reliable price collapse still happens one tier down, where Flash-0731 delivers 50 index points at $0.28 output. The gap that should drive your architecture is therefore not open-versus-closed but frontier-versus-sufficient: identify which of your workloads genuinely require the top band, pay for it there (closed or K3), and route everything else to the tier where intelligence is a commodity. That routing decision is exactly what section 14 operationalizes.
12. Licenses in August 2026: The Fine Print Is Coming Back
The July edition described current open-model licensing as the cleanest in the field's history, a top tier running almost entirely on MIT and Apache 2.0. One month later that statement needs an asterisk, because the new intelligence leader introduced the most commercially significant license terms the open frontier has seen. The direction of travel matters more than any single clause: as open models became genuinely frontier-class, the licenses started acquiring teeth, and diligence habits that atrophied during the clean-license era need to come back.
The first-principles logic behind the shift is straightforward once stated. A lab open-sources aggressively when the weights are a distribution instrument: free capability buys developers, mindshare, and a seat at the standard-setting table, and the foregone revenue is small because the model is not the best available anyway. The moment a lab's open model becomes the best available, the calculus inverts, because now the weights are the product, and every inference provider serving them is monetizing an asset the lab paid nine figures to train. Moonshot's response was not to close K3 but to meter the serving layer, which is a genuinely new position on the open-closed spectrum: open to users, licensed to resellers. Expect the pattern to recur whenever an open lab holds the frontier, and read every future "open" release by asking who, exactly, is being asked to pay.
| License | Models (August 2026) | Commercial use | The fine print |
|---|---|---|---|
| MIT | DeepSeek V4 Flash 0731, GLM-5.2, MiMo-V2.5-Pro | Unrestricted | None: use, modify, resell, no attribution UI required |
| Apache 2.0 | Hy3, Inkling, gpt-oss-120b | Unrestricted | Patent grant + NOTICE file; Hy3 notably dropped its April geo-restrictions |
| Kimi K3 License | Kimi K3 | Gated at scale | MaaS revenue over $20M/year needs a Moonshot agreement; display "Kimi K3" above 100M MAU or $20M/month |
| Custom community | Gemma 4, Llama 4 (legacy) | Allowed with conditions | Use policies and redistribution terms; legal must actually read |
The Kimi K3 license deserves precision because it is genuinely novel. K2 shipped under Modified MIT with a single cosmetic attribution clause at consumer mega-scale. K3 keeps that clause and adds a structurally different one: any model-as-a-service business earning more than $20 million a year from serving K3 must negotiate a separate commercial agreement with Moonshot - Roo. That is not branding; it is a claim on the serving layer's revenue, aimed at the OpenRouters and Fireworks of the world, and it makes K3 the first leading open model whose economics are partially enclosed by design. For a team running K3 on its own infrastructure for its own product, neither clause is likely to bind. For anyone reselling inference, the calculus changed.
Two July events cut the other way, and the balance is worth stating fairly. Tencent released Hy3 under Apache 2.0 with zero carve-outs, explicitly reversing its April preview's exclusion of the EU, UK, and South Korea, proof that license terms move toward openness under competitive pressure as well as away from it - Digital Applied. And Thinking Machines put a 975B-parameter American flagship under plain Apache 2.0. The practical hierarchy for buyers is unchanged in shape: MIT and Apache 2.0 first (and note from our own errata that GLM-5.2 is MIT per its Hugging Face tag, which we previously misreported), custom-but-cosmetic second, revenue-gated licenses third with counsel involved, and always check the specific checkpoint's license page rather than the family's reputation. The new habit August adds: model the license against your five-year revenue plan, not your current size, because the K3 clause is the kind of term that is irrelevant until the day it is the only thing that matters in your acquisition diligence.
13. Pricing Economics: Cache Hits, Surge Hours, and the $15 Output Token
Open-model pricing in August 2026 spans two orders of magnitude, and the spread is the story. At the floor, DeepSeek V4 Flash 0731 charges $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit, and $0.28 for output - DeepSeek. At the ceiling, Kimi K3 charges $3.00 per million input on a miss and $15.00 for output, within sight of Claude Opus 5's closed-frontier $5 and $25 - Anthropic. Between them sit GLM-5.2 at $0.60/$1.50 and MiniMax M3 at $0.24/$0.96 on current promotional rates. The era in which "open" reliably meant "two orders of magnitude cheaper than closed" ended on July 16; open pricing is now a spectrum you must navigate, not a discount you can assume.
The metric that actually decides deployments is cost per completed task, and cache mechanics dominate it for agents. An autonomous agent re-reads its context on every step: a 20-step session over a 200K-token codebase processes roughly 4M input tokens of which the overwhelming majority are repeated prefix. At Flash's cache-hit rate those re-reads cost about a penny in total; the same repetition at K3's $0.30 cache-hit rate costs around a dollar; and output tokens, which caching never discounts, then dominate the K3 bill, since a verbose reasoning model at $15 per million output can spend a quarter on a single SVG, as Simon Willison measured on launch day - Simon Willison. OpenRouter's own serving data reinforces how central caching has become, noting effective customer prices routinely land 60-80% below list on repeated-context workloads - OpenRouter. The full framework for this arithmetic, including failure-rate adjustment (a cheap model that fails twice is not cheap), lives in our true cost of LLM inference guide and the companion cost-cutting playbook.
A worked example makes the spread concrete, using only prices verified above and round arithmetic. Take a research-agent task: 30 steps over a session that accumulates 150K tokens of context, re-read each step, producing 20K output tokens total. Input volume is roughly 4.5M tokens, of which perhaps 4.3M are cached prefix after the first pass. On Flash-0731, that is about $0.03 of cache-miss input, a penny of cache-hit input, and $0.006 of output: call it four cents per task. On GLM-5.2 at list-ish rates the same shape lands around $0.60. On K3 it is roughly $0.65 of cache-miss input, $1.29 of cache-hit re-reads, and $0.30 of output before accounting for K3's verbose max-effort reasoning, so $2 to $3 per task in practice. On Opus 5 at $5/$25 flat, the session costs over $20 unless the closed provider's own caching discounts apply. Now multiply by ten thousand tasks a month: $400 versus $6,000 versus $25,000 versus $200,000. The ranking of models by monthly bill is not the ranking by intelligence, and the delta between adjacent tiers funds an engineer. This is why the correct unit of decision is the task profile, not the model: the same arithmetic with an output-heavy profile (code generation, long drafting) punishes K3 and Opus far harder, while a vision-judgment profile with tiny outputs closes the gap dramatically.
The newest variable is time. DeepSeek's announced peak/off-peak policy will double all billing items during Beijing business hours (9:00-12:00 and 14:00-18:00 UTC+8) once it takes effect, making it the first LLM API priced like an electricity grid - DeepSeek. For interactive products serving live users, surge hours are simply a cost you absorb or route around. For the growing share of agent work that is batchable (overnight research runs, scheduled report generation, bulk data processing), time-shifting into off-peak windows is a free 50% discount on the cheapest frontier-adjacent tokens in the market. Pricing strategy, cache strategy, and scheduling strategy have merged into one discipline, and the teams that treat them as one line item labeled "API costs" are leaving half the open-model advantage on the table.
14. What an Agent Platform Actually Routes Where
Everything above is what the models are. This section is what we do about it, because O-mega runs production agent workforces and every July release either did or did not change our routing table with real money attached. This is the section an aggregator cannot write, and we will be specific enough to be falsifiable, including about what we do not use. The framing that matters: in an agent platform, "which model is best" is a malformed question. Agents are roles, roles have cost-of-error and token-volume profiles, and the routing table maps profiles to models, a discipline we covered from the buyer's side in our best LLM for AI agents ranking.
High-volume worker agents (research sweeps, data extraction, bulk drafting, sub-agent swarms) route to DeepSeek V4 Flash 0731, and July strengthened that call twice: the 0731 re-post-train raised per-step reliability at zero price change, and cache-hit economics remain unmatched for the re-read-heavy loops swarms generate. Long-document research agents, the ones holding entire data rooms or hours of transcript in context, route to MiniMax M3 for the 1M multimodal window at bottom-tier prices. Agentic coding roles route to GLM-5.2, whose SWE-bench Verified and DeepSWE leads among open weights match exactly the failure mode that costs us most (a wrong patch is expensive; a slow one is not), with the calibration detail that we treat closed models as the ceiling for the hardest tickets, per our open-source AI coders ranking. Screen-operating browser agents are the one role where we route selectively to Kimi K3: native vision plus frontier judgment in a single loop genuinely reduces multi-step navigation failures, and because browser sessions are output-light relative to research or coding sessions, the $15 output rate hurts less where K3 helps most.
Two more roles complete the table. Private and on-premise deployments, the workloads where customer data may not leave controlled infrastructure, route to gpt-oss-120b on single-GPU hardware today, with Hy3 under FP8 as the candidate upgrade for teams that can field a sub-300GB single node and want materially more capability inside the same Apache 2.0 comfort zone; quantization quality is the gating question there, and our TurboQuant compression guide covers how modern quantization trades footprint against fidelity. And drafting and content roles, which are output-heavy by definition, stay pinned to the cheap-output tier (Flash-0731, M3) no matter how tempting frontier quality looks, because output tokens are the one cost that caching never rescues and drafting is the profile that generates the most of them. The general rule underneath all of these assignments: match the model to the role's dominant cost term, whether that is input re-reads, output volume, error remediation, or compliance constraints, and the model choice usually makes itself.
Three verdicts on the other side of the ledger, stated plainly. We do not route production traffic to Nemotron 3 Ultra: at 38 on the current index with no cost advantage we can verify against the Chinese pack, it wins no role in our table, and NVIDIA-stack alignment is not a workload. We do not yet route to Inkling or Hy3: Inkling because three weeks is too young for load-bearing roles (its audio-native design has us running it in evaluation for voice agents now), and Hy3 because its headline search scores are still vendor-reported; both are the likeliest promotions in our next revision, which is a different thing from a promotion today. And we do not use K3 as a general workhorse despite it being the smartest open model ever shipped, for the arithmetic reason section 13 established: at $15 per million output tokens, routing bulk work to K3 recreates the closed-frontier bill that open routing exists to avoid.
Two portfolio-level habits make the table durable rather than heroic. First, the eval set is the asset and models are components: every role has a golden-task suite of real workload samples, and July's churn (a retirement, a re-post-train, a new frontier) was absorbed by re-running suites, not by debates. Teams that validated by vibes in July are still arguing; teams with suites migrated in days. Second, abstract the model name and the reasoning mode: the deepseek-reasoner retirement burned exactly the teams that had hardcoded a model string where a capability flag belonged, since reasoning is now a request parameter on deepseek-v4-flash. For readers building this stack themselves, our primer on making LLMs autonomous covers the scaffolding, and for readers who would rather have the routing treadmill absorbed for them, that is the operating layer O-mega sells; both paths get cheaper every quarter the open frontier advances.
15. Conclusion: Choosing Your Model in August 2026
The durable facts of August 2026, separated from the fast-rotting ones. Durable: an open model has now passed the previous generation's closed champion, confirming the 3-6 month lag as the planning assumption; the open frontier is no longer uniformly cheap, so open-versus-closed has been replaced by frontier-versus-sufficient as the decision that matters; licenses at the top are re-acquiring commercial teeth; and cache, surge, and scheduling mechanics now move real budgets more than list prices do. Fast-rotting: every version number and index score in this guide, which is why each one carries a source you can re-check in one click, and why this edition begins with a list of its predecessor's failures rather than a claim of finality.
In prose: pick DeepSeek V4 Flash 0731 as the default, because intelligence, price, MIT terms, and ecosystem compound there better than anywhere else, and July's upgrade made the default ten index points smarter for free. Pick GLM-5.2 when agentic coding is the job. Pick Kimi K3 when a task is worth frontier intelligence and you have priced its output tokens honestly. Pick MiniMax M3 for long-document and mixed-media research, Hy3 if you are building search-heavy agents on self-hosted Apache 2.0 weights and can wait for independent benchmark confirmation, Inkling if you are betting on audio-native agents or American-substrate fine-tuning, and gpt-oss-120b or Gemma 4 when the constraint is hardware you already own. Then schedule the re-check: this market rewrote its own top three inside 25 days, and the Artificial Analysis open-weights index will tell you within a minute whether this guide's photograph has aged.
This guide was researched and written by Yuma Heymans (@yumahey), founder of O-mega and co-founder of HeroHunt.ai, whose production routing tables had to absorb every release documented above, which is a more motivating fact-checking regime than any editorial calendar.
This edition reflects the open-weight model landscape as of August 5, 2026, with every benchmark score, price, and download count verified against the linked primary sources on that date. This market invalidated our previous edition in under three weeks; assume it will try to do the same to this one, and re-verify against the linked sources before committing budget.