The tracked ranking of agent skills: what to install, what just overtook what, and what the leaderboard's own numbers cannot tell you.
grill-me, a skill whose entire job is interrogating you before your agent writes a line of code, just overtook Anthropic's frontend-design as the most-installed non-bundled skill in the world. It sits at 756.3K installs against frontend-design's 742.3K on the live leaderboard - skills.sh. Four weeks ago that gap ran the other way by more than 150K installs. In the same four weeks, Matt Pocock's behavioral skills went from four entries in the global top 15 to five of the top nine, ByteDance's Lark enterprise skills entered the top 10 from nowhere, and the registry grew from roughly 670,000 skills to 1,131,746.
This guide is a ranking, but it is deliberately more than a snapshot of a leaderboard, because we have learned the hard way that snapshots of this leaderboard have a half-life of about four weeks. We published a ranking in January built on marketplace access counts; the ground truth moved under it. We rebuilt it in July on verified install data; the order rotated again within a month. What we hold now that nobody else publishing in this space holds is three timestamped observations of the same ecosystem: January's access-count era, July 8 install data, and a full re-verification of every number against live sources on August 5, 2026. That longitudinal record lets this guide show you something the leaderboard itself does not display: install velocity, which skills are accelerating, which are stalling, and which quietly fell out of the top 15 entirely.
So here is the deal this page makes with you: the top 10 ranked with live data, first, with scores you can argue with. Then the four-week changelog showing exactly what moved. Then what skills are and how to install them if you are new, profiles of every ranked skill, the security wave you cannot ignore, and the September 2026 model lineup actually running these skills. We maintain the long-tail companion in our top 100 Claude Code skills ranking if you want depth beyond the head of the distribution.
Contents
- The Top 10, Ranked
- The Four-Week Changelog: What Moved and How Fast
- Agent Skills in Brief: The Format and the Open Standard
- grill-me: The New Number One
- frontend-design: The Design-Taste Transplant
- agent-browser: Hands for the Web
- The Pocock Behavioral Stack: tdd, grill-with-docs, improve-codebase-architecture
- vercel-react-best-practices: Framework Judgment, Inlined
- Superpowers: The Methodology
- web-design-guidelines: The Hundred-Point Audit
- find-skills and the Bundler Problem
- The Enterprise Land-Grab: Microsoft, and Now ByteDance
- What Fell Out: remotion, Ralph, and the January Casualties
- Security: The Malicious-Skill Supply Chain
- Skills Beyond Coding: Cowork and Document Work
- Skills vs MCP vs Plugins vs AGENTS.md
- How to Ship and Measure Your Own Skill
- The Models Running Your Skills in September 2026
- Outlook and Decision Framework
1. The Top 10, Ranked
Every ranking encodes an editorial stance, and ours is stated up front: raw install counts are partly a distribution artifact, and a ranking that repeats them uncritically is repeating the artifact. The two loudest examples on the current leaderboard are find-skills at 2.8M installs and setup-matt-pocock-skills at 544.6K - skills.sh. Both are bundlers: skills whose function is installing other skills, which means their counts measure onboarding funnels, not adoption on merit. Our scoring discounts that, which no scraper-built listicle does, because a scraper cannot tell plumbing from product.
Four criteria, weights summing to 100%. Practical impact (30%) measures how much the skill changes the quality or scope of what an agent produces, because a skill that transforms output beats a popular skill that saves a minute. Organic adoption (25%) is install rank adjusted for bundler and default-install distortion. Trust and maintenance (25%) weighs publisher identity, maintenance activity, and auditability, which the supply chain attacks of this year (section 14) turned from nicety into necessity. Momentum (20%) is the new criterion this refresh adds: four-week install velocity, computed from our July 8 observations against the August 5 live leaderboard, a delta the leaderboard itself does not show. All install figures were read from the skills.sh leaderboard and all star counts from the GitHub API on August 5, 2026.
| # | Skill | What It Does | Impact (30%) | Adoption (25%) | Trust (25%) | Momentum (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | grill-me | Interrogates you before building, 756.3K installs | 8 - kills wrong-direction builds early | 9 - top organic skill globally | 9 - Pocock repo, 203.9K stars | 10 - +58% in four weeks | 8.9 |
| 2 | frontend-design | Anthropic's design-quality skill, 742.3K installs | 9 - most visible output transform | 9 - #2 organic, was #1 | 10 - Anthropic official repo | 6 - +16%, half the leaders' pace | 8.7 |
| 3 | agent-browser | Ref-based browser control, 629.0K installs | 9 - closes the verification loop | 8 - #5 globally | 9 - Vercel Labs, 39.9K stars | 8 - +20%, strong absolute adds | 8.6 |
| 4 | tdd | Test-first sequencing enforced, 595.4K installs | 8 - converts bluffing into feedback | 7 - #8 globally | 9 - Pocock repo, active | 10 - +59% in four weeks | 8.4 |
| 5 | grill-with-docs | grill-me against your own docs, 641.9K installs | 7 - overlaps grill-me's value | 8 - #4 globally | 9 - Pocock repo, active | 10 - +61%, fastest tracked | 8.4 |
| 6 | improve-codebase-architecture | Structural refactor advisor, 617.5K installs | 7 - advisory, payoff varies by repo | 8 - #6 globally | 9 - Pocock repo, active | 10 - +57% in four weeks | 8.4 |
| 7 | vercel-react-best-practices | React/Next.js perf rules, 606.7K installs | 9 - catches real regressions | 8 - #7 globally | 9 - Vercel official, maintained | 6 - +14%, slowest of top cohort | 8.2 |
| 8 | Superpowers | Full dev methodology, 266.8K GitHub stars | 10 - changes the whole dev loop | 8 - most-starred skills repo | 7 - solo-led, deep behavior changes | 6 - +7% stars, cooling | 8.0 |
| 9 | web-design-guidelines | 100+ a11y/UX audit rules, 516.5K installs | 8 - systematic a11y coverage | 7 - #32, was #6 in July | 9 - Vercel official | 6 - +16%, below cohort | 7.6 |
| 10 | find-skills | Registry discovery meta-skill, 2.8M installs | 7 - plumbing, not output | 7 - #1 raw, bundler-discounted | 9 - Vercel Labs maintained | 7 - +17% on a huge base | 7.5 |
Read the table against the leaderboard and the editorial stance becomes concrete. find-skills is #1 by raw installs and #10 here, because being the default first install is a distribution fact, not a quality signal. grill-me is #2 raw and #1 here, because its installs were earned one convinced user at a time and its four-week growth rate is the strongest evidence on the board that the convincing is still happening. Two skills that would have made a raw-install top 10, lark-approval (#10 raw) and setup-matt-pocock-skills (#9 raw), are excluded: the first is scoped to ByteDance's Feishu suite and covered in section 12, the second is a pure bundler. microsoft-foundry, at #36 raw, just misses the cut at 7.3 and is also covered in section 12.
The compression at the top of this chart is the story. In July, frontend-design led the organic field by more than 100K installs. Today the top seven organic skills sit within a 161K-install band of each other, four of them from a single author's repository. When a ranking compresses like that, order stops being the interesting fact and velocity starts. That is what the next section measures.
2. The Four-Week Changelog: What Moved and How Fast
This section is the reason this guide exists in its current form. A ranking you can scrape is a commodity; a ranking with a memory is not. Because we recorded the full leaderboard on July 8 and re-verified it against the live site on August 5, we can publish the install deltas that the leaderboard, which only shows current totals and a small weekly sparkline, does not surface. Four weeks is also, empirically, about one full rotation of this ecosystem's attention, which makes the window meaningful rather than arbitrary.
The headline movement: grill-me grew 58% in four weeks, from 479.9K to 756.3K installs, and overtook frontend-design, which grew 16% over the same window - skills.sh. That is not an isolated spike. The whole Pocock cluster moved together: grill-with-docs +61% (398.1K to 641.9K), tdd +59% (375.5K to 595.4K), improve-codebase-architecture +57% (393.5K to 617.5K). Meanwhile the official vendor cohort grew at a much flatter 14-20%: agent-browser +20%, frontend-design and web-design-guidelines +16%, vercel-react-best-practices +14%, microsoft-foundry +15%. Two growth regimes, cleanly separated, visible only if you kept the July numbers.
Beyond the top cohort, three changelog entries carry real information. First, remotion-best-practices fell out of the top 15: it ranked #8 globally in our July data at 414.5K installs and now sits at #55 with 464.2K - skills.sh/remotion-dev/skills. It still grew 12%, but the field grew past it, and that demotion of the only serious video-creation skill says the leaderboard's center of gravity is consolidating on coding discipline and enterprise procedure, not creative output. Section 13 covers what that means if you build video with agents.
Second, ByteDance arrived. Skills published under open.feishu.cn, the developer platform of ByteDance's Lark/Feishu office suite, now hold 22 of the global top 31 slots: lark-approval at #10 (529.4K installs) with the next 21 leaderboard positions stacked behind it, plus lark-okr (#34, 509.2K), lark-markdown, lark-vc-agent, and lark-apps further down - skills.sh. None of these existed in the top ranks four weeks ago, and no English-language ranking we can find has covered them at all. This is the first non-Western enterprise vendor to storm the leaderboard, and section 12 argues it is the strongest evidence yet for the enterprise land-grab thesis. Third, the registry itself grew 69% in about two months: from approximately 669,670 skills in June - rywalker.com - to 1,131,746 on the live counter today - skills.sh.
The star side of the ecosystem tells a matching story with one twist. mattpocock/skills grew 27% in four weeks to 203,962 stars, overtaking the Karpathy behavioral-rules repository (199,699) that had been the darling of the spring - mattpocock/skills. Superpowers remains the most-starred skills repository at 266,828 stars, but its four-week growth of about 7% is a third of its spring pace - obra/superpowers. Stars endorse; installs adopt. Pocock now leads on both, which has not been true of any single author since the format existed. Five of the global top nine skills by installs come from his one repository (grill-me, grill-with-docs, improve-codebase-architecture, tdd, and setup-matt-pocock-skills), with handoff at #33 and triage at #35 close behind.
For readers who followed this guide's earlier editions, the longer arc in one paragraph: our January ranking was built on marketplace access counts and was almost entirely falsified when Vercel's install-tracked registry launched; six of those ten skills (Prompt Lookup, SkillsMP, Dify Frontend Tester, Electron Upgrade Advisor, Next.js Cache Optimizer, Skill Writer) are nowhere near today's top ranks, three survived (react-best-practices, web-design-guidelines, remotion), and one, Ralph, was promoted into Claude Code itself. Our July ranking got the source of truth right and the order wrong within a month, because it treated a fast-moving distribution as a still photograph. The correction this edition makes is structural: dated observations, published deltas, and an explicit warning that any specific number on this page is a timestamped reading, not a permanent fact. That is what "tracked ranking" means.
3. Agent Skills in Brief: The Format and the Open Standard
If you already run skills daily, skip to section 4. For everyone else: an agent skill is a folder. A directory containing a SKILL.md file with YAML frontmatter (a name and a one-line description), a body of instructions, and optionally scripts and reference files. When an agent meets a task matching the description, it loads the instructions and follows them. The load-bearing idea is progressive disclosure: at rest, the agent sees only each skill's name and description, a few dozen tokens; the full body loads only on relevance; bundled files load only on need - Anthropic Engineering. That is why an agent can carry hundreds of skills without drowning its context window.
The reason this one convention now matters commercially is that it stopped being one vendor's convention. Anthropic published Agent Skills as a formal open standard on December 18, 2025, with the spec and SDK at agentskills.io - The Decoder. The spec repository holds 23,864 stars - agentskills/agentskills, and the adopter list names more than 40 client integrations: Claude Code and Claude.ai, OpenAI Codex, Gemini CLI, GitHub Copilot, VS Code, Cursor, JetBrains Junie, Goose, Letta, and Databricks Genie Code among them. A skill authored once runs in Codex, where it is invoked with a $skill-name mention or selected automatically - OpenAI Codex docs, and in GitHub Copilot, which documents agent skills as a first-class concept - GitHub Docs. Write once, run everywhere is not aspiration; it is the shipped state of the ecosystem.
Distribution consolidated as fast as the format did. Vercel launched the skills CLI and the skills.sh directory on January 20, 2026, initially supporting 17 agents and now more than 20 - Vercel changelog. Installing any published skill is one command:
npx skills add mattpocock/skills/grill-me
# or search interactively
npx skills find "requirements gathering"
The host platforms matured alongside the format, and two platform features have become load-bearing for skill users. Claude Code's changelog through the summer added a disableBundledSkills setting, nested .claude/skills directory loading with clash resolution, stacked slash-skill invocation letting a single message chain up to five skills, and, most importantly, a --safe-mode flag that starts a session with every customization (CLAUDE.md, plugins, skills, hooks, MCP servers) disabled - Claude Code changelog. That last one matters beyond convenience: when an agent starts behaving strangely, --safe-mode is the fastest way to establish whether an installed skill is steering it, which makes it the triage tool the security section leans on. If you are choosing a first client for skill-heavy work, the maturity of these controls is a real selection criterion, not a footnote.
The registry that command feeds is the closest thing this ecosystem has to ground truth, which is why this guide leans on it while also refusing to take it at face value. It counts something real: a CLI action that placed files on a machine. It does not count usefulness, safety, or whether the install arrived through a bundler's funnel. Those judgments are the ranking's job, and they start with the new number one.
4. grill-me: The New Number One
What it is. grill-me inverts the default agent interaction. Instead of taking your request and building immediately, the agent interrogates you first: what is this for, who uses it, what are the constraints, what does done mean, what should explicitly not happen. Implementation starts only when the requirements survive the grilling. It is the flagship of Matt Pocock's "Skills for Real Engineers" repository, sits at 756.3K installs, and after growing 58% in four weeks it is the highest-ranked skill on this list and the top organic skill on the leaderboard - skills.sh/mattpocock/skills.
Why a questioning skill outranks capability skills. The most expensive failure mode in agent work is not bad code; it is a confident build in the wrong direction. Frontier models will competently construct whatever you underspecified, and you discover the misunderstanding only after reviewing a finished, wrong artifact. grill-me moves that discovery to minute one, where it costs a conversation instead of a rebuild. The economics compound with model strength: the better the model builds, the more a wrong-direction build costs before anyone notices, which is our best explanation for why this skill's growth accelerated precisely as the strongest model generation yet (section 18) rolled out. The install curve is the market pricing in a truth experienced operators already knew: specification, not capability, is the binding constraint on agent output in 2026.
How to use it well. Selectively and honestly. Selectively: prefix it to anything ambiguous, novel, or expensive to redo, and skip it for mechanical tasks with a genuinely complete spec. Honestly: the skill works only if you answer the questions rather than brushing them off, because your answers become the working spec. Teams get a second-order benefit they rarely anticipate: the grilling transcript is documentation, capturing the requirements conversation that normally evaporates. We see the same dynamic on the O-mega platform from the operator side: sessions that begin with an explicit requirements exchange produce dramatically fewer abandoned-and-retried runs than sessions that begin with a bare instruction, across every session type we operate.
A concrete session shape makes the value tangible. Ask an agent to "add an export feature to the dashboard" without grill-me and you typically get an immediate build: CSV export of whatever table the agent found first, plausible and useless if what you actually needed was a scheduled PDF report for a board pack. Under grill-me, the same request first produces the questions that would have exposed the gap: who consumes the export, in what format, triggered how, covering which data range, with what handling for the columns containing personal data. Five questions, ninety seconds, and the wrong build never starts. The skill's install curve is hundreds of thousands of individual versions of that ninety seconds, which is why we weight its momentum as the most meaningful signal on the board.
Limitations. It adds latency and friction by design, and on tasks that did not need it the friction is pure cost. It cannot surface constraints that neither you nor the agent knows exist. And its value depends on the model's questioning quality, which is high on current frontier models and degrades on small ones, which ask generic checklist questions instead of load-bearing ones.
5. frontend-design: The Design-Taste Transplant
What it is. frontend-design is Anthropic's official skill for making agent-built interfaces look professionally designed rather than obviously generated. It holds 742.3K installs, ranks #3 raw and #2 organic, and ships in the official anthropics/skills repository, now at 166,356 stars. It encodes concrete design judgment: typography scale and pairing, spacing systems, color discipline, layout composition, states and affordances, the dozens of small decisions separating a credible product UI from bootstrap-flavored output.
Why it holds #2 despite losing the top spot. Nothing about the skill got worse; the competition got faster. Its impact score remains the highest of any knowledge skill on this list because it produces the most visible before-and-after in the ecosystem: ask an agent for a landing page without it and you get competent, generic markup; ask with it and you get deliberate typography, a coherent spacing rhythm, restrained color. Its trust score is a 10, maintained inside Anthropic's own repository. What dropped is momentum: +16% in four weeks against the Pocock cluster's +57-61%. Our reading, and it is a reading rather than a measurable fact, is saturation rather than rejection: frontend-design was the ecosystem's default second install for two quarters, and the population that wants it largely has it.
How to use it well. The skill supplies design-system thinking; you still own the product intent: audience, tone, one or two reference points. The most effective pattern remains the iteration loop, generate, screenshot, critique against the skill's own principles, regenerate. On platforms where agents build web properties as a managed capability, this kind of design-guideline knowledge increasingly ships as standing equipment rather than a per-session install, which is how O-mega wires it and one reason install counts undercount design-skill usage generally: platform-embedded copies never register on the leaderboard.
Limitations. It cannot rescue a confused product brief, and it biases toward a recognizable contemporary aesthetic: clean, spacious, type-led. If your brand needs something deliberately unusual, the skill's defaults become a gravity well to pull against. It concentrates on visual and interaction quality, not information architecture: it will not tell you the page answers the wrong question.
6. agent-browser: Hands for the Web
What it is. agent-browser gives coding agents a controllable browser through a ref-based command interface: the agent snapshots a page, receives stable element references from the accessibility tree, and issues click, type, and navigate commands against them. Launched January 11, 2026 by Vercel Labs, it stands at 629.0K installs (#5 globally) and 39,946 GitHub stars - vercel-labs/agent-browser. The design insight is economy: structured snapshots instead of screenshot-driven vision loops, which is faster, cheaper, and far less flaky.
Why it matters. An agent without a browser cannot verify its own work: it can write a UI but not see it render, produce a checkout flow but not walk through it. agent-browser closes that loop, and closing the verification loop is the single biggest reliability upgrade available to agent-built software. Its +20% four-week growth, the strongest in the official-vendor cohort, tracks the broader shift toward agents that check their work rather than merely produce it. The skill also unlocks the wider category of web work: form automation, scraping, end-to-end testing, competitive research. For a comparison of the full landscape of browser-driving agents and where ref-based automation wins and loses against vision-based approaches, our browser-use agents review covers the field in depth.
How to use it well. The highest-value pattern is self-verification: after the agent builds or changes a flow, require it to drive that flow in the browser and report observations before declaring success. Second is scheduled journey checks: agents that periodically walk critical user paths and file issues when something breaks. A first-hand note from operating browser automation at platform scale: which public browser skills matter depends heavily on whether your agent stack executes browser work as a managed session type. O-mega runs browser automation as a built-in session type with its own controlled environment, so for our agents the leverage sits in site-specific procedure skills layered on top, not in the browser-control primitive itself. If you self-host your agents, the primitive is exactly what agent-browser provides, and it belongs in your first five installs.
Limitations. It operates a real browser, so it inherits the web's hostility: login walls, CAPTCHAs, bot detection, and rate limits are its weather. It deliberately does not solve authentication vaulting or IP rotation, which production scraping needs. And handing an agent a browser widens the security surface: a malicious page's content becomes agent input, a prompt-injection vector covered alongside the rest of the threat model in our prompt injection defense guide.
7. The Pocock Behavioral Stack: tdd, grill-with-docs, improve-codebase-architecture
Ranks four through six belong to one repository, and that concentration is itself the finding. Beyond grill-me, Matt Pocock's skills repo places tdd at #8 globally (595.4K installs, +59% in four weeks), grill-with-docs at #4 (641.9K, +61%, the fastest four-week growth we tracked), and improve-codebase-architecture at #6 (617.5K, +57%) - skills.sh. Add the bundler at #9, and one author holds five of the global top nine, with handoff (#33) and triage (#35) close behind. The repository's 27% star growth in the same window, to 203,962, means the endorsement curve is keeping pace with the adoption curve.
What the three ranked skills share is the property section 1 weights heaviest in spirit: they are procedural interventions, not knowledge payloads. tdd forces a failing test before implementation, converting the model's confident-but-wrong tendencies into fast, visible feedback. grill-with-docs runs the grill-me interrogation against your own documentation, so the questions come armed with what your project already claims to be true, catching the specifically painful failure where an agent builds something your own docs prohibit. improve-codebase-architecture structures the refactor conversation: instead of "clean this up," it drives an assessment of coupling, boundaries, and dependency direction before any code moves. None of them teach the model a fact it lacks. All of them constrain how it behaves when it already knows enough to be dangerous.
Why is the market converging on discipline over knowledge this quarter? From first principles: as base models strengthened through 2026, the binding constraint on agent output moved. In 2025, agents failed because they could not do things, so the market bought capability. By mid-2026, frontier models rarely fail on capability and routinely fail on judgment defaults: they overbuild, under-ask, skip verification, plow past ambiguity. Knowledge skills cannot fix judgment defaults; procedural constraints can. The Karpathy behavioral-rules repository, four rules distilled from Andrej Karpathy's observations about LLM coding pitfalls, rode the same insight to 199,699 stars - multica-ai/andrej-karpathy-skills, and this quarter Pocock's install-verified versions of the same philosophy overtook it on stars too. The highest-leverage text you can put in front of a strong model is not information. It is discipline. This matches what Claude Code's own internals reveal, where the system prompt is mostly behavioral constraint, as we documented in our leaked source analysis.
The practical guidance for stack assembly follows directly: install behavioral skills before capability skills. A disciplined agent with two knowledge skills outperforms an undisciplined agent with twenty, because the behavioral layer decides whether knowledge gets applied at the right moment. The recurring stack among heavy users is one requirements skill (grill-me or grill-with-docs, rarely both), one sequencing skill (tdd), one restraint layer (Karpathy rules or equivalent), then domain knowledge as work demands. A caution that applies to the whole cluster: these skills change agent behavior deeply, and the repository moves fast. Pin versions for team use and review updates before upgrading, for the reasons section 14 makes vivid.
8. vercel-react-best-practices: Framework Judgment, Inlined
What it is. vercel-react-best-practices packages Vercel's accumulated React and Next.js performance judgment into rules the agent applies while writing or reviewing code: parallelizing data fetches instead of waterfalling, keeping client bundles lean, respecting server/client component boundaries, avoiding render cascades, each with concrete bad-versus-good examples. It stands at 606.7K installs (#7 globally), in the officially maintained vercel-labs/agent-skills collection at 29,761 stars.
Why it earns 8.2. Impact stays a 9 because the failure modes it catches are the dominant real-world causes of slow React apps, and agents without this knowledge reproduce them constantly, since so much training data does. Trust stays a 9: a year of framework releases and issue reports has hardened the rule set, and the vendor whose framework you are optimizing maintains it. What costs it position is momentum, +14% in four weeks, the slowest of the top organic cohort. The plausible read parallels frontend-design: the React-building population that wants it has it, and the frontier of new installs has moved to stack-agnostic behavioral skills. Scope is the other cap: it is worth a 9 or 10 to a React/Next.js team and near zero to everyone else.
How to use it well. Treat it as a standing reviewer, not an occasional linter: agents that load it during code generation avoid the regressions; agents that only audit afterward inherit them. The transferable lesson for non-React teams is the pattern, not the skill: find the skill in which your framework's vendor encoded their performance review culture, and if it does not exist, that gap is your authoring opportunity (section 17).
Limitations. The skill enforces Vercel's opinions, well-earned but not neutral: it steers architectures toward patterns that shine on Vercel's platform. Most advice is platform-agnostic, but keep the provenance in mind for deployment-adjacent suggestions. And like every review skill, it identifies problems more reliably than it sizes them: twelve flagged issues may contain two that matter for your traffic profile, and you still make that call.
9. Superpowers: The Methodology
What it is. Superpowers remains the most-starred skills repository on GitHub at 266,828 stars and 23,851 forks, created by Jesse Vincent in October 2025 - obra/superpowers. It is not one skill but an integrated development methodology delivered as a skill framework: test-driven development discipline, systematic debugging procedure, subagent orchestration patterns, planning and brainstorming protocols, and meta-skills that teach the agent to acquire further skills. Installing it changes the agent's entire working posture, from "generate code when asked" to "operate a development process."
Why it slid from the podium. Its impact score is still the only 10 on this list: nothing else changes as much about how an agent works. Two forces cost it rank in a momentum-weighted scoring. First, its growth cooled to +7% stars in four weeks, roughly a third of its spring pace, while the Pocock cluster compounded at +57-61% installs. Second, the market is unbundling it: grill-me, tdd, and improve-codebase-architecture each deliver one slice of the Superpowers philosophy as a single cheap install, and the install data says most adopters now assemble the methodology piecewise rather than committing to the whole framework. Superpowers is the cathedral; the market is currently buying bricks. For teams that do commit, the compounding is real: each phase's output disciplines the next, and the fork count says thousands of teams are adapting the methodology to their own conventions rather than merely endorsing it.
How to use it well. Commit or skip; cherry-picking gets you context load without the payoff. Run one real feature end-to-end under the full loop (brainstorm, plan, test-first implementation, systematic verification) before judging. It pairs naturally with loop-style autonomous execution, where a methodology plus fresh-context iteration is what makes overnight unattended work trustworthy, mechanics we cover in our guide to writing loops for AI coding agents and, for the long-horizon case, in our long-running coding agents guide.
Limitations. Trust scores a 7, not from any incident but from structure: a fast-moving, opinionated, solo-led project whose updates change agent behavior more deeply than any rules list. Pin a version for team use. It is heavyweight for quick edits, and the deepest integration remains Claude-first even as ports spread, which costs portability against pure SKILL.md entries.
10. web-design-guidelines: The Hundred-Point Audit
What it is. web-design-guidelines is Vercel's audit skill: over 100 concrete rules covering accessibility (ARIA usage, alt text, focus order, contrast), form behavior, keyboard navigation, responsive layout, typography, dark mode, touch targets, and internationalization. It stands at 516.5K installs, having slipped from #6 in our July data to #32 on today's leaderboard - skills.sh. Where frontend-design shapes what the agent creates, this skill inspects what exists: point it at a page or component tree and it produces findings with the violated rule and the fix.
Why it stays ranked despite slipping. Audit checklists are the perfect content type for the skill format: voluminous, precise, dull, exactly the profile of expertise humans skip under deadline and agents execute perfectly every time. The rules are objective enough that findings are mostly verifiable on inspection, which keeps trust high. And the regulatory floor keeps rising: with the European Accessibility Act's obligations applying to most consumer-facing digital services, a systematic accessibility pass moved from virtue to compliance requirement, and this skill remains the cheapest systematic pass available. Its slide down the leaderboard reflects the field's growth, not decay; +16% in four weeks is healthy for a mature audit tool.
How to use it well. Wire it into the definition of done rather than running occasional cleanups: an agent that audits every new page before marking work complete catches issues while they are one-line fixes. Pair it with agent-browser so findings come from the rendered DOM rather than static reading, which catches runtime issues (focus traps, contrast in computed colors) that code review misses.
Limitations. It measures conformance, not experience: a page can pass every rule and still confuse users, because task-flow quality is not rule-checkable. It produces false positives on intentional patterns (decorative images, deliberate focus handling), so findings need a human accept/reject pass. Scope is web only. And there is a portfolio-level caution its slide illustrates: audit skills sit downstream of creation skills in the value chain, so when budget or attention forces a choice, teams keep the skill that shapes output (frontend-design) and defer the one that inspects it. If accessibility is a compliance requirement for you rather than a preference, resist that ordering; the audit is the layer with legal weight.
11. find-skills and the Bundler Problem
What it is. find-skills is the discovery meta-skill: install it once and your agent can search the registry, evaluate candidates, and install what it needs from inside the conversation. At 2.8M installs, up roughly 400K in four weeks, it holds the #1 raw slot by nearly a factor of four - skills.sh. It is published by Vercel Labs alongside the skills CLI (28,072 stars) - vercel-labs/skills.
Why we discount it, and what the discount reveals. find-skills is the default first install: Vercel's onboarding starts with it, most tutorials start with it, most agent-platform setup guides start with it. Its count therefore measures the ecosystem's onboarding funnel more than any judgment about the skill itself. The August leaderboard supplied a perfect natural experiment proving the category: setup-matt-pocock-skills, a bundler whose sole function is installing the Pocock collection, entered the global top 10 at #9 with 544.6K installs - skills.sh. Nobody chose it on merit; they chose the collection, and the bundler's count is an echo of that choice. A ranking that cannot distinguish an echo from a voice will increasingly mis-rank the head of this leaderboard as more distribution flows through bundles. That is the argument for the organic-adoption adjustment in section 1, and as far as we can tell no aggregator ranking makes it.
How to use it well. Constrain the source: instruct the agent to prefer publishers you already trust (anthropics, vercel-labs, mattpocock, microsoft, your own org) over raw popularity. And review before activation: have the agent print the SKILL.md of anything it proposes to install, and read it. That thirty-second habit is your primary defense against the supply-chain mechanics in section 14, because discovery skills are precisely the mechanism that turns one careless session into an installed backdoor.
Limitations. It improves no output directly: indispensable plumbing, but plumbing. It is only as good as the registry's hygiene, and skills.sh verifies installs, not safety. And its 2.8M installs make it the single most attractive typosquatting and dependency-confusion target in the ecosystem, which is exactly why the vetting habit matters.
12. The Enterprise Land-Grab: Microsoft, and Now ByteDance
The July edition of this guide argued that Microsoft's azure-skills sprint validated the SKILL.md format for enterprises: a repository created February 26, 2026 that placed four skills in the top 15 of our July reading within four months. The August data both confirms the thesis and internationalizes it. microsoft-foundry now holds 506.1K installs at #36, with thirteen more Azure skills filling ranks 37 through 49 directly behind it, and azure-messaging, azure-rbac, azure-compute, azure-cloud-migrate, and azure-hosted-copilot-sdk continuing through the top 60 - skills.sh. Microsoft's cohort grew a steady 15% over four weeks: the flat, durable growth of procurement-driven adoption rather than viral discovery.
The new fact is ByteDance. Skills published under open.feishu.cn, the developer platform of ByteDance's Lark (Feishu) office suite, went from absent to 22 of the global top 31 slots in four weeks: lark-approval at #10 with 529.4K installs and the next 21 leaderboard positions behind it, lark-okr at #34 with 509.2K, then lark-markdown, lark-vc-agent, and lark-apps - skills.sh. These teach agents to operate Lark's approval workflows, OKR system, documents, and video conferencing. Their absence from every English-language ranking we can find is partly a language artifact and partly aggregator laziness, but the strategic fact is unambiguous: the first non-Western enterprise vendor has decided the agent-readable version of its platform belongs on the global open-standard registry, and its install velocity in one month matched what took Microsoft a quarter. Lark claims hundreds of millions of users across Asia's enterprise market; if its skills sustain anything like this pace, the leaderboard's top 10 will not stay a Western artifact for long. We score lark-approval at 7.3, just off the top 10: impact capped by Feishu-scoped value, trust capped by newness and thin English documentation, momentum a 10 by definition.
The strategic logic driving both vendors is the same, and it predicts the next entrants. For a platform vendor, an official skill is simultaneously documentation, onboarding, and moat. Documentation, because it is the platform's procedures in executable form, always fresher than training data. Onboarding, because an agent with the vendor's skill makes a developer productive on the platform in minutes. Moat, because once an organization's agents fluently operate Azure, or Lark, or Stripe, the switching cost of retraining workflows grows quietly with every automated procedure. This is the developer-relations playbook compressed into a folder of markdown, which is why official vendor skills carry such a high quality bar: they are competing artifacts, not community contributions. Watch for the remaining enterprise majors (SAP, Salesforce, ServiceNow are conspicuously absent from the leaderboard) because the pattern says their arrival is a when, not an if.
There is a tension the open standard only partially resolves, and it belongs in your architecture thinking. The format is portable: microsoft-foundry runs identically in Claude Code, Codex, or Gemini CLI, so no agent vendor locks you in. But the content steers: each vendor skill encodes its publisher's preferred architecture, and a stack assembled entirely from vendor skills will build you a system shaped by those vendors' interests. The counterweight is deliberate: pair vendor skills with behavioral skills that enforce your own review discipline, and maintain private skills encoding your architecture decisions so the agent weighs your constraints against vendor defaults. The stabilizing enterprise pattern is a three-layer stack: open-standard behavioral skills at the base, vendor skills for platform procedure, private organizational skills on top, with the organizational layer holding the final word.
Read this chart against the install chart in section 1 and a split emerges: stars and installs measure different communities. The star leaders are community methodology repos that developers evaluate and endorse; the install leaders skew toward official and enterprise skills pulled during setup. The August twist is that Pocock's repo now leads the community cohort on both measures simultaneously, which no author has done before. A skill stack built from only one chart misses half the ecosystem.
13. What Fell Out: remotion, Ralph, and the January Casualties
An honest tracked ranking reports its exits, because a demotion is information the leaderboard's current view hides. The most significant August exit is remotion-best-practices, the programmatic-video skill we ranked in the top 10 in both January and July. It now sits at #55 globally with 464.2K installs - skills.sh/remotion-dev/skills. Note what the number says and does not say: it grew 12% in four weeks, its 28 modular rule files remain the reference example of well-structured skill authoring - Remotion docs, and it is still the only serious entry in its category. It did not get worse; the leaderboard's center of gravity moved to coding discipline and enterprise procedure, and a category with one mature player generates no competitive install churn. If you build video with agents, install it exactly as before. If you are reading the leaderboard as a market map, its slide is a datum: creative-output skills are not where this quarter's growth lives.
Ralph exited the ranking in the opposite direction: upward, into the platform. The technique of running an agent in a persistent loop against a goal until the work is verifiably done, pioneered by Geoffrey Huntley as a bash while-loop, became an official ralph-wiggum plugin shipped inside the Claude Code repository, implemented with a Stop Hook that intercepts the agent's attempts to finish and re-feeds the goal - anthropics/claude-code, alongside native /loop, /goal, and /batch commands. HumanLayer's history documents the full arc from hack to adoption - HumanLayer. The documented outcomes remain remarkable: The Register's January coverage describes Y Combinator teams shipping six repositories overnight for $297 in API costs, and Huntley's three-month continuous loop built a programming language, compiler included - The Register. We rank installable skills, and Ralph is now a platform primitive rather than an installable skill, so it leaves the table while remaining essential practice; the preconditions (machine-checkable goals, external memory, spend caps) are covered in our loops guide.
The January casualties complete the record, compressed because the lesson matters more than the obituaries. Of our original January top 10, built on marketplace access counts, six are nowhere near today's top ranks: Prompt Lookup (absorbed by stronger base models), Skill Installer/SkillsMP (superseded by npx skills add and find-skills), Dify Frontend Tester and Electron Upgrade Advisor (niche audiences that never converted to installs), Next.js Cache Optimizer (absorbed into Vercel's official skills), and Skill Writer (superseded by Anthropic's skill-creator and the published spec). The survivors, react-best-practices, web-design-guidelines, and remotion, share one property: each encodes a large, maintained body of expert judgment a base model cannot regenerate reliably on demand. The casualties were thin utility wrappers whose function the platform or the models absorbed. That selection pressure operates on all 1.13 million registry entries, and it is the durable test to apply before depending on any skill: would this survive its own function being absorbed into the next model generation?
14. Security: The Malicious-Skill Supply Chain
The defining skills story of 2026 still is not on any leaderboard. In early February, Palo Alto Networks' Unit 42 analyzed skills on the OpenClaw marketplace and found roughly 17% carried malicious payloads in the platform's first weeks, with Koi Security's associated ClawHavoc disclosure documenting 341 malicious skills at the time of discovery - Unit 42. Those figures are dated events now, but nothing structural has changed since: the registry doubled, the attack surface doubled with it, and install-verified popularity remains no proxy whatsoever for safety. If you install skills, you are running a package manager for your AI's brain, and this year proved the packages get poisoned. The parallel ecosystem's story is told in our OpenClaw workforce guide.
The attack anatomy matters because half of it is genuinely novel. A large-scale 2026 empirical study of malicious skills documents two distinct classes - arXiv. Code-level attacks are the familiar kind: a skill bundles scripts, the agent executes them, the script exfiltrates credentials or establishes persistence; your existing security intuitions roughly apply. Instruction-level attacks are the new kind: the SKILL.md itself contains directives steering the agent to quietly include data in outbound requests, weaken code it writes, or persist instructions into other files. No executable payload exists to scan. The weapon is prose, the execution engine is the model's own instruction-following, and the attack activates only in the semantic context the attacker chose. The academic response (MalSkillBench for benchmarking detection, PhantomSkill demonstrating attacks that stay dormant through review, SkillSieve on the defense side, at least eight arXiv papers in the first half of 2026) maps a threat class that is detectable in bulk but not reliably in any single case. The gap between those two clauses is where your vetting habits live.
A first-hand data point from the curation side, because O-mega operates a curated skill library and the rejection pile is instructive. The skills we decline to index are rarely cartoonish malware; they are scope creepers: a documentation skill whose instructions also direct the agent to fetch a remote URL "for updates" on every activation, a formatting skill that asks the agent to write persistent instructions into project config files, bundled scripts with network calls nowhere implied by the skill's description. Each has an innocent explanation and each is indistinguishable in form from an instruction-level attack, which is precisely the point: the safe-versus-unsafe line runs through undeclared behavior, not stated intent. Our operating rule, and the one worth stealing for your own vetting, is that a skill earns trust when everything it does is implied by its one-line description, and loses it the moment an instruction serves the publisher rather than the user. Executing skills inside controlled, sandboxed sessions rather than on user machines, which is how the O-mega platform runs them, bounds the blast radius; it does not remove the need for the judgment.
The defense posture for self-hosted stacks is not optional hygiene; it is the cost of participation. Concretely:
- Read before you run: open the SKILL.md and every bundled script; if you cannot explain what an instruction is for, do not install it
- Pin trusted publishers: prefer anthropics, vercel-labs, mattpocock, microsoft, remotion-dev, and your own org; treat unknown publishers as untrusted code-review subjects
- Use the kill switches: Claude Code's
--safe-modestarts a session with every skill, hook, plugin, and MCP server disabled, the fastest way to test whether a misbehaving agent is being steered by something you installed - Claude Code changelog - Scope credentials: agents with skills installed should hold least-privilege tokens, so a compromised session has a bounded blast radius
- Audit periodically: skills update; a publisher compromise can turn yesterday's safe install into today's payload, exactly like npm
The checklist is short because the habit matters more than the length. The one-sentence version: treat every skill as code you are deploying to your most privileged machine, because that is literally what it is. This year supplied hundreds of documented reasons, and the doubling registry supplies more every month.
15. Skills Beyond Coding: Cowork and Document Work
Every skill profiled so far touches software, and that scoping is now the minority use case for the format. Claude Cowork, Anthropic's agent for non-coders, shipped as a desktop app in January 2026 and expanded to web and mobile with cloud execution on July 7, 2026 for Max subscribers, and the number that reframes this entire guide came with it: more than 90% of Cowork usage is non-software work, roughly half of it business operations or content creation - TechCrunch. The same SKILL.md format that teaches a coding agent React performance teaches an office agent how to process invoices, and the office is the larger market.
The bridge is Anthropic's own document skills: the official anthropics/skills repository bundles docx, pdf, pptx, and xlsx creation and manipulation, available to paid Claude.ai plans. Add the partner skills announced with the open standard (Canva, Stripe, Notion, and Zapier among the launch cohort) - The Decoder, and now ByteDance's Lark suite from section 12, and the shape of the non-developer skills economy is visible: vendors publish skills that make their products operable by agents, and knowledge workers install procedures the way they once installed apps. The Lark entrance is best understood through this lens too: approval workflows and OKR updates are exactly the recurring office procedures the format was built to encode.
A worked example makes it concrete. A finance operator processes vendor invoices weekly: extract fields from PDFs, check amounts against purchase orders, flag mismatches over a threshold, produce a summary spreadsheet. Encoded as a private skill, that becomes a SKILL.md describing extraction fields and tolerance rules, referencing the pdf and xlsx skills for mechanical work, with the company's escalation thresholds in a reference file. Writing it costs an hour once. Every subsequent week the procedure runs in minutes, identically, including the edge-case handling that previously lived in one employee's head. Multiply that shape across the recurring workflows of any operations, marketing, or HR function and you have the mechanism behind the 90% figure: not exotic capabilities, but ordinary procedures finally having a durable, executable home. Consumer pricing context, re-verified this week: Claude Pro at $17/month annual ($20 monthly), Max from $100/month, Team at $20 per seat annual with premium seats at $100, all bundling Claude Code and Cowork - claude.com/pricing. Setup and platform mechanics are in our Cowork starter guide, with the ecosystem view in our Cowork pricing and agent ecosystem analysis.
This is also where we will state the pattern O-mega sees most clearly from operating both halves of a skill economy: public utility skills commoditize; private organizational skills compound. Public skills race to the same ceiling, because every improvement is instantly available to everyone, including your competitors; the leaderboard churn documented in section 2 is that race made visible. Private skills, your organization's own procedures encoded once and executed by every agent thereafter, appreciate with each edge case they absorb, and they are the only part of a skill stack competitors cannot install. On the O-mega platform, private skills are stored as files with semantic discovery, searched in the same unified index as the public library, so an agent reaching for procedure finds your version of the workflow before a generic one. Whatever platform you use, that division, generic capability rented from the commons, differentiating procedure owned privately, is the stable end-state the last six months of data point to.
16. Skills vs MCP vs Plugins vs AGENTS.md
Mid-2026's agent stack has four extension mechanisms that get conflated constantly, and choosing wrong wastes context, money, or both. The clean disambiguation: skills are procedures, knowledge loaded on demand that changes how the agent performs a task. MCP is connectivity, live tool connections through the Model Context Protocol that let the agent call external systems at runtime. Plugins are packaging, Claude Code's bundle format shipping skills, hooks, slash commands, and MCP configuration as one installable unit, which is how ralph-wiggum ships. AGENTS.md is standing context, a per-repository instruction file read automatically by Codex and a growing list of agents, holding the always-relevant conventions of one codebase.
The composition rules follow from what each layer is. A skill can instruct the agent to use an MCP tool: the procedure says "query the warehouse via the MCP connection," which is the correct division of labor, MCP providing the verb and the skill providing the judgment about when and how. AGENTS.md holds only what applies to every task in the repo, because it loads unconditionally and every token in it taxes every session, where a skill's tokens are spent only when triggered. Plugins wrap any of the above for distribution. The test that resolves most cases: is the thing you are adding knowledge (skill), access (MCP), repo law (AGENTS.md), or a shippable combination (plugin)?
The frequent practical mistake runs in both directions. Teams reimplement API documentation as MCP servers when a skill plus a plain HTTP call would do, paying server maintenance for what is really just procedure. And teams stuff giant procedure documents into AGENTS.md or MCP tool descriptions, paying unconditional context tax for conditional knowledge. A worked decision from this month's leaderboard makes the layering concrete: ByteDance's lark-approval is correctly a skill, not an MCP server, because the hard part of automating an approval workflow is procedural (which approval chain applies, what fields the request needs, how exceptions escalate), while the API calls themselves are ordinary HTTP the agent can already make. Had ByteDance shipped it as a server, every session would carry tool definitions it rarely needs; as a skill, the whole procedure costs approximately thirty tokens until the moment someone asks for an approval. The protocol side has its own current-events layer this year, with the MCP specification's move toward stateless operation and long-running tasks reshaping what belongs in a server at all, which we cover in our MCP 2026 spec guide, and the protocol-selection question, MCP against agent-to-agent alternatives, in our MCP vs A2A comparison.
17. How to Ship and Measure Your Own Skill
A 1.13-million-skill registry can read as "everything is taken," but the install distribution says the opposite: a tiny head, most of it now held by a handful of publishers, and an enormous tail of thin coverage. The head keeps rotating (this article has now documented two full rotations in seven months), and the breakout path is visible in the data: grill-me proved a solo author with a well-shaped idea can reach the top of the global leaderboard, and the Lark suite proved an organization can move from zero to the top 15 in a month by encoding procedures a real user base runs daily. The opportunity in late 2026 is specificity: your framework's migration path, your industry's compliance procedure, your niche tool's best practices, the expertise that is scarce rather than general.
The craft is editorial, and three principles carry most of it. Write the description line for retrieval: it is the only text the agent sees at discovery time, so it must state precisely when the skill applies, in the vocabulary a user's request would contain. Structure for progressive disclosure: a lean SKILL.md holding the core procedure, depth pushed into reference files loaded on demand, the pattern remotion's 28-file structure demonstrates at production scale. And test against the failure you are fixing: run the same tasks with and without the skill installed, and keep only instructions that change behavior, because every sentence that does not is context tax on every future activation. The format itself is documented at agentskills.io/specification, and Anthropic's skill-creator scaffolds the structure interactively.
The description line deserves one concrete illustration because it is where most first-time authors fail:
# Weak: vague, matches everything and nothing
name: database-helper
description: Helps with database tasks.
# Strong: states the trigger conditions in the user's vocabulary
name: postgres-migration-review
description: Reviews PostgreSQL schema migrations for locking hazards,
irreversible operations, and missing indexes before they run against
production. Use when writing, reviewing, or debugging a migration file.
The weak version either never activates or activates constantly; the strong version activates exactly when a migration is in play, because its description contains the words a real request contains. Discovery-time matching is the only mechanism deciding whether your carefully written procedure ever loads, so the description is not metadata, it is the skill's API. Test activation if you test nothing else: five requests that should trigger it, five that should not, checked in both directions.
Publishing and measurement run through the same pipe as consumption: push to a public GitHub repository and the skill is installable immediately via npx skills add your-org/your-skill, with skills.sh tracking installs from that moment - Vercel changelog. Be clear-eyed about what the scoreboard signals after this month's data: raw counts now mix organic adoption with bundler distribution, so watch your install velocity and retention of activation rather than the absolute number, and for a niche skill a few thousand installs can mean complete saturation of the audience that matters. The strategic choice is the one section 15 framed: publish what markets your expertise; keep private what constitutes it. The procedures that differentiate your business belong in a private skill layer, whether that is Claude enterprise skill distribution or a platform-managed library, not on a public leaderboard.
18. The Models Running Your Skills in September 2026
Skills are instructions, and instructions are only as good as the model following them, so a skill ranking silently assumes a model lineup. Here is the September 2026 lineup, verified against the live model documentation this week. Anthropic's current family: Claude Fable 5.1 (claude-fable-5-1, $10 input / $50 output per million tokens, 1M-token context, 128K max output, released September 1, 2026), Claude Opus 5 (claude-opus-5, $5/$25, 1M context, 128K output), Claude Sonnet 5 ($2/$10, 1M context), and Claude Haiku 4.5 ($1/$5, 200K context), with Claude Mythos 5.1 sharing Fable 5.1's specs but remaining invitation-only within Project Glasswing - Claude model docs. Two facts change the practical calculus. First, Anthropic's documentation names Opus 5 as the recommended starting point for most workloads, complex agentic coding included, with the Opus 4.x line (4.8, 4.7, 4.6) and now Fable 5 itself sitting on the legacy list at unchanged prices, a generational handoff we benchmark in our Opus 5 vs 4.8 comparison. Second, Claude Opus 4.1 retired on August 5, 2026, the first skills-era Claude model to complete the full deprecation cycle: a reminder that model references inside your own skills and scripts rot on a schedule.
The number that matters most to anyone carrying a large skill library arrived with that handoff. Fable 5.1 cut cache reads by 75%, to $0.25 per million tokens, billing them at 0.025x base input where every other Claude model charges 0.1x, and Anthropic puts the resulting bill at around 25% lower on typical workloads and up to roughly 45% lower on complex coding and highly agentic tasks - Anthropic. That lands directly on the economics of this guide, because a skill library is a cached prefix: the cost of carrying fifty SKILL.md descriptions through a long agent loop is one cache read repeated on every turn, and on the top model it is now a quarter of what it was.
On the OpenAI side, the July edition of this guide went stale within a day of publishing, which is its own lesson. The GPT-5.6 family (Sol, Terra, Luna) launched publicly across ChatGPT, Codex, and the API in the week of July 9-10, 2026, with Sol posting 53.6 on Agents' Last Exam and state-of-the-art results on Terminal-Bench 2.1 - Latent Space. Two rounds of cuts have repriced the family since: Luna fell 80% and Terra 20% on July 30, then Sol dropped 20% on input and 33% on output on August 21, leaving Sol at $4/$20, Terra at $2/$12, and Luna at $0.20/$1.20 per million tokens. Sol is no longer the top of that stack either. GPT-6 Astra shipped on September 3, 2026 as OpenAI's most capable model, priced at $10/$50 with a 1.05M-token context and 128K max output - OpenAI API pricing. Full benchmark and pricing detail lives in our GPT-5.6 breakdown, and the head-to-head agentic comparison in our GPT-5.6 vs Claude Opus 5 analysis.
The model-skill interaction has a shape worth internalizing: stronger models extract more from the same skill. A behavioral skill like grill-me asks sharper questions on a frontier model than a small one; a dense rule set gets applied with better judgment about which rules bind in context. This cuts against the intuition that skills are a crutch for weak models; in practice skills are leverage, and leverage scales with what it is applied to, which is one reason the behavioral-skill boom coincided with the strongest model generation yet. The economics follow: 1M-token contexts on Fable 5.1, Opus 5, and Sonnet 5 make large skill libraries nearly free to carry at discovery stage, Sonnet 5's now-permanent $2/$10 rate makes bulk skill work cheap year-round rather than for one closing month, and methodology or loop workloads want the strongest model available because judgment errors compound across iterations, affordable on subscription rather than API metering as detailed in our Claude Code cost breakdown. For the cross-vendor selection question, our August 2026 LLM ranking for agents scores the full field on agentic criteria.
19. Outlook and Decision Framework
Standing on three timestamped observations rather than one, four trajectories look structural. Behavioral skills keep compounding: two consecutive four-week windows of 57-61% growth for procedural-discipline skills, against 14-20% for knowledge skills, is a trend with a mechanism behind it (stronger models raise the cost of bad judgment faster than the cost of missing knowledge), and trends with mechanisms persist. The leaderboard internationalizes: ByteDance's one-month sprint to 22 of the top 31 slots ends the era of reading skills.sh as a Western developer artifact, and the obvious followers are Asia's other enterprise platforms. Bundler distortion grows: as more distribution flows through collection installers like setup-matt-pocock-skills, raw install rank will diverge further from organic adoption, and rankings that do not adjust will quietly become measurements of onboarding funnels. And private skill layers become the moat: public procedure is commoditizing at exactly the pace the changelog documents, which leaves organization-specific skill libraries as the only compounding, non-installable asset in the stack, a bet every serious platform from Claude enterprise plans to O-mega is making with first-class private-skill infrastructure. The broader platform context for these bets is mapped in our Anthropic ecosystem guide.
The decision framework, compressed. If you are new to skills: install grill-me, frontend-design, and find-skills, build one real thing, and add nothing else until you feel a specific gap. If you run a development team: adopt a behavioral base (grill-me plus tdd, or full Superpowers if you will commit to the whole loop), add your stack's official vendor skills, mandate the section 14 vetting habit, and pin publishers and versions. If you are responsible for security: treat skills as a package-manager surface, make --safe-mode triage standard procedure, and inventory what is already installed across your org, because the base rate of malicious packages is a measured quantity, not a hypothetical. If you run a business function outside engineering: start with Anthropic's document skills in Cowork and encode one recurring workflow as a private skill this quarter; compounding starts when the second workflow reuses patterns from the first. If you want the workforce without the assembly: managed platforms bundle curated skills, controlled execution, and the security perimeter, the trade O-mega and its peers offer against the self-assembled stack this guide equips you to build.
The through-line across seven months of tracking this ecosystem: the ranking authority changed once, the top of the leaderboard rotated twice, an attack wave arrived, a community hack became a platform primitive, a second superpower's enterprise vendor showed up, and the addressable market quietly shifted from developers to everyone. Any top-10 list here is a timestamped reading, and this guide now says so on its face, because pretending otherwise is how rankings rot. The durable skill remains the meta one: check the leaderboard, read the SKILL.md, measure the before-and-after, and keep your own procedures encoded and versioned. Agents are becoming the workforce; skills are the training program. Run it deliberately, and check back for the next changelog.
Written by Yuma Heymans (@yumahey), founder and CEO of O-mega and co-founder of HeroHunt.ai, who spends his weeks inside the skill economy this guide measures: curating O-mega's public skill library and building the private-skill infrastructure its agent workforces run on.
This guide is a tracked ranking: every install count and star count was read from live sources on August 5, 2026, the model names and prices were re-verified on September 8, 2026, and the four-week deltas come from our own July 8 observations. These numbers move fast; treat any figure here as a dated reading and verify before purchasing or architecture decisions.