Skills registry

The leaderboard is public. So are the evals.

Skills ranked by measured lift — not download counts. Every score is backed by evals you can inspect, case by case.

← Why rank by lift? See how SkillScore worksBuild your own — author, eval, govern ↓
AllCodingData & SQLSupportSecurityResearch & webOps & infra✓ Benchmarked only
Top skillsby SkillScore · all categories
lift measured on gemini-3.5-flash · 21+ cases per skill · re-run any skill on your own model with the open runner
#SkillSkillScoreLiftTurns ΔTokens ΔRatingTeams
1st
sql-result-contract ✓ benchmarked
data & sql
91+66%−5%+2% 4.9
2nd
secret-scanner ✓ benchmarked
security
89+52%−4%−2% 4.8
3rd
commit-conventions ✓ benchmarked
coding
87+35%−12%−8% 4.8
04
error-triage-protocol ✓ benchmarked
support
86+28%−14%−5% 4.7
05
postgres-migration-guard ✓ benchmarked
ops & infra
85+40%−9%−4% 4.7
06
source-citation-guard ✓ benchmarked
research & web
84+44%−10%+6% 4.7
07
query-plan-reviewer ✓ benchmarked
data & sql
83+38%−8%−2% 4.7
08
pr-review-checklist ✓ benchmarked
coding
82+30%−9%−4% 4.6
09
incident-postmortem ✓ benchmarked
ops & infra
82+34%−13%+7% 4.6
10
refund-policy-writer ✓ benchmarked
support
81+30%−11%−4% 4.6
11
prompt-injection-shield ✓ benchmarked
security
81+33%−2%−1% 4.6
12
json-strict-output ✓ benchmarked
coding
80+22%−6%−3% 4.5
13
search-query-planner ✓ benchmarked
research & web
80+29%−8%+4% 4.5
14
pii-redactor ✓ benchmarked
security
79+31%0%−3% 4.5
15
terraform-plan-review ✓ benchmarked
ops & infra
79+32%−7%−3% 4.5
16
csv-schema-inference ✓ benchmarked
data & sql
78+27%−6%+4% 4.4
17
ticket-summarizer ✓ benchmarked
support
78+25%−16%−7% 4.5
18
test-case-generator ✓ benchmarked
coding
77+24%−7%+3% 4.4
19
schema-linter ✓ benchmarked
data & sql
77+26%−7%−2% 4.4
20
page-summary-extractor ✓ benchmarked
research & web
77+23%−9%−5% 4.4
21
dependency-audit ✓ benchmarked
security
76+25%−6%+5% 4.3
22
k8s-manifest-lint ✓ benchmarked
ops & infra
76+24%−5%−2% 4.3
23
support-triage ✓ benchmarked
support
75+21%−18%−6% 4.6
24
rag-chunk-strategy ✓ benchmarked
research & web
75+20%−4%−8% 4.3
25
dbt-model-conventions ✓ benchmarked
data & sql
74+18%−5%−3% 4.2
26
authz-boundary-check ✓ benchmarked
security
74+23%−5%−1% 4.3
27
tone-consistency-guard ✓ benchmarked
support
73+17%−3%+2% 4.2
28
web-scraper-toolkit caution
research & web
73+22%−3%+5% 4.2
29
changelog-writer ✓ benchmarked
coding
72+19%−4%−9% 4.1
30
runbook-writer ✓ benchmarked
ops & infra
71+16%−6%−11% 4
1st
sql-result-contract ✓ benchmarked
data & sql · ★ 4.9
91
+66%
2nd
secret-scanner ✓ benchmarked
security · ★ 4.8
89
+52%
3rd
commit-conventions ✓ benchmarked
coding · ★ 4.8
87
+35%
04
error-triage-protocol ✓ benchmarked
support · ★ 4.7
86
+28%
05
postgres-migration-guard ✓ benchmarked
ops & infra · ★ 4.7
85
+40%
06
source-citation-guard ✓ benchmarked
research & web · ★ 4.7
84
+44%
07
query-plan-reviewer ✓ benchmarked
data & sql · ★ 4.7
83
+38%
08
pr-review-checklist ✓ benchmarked
coding · ★ 4.6
82
+30%
09
incident-postmortem ✓ benchmarked
ops & infra · ★ 4.6
82
+34%
10
refund-policy-writer ✓ benchmarked
support · ★ 4.6
81
+30%
11
prompt-injection-shield ✓ benchmarked
security · ★ 4.6
81
+33%
12
json-strict-output ✓ benchmarked
coding · ★ 4.5
80
+22%
13
search-query-planner ✓ benchmarked
research & web · ★ 4.5
80
+29%
14
pii-redactor ✓ benchmarked
security · ★ 4.5
79
+31%
15
terraform-plan-review ✓ benchmarked
ops & infra · ★ 4.5
79
+32%
16
csv-schema-inference ✓ benchmarked
data & sql · ★ 4.4
78
+27%
17
ticket-summarizer ✓ benchmarked
support · ★ 4.5
78
+25%
18
test-case-generator ✓ benchmarked
coding · ★ 4.4
77
+24%
19
schema-linter ✓ benchmarked
data & sql · ★ 4.4
77
+26%
20
page-summary-extractor ✓ benchmarked
research & web · ★ 4.4
77
+23%
21
dependency-audit ✓ benchmarked
security · ★ 4.3
76
+25%
22
k8s-manifest-lint ✓ benchmarked
ops & infra · ★ 4.3
76
+24%
23
support-triage ✓ benchmarked
support · ★ 4.6
75
+21%
24
rag-chunk-strategy ✓ benchmarked
research & web · ★ 4.3
75
+20%
25
dbt-model-conventions ✓ benchmarked
data & sql · ★ 4.2
74
+18%
26
authz-boundary-check ✓ benchmarked
security · ★ 4.3
74
+23%
27
tone-consistency-guard ✓ benchmarked
support · ★ 4.2
73
+17%
28
web-scraper-toolkit caution
research & web · ★ 4.2
73
+22%
29
changelog-writer ✓ benchmarked
coding · ★ 4.1
72
+19%
30
runbook-writer ✓ benchmarked
ops & infra · ★ 4
71
+16%
illustrative sampleSee the live ranking →
Browse the catalog40k+ skills · card view
sql-result-contract ✓ benchmarked
data & sql
91SkillScore
Validate query results against a typed contract before they return.
+66% lift4.9view evals →
secret-scanner ✓ benchmarked
security
89SkillScore
Catch credentials in code and config before they reach a commit.
+52% lift4.8view evals →
commit-conventions ✓ benchmarked
coding
87SkillScore
Enforce your repo’s commit format — consistent, parseable messages.
+35% lift4.8view evals →
error-triage-protocol ✓ benchmarked
support
86SkillScore
Triage a failing run to the right owner with a consistent, reviewable protocol.
+28% lift4.7view evals →
postgres-migration-guard ✓ benchmarked
ops & infra
85SkillScore
Catch unsafe schema migrations before they hit production.
+40% lift4.7view evals →
source-citation-guard ✓ benchmarked
research & web
84SkillScore
Every claim carries a source that resolves — or it doesn’t ship.
+44% lift4.7view evals →
query-plan-reviewer ✓ benchmarked
data & sql
83SkillScore
Read the plan before the query ships — catch the scan that has no index behind it.
+38% lift4.7view evals →
pr-review-checklist ✓ benchmarked
coding
82SkillScore
Review a diff against the checks your team actually cares about.
+30% lift4.6view evals →
incident-postmortem ✓ benchmarked
ops & infra
82SkillScore
Write the postmortem from the timeline, with causes separated from symptoms.
+34% lift4.6view evals →
refund-policy-writer ✓ benchmarked
support
81SkillScore
Answer refund requests against your actual policy, edge cases included.
+30% lift4.6view evals →
prompt-injection-shield ✓ benchmarked
security
81SkillScore
Detect and block prompt-injection attempts in tool inputs.
+33% lift4.6view evals →
json-strict-output ✓ benchmarked
coding
80SkillScore
Lint model output against a strict JSON schema before anything downstream reads it.
+22% lift4.5view evals →
search-query-planner ✓ benchmarked
research & web
80SkillScore
Plan the searches before running them, so one angle doesn’t stand in for all of them.
+29% lift4.5view evals →
pii-redactor ✓ benchmarked
security
79SkillScore
Strip PII from traces before they leave your workspace.
+31% lift4.5view evals →
terraform-plan-review ✓ benchmarked
ops & infra
79SkillScore
Read a plan diff for the destroy you didn’t mean to approve.
+32% lift4.5view evals →
csv-schema-inference ✓ benchmarked
data & sql
78SkillScore
Infer and pin a column schema from messy CSVs instead of guessing per file.
+27% lift4.4view evals →
ticket-summarizer ✓ benchmarked
support
78SkillScore
Compress a long thread into what the next agent needs to act.
+25% lift4.5view evals →
test-case-generator ✓ benchmarked
coding
77SkillScore
Turn a function’s edge cases into runnable tests, not prose about tests.
+24% lift4.4view evals →
schema-linter ✓ benchmarked
data & sql
77SkillScore
Lint output schemas for missing or drifted required fields.
+26% lift4.4view evals →
page-summary-extractor ✓ benchmarked
research & web
77SkillScore
Pull the claim, the number, and the date out of a page — skip the nav.
+23% lift4.4view evals →
dependency-audit ✓ benchmarked
security
76SkillScore
Read an advisory against your lockfile and say whether you’re actually reachable.
+25% lift4.3view evals →
k8s-manifest-lint ✓ benchmarked
ops & infra
76SkillScore
Flag the missing limits, probes, and rollout settings before apply.
+24% lift4.3view evals →
support-triage ✓ benchmarked
support
75SkillScore
Route support conversations to the right resolution path.
+21% lift4.6view evals →
rag-chunk-strategy ✓ benchmarked
research & web
75SkillScore
Chunk on structure instead of character count, then check what retrieval returns.
+20% lift4.3view evals →
dbt-model-conventions ✓ benchmarked
data & sql
74SkillScore
Keep model names, tests, and materializations to one house style.
+18% lift4.2view evals →
authz-boundary-check ✓ benchmarked
security
74SkillScore
Check every new route for the tenancy filter it’s supposed to carry.
+23% lift4.3view evals →
tone-consistency-guard ✓ benchmarked
support
73SkillScore
Hold replies to one voice across a team of agents and humans.
+17% lift4.2view evals →
web-scraper-toolkit caution
research & web
73SkillScore
Structured extraction helpers — listed with a caution note.
+22% lift4.2view evals →
changelog-writer ✓ benchmarked
coding
72SkillScore
Draft release changelogs from your merged PRs.
+19% lift4.1view evals →
runbook-writer ✓ benchmarked
ops & infra
71SkillScore
Turn a fix you just made into steps the next on-call can follow.
+16% lift4view evals →
SkillScore

One number, four signals — and one of them can't be bought.

Benchmark lift at 32% (pass rate with the skill minus without), live pass rate from real runs at 32%, an AI-judged quality review at 16%, and adoption at 20% — which grows logarithmically and is capped by org diversity, so volume from a single workspace can’t buy a rank. A missing signal isn’t scored as zero; the remaining weights rescale.

SkillScore · commit-conventions
Benchmark lift · 32%+35%
Live pass rate · 32%86%
AI rating · 16%A
Adoption · 20%
SkillScore87
no adoption signal yet — the other three weights rescale to 100%
SkillSafety

Nothing lists without clearing the gate.

Every submission runs a three-stage pipeline — static scan, AI security review, content check — and carries its band. Blocked skills never appear.

SkillSafety bands
passedcleared all three stages
cautionlisted with the reason shown
blockednever lists
Skill router

Know which skills your agent actually uses.

smart_route() ranks a shortlist, then compares what you offered against what the agent activated. Under 40% flags menu bloat — and shows which skills to cut.

Offered vs activated · last 7d
Offered12 skills
Activated5 skills
menu bloat42% activation — trim 7
How a skill earns its badge
01Scanned02Security review03Content check04Benchmarked05Band + score published
Build on it

Your team’s skills, not just everyone else’s.

Fork a public skill or write your own, prove it works on your cases, and ship it to your team — on the same rails that rank the public registry. Authoring, evals, versioning, access: skills governance without building the system.

Your eval · refund-policy-writer
Pass rate without54%
Pass rate with84%
Lift+30 pts
workspace-only · illustrative — run it with the open eval runner
From your draft to your team’s router
01Author or forkStart from a benchmarked skill or a blank one. The skill format, linter, and local safety scan are built in.
02Eval before you shipRun the same harness that scores public skills against your own cases — pass rate before and after, and the exact lift.
03Share with your teamPublish to your workspace. Your agents pick it up via smart_route() — and real usage feeds back into the score.
04Governed by defaultThe same safety gate and pinned versions as the public registry. Governance runs on the rails — not as a process you police.
Skills governance

Everything you’d build internally — already here.

Keeping a team’s skills authored, evaluated, safe, versioned, and shared is real work — skills governance. In-house it’s a stack you build and maintain. Here, it’s what the registry already does.

Skills governanceBuild it in-houseOn DecimalAI
Eval harness — with / without runsyou build it✓ the same harness that ranks the registry
Safety review on every versionyou build it✓ SkillSafety gate included
Quality ranking nobody gamesyou build it✓ SkillScore — measured lift
Versioning & pinningyou build it✓ pinned versions built in
Sharing & access controlyou build it✓ workspace publishing
the upkeepweeks of engineering · yours to maintainpip install decimalai · today
Publish your first skill →
FORK & EVAL ON FREE · TEAM PUBLISHING FROM CORE $49
One engine

The engine that ranks skills also versions your whole agent.

Explore Agent versioning →

Adopt the skills that are proven to help.

Browse free — no sign-up. Two lines to instrument your agent.

Browse the registry