把工作流变成可研究、可评测、可发布的乔木 Agent Skill | Turn workflows into researched, tested, release-ready agent skills.
-
Updated
Aug 4, 2026 - Python
把工作流变成可研究、可评测、可发布的乔木 Agent Skill | Turn workflows into researched, tested, release-ready agent skills.
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
Open-source self-hosted web tool for evaluating Agent Skills with rubric scores, Deep Review, and improvement suggestions.
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
The skill OS for Codex, Claude Code, and Gemini CLI. One pool, one router, one feedback loop — across all three hosts. Per-turn semantic top-K with dynamic context sizing, session-end self-evaluation, and evidence-blended re-ranking that gets better the more you use it.
OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
Your AI agent skill doctor - 5-dimension scoring + security gate + leaderboard
A complete, open-source guide to Agent Skill design: from cognition to production. 28 chapters + 6 appendices.
Evidence-first evaluation and creation for Agent Skills.
opencode 技能工厂 · Create, evaluate & optimize agent skills via DB physical-evidence eval engine. Cross-platform packaging (trae/claude/generic) + governance. 用 opencode.db 物理证据消除 LLM 自评偏差。
MCP server for Claude Code: Anna's Archive search/download + Gemini methodology extraction → audited Claude Code SKILL.md. One tool call, end-to-end.
判断一个skill适不适合你用的skill——啥好用。它可以根据你的需求,把候选 Skill 的用途、安装代价、风险和需求匹配度做成清晰易懂的 HTML 说明书,提升你挑选skill的效率,省点时间和token。
Convert methodology books from Anna's Archive into Claude Code skills using the Model Context Protocol.
Triage-trainer:从零为您的个人助手构建定制化的导诊 Skill,赋予精准的就诊科室推荐能力
Claim-first repo 评测框架。in target repo + claim map → out bilingual verdict dossier + all-evals dashboard
Codex skill review / skill 检验 tool for evaluating necessity, baseline lift, rubric score, and keep/simplify/merge/rewrite/delete/patch decisions.
Evidence-backed release audits for Agent Skills, covering discovery, behavioral uplift, safety, and package integrity. 为 Agent Skill 做有证据的上线审查,验证能否被正确触发、真实提升结果,并安全安装与复现。
Open-source framework for reproducible A/B evaluation of agent skills and instructions
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.
Add a description, image, and links to the skill-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the skill-evaluation topic, visit your repo's landing page and select "manage topics."