
Official Anthropic Agent Skills repo with document and example packs
Set up evals for LLM apps with automated metrics, human feedback, and benchmarks instead of vibes alone.
Clone or download it from github.com, then drop the folder into ~/.claude/skills/ (Claude Code) or your agent's skills directory.
Share this product's name and rating in your README or on your website.
Updates may be delayed by image caching.
From the wshobson/agents plugin collection.
Shipping an AI feature is easy; knowing whether it got better or worse is the hard part. This skill helps an agent design an evaluation setup for an LLM application, mixing automated metrics, human review, and benchmark runs, so prompt or model changes come with evidence.
Used it? Write the first review.

Official Anthropic Agent Skills repo with document and example packs

Official Browserbase skills for cloud browser automation

Production agent skills from Matt Pocock's .agents directory

Official OpenAI Agent Skills for Codex and ChatGPT workflows

Official Claude document skills for Word, PDF, PowerPoint, Excel

Curated collection of 1000+ SKILL.md packs for coding agents