vibehacker

LLM Evaluation

Skill

Set up evals for LLM apps with automated metrics, human feedback, and benchmarks instead of vibes alone.

by wshobsonAgents & AutomationData & Analytics Listed Jul 24, 2025

Clone or download it from github.com, then drop the folder into ~/.claude/skills/ (Claude Code) or your agent's skills directory.

Product badge

Share this product's name and rating in your README or on your website.

LLM Evaluation: rating on VibeHacker

Updates may be delayed by image caching.

LLM Evaluation cover

About LLM Evaluation

From the wshobson/agents plugin collection.

Shipping an AI feature is easy; knowing whether it got better or worse is the hard part. This skill helps an agent design an evaluation setup for an LLM application, mixing automated metrics, human review, and benchmark runs, so prompt or model changes come with evidence.

Used LLM Evaluation?

Log in to write a review.

No reviews yet

Used it? Write the first review.

Similar skills

View all