Research: eval-driven-dev

Cached research evidence for eval-driven-dev (not authority).

Back to catalog page

awesome-copilot skill for eval-driven development: build automated evaluation pipelines that run the AI app end-to-end with real inputs, score outputs with evaluators, produce pass/fail via test frameworks (pixie etc).

Target agents: antigravity, claude-code, codex, crush, cursor, gemini-cli, github-copilot, grok, opencode.

trust_tier=needs-inspection; status=inspect-then-install; provenance=verified-install-command; GitHub curated; powerful for regression but involves executing the app under test and scoring (potential cost, flakiness, data needs); inspect eval harness integration and data sources.

Install: npx skills add github/awesome-copilot --skill eval-driven-dev; status=inspect-then-install; selector=named; policy=Inspect source, hooks, scripts, credentials, and dedupe before install.

github/awesome-copilot (GitHub curated)

arize-evaluator, phoenix-evals, agentic-eval; general property-based or LLM eval skills.

> Evidence synthesized from public web sources (GitHub repos, official docs, skill registries); confidence reflects source reputation and public signals only. Not an endorsement.