Research: agentic-eval

Cached research evidence for agentic-eval (not authority).

Back to catalog page

awesome-copilot skill for evaluating agentic (tool-using, multi-step) AI systems: success criteria, trajectory eval, tool use correctness.

Target agents: antigravity, claude-code, codex, crush, cursor, gemini-cli, github-copilot, grok, opencode.

trust_tier=needs-inspection; status=inspect-then-install; provenance=verified-install-command; GitHub curated; pairs well with phoenix/arize eval skills; abstract so customize metrics.

Install: subset of github/awesome-copilot bundle; status=inspect-then-install; selector=named; policy=Inspect source, hooks, scripts, credentials, and dedupe before install.

github/awesome-copilot (GitHub curated)

arize-evaluator, phoenix-evals, eval-driven-dev; general LLM-as-judge skills.

> Evidence synthesized from public web sources (GitHub repos, official docs, skill registries); confidence reflects source reputation and public signals only. Not an endorsement.