Research: agentic-eval
Cached research evidence for agentic-eval (not authority).
Purpose
Section titled “Purpose”awesome-copilot skill for evaluating agentic (tool-using, multi-step) AI systems: success criteria, trajectory eval, tool use correctness.
Harness Coverage
Section titled “Harness Coverage”Target agents: antigravity, claude-code, codex, crush, cursor, gemini-cli, github-copilot, grok, opencode.
Trust And Risks
Section titled “Trust And Risks”trust_tier=needs-inspection; status=inspect-then-install; provenance=verified-install-command; GitHub curated; pairs well with phoenix/arize eval skills; abstract so customize metrics.
Install Prerequisites
Section titled “Install Prerequisites”Install: subset of github/awesome-copilot bundle; status=inspect-then-install; selector=named; policy=Inspect source, hooks, scripts, credentials, and dedupe before install.
Upstream Maintainer
Section titled “Upstream Maintainer”github/awesome-copilot (GitHub curated)
Comparable Alternatives
Section titled “Comparable Alternatives”arize-evaluator, phoenix-evals, eval-driven-dev; general LLM-as-judge skills.
> Evidence synthesized from public web sources (GitHub repos, official docs, skill registries); confidence reflects source reputation and public signals only. Not an endorsement.
