Research: phoenix-evals
Cached research evidence for phoenix-evals (not authority).
Purpose
Section titled “Purpose”From Arize-ai/phoenix: skill to build and run evaluators for AI/LLM applications (RAG, agents) using Phoenix evals package. Includes relevance, hallucination, QA, custom evaluators and experiment tracking.
Harness Coverage
Section titled “Harness Coverage”Target agents: antigravity, claude-code, codex, crush, cursor, gemini-cli, github-copilot, grok, opencode.
Trust And Risks
Section titled “Trust And Risks”trust_tier=needs-inspection; status=inspect-then-install; provenance=verified-install-command; Arize official (phoenix repo 10k+ stars, MIT); open source evals platform; inspect for which models used as judges and data exported.
Install Prerequisites
Section titled “Install Prerequisites”Install: npx skills add Arize-ai/phoenix or via skills selectors; status=inspect-then-install; selector=named; policy=Inspect source, hooks, scripts, credentials, and dedupe before install.
Upstream Maintainer
Section titled “Upstream Maintainer”Arize-ai/phoenix (official Arize)
Comparable Alternatives
Section titled “Comparable Alternatives”arize-evaluator, eval-driven-dev, agentic-eval; Phoenix tracing sibling phoenix-tracing.
> Evidence synthesized from public web sources (GitHub repos, official docs, skill registries); confidence reflects source reputation and public signals only. Not an endorsement.
