Research: phoenix-evals

Cached research evidence for phoenix-evals (not authority).

Back to catalog page

From Arize-ai/phoenix: skill to build and run evaluators for AI/LLM applications (RAG, agents) using Phoenix evals package. Includes relevance, hallucination, QA, custom evaluators and experiment tracking.

Target agents: antigravity, claude-code, codex, crush, cursor, gemini-cli, github-copilot, grok, opencode.

trust_tier=needs-inspection; status=inspect-then-install; provenance=verified-install-command; Arize official (phoenix repo 10k+ stars, MIT); open source evals platform; inspect for which models used as judges and data exported.

Install: npx skills add Arize-ai/phoenix or via skills selectors; status=inspect-then-install; selector=named; policy=Inspect source, hooks, scripts, credentials, and dedupe before install.

Arize-ai/phoenix (official Arize)

arize-evaluator, eval-driven-dev, agentic-eval; Phoenix tracing sibling phoenix-tracing.

> Evidence synthesized from public web sources (GitHub repos, official docs, skill registries); confidence reflects source reputation and public signals only. Not an endorsement.