Research: arize-evaluator

Cached research evidence for arize-evaluator (not authority).

Back to catalog page

Arize skill (from Arize-ai/arize-skills or phoenix) for building/running evaluators for LLM apps (RAG relevance, answer relevance, hallucination, groundedness etc). Integrates with Phoenix observability.

Target agents: antigravity, claude-code, codex, crush, cursor, gemini-cli, github-copilot, grok, opencode.

trust_tier=needs-inspection; status=inspect-then-install; provenance=verified-install-command; source=Arize AI (phoenix/arize-skills repos); 10k+ stars on phoenix; open source (MIT); pairs tracing + evals; inspect for evaluator model costs and data sent to Arize/Phoenix.

Install: subset of github/awesome-copilot or direct arize-ai/arize-skills; or npx skills add Arize-ai/phoenix; status=inspect-then-install; selector=named; policy=Inspect source, hooks, scripts, credentials, and dedupe before install.

Arize-ai/arize-skills and Arize-ai/phoenix (official Arize)

phoenix-evals, agentic-eval, eval-driven-dev; LangSmith evals or other LLM eval harness skills.

> Evidence synthesized from public web sources (GitHub repos, official docs, skill registries); confidence reflects source reputation and public signals only. Not an endorsement.