open source · local-first · MCP

Evals for AI systems,
as an MCP.

retriEVAL is a local-first LLM-evaluation server you connect to Claude, Cursor, or CI. It scores RAG by stage, telling retriever failures apart from generator failures, authors metrics from plain language, and tracks every run with charts and history. Bring your own model: run free on Ollama, or use your Claude, OpenAI, or any LLM key.

The loop: score your AI's outputs, track them, improve
01 · Collect
Outputs
from your app or RAG
02 · Score
By stage
retriever vs generator
03 · Track
Charts + history
on your dashboard
04 · Improve
Iterate
gate quality over time

What makes it different

The metrics aren't new; DeepEval and Ragas cover those. The shape is: a standalone, local-first MCP an agent can call mid-workflow, with RAG scored by stage.

stage-separated

Retriever vs generator

contextual_recall flags retrieval misses; faithfulness flags hallucinations. Know which half to fix, not just "the answer is bad."

in your editor

Runs in Cursor

Build custom skills and QE pipelines around it to keep improving evals on your AI products. retriEVAL is the scoring engine your workflows reach for.

bring your own model

Your model, your key

Run free on a local Ollama model, or plug in your own Claude, OpenAI, or any OpenAI-compatible LLM key. No platform account, no data leaving your box.

history

Charts + run history

Every run is saved (file or Supabase). Trend by model, compare runs, drill into the case that regressed, on a dashboard you own.

Or just run evals by chatting

No pipeline required. Connect retriEVAL in Claude or another chat UI, toggle it on, and a manual tester can score outputs, author metrics, and pull up charts in plain language.

Toggle it on, like any connector
Connectors
Web search
retriEVAL
Google Drive

Illustration. Uses the hosted server, since chat UIs connect to remote MCPs.

Evals appear right in the chat
Chat
Score this answer against the refund policy: faithfulness and answer relevancy.
Ran 3 metrics on 12 cases  1 FAILING
faithfulness.92PASS
answer_relevancy.88PASS
contextual_recall.58FAIL
plot_metric_trend · faithfulnesslast 6 runs
1.0 0.5 0.0 0.70

Try it live

Run a real eval right here, no signup. Pick a sample (or paste your own input, context, and output), choose a free judge model, and watch retriEVAL score it. Nothing is saved.

1 · Choose a sample
2 · Metric
3 · Judge model

Runs on a free model · nothing is saved · limited per day

Scoring with the judge model…

Quickstart

Run it locally over stdio, or connect the hosted server from anywhere.

Local (Claude Desktop / Cursor)
# install
pip install retrieval-mcp
export ANTHROPIC_API_KEY=sk-ant-...

# add to claude_desktop_config.json
{
  "mcpServers": {
    "retrieval": {
      "command": "retrieval-mcp"
    }
  }
}
Hosted (any client, incl. Claude.ai)
# Claude.ai → Settings → Connectors
Add custom connector:
  URL: https://YOUR-APP.up.railway.app/mcp
  Auth: Bearer <token>

# then just ask Claude
"Load my golden set and run
 faithfulness + answer_relevancy."

Pick your judge: free local Ollama (RETRIEVAL_JUDGE_BACKEND=ollama), your Anthropic key, your OpenAI key, or any OpenAI-compatible endpoint via OPENAI_BASE_URL. Full deploy guide (Supabase + Railway + Vercel) is in the repo's DEPLOY.md.

See it in action

Walkthroughs and demos: connecting retriEVAL, authoring a metric, and building QE pipelines in Cursor.

▶ video coming soon
Quickstart: connect + first eval
▶ video coming soon
Build a QE pipeline in Cursor

YouTube demos will be embedded here.

How it compares

Honest positioning: not "better metrics," but a different shape: standalone, local, and usable as a tool inside an agent.

 retriEVALHosted eval platformsEval libraries
Works as an MCP tool inside an agentyesplatform-gatedno
Runs local, no accountyesnoyes
Bring your own model / key (Ollama, Claude, OpenAI, …)yeslimitedyes
RAG scored by stage, built inyesvariesmanual
Own dashboard + historyyesyesno