Multi-vendor LLM evaluation

Choose your model vendor with evidence, not marketing.

Run every major model — Anthropic, OpenAI, Google, and a dozen others, by API key or the Claude / ChatGPT subscription you already pay for — against your own policy, your own tickets, and a scorecard weighted to what your business cares about.

Self-hosted by design. Runs inside your own environment, on your own API keys — your tickets and your policy never leave it.

Works with Anthropic · OpenAI · Google · Meta · xAI · Z.ai · +14 more
19
providers supported
60
reference tickets (20 held-out), 12 adversarial

Real run output, not a mockup

Watch the evaluation work, on one real ticket

Reference results are being refreshed at matched reasoning effort. This section fills in automatically, with one real ticket and every candidate's scored reply, when the run completes and passes validation.

The gap in procurement today

What your vendor's benchmark can't tell you

Four questions that come up before every model contract gets signed — and that a vendor-supplied leaderboard was never built to answer.

“Vendor benchmarks are marketing.”

None of them are scored against your policy, your tickets, or your definition of a good answer.

“A model swap can silently break production.”

Quality drops. Nobody notices until a customer complains — there was no gate that would have caught it first.

“Nobody's testing for the failure that matters.”

A model that's cheaper but leaks its system prompt isn't a saving — it's a liability with a lower price tag.

“$/1M tokens isn't a budget line.”

Nobody's translated the rate card into $/month at your actual volume before the contract is signed.

What it's actually for

Built around four decisions, not one demo

Vendor selection & procurement

The model you're being pitched, against every other vendor's flagship, on the same tickets — ranked by quality, latency, and cost weighted to your workload.

scorecard.weights: quality / latency / cost

Safety & compliance red-teaming

A prompt-injection, a system-prompt leak attempt, and an over-refusal trap ride along in every run. A violation is a hard gate, never averaged into the score.

critical_violation → hard gate, not a discount

Cost governance at scale

Set your monthly ticket volume once; every candidate's rate card becomes a $/month figure finance can actually act on.

scorecard.monthly_volume → $/month projection

Regression protection on model swaps

Run the new model version against the same tasks and a saved baseline; the gate fails the deploy if quality drops or violations rise — enforceable in CI.

--baseline / --regression-threshold → exit 1

Case study

The reference workload: support-ticket triage

Ships with a support inbox for Northwind Cloud — a synthetic reference company and a real refund/priority policy. Results are being refreshed at matched reasoning effort; the scorecard appears here automatically once a run completes and passes validation.

Why the number is trustworthy

Four rules the scoring never breaks

i.

Graded exactly, where exactness is possible

Routing, priority, and the refund, escalation, and retention actions the agent declares are checked against gold labels — no judge opinion where there's a right answer.

ii.

Judged only where judgment is required

Policy adherence, resolution, and tone are scored by an LLM judge against the exact policy the candidate saw.

iii.

Every run reports how it was judged

Judge coverage, reasoning effort, model IDs and CLI versions are recorded with every run, and results with too many unjudged answers are never published.

iv.

A violation is pass/fail, never a discount

critical_violation is a hard gate — a forbidden refund promise or a leaked prompt fails outright, never smoothed into an average.

Provider coverage

No lock-in built into the tool that measures lock-in

Frontier Anthropic·OpenAI·Google·Z.ai·xAI·Meta
Open-weight Moonshot (Kimi)·Alibaba (Qwen)·DeepSeek·NVIDIA (Nemotron)·Meta (Llama)
Aggregators OpenRouter·Together·Groq·Fireworks·DeepInfra
Subscriptions Claude Code CLI·Codex CLI·Gemini CLI — no per-token key required
Local & neutral Ollama (self-hosted)·Nous Research Hermes (no stake in the outcome, used as judge)

Most of these speak the same OpenAI-compatible wire protocol underneath — a new adapter is typically about eight lines of code, so coverage grows without a rewrite each time a vendor ships something new.

Run it against your own policy.

Clone it, point it at your own tickets and your own rulebook, or start from the Northwind Cloud reference workload above.

View on GitHub →
git clone https://github.com/gpatwa/multi-agent-eval cd multi-agent-eval python main.py --config config.triage.yaml \ --tasks tasks.triage.yaml --out results/