Certified takeover of LLM judgment calls

Most of your LLM judgment calls don’t need an LLM.

Sundr learns the classifications, routing and rubric scores your LLM makes today. It splits each decision into narrow questions for decision models like Jev and Clef, then answers the calls it can, inside an error budget it certifies on your own traffic. Everything else still goes to your LLM.

99.2%agreement, phishing benchmark
80%of toxicity calls certified at 1%
$0.16per 1,000 decisions
One call email #48,112 Narrow questions Link ≠ sender domain 0.97 Asks to sign in 0.92 Urgent pressure 0.71 From a known vendor 0.06 Free-hosting link 0.88 Verdict phishing p = 0.994 certified ≤ 1% error
Illustrative: one email and five of the questions Sundr asks about it.
Drop-in for OpenAI and Anthropic APIs Decision models TypeSafe Jev · Cloudflare Clef Imports Langfuse · Helicone · LangSmith · Braintrust
01 · Results

Ask narrow questions, and a decision model agrees with the judge.

Asked one broad question, a decision model is a weak judge. Asked a dozen narrow ones, combined by a small calibrated model, it matches the reference far more often, and knows when it doesn’t.

+6.8 pts
average agreement gained by splitting the question, on the six benchmarks that fit
4 of 1,609
certified test splits came in significantly over budget (0.25%; the guarantee allows 5%)
7
public benchmarks, 100 random train / calibrate / test splits each
Agreement with the reference label
One direct question vs. Sundr’s split, Clef-flash decision model, held-out test calls
60%70%80%90%100%Phishing emailsjev-phishing-bench, 1,934 emails99.2%Contract audit rightsLegalBench CUAD, 1,216 clauses98.6%Definitions in opinionsLegalBench, 1,337 sentences96.6%Overruling languageLegalBench, 2,394 sentences96.3%Banking intents (77 classes)BANKING77, 13,083 messages95.3%Toxic commentsCivil Comments, 20,000 comments94.6%Chatbot preferenceMT-Bench human votes, 3,35566.2%
Chatbot preference is the honest miss: human votes disagree with each other too often to certify. Benchmarks run Oct 9, 2026
View as table
BenchmarkOne questionSundrCertified at 1%Certified at 2%
Phishing emails
jev-phishing-bench, 1,934 emails
77.4%99.2%74%99%
Contract audit rights
LegalBench CUAD, 1,216 clauses
94.6%98.6%37%91%
Definitions in opinions
LegalBench, 1,337 sentences
94.3%96.6%2%47%
Overruling language
LegalBench, 2,394 sentences
92.5%96.3%12%69%
Banking intents (77 classes)
BANKING77, 13,083 messages
95.4%95.3%78%91%
Toxic comments
Civil Comments, 20,000 comments
85.9%94.6%80%87%
Chatbot preference
MT-Bench human votes, 3,355
64.1%66.2%0%0%
02 · How it works

Four steps, and your users never notice the first three.

Nothing changes for your product until a threshold is certified on your own traffic. After that, the calls Sundr can answer stop costing you an LLM call.

01

Shadow

Point your judgment calls at Sundr with one base URL. Every call still goes to your LLM; Sundr records the verdict.

02

Split

Claude drafts 30–60 narrow questions for the task. A linter strips the patterns decision models get wrong: negation, counting, “unsure” options.

03

Certify

A small combiner learns your judge. Learn then Test picks a threshold whose error on held-out calls stays under your budget with 95% confidence.

04

Take over

Calls above the threshold are answered in about half a second. An audit slice keeps checking, and if the error drifts, traffic hands back to your LLM automatically.

One line of integration.

Keep your client, your prompts and your response schema. Sundr answers in exactly the shape your code already parses, including forced tool calls. Calls without a task header, and streaming calls, pass straight through.

The proxy is in design-partner preview. The audit is open to everyone.

from openai import OpenAI

client = OpenAI(
    base_url="https://proxy.sundr.ai/openai/v1",
    default_headers={
        "X-Sundr-Key": SUNDR_KEY,
        "X-Sundr-Task": "ticket-router",
    },
)
# Same call, same JSON schema. Sundr answers what it can certify;
# everything else is forwarded to your model untouched.
03 · The guarantee

A threshold tuned to 1% misses it half the time. Ours doesn’t.

The usual way to automate with a confidence score is to pick the loosest threshold whose error on a labelled sample is under budget. It stops exactly where luck made the sample look best, so on new traffic it overshoots: one independent evaluation tuned for 1% and measured 1.6%.

Sundr certifies instead. It walks thresholds from strict to loose with a binomial test at each step (Learn then Test) and stops at the first one it can’t vouch for. The bound holds whatever the model’s calibration, as long as new traffic resembles the sample, and in production a live audit slice checks that it still does.

Error ≤ your budget on the calls Sundr answers.95% confidence at certification, re-checked continuously, rolled back automatically.

The price is data: about 120 ÷ budget calibration calls, so roughly 12,000 for a 1% budget. Below that, the audit tells you which budget it can promise.

Trials where true error exceeded a 1% budget
300 simulated deployments per calibration size, slightly overconfident model
0%20%40%60%56%0%12062%0%40055%1%1,00058%2%3,00050%3%12,000calibration calls the threshold was tuned on
View as table
Calibration callsNaive over budgetSundr over budget
12056%0%
40062%0%
1,00055%1%
3,00058%2%
12,00050%3%
04 · Pricing

You pay from what Sundr saves you.

Savings are measured call by call against what your LLM would have charged, so the invoice is a share of a number you can audit.

Savings Audit
Free
on up to 100,000 logged calls
  • Certified coverage at 0.5–5% budgets
  • Monthly savings estimate
  • The questions Sundr would ask
  • Logs deleted when the report is ready
Run an audit
Enterprise
Annual
in your VPC on open-weight models
  • Self-hosted proxy and monitor
  • Zero data retention
  • SSO and certificate reports
  • Custom error budgets per route
Talk to us
05 · Free audit

See how much of last week you could have handed over.

Upload a log export of your judgment calls. You’ll get the share Sundr could answer at each error budget, certified on your own data, and what it would have saved.

  • Exports from Langfuse, Helicone, LangSmith, Braintrust, or raw OpenAI and Anthropic logs
  • About 30,000 calls certifies a 1% budget; 300 is the minimum
  • Encrypted at rest, deleted as soon as the report is written

Or let your coding agent find the calls and export them

Field names and volume

Dotted paths reach into nested JSON, e.g. request.messages.

Your logs are sent only to the decision-model and Claude APIs that run the audit. Privacy

06 · Questions
What counts as a judgment call?

Any LLM call whose answer is one of a fixed set: a label, a route, a yes or no, a 1–5 score. Ticket triage, moderation, intent detection, policy checks, LLM-as-judge evals and agent “is this done?” checks all qualify. Open-ended generation doesn’t.

Which models answer the calls?

Decision models: TypeSafe’s Jev and Cloudflare’s Clef today, routed per task, with open-weight models for self-hosted deployments. They return calibrated probabilities rather than text, in well under a second, for a fraction of a cent per thousand calls.

Is agreement with my LLM the same as accuracy?

No. By default Sundr certifies agreement with the model you use today, because its past verdicts are free labels. Send a few hundred human-labelled calls and the report states your LLM’s own error rate beside Sundr’s, or certify against the human labels directly.

How much traffic do I need?

Certifying a budget reliably takes about 120 ÷ budget held-out calls: 12,000 for 1%, 6,000 for 2%. With less, the audit shows the tightest budget it can honestly promise.

What happens if accuracy drifts?

A slice of answered calls, weighted toward the least certain, still goes to your LLM. Sundr estimates live error from it with a valid confidence interval. If the bound crosses your budget, the task switches back to your LLM automatically and you’re alerted.

What happens to my data?

Audit uploads are encrypted at rest and deleted as soon as the report is written; reports are deleted after 30 days. See the privacy notice for the providers involved.