Most of your LLM judgment calls don’t need an LLM.
Sundr learns the classifications, routing and rubric scores your LLM makes today. It splits each decision into narrow questions for decision models like Jev and Clef, then answers the calls it can, inside an error budget it certifies on your own traffic. Everything else still goes to your LLM.
Ask narrow questions, and a decision model agrees with the judge.
Asked one broad question, a decision model is a weak judge. Asked a dozen narrow ones, combined by a small calibrated model, it matches the reference far more often, and knows when it doesn’t.
View as table
| Benchmark | One question | Sundr | Certified at 1% | Certified at 2% |
|---|---|---|---|---|
| Phishing emails jev-phishing-bench, 1,934 emails | 77.4% | 99.2% | 74% | 99% |
| Contract audit rights LegalBench CUAD, 1,216 clauses | 94.6% | 98.6% | 37% | 91% |
| Definitions in opinions LegalBench, 1,337 sentences | 94.3% | 96.6% | 2% | 47% |
| Overruling language LegalBench, 2,394 sentences | 92.5% | 96.3% | 12% | 69% |
| Banking intents (77 classes) BANKING77, 13,083 messages | 95.4% | 95.3% | 78% | 91% |
| Toxic comments Civil Comments, 20,000 comments | 85.9% | 94.6% | 80% | 87% |
| Chatbot preference MT-Bench human votes, 3,355 | 64.1% | 66.2% | 0% | 0% |
Four steps, and your users never notice the first three.
Nothing changes for your product until a threshold is certified on your own traffic. After that, the calls Sundr can answer stop costing you an LLM call.
Shadow
Point your judgment calls at Sundr with one base URL. Every call still goes to your LLM; Sundr records the verdict.
Split
Claude drafts 30–60 narrow questions for the task. A linter strips the patterns decision models get wrong: negation, counting, “unsure” options.
Certify
A small combiner learns your judge. Learn then Test picks a threshold whose error on held-out calls stays under your budget with 95% confidence.
Take over
Calls above the threshold are answered in about half a second. An audit slice keeps checking, and if the error drifts, traffic hands back to your LLM automatically.
One line of integration.
Keep your client, your prompts and your response schema. Sundr answers in exactly the shape your code already parses, including forced tool calls. Calls without a task header, and streaming calls, pass straight through.
The proxy is in design-partner preview. The audit is open to everyone.
from openai import OpenAI client = OpenAI( base_url="https://proxy.sundr.ai/openai/v1", default_headers={ "X-Sundr-Key": SUNDR_KEY, "X-Sundr-Task": "ticket-router", }, ) # Same call, same JSON schema. Sundr answers what it can certify; # everything else is forwarded to your model untouched.
A threshold tuned to 1% misses it half the time. Ours doesn’t.
The usual way to automate with a confidence score is to pick the loosest threshold whose error on a labelled sample is under budget. It stops exactly where luck made the sample look best, so on new traffic it overshoots: one independent evaluation tuned for 1% and measured 1.6%.
Sundr certifies instead. It walks thresholds from strict to loose with a binomial test at each step (Learn then Test) and stops at the first one it can’t vouch for. The bound holds whatever the model’s calibration, as long as new traffic resembles the sample, and in production a live audit slice checks that it still does.
The price is data: about 120 ÷ budget calibration calls, so roughly 12,000 for a 1% budget. Below that, the audit tells you which budget it can promise.
View as table
| Calibration calls | Naive over budget | Sundr over budget |
|---|---|---|
| 120 | 56% | 0% |
| 400 | 62% | 0% |
| 1,000 | 55% | 1% |
| 3,000 | 58% | 2% |
| 12,000 | 50% | 3% |
You pay from what Sundr saves you.
Savings are measured call by call against what your LLM would have charged, so the invoice is a share of a number you can audit.
- Certified coverage at 0.5–5% budgets
- Monthly savings estimate
- The questions Sundr would ask
- Logs deleted when the report is ready
- 30-day free shadow run first
- Certified threshold per task
- Live error monitor and automatic rollback
- Jev and Clef, routed per task
- Self-hosted proxy and monitor
- Zero data retention
- SSO and certificate reports
- Custom error budgets per route
See how much of last week you could have handed over.
Upload a log export of your judgment calls. You’ll get the share Sundr could answer at each error budget, certified on your own data, and what it would have saved.
- Exports from Langfuse, Helicone, LangSmith, Braintrust, or raw OpenAI and Anthropic logs
- About 30,000 calls certifies a 1% budget; 300 is the minimum
- Encrypted at rest, deleted as soon as the report is written
What counts as a judgment call?
Any LLM call whose answer is one of a fixed set: a label, a route, a yes or no, a 1–5 score. Ticket triage, moderation, intent detection, policy checks, LLM-as-judge evals and agent “is this done?” checks all qualify. Open-ended generation doesn’t.
Which models answer the calls?
Decision models: TypeSafe’s Jev and Cloudflare’s Clef today, routed per task, with open-weight models for self-hosted deployments. They return calibrated probabilities rather than text, in well under a second, for a fraction of a cent per thousand calls.
Is agreement with my LLM the same as accuracy?
No. By default Sundr certifies agreement with the model you use today, because its past verdicts are free labels. Send a few hundred human-labelled calls and the report states your LLM’s own error rate beside Sundr’s, or certify against the human labels directly.
How much traffic do I need?
Certifying a budget reliably takes about 120 ÷ budget held-out calls: 12,000 for 1%, 6,000 for 2%. With less, the audit shows the tightest budget it can honestly promise.
What happens if accuracy drifts?
A slice of answered calls, weighted toward the least certain, still goes to your LLM. Sundr estimates live error from it with a valid confidence interval. If the bound crosses your budget, the task switches back to your LLM automatically and you’re alerted.
What happens to my data?
Audit uploads are encrypted at rest and deleted as soon as the report is written; reports are deleted after 30 days. See the privacy notice for the providers involved.