Skip to content
New on arXiv: a judge's uncertainty decomposes, so expert labels go exactly where they remove error
The platform

Self-improving evals.

An evaluation agent that learns your experts' definition of good and reviews every output in production.

It plugs into the logs you already have, self-hosted if you're regulated, and it's running in about two weeks.

The problem

You don't know what your AI is getting wrong right now.

01

Your test suite was accurate the week you wrote it. Your LLM-as-judge gives the same scores on day 100 as day 1.

02

Your AI is handling things differently than you expect - and nobody notices until a customer complains.

03

Every team we've worked with discovers failure patterns in week one they had no idea existed. Not generic "hallucination" - specific failures that matter for your domain.

The loop

Discover, evaluate, learn.

01

Discover failure modes

Point Composo at your production traces and it finds and names the failure clusters nobody wrote criteria for. One click from a discovered failure mode to a live, calibrated evaluator.

02

Evaluate every output

One API call: one output, one criterion, and you get back a continuous score with reasoning and cited sources. Fast and cheap enough to run on everything, not a sample - and to block before a customer sees it.

03

Learn from your experts

Your reviewers' corrections compound: fix one case and similar cases improve automatically. The judge knows which cases it shouldn't decide alone, and raises exactly those - so the review queue stays small.

What happens

What the first four weeks look like.

Week 1

The failure report

We connect to your production traces and run our engine. You get a failure report - every failure categorised by type, severity, and frequency. This is usually the "oh shit" moment.

Weeks 2-3

Your experts calibrate

Your domain experts review what we flagged and correct where we're wrong. Every correction makes the system smarter - similar cases improve automatically. We build out guardrails for the worst patterns.

Week 4

Handover

You own everything. The evaluation criteria, the failure taxonomy, the guardrail rules, all correction data. The system works without us.

Ongoing

It gets smarter

Platform maintenance, upgrades, and tuning as your product evolves. Optional - the system works without us.

Your team commits ~10 hours over 4 weeks. We handle everything else. You own everything at the end.

The engine

No judgement from the output alone.

The rubric is assembled fresh for every judgement - from your standards, your documents and your experts' past corrections. In the middle sits a dynamic ensemble of judge models, sized by task complexity.

Diagram: five sources - your standards and SOPs, your existing logging, every prior judgement, your reviewers' corrections, and lookup tools - all feed the evaluation agent, which returns a verdict with its reasoning and sources attached.
On the record

Every verdict shows its working.

A judgement is a record: the rubric it was scored against, the sources it consulted, the reasoning behind the score - so a reviewer can check the judge as easily as the output.

A single evaluation run drawn as a five-stage record: building the evaluation rubric from customer standards, searching evaluation memories and expert annotations, combining sub-agent results, and the completed 0.34 score with its reasoning and where it lands on the rubric scale.
Escalation

The judge knows when it is not sure.

Uncertain cases escalate to your experts; everything auto-resolved carries a stated confidence. The method is published - arXiv:2609.06444.

Curve of auto-resolved accuracy against the share of outputs judged automatically: at the operating point, 74% of outputs are judged automatically at roughly 95% accuracy, and the least-confident 26% escalate to your experts.
Under the hood

Built from production, not a template

01

Custom failure taxonomy for your domain

A failure taxonomy specific to your use case - learnt from your traces and your experts, and informed on day one by patterns from many production deployments across healthcare, fintech, CX, legal, and multi-agent systems.

02

Learns from your traces and experts

Your production traces and expert corrections build a memory of what quality means for your domain. Month-1 corrections still improve month-6 evaluations. The system gets smarter every week without retraining.

03

Dynamic ensemble of agents

Multiple specialised agents work together - blending fast and deep evaluation intelligently. Beats any single model alone. Fast enough to block, cheap enough to run on everything.

04

Built for the agent era

The trace endpoint evaluates multi-agent workflows agent by agent, and the MCP surface means your AI agents can query eval data directly today.

The failure taxonomy

An engine that has already seen your type of failure.

Every engagement adds to a structured library of AI failure patterns - categorised by type, severity, and domain. Hallucinated medications in healthcare. Unsupported conclusions in legal. Confident wrong answers in customer support.

When we deploy into your stack, the engine already knows what to look for. Your expert corrections make it specific to your domain, and the taxonomy grows with every engagement - anonymised, cross-customer, compounding. This is the thing that takes 6 months to build internally and starts from zero every time.

See it in action

See what we find in a real clinical AI output

See how Composo evaluates a real clinical AI output - with analysis, source citations, and expert corrections that compound over time.

Deployment

Built for regulated industries and security-conscious enterprises.

Your region

Separate EU and US endpoints, both live - so European data stays in Europe.

Your cloud

Or self-hosted in your own Azure, AWS or GCP, where nothing leaves your environment at all.

Audited

SOC 2 Type II certified, pen-tested, HIPAA/BAA-ready. Details on the security page.

What you get

Your system. Your data. Your rules.

Calibrated evaluators

Specific to your domain and use case

Dynamic failure taxonomy

Every pattern categorised and severity-ranked

Guardrail rules and thresholds

Running in your stack, sub-second latency

All annotation data

From your domain experts, compounding over time

Deployment runbook

Complete documentation and configuration

Self-hosted option

Deploy in your Azure, AWS, or GCP, or use our separate EU and US endpoints.

See it on your data

Send us a handful of production traces. We'll deliver scored results with a failure report - what's going wrong, how often, and how severe. Takes under a week.