Self-improving evals.
An evaluation agent that learns your experts' definition of good and reviews every output in production.
It plugs into the logs you already have, self-hosted if you're regulated, and it's running in about two weeks.
You don't know what your AI is getting wrong right now.
Your test suite was accurate the week you wrote it. Your LLM-as-judge gives the same scores on day 100 as day 1.
Your AI is handling things differently than you expect - and nobody notices until a customer complains.
Every team we've worked with discovers failure patterns in week one they had no idea existed. Not generic "hallucination" - specific failures that matter for your domain.
Discover, evaluate, learn.
Discover failure modes
Point Composo at your production traces and it finds and names the failure clusters nobody wrote criteria for. One click from a discovered failure mode to a live, calibrated evaluator.
Evaluate every output
One API call: one output, one criterion, and you get back a continuous score with reasoning and cited sources. Fast and cheap enough to run on everything, not a sample - and to block before a customer sees it.
Learn from your experts
Your reviewers' corrections compound: fix one case and similar cases improve automatically. The judge knows which cases it shouldn't decide alone, and raises exactly those - so the review queue stays small.
What the first four weeks look like.
The failure report
We connect to your production traces and run our engine. You get a failure report - every failure categorised by type, severity, and frequency. This is usually the "oh shit" moment.
Your experts calibrate
Your domain experts review what we flagged and correct where we're wrong. Every correction makes the system smarter - similar cases improve automatically. We build out guardrails for the worst patterns.
Handover
You own everything. The evaluation criteria, the failure taxonomy, the guardrail rules, all correction data. The system works without us.
It gets smarter
Platform maintenance, upgrades, and tuning as your product evolves. Optional - the system works without us.
Your team commits ~10 hours over 4 weeks. We handle everything else. You own everything at the end.
No judgement from the output alone.
The rubric is assembled fresh for every judgement - from your standards, your documents and your experts' past corrections. In the middle sits a dynamic ensemble of judge models, sized by task complexity.
Every verdict shows its working.
A judgement is a record: the rubric it was scored against, the sources it consulted, the reasoning behind the score - so a reviewer can check the judge as easily as the output.
The judge knows when it is not sure.
Uncertain cases escalate to your experts; everything auto-resolved carries a stated confidence. The method is published - arXiv:2609.06444.
Built from production, not a template
Custom failure taxonomy for your domain
A failure taxonomy specific to your use case - learnt from your traces and your experts, and informed on day one by patterns from many production deployments across healthcare, fintech, CX, legal, and multi-agent systems.
Learns from your traces and experts
Your production traces and expert corrections build a memory of what quality means for your domain. Month-1 corrections still improve month-6 evaluations. The system gets smarter every week without retraining.
Dynamic ensemble of agents
Multiple specialised agents work together - blending fast and deep evaluation intelligently. Beats any single model alone. Fast enough to block, cheap enough to run on everything.
Built for the agent era
The trace endpoint evaluates multi-agent workflows agent by agent, and the MCP surface means your AI agents can query eval data directly today.
An engine that has already seen your type of failure.
Every engagement adds to a structured library of AI failure patterns - categorised by type, severity, and domain. Hallucinated medications in healthcare. Unsupported conclusions in legal. Confident wrong answers in customer support.
When we deploy into your stack, the engine already knows what to look for. Your expert corrections make it specific to your domain, and the taxonomy grows with every engagement - anonymised, cross-customer, compounding. This is the thing that takes 6 months to build internally and starts from zero every time.
See what we find in a real clinical AI output
See how Composo evaluates a real clinical AI output - with analysis, source citations, and expert corrections that compound over time.
Built for regulated industries and security-conscious enterprises.
Your region
Separate EU and US endpoints, both live - so European data stays in Europe.
Your cloud
Or self-hosted in your own Azure, AWS or GCP, where nothing leaves your environment at all.
Audited
SOC 2 Type II certified, pen-tested, HIPAA/BAA-ready. Details on the security page.
Your system. Your data. Your rules.
Calibrated evaluators
Specific to your domain and use case
Dynamic failure taxonomy
Every pattern categorised and severity-ranked
Guardrail rules and thresholds
Running in your stack, sub-second latency
All annotation data
From your domain experts, compounding over time
Deployment runbook
Complete documentation and configuration
Self-hosted option
Deploy in your Azure, AWS, or GCP, or use our separate EU and US endpoints.
Calibrated to your domain.
Healthcare AI
Clinical AI quality. Catch hallucinated medications, omitted findings, and diagnostic leaps before they reach patients.
Financial services
Quality infrastructure for agentic financial AI. Catch silent failures in FX, hedging, fraud detection, and financial advice.
Enterprise AI
Domain-specific quality for enterprise AI systems at scale. Deploy a quality layer calibrated to your business logic in 2-4 weeks.
Legal tech
Evaluation for document extraction, contract review, and legal reasoning. Built on domain-expert calibration, not text similarity.
See it on your data
Send us a handful of production traces. We'll deliver scored results with a failure report - what's going wrong, how often, and how severe. Takes under a week.