Skip to content
New on arXiv: a judge's uncertainty decomposes, so expert labels go exactly where they remove error
Customers

Trusted by teams where quality isn't optional.

In production since early 2025, across customer service, enterprise software, healthcare, fintech and legal - with over a million evaluations processed on real production traffic.

  • Accenture
  • Instrumentl
  • MIT
  • SentiSum
  • Palantir
  • Leena AI
  • ETH Zurich
  • DigitalGenius
  • Qapitol QA

Used by teams at the companies above - customers and partners, current and past.

Inside their products

The evaluation layer inside other companies' products.

The strongest proof we have: customers who took Composo beyond their own QA and shipped it to their customers.

A leading customer-service AI platform

Ships Composo-powered evaluation as a customer-facing feature inside its own product - its enterprise customers see our scores on their conversations, presented as the platform's own quality layer.

An enterprise HR AI platform

Composo scores sit in the dashboards its enterprise customers use every day. The usage grew on the customer's own initiative - evidence their customers can read, inside the product.

Diagram of the calibration loop in three stages: discover failure modes in the traces, judge every output against a rubric assembled for it, and learn as the customer's experts review the judge - every correction becomes calibration data.
The loop behind every deployment: corrections from the customer's own experts compound into the judge.
See it in action

See what we find in a real clinical AI output

See how Composo evaluates a real clinical AI output - with analysis, source citations, and expert corrections that compound over time.

In their words

“We cut our QA cycle time by 70%. Instead of relying purely on human review, now we instantly know which prompts are failing and why.”

Head of AI Engineering Enterprise SaaS platform
We embedded Composo into our AI Workers from day one - best decision we've made on testing. As an early stage start-up, we can't afford to waste time on manual evals or debugging. They provide peace of mind for us and our customers. No brainer.
Fehmi Sener CTO, 5u.ai
For the first time, we can ship with complete confidence knowing exactly what our AI quality looks like at scale.
Senior Software Engineer Instrumentl
LLM as a Judge was far too unreliable. Composo gave us the deterministic scoring we needed to actually track improvements.
Senior ML Engineer Fortune 500 Financial Services
The numbers 2026
1M+ Evaluations processed on real production traffic.
10K+ Evaluations a day at the largest customers.
90% Agreement with domain experts on flagged failures.
2-4 weeks To production, against a 3-6 month internal build.

The full stories, anonymised.

Problem, deployment and result for 5 production engagements - across clinical notes, financial planning, customer support and legal work.