# Composo > AI evaluation that learns your standards. Composo finds the failures your evals miss and fixes them before they reach customers. Composo is an AI evaluation platform for teams shipping LLM applications to production — RAG systems, agents, classifiers, and domain-specific AI in clinical, financial, legal, and enterprise settings. Our generative reward model achieves ~95% agreement with expert evaluators (vs ~70% for LLM-as-judge) with deterministic, repeatable scoring that learns your team's standards over time. ## Docs - [Homepage](https://www.composo.ai/): Composo product overview and positioning. - [Use cases](https://www.composo.ai/use-cases): Hub describing who Composo is built for and how teams deploy it. - [Healthcare](https://www.composo.ai/healthcare): Clinical AI evaluation — ambient scribes, clinical decision support, and medical record extraction. - [Fintech](https://www.composo.ai/fintech): Evaluation for financial AI — agents, document workflows, and customer-facing assistants. - [Legal](https://www.composo.ai/legal): Evaluation for legal-tech AI workflows. - [Enterprise](https://www.composo.ai/enterprise): Enterprise AI quality, reliability, and QA infrastructure. - [Guardrails](https://www.composo.ai/guardrails): Domain-specific guardrails for blocking bad LLM outputs at inference time. - [Blog](https://www.composo.ai/blog): Index of all Composo blog posts on LLM evaluation, agents, RAG, clinical AI, guardrails, and AI quality engineering. Follow links from this page for individual posts. - [Case studies](https://www.composo.ai/case-studies): Index of customer case studies across healthcare, B2B SaaS, financial services, enterprise, and legal tech. Follow links from this page for individual studies. - [Book a demo](https://www.composo.ai/book-a-demo): Schedule a product walkthrough. - [Security](https://www.composo.ai/security): Security posture, certifications, and data handling. ## Selected blog posts - [What We Found Inside Clinical AI Systems That Were Passing Every Eval](https://www.composo.ai/post/clinical-ai-failure-modes): Findings from clinical AI engagements — actual failure patterns from production clinical AI, categorised by type, with real examples. Discussions becoming decisions, dangerous omissions, dosage errors, and diagnostic leaps. - [Improving LLM Judges With Experiments, Not Vibes](https://www.composo.ai/post/llm-judge-criteria-ensembling): Our open-source research on RewardBench 2 shows that two simple techniques — task-specific criteria injection and ensembling — improve LLM judge accuracy by up to 13.5pp (71.7% baseline → 85.8% with Claude Haiku 4.5), with the same pattern holding across OpenAI GPT-5.4 and Anthropic Claude families. ## Compare - [Composo vs LangSmith](https://www.composo.ai/vs/langsmith): How Composo compares to LangSmith for LLM evaluation. - [Composo vs Langfuse](https://www.composo.ai/vs/langfuse): How Composo compares to Langfuse for LLM evaluation. - [Composo vs Braintrust](https://www.composo.ai/vs/braintrust): How Composo compares to Braintrust for LLM evaluation. - [Composo vs Arize](https://www.composo.ai/vs/arize): How Composo compares to Arize for LLM evaluation. - [Composo vs Galileo](https://www.composo.ai/vs/galileo): How Composo compares to Galileo for LLM evaluation. - [Composo vs build-it-yourself](https://www.composo.ai/vs/build-it-yourself): When to build evaluation infrastructure in-house vs buy. ## Optional - [Ground Truth newsletter](https://www.composo.ai/ground-truth): Subscribe to our newsletter on real-world AI evaluation. - [Privacy policy](https://www.composo.ai/privacy-policy) - [Terms of service](https://www.composo.ai/terms) - [Cookie policy](https://www.composo.ai/cookie-policy)