27 Aug 2026
Demo: Insights & Dashboards
A short walkthrough of the Composo dashboard: overall score, what needs attention, and what regressed since last period.
Luke Markham
Technical articles on AI evaluation, failure modes, and what we're learning from production deployments.
For our own experiments, benchmarks and findings, see Composo Research.
27 Aug 2026
A short walkthrough of the Composo dashboard: overall score, what needs attention, and what regressed since last period.
Luke Markham
19 Aug 2026
A short walkthrough of the failure map: your failing traces grouped into named failure modes, sized, and tracked week to week.
Luke Markham
26 May 2026
Domain-calibrated guardrails for AI scribes. Catches hallucinated medications, fabricated values, and unsupported clinical inference at the inference boundary, with sub-second latency.
Ryan Lail
A short walkthrough of the Composo dashboard: overall score, what needs attention, and what regressed since last period.
A short walkthrough of the failure map: your failing traces grouped into named failure modes, sized, and tracked week to week.
Domain-calibrated guardrails for AI scribes. Catches hallucinated medications, fabricated values, and unsupported clinical inference at the inference boundary, with sub-second latency.
Build or buy eval infrastructure? The six layers after a baseline judge (ensemble through meta-eval), why a first slice of scoring is not the same milestone as release-grade depth, and how we benchmark quality (71% baseline judge to 83.6% on our internal set).
Findings from clinical AI engagements - actual failure patterns from production clinical AI, categorised by type, with real examples. Discussions becoming decisions, dangerous omissions, dosage errors, and diagnostic leaps.
Introducing Composo: AI evaluation that learns your standards. Not manual review, not LLM-as-judge -- a third option that gets better the more you use it.
A deep dive into Composo's generative reward model architecture that achieves 95% agreement with expert evaluators, compared to ~70% for LLM-as-judge approaches.
A practical guide to evaluating LangGraph multi-agent workflows using Composo's agent evaluation framework with quantitative scoring across 5 key dimensions.
Published AI scribe failure rates are higher than vendor marketing suggests. Real lawsuits have started. Here is what the specific failure patterns look like and what clinical evaluation needs to catch.
A component-based evaluation framework for agentic LLM systems covering tool call formulation, tool choice, response integration, reasoning evaluation, and system-level analysis.
A practical guide to evaluating clinical AI in production: specific failure modes to catch, why generic evaluation misses them, regulatory context, and what good quality infrastructure looks like.
A comprehensive guide to evaluating LLM classification quality, covering supervised metrics, generative reward models, and LLM-as-judge approaches.
A comprehensive guide to evaluating RAG applications, covering generation metrics, retrieval assessment, and advanced CAG-based oracle evaluation techniques.
Composo Align achieves 95% agreement with expert preferences vs 72% for LLM-as-judge, with 100% score consistency through its deterministic generative reward model.
Composo Align uses a generative reward model architecture to provide deterministic, consistent scoring for LLM evaluation, achieving 95% agreement with expert preferences.
A plain-language explanation of LLM evaluation: what it is, why LLM-as-judge plateaus, what production AI quality actually requires, and how to think about build vs buy.
LLM evaluations that worked last quarter can silently stop working this quarter. Evaluation drift is a first-class failure mode of production AI quality systems. Here is how to detect it and what to do about it.
An honest analysis of the real cost of building an LLM evaluation pipeline internally vs buying a deployed quality layer. Real numbers, real timelines, real trade-offs.
A structured guide to evaluating LLM applications, covering common challenges with human vibe checks and LLM-as-judge, and key steps to building a reliable evaluation framework.
A practical guide to AI guardrails for production LLM systems. What they catch, what they miss, and how to deploy domain-specific guardrails that actually work at inference time.