Skip to content
New on arXiv: a judge's uncertainty decomposes, so expert labels go exactly where they remove error
← Back to Research

Inside 847 Production Clinical AI Notes

Seb Fox · CEO & Co-founder ·

Seb gave this talk at the AI Engineer World’s Fair in San Francisco: a teardown of production clinical notes from deployed AI scribes, and why the dangerous failures are the ones that read completely fine.

In the video: a note that looks complete until you know what the patient actually said, the failure patterns across three commercial scribes and why they happen even with a perfect transcript, the reveal that a carefully engineered LLM judge passed the failing notes too, and the Discover, Capture, Calibrate loop that catches what a rubric can’t.

The full verified study is now on arXiv as two papers: One note in three, our census of three deployed scribes, and LLM Judges Verify Presence, Not Absence, on why judges miss omissions and what recovers detection. The benchmark behind them is open on Hugging Face, with code on GitHub.