Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
AI Engineer World's Fair 2026 · 19:48
AI evaluation and quality infrastructure
Composo builds AI evaluation systems that help teams find and control failures in production applications. Its platform connects to application traces, discovers recurring failure patterns and evaluates outputs against domain-specific criteria. Teams can use the results to block failing outputs, monitor quality and route uncertain cases to experts. The evaluation API returns a score between 0 and 1, an explanation and supporting sources.
Composo’s co-founders are Seb Fox, CEO, a medical doctor who previously led AI engagements at McKinsey’s QuantumBlack, and Luke Markham, CTO, whose machine-learning engineering experience includes Graphcore and Tessella. The evaluation engine assembles a fresh rubric from customer standards, documents and past expert corrections, using an ensemble of judge models sized to task complexity. Its clinical evaluation research examines why judges miss omitted information: checking a note against facts extracted from the underlying consultation helps recover errors that ordinary judging overlooks.
Composo serves healthcare, financial-services, legal and enterprise AI workflows. In 2026, the company reported more than one million production outputs evaluated and use by over 30 AI teams. Customers can use separate EU and US endpoints or self-host in their own cloud. They retain ownership of evaluation criteria, failure taxonomies, guardrail rules and expert correction data; ongoing platform maintenance and tuning are optional.
Composo’s supplied archive contains one recording. It offers a focused path through clinical AI failure analysis and expert-informed evaluation. The findings and proposed methods below are claims presented in the recording, not independently verified current facts or statements about Composo’s current products.
Watch Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo for concrete examples of how plausible clinical notes can omit important information or introduce errors. Fox distinguishes transcription mistakes from generation failures and reports error rates from a real-world study. Readers assessing clinical-note systems can begin with these examples and follow his discussion of why an accurate transcript does not guarantee a faithful note.
For evaluation design, continue to his account of an automated judge that still missed serious errors, then his proposed approach: capture experts’ reasoning and corrections, and retrieve relevant judged cases, references, and guidelines for each evaluation. The useful question running through the talk is how to determine which differences matter in context. Fox presents dataset-specific improvements from this approach; the recording does not establish its effectiveness across other applications or its current availability as a product.
AI Engineer World's Fair 2026 · 19:48
Affiliations reflect their AIE appearances, not necessarily current employment.