Automated ROTEM interpretation.
Turning physician interpretation logic into explicit, reproducible software for rotational thromboelastometry.
Deterministic Python prototype v0.1 built; retrospective validation in development. Not clinically validated or deployed for autonomous reporting.
Clinical knowledge engineering: interpretation rules, assay dependencies, exception handling, and standardized report design
The clinical problem
ROTEM studies require physician review and a narrative interpretation. Much of the routine reporting follows repeatable logic, but the relationships among assays make simple templating insufficient. The goal is to automate the repetitive work while preserving physician oversight and making the reasoning inspectable.
From expert judgment to executable rules
I used an LLM as an interactive prototyping environment to formalize interpretation logic, work through real and synthetic patterns, and refine report language. That specification has now moved into a deterministic Python package with versioned rules and reference intervals, clinical-state classification, provenance, and constrained narrative generation.
The clinical interpretation is rule-based. The LLM helped elicit and organize the knowledge; it is not the intended source of clinical truth or the runtime decision-maker. The work separates input normalization, clinical classification, and narrative rendering so that each can be reviewed independently.
Reproducibility is the central requirement
Same inputs + same reference configuration + same rule-set version = the same clinical classification and report.
The aim is a report whose statements can be traced to explicit rules, rather than a free-form answer that changes with conversational context. Laboratory reference intervals belong in versioned configuration, separate from interpretation logic.
Interpret what is actually known
The specification distinguishes coagulation initiation, heparin effect, fibrin-based clot strength, platelet contribution, and fibrinolysis. These domains have dependencies: a reduced composite clot-firmness measurement does not, by itself, establish a platelet abnormality.
Assay availability is part of the input. The system must preserve uncertainty when required information is absent, and distinguish missing, invalid, unmeasurable, and physiologically zero results.
Explore the assay dependencies and edge cases
Paired assays: INTEM and HEPTEM are evaluated together to characterize correction after heparin neutralization, with residual abnormality considered separately. Percent correction is reported explicitly. The model avoids subjective labels such as “slight”; correction that does not meet a validated threshold leaves heparin effect not established.
Dependent interpretation: FIBTEM informs whether platelet contribution can be assessed independently from composite clot firmness. When FIBTEM is abnormal, the current rules avoid inferring independent platelet adequacy. Without FIBTEM, the report retains the combined platelet/fibrinogen uncertainty. Elevated composite MCF and CFT findings are displayed and classified without adding unsupported physiologic inferences.
Analytic states: An unmeasurable FIBTEM is a distinct state, not a numeric zero or an omitted observation. Rules and report templates must handle it deliberately.
Panel awareness: Bleeding, comprehensive, and heparin panels expose different observations. The engine should never infer a result from an assay that was not performed.
A clinical state before a sentence
The architecture classifies the observations into an intermediate phenotype before rendering the report. That representation records factor-activity findings, heparin effect, fibrinogen and platelet assessments, and fibrinolysis—including indeterminate states.
Constrained templates then present findings in a consistent order. Common phenotypes can use a consolidated narrative rather than a series of repetitive sentences. This separation supports machine-readable output and independent review of the logic and wording.
What an auditable report should retain
Each statement should connect back to the original measurement and its source, the reference configuration, the rule evaluated, the resulting clinical state, and the narrative template used. The question is: “Why did the system say this?” The answer should be a reproducible provenance chain.
Input source → reference-set version → rule and result → clinical state → template and report.
Boundaries around images and recommendations
Structured instrument or LIS observations are the preferred input. Screenshot ingestion is a separate development path: extract measurements, display them for human confirmation, and only then apply the deterministic rules. An image model should not directly generate the clinical interpretation.
Management recommendations are also a separate layer. A separate recommendations module has been implemented. Recommendations remain conservative and require clinical context and institutional governance; the current approach does not recommend reversal for a heparin-effect finding.
Implementation and testing
The v0.1 Python implementation is maintained in a private source repository, with continuous integration configured for Python 3.10–3.13. The current development suite contains 18 passing deterministic tests, including protections against historical interpretation failures and tests of the agreed clinical rules.
These are software regression tests. They establish expected behavior for specified cases; they do not establish clinical performance or patient benefit.
Retrospective validation in development
Six historical cases have been recovered and normalized into a separate review corpus. Cases with unresolved specification conflicts, historical failures, or pending expert review remain separate from the accepted reference set. Qualitative observations are preserved without inventing numeric measurements.
The next phase is retrospective comparison against reviewed clinical cases. Cases with raw measurements and historical physician interpretations can support concordance assessment; raw-number-only cases can help characterize patterns and identify gaps in test coverage.
Clinical review, local reference configuration, and workflow validation remain prerequisites to clinical deployment. Data access and the validation corpus are still being developed.
What this work demonstrates
Domain expertise becomes more useful when its assumptions, dependencies, and limits are explicit. This project uses AI to accelerate knowledge engineering while placing the clinical reasoning in software that can be tested, versioned, and audited.
From clinical need to institutional action.
These projects connect clinical expertise, usable systems, and accountable decisions—the same concerns that shape my approach to clinical informatics.