Clinical AI Insight

Outcome Ascertainment Bias in Clinical AI: Why Prediction Models Learn Who Gets Tested

A prediction model cannot learn a clinical outcome that the healthcare system never observes. When testing, surveillance, diagnosis, or documentation determines which outcomes become visible, the model may learn patterns of observation and care rather than biology alone.

By Mohamed Khair Ibraheem, MD, MS September 8, 2026 Critical Care · Clinical AI · Prediction Models
What is outcome ascertainment bias in clinical AI? Outcome ascertainment bias occurs when the outcome used to train or evaluate a clinical prediction model is observed differently across patients because testing, surveillance, diagnosis, or documentation is selective. The recorded label can therefore reflect both the patient state and the healthcare process that determined whether that state became visible.

Clinical machine learning usually begins with an apparently simple pair: predictors and an outcome. The predictors may include vital signs, laboratory values, medications, comorbidities, imaging, or physiologic signals. The outcome may be sepsis, acute kidney injury, infection, bleeding, thrombosis, organ failure, or another event that matters clinically.

The difficulty is that many clinical outcomes are not passively waiting in the electronic health record. They become known only after somebody looks for them. A blood culture has to be ordered. A troponin has to be measured. Imaging has to be performed. A complication has to be recognized and documented. Follow up has to occur long enough for an event to be captured.

This creates a fundamental problem for clinical AI. The model is trained on what was observed, but what was observed is partly determined by clinical suspicion, local workflow, resource availability, practice style, access to care, and prior decisions. In that setting, the label is not simply a biological truth. It is a biological truth filtered through a process of ascertainment.

Why selective testing changes the label

Suppose a model is developed to predict a condition that is confirmed only when a diagnostic test is performed. Patients with a positive test are labeled positive. Patients with a negative test are labeled negative. The difficult group is the patients who were never tested.

A common analytic shortcut is to treat an untested patient as though the condition were absent. That choice may be convenient, but it creates a strong assumption: no test means no disease. In real clinical care, no test may instead mean that the clinician did not suspect the condition, the presentation was atypical, another diagnosis dominated attention, testing was unavailable, or the patient was discharged before the outcome became visible.

Chang, Sjoding, and Wiens described this problem as disparate censorship and showed how differences in testing can create label bias when untested patients are effectively treated as negative. Their work demonstrates an important principle: when testing rates differ among patients with comparable underlying risk, the resulting labels can carry systematic error into model development.

The central question is not only whether the recorded outcome is correct.

It is whether every patient had a comparable opportunity for that outcome to be detected.

Outcome ascertainment bias is not ordinary missing data

It is tempting to frame this as a missing data problem, but that is incomplete. Missingness usually focuses attention on absent variables. Outcome ascertainment concerns the mechanism that determines whether the target itself becomes known.

The distinction matters because the observation process can be clinically informative. A clinician orders a test for a reason. Surveillance is intensified for a reason. Some patients receive more imaging, more laboratory testing, longer monitoring, or closer follow up because their perceived risk is higher. If the model uses labels created downstream of those decisions, it can inherit the decision process itself.

This connects outcome ascertainment bias with older concepts such as verification bias and workup bias in diagnostic research. Decades before modern clinical AI, prediction researchers recognized that estimates could be distorted when the decision to apply a reference standard depended on the clinical findings being studied. Machine learning does not remove that problem. Large electronic health record datasets can scale it.

Why critical care is especially vulnerable

Critical care and acute illness are natural settings for this problem because observation is intense but uneven. Two patients with similar underlying disease may receive very different diagnostic trajectories depending on their initial presentation, location of care, clinician suspicion, hemodynamic stability, local protocols, and timing.

Consider sepsis. A model may rely on outcomes defined by cultures, organ dysfunction measures, treatment patterns, or combinations of clinical observations. Yet culture collection is selective. Laboratory measurement frequency is selective. ICU transfer is selective. Antibiotic treatment is selective. Even the timing of documentation is influenced by workflow.

The result is that a model can appear to learn sepsis risk while partly learning who is likely to trigger the clinical actions through which sepsis becomes observable.

This is relevant to interpretability as well. In our 2026 systematic review and meta analysis of interpretable AI for sepsis and severe infection, the literature showed promising discrimination for sepsis detection but substantial methodological limitations. Fourteen of fifteen included studies were judged to have high overall risk of bias, emphasizing that understandable model outputs do not eliminate weaknesses in outcome definition, study design, validation, or clinical implementation.

A high AUC does not prove that the model learned the right phenomenon

Discrimination metrics answer an important but narrow question: can a model rank patients according to the recorded outcome? They do not prove that the recorded outcome is an unbiased representation of the clinical state we actually care about.

A model can therefore perform well against an imperfect label. If testing behavior is stable within the development dataset, the model may accurately reproduce that behavior. Cross validation can reward it. Internal validation can confirm it. Feature importance methods can explain it. None of those steps guarantees that the target was ascertained equally across patients.

This is one reason that label bias can be difficult to detect through routine performance measures. The GUIDE framework for predictive information in healthcare specifically highlights outcome ascertainment as a source of label bias when the observed outcome differs systematically from the ideal outcome that should guide clinical decision making.

Proxies can create the same problem in another form

Sometimes the problem is not whether an outcome was measured but whether the chosen label represents the desired clinical construct. A well known example comes from an algorithm that used healthcare cost as a proxy for healthcare need. Because spending was not equivalent to illness burden across patient groups, a seemingly reasonable target encoded a structural difference in care rather than the clinical need the system intended to identify.

The lesson is broader than that single application. Clinical AI developers should ask whether the target is the clinical phenomenon itself, a proxy for that phenomenon, or a healthcare process correlated with it. Those are not interchangeable.

External validation helps, but it does not automatically solve ascertainment bias

External validation is essential for assessing generalizability, but an external dataset can reproduce the same observation mechanism. Two institutions may both test high risk patients more aggressively. Two databases may both define an outcome using the same ordered laboratory test. A model can therefore validate across datasets while the same underlying ascertainment process remains embedded in both.

Changes in clinical policy can also produce dataset shift. If a health system changes when it orders cultures, imaging, biomarkers, or surveillance tests, the relationship between patient state and recorded label may change even when the biology does not. Work on dataset shift in health AI has emphasized that changes in measurement and clinical policy can degrade model reliability after deployment.

A practical framework for researchers

Outcome ascertainment should be treated as part of the data generating process, not merely as a limitation mentioned at the end of a manuscript. A practical evaluation can begin with seven questions.

1. Define the clinical outcome before defining the database label

State the biological or clinical event the model is intended to predict. Then describe exactly how the electronic record operationalizes that event. The gap between those two definitions is where ascertainment problems often begin.

2. Map the pathway through which the outcome becomes visible

Identify every decision required for detection. Does a clinician have to order a test? Does the patient need imaging? Is specialist review required? Is follow up necessary? Does documentation depend on billing or coding?

3. Quantify who gets tested

Testing frequency should be described across relevant clinical subgroups, care settings, severity levels, and time periods. Large differences do not automatically prove bias, but they reveal where unequal observation may exist.

4. Do not automatically equate unobserved with negative

Whenever possible, distinguish confirmed negative outcomes from outcomes that were never assessed. Sensitivity analyses can explore how different assumptions about the unobserved group change model performance.

5. Evaluate the testing process as well as the prediction process

If testing probability is strongly predictable from the same variables used by the outcome model, that is important information. It suggests that part of the apparent signal may be related to clinical suspicion or workflow.

6. Validate across different surveillance regimes

A useful external validation setting is not merely geographically different. It should also differ in practice patterns when possible. Robust performance across different testing policies provides stronger evidence that the model is learning more than one institution's observation process.

7. Report ascertainment explicitly

Readers should know how the outcome was detected, how many patients were never assessed for it, how observation varied across groups, and which assumptions converted clinical records into model labels.

Deployment can turn ascertainment bias into a feedback loop

The problem becomes more important once a model influences care. If a model trained on historical testing behavior predicts that certain patients are low risk, clinicians may test those patients less often. Fewer outcomes are then detected in that group. Those new data can later reinforce the original prediction pattern.

This means the model is no longer simply describing a historical dataset. It is participating in the mechanism that generates future labels.

For clinical AI governance, monitoring should therefore include more than discrimination and calibration. Health systems should watch whether deployment changes testing rates, surveillance intensity, treatment, escalation decisions, and the probability that outcomes are observed.

The broader principle

Clinical AI is often described as learning from real world data. That description is correct but incomplete. Real world clinical data are produced by a healthcare system. They reflect disease biology, but they also reflect suspicion, access, clinician behavior, resource allocation, policy, and documentation.

For high consequence prediction models, especially in acute and critical illness, the reliability of the label deserves the same scrutiny as the architecture of the algorithm.

A model cannot learn an outcome that the healthcare system never sees.

Before asking whether a clinical AI model predicts well, we should ask how the outcome became observable, who had the opportunity to be diagnosed, and whether the label represents biology or the process of looking for it.

Mohamed Khair Ibraheem, MD, MS

Mohamed Khair Ibraheem, MD, MS

Physician scientist focused on critical care, clinical artificial intelligence, acute illness, sepsis, and transplantation.

Research program · Publications · ORCID

Selected references

  1. Chang T, Sjoding MW, Wiens J. Disparate Censorship & Undertesting: A Source of Label Bias in Clinical Machine Learning. Proceedings of Machine Learning Research. 2022. PubMed
  2. GUIDE investigators. Guidance for unbiased predictive information for healthcare decision making and equity: considerations when race may be a prognostic factor. npj Digital Medicine. 2024. Nature
  3. Panzer RJ, Suchman AL, Griner PF. Workup bias in prediction research. Medical Decision Making. 1987. PubMed
  4. Subbaswamy A, Saria S. From development to deployment: dataset shift, causality, and shift stable models in health AI. Biostatistics. 2020. Oxford Academic
  5. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019. PubMed
  6. Ibraheem M, Khalil M, Khalil S, Contreras H, Azar J. Early clinical decision support using interpretable artificial intelligence in acute illness (sepsis and infection): A systematic review and meta analysis. International Journal of Medical Informatics. 2026. PubMed

This article is an independent academic perspective for educational and scholarly purposes. It is not a peer reviewed publication and does not provide individual medical advice.