Illustration of a digestive system suffering from irritable bowel syndrome.
Gastroenterology & GI Surgery

Machine Learning Improves Accuracy of Inflammatory Bowel Disease Identification Beyond ICD-10 Codes

Houston Methodist researchers developed a multi-class machine learning approach to identify true cases of inflammatory bowel disease from electronic medical records, to improve patient cohort accuracy and enable more reliable clinical research.

Electronic medical records (EMRs) have transformed clinical research by providing digital access to large amounts of patient data. However, accurately identifying patients with specific diseases remains a major challenge. Researchers often rely on International Classification of Diseases (ICD-10) codes to identify patient populations, but these billing codes are not always reliable for confirming a diagnosis. Misclassification can introduce bias into research studies, reduce the accuracy of findings and limit the development of high-quality patient cohorts.

To address this challenge, Dr. Bincy Abraham, Professor of Clinical Medicine at Houston Methodist, and her research team developed a machine learning (ML) model that can distinguish patients with Crohn's disease, ulcerative colitis and those without inflammatory bowel disease (IBD) using information in EMRs. This research study was recently accepted for publication in the journal Digestive Diseases and Sciences.

By integrating multiple clinical features rather than relying solely on diagnosis codes, the model substantially improved the accuracy of identifying true IBD cases. This ML model can also be applied to other conditions to improve the quality of clinical research.

“Our multi-class machine learning approach substantially improved IBD classification compared to ICD-10 codes, increasing the positive predictive value from 65% to 90%. Our framework provides a scalable, accurate, and interpretable method for EMR-based research.”


Bincy Abraham, MD, MS

Unlike traditional binary models, the algorithm can classify patients as having Crohn's disease, ulcerative colitis or no confirmed IBD.

Why accurate IBD identification matters

IBD, which includes Crohn's disease and ulcerative colitis, is a chronic inflammatory condition of the gastrointestinal tract requiring long-term medical management. IBD symptoms often overlap with those of other gastrointestinal disorders such as irritable bowel syndrome (IBS), infectious colitis and diverticular disease. Overlapping symptoms, such as chronic diarrhea, abdominal pain, fatigue and unintended weight loss, can result in patients receiving multiple or evolving diagnoses.

Accurate identification of IBD is vital to prevent irreversible intestinal damage, guide targeted subtype treatment and significantly improve long-term patient outcomes.

ICD-10 codes alone are not sufficiently reliable

Researchers frequently use ICD-10 codes to identify patients with IBD for retrospective studies. However, Dr. Abraham’s research found that relying on one or more ICD-10 codes correctly identified true IBD in only about 65% of cases, highlighting the limitations of using administrative coding alone.

Diagnostic accuracy is crucial in retrospective research because validated patient populations reduce misclassification bias, improve study reliability and provide stronger foundations for future clinical studies.

Building a smarter prediction model

To improve case identification, Dr. Abraham and her team first performed detailed manual chart reviews to establish a validated cohort of patients with confirmed Crohn's disease, ulcerative colitis or no IBD. Of nearly 35,000 patients with ICD-10 codes, the team manually validated 1,200. This carefully curated dataset served as the foundation for training several machine learning algorithms.

The team evaluated three commonly used machine learning approaches:

  • Logistic regression

  • Random forest

  • Extreme gradient boosting (XGBoost)

The models incorporated 198 clinical variables extracted from the EMRs, including diagnosis codes, medication history, imaging findings, pathology reports, specialist involvement and procedural information.

Researchers also applied recursive feature elimination with cross-validation (RFECV), a feature-selection technique that systematically removes less informative variables while maintaining overall model performance.

“It is important to ensure that the cohort you are working with is validated before you look at either retrospective or prospective outcomes. Our goal was to narrow down the patients that we know for sure to have IBD, so that when we look at their outcomes and comorbidities, we are certain that they actually have the disease,” says Dr. Christopher Fan, Assistant Professor of Medicine at Houston Methodist, who served as the lead author of this study.

Random forest delivers the strongest performance

Among the ML approaches tested, tree-based models consistently outperformed traditional logistic regression. The random forest model achieved the highest overall performance in distinguishing Crohn's disease, ulcerative colitis and patients without IBD. Interestingly, reducing the number of input variables by nearly 80% using RFECV had little effect on predictive performance. This finding suggests that highly efficient models can be developed without sacrificing diagnostic accuracy, reducing computational demands and simplifying future implementation across healthcare systems.

Making artificial intelligence explainable

One of the most important aspects of implementing artificial intelligence in healthcare is understanding why a model reaches a particular prediction. Rather than functioning as a "black box," the investigators incorporated explainable AI using SHapley Additive exPlanations (SHAP), allowing clinicians to visualize how individual clinical features influenced each prediction.

The model identified clinically meaningful predictors that aligned with real-world medical practice.

Diagnosis codes for Crohn's disease and ulcerative colitis strongly predicted their respective diseases. In contrast, codes for competing gastrointestinal disorders (IBS, infectious gastroenteritis and diverticular disease) were more strongly associated with patients ultimately classified as not having IBD.

Medication history also proved highly informative. Use of IBD-specific therapies was associated with both Crohn's disease and ulcerative colitis, while aminosalicylate medications demonstrated a stronger relationship with ulcerative colitis, reflecting current treatment guidelines.

By providing transparent explanations for its predictions, the model increases clinician confidence and supports responsible integration of artificial intelligence into clinical research.

“Accurate patient identification is fundamental to both clinical research and patient care. By combining machine learning with explainable artificial intelligence, we can move beyond ICD-10 codes alone to create more reliable patient cohorts for future research and ultimately improve the quality of evidence generated from electronic health records," adds Dr. Abraham.

Advancing research through multi-class machine learning

Unlike previous machine learning approaches that required separate models for Crohn's disease and ulcerative colitis, the Houston Methodist team developed a single multi-class model that can simultaneously classify patients into one of three categories: Crohn's disease, ulcerative colitis or no IBD.

This unified framework simplifies implementation while better reflecting the diagnostic uncertainty commonly encountered in clinical practice. Patients may carry both Crohn's disease and ulcerative colitis diagnosis codes before a definitive diagnosis is established, making a multi-class approach particularly valuable.

Looking ahead

Although the study was conducted within a single healthcare system and relied primarily on structured EMR data, the findings show ML's potential to improve diagnostic accuracy for large-scale clinical research. As healthcare increasingly embraces data-driven medicine, explainable machine learning models offer an opportunity to build more accurate research cohorts, reduce diagnostic uncertainty and strengthen both retrospective and prospective studies.

Subscribe to our Newsletter
Please enter an email
Please enter a valid email
Related Articles