2026-09-09
The deck
The story
When does this patient go home?
A quality improvement study compared two ways of answering it at Houston Methodist, and reported which one was closer.
The hospital's own staff, and a commercial clinician-assisted tool integrated with the EHR, over inpatient encounters discharged in the window below.
At admission the case managers and the tool sat within a fraction of a day of each other on error.
The case managers did match the exact date more often, 23.6% against 15.3%.
The authors say this is the point where an accurate date counts for most in bed management.
One improved. One got worse.
Mean absolute error in days. The two never meet.
The strong admission correlation did not persist nearer discharge and should not be interpreted as evidence that the tool captured the same information clinicians use throughout the stay.
Estimates were part of usual care and visible to care teams. They may be partly self-fulfilling. The authors' own word for it is advantaged nearer discharge.
Accessible in the EHR but not enabled automatically.
No baseline patient demographic data were available.
The authors say what has to happen before anyone decides the tool's extra accuracy in some subgroups justifies the complexity of putting it in.
Warrants prospective, outcome-based evaluation.
What was verified
A quality improvement study at Houston Methodist Hospital examined inpatient encounters discharged between August 1st, 2023 and February 28th, 2024.
This quality improvement study included inpatient encounters at Houston Methodist Hospital with discharge dates between August 1, 2023, and February 28, 2024.
A day before discharge, case managers placed far more of their estimates within one day of the real date than the AI tool did.
At 24 hours, 79.5% of case manager estimations fell within 1 day vs 37.9% for AI estimations
Closer to discharge the case managers beat the tool on mean absolute error at both the 48 hour and the 24 hour mark.
Nearer to discharge, case managers vs AI were more accurate (48 hours: MAE, 1.29 [95% CI, 1.26-1.32] vs 1.59 [95% CI 1.57-1.61] days; 24 hours: MAE, 0.98 [95% CI, 0.96-1.01] vs 1.93 [95% CI, 1.91-1.95] days) (both P < .001).
At admission the two methods sat close together on error, and the case managers still hit the exact date and landed within one day more often than the tool.
At admission, accuracy was similar between methods (MAE, 4.20 [95% CI, 4.08-4.31] vs 4.27 [4.15-4.38] days; P = .001), although case managers matched the exact date more often (23.6% vs 15.3%) and fell within 1 day more often (46.3% vs 40.8%).
The dates came from a commercial clinician assisted tool wired into the electronic health record.
Artificial intelligence–estimated discharge dates were generated by a commercial clinician-assisted tool integrated with the EHR.
The study found the AI estimate matched usual care at admission. It lost to case managers as discharge approached, which is when an accurate date is most useful for managing beds.
This quality improvement study found that AI-estimated discharge was comparable to usual care at admission but outperformed by case managers as discharge approached (when accurate estimations are most actionable for bed management).
The authors warn that the strong correlation at admission can't be read as evidence the tool captured what clinicians know across the stay.
The strong admission correlation did not persist nearer discharge and should not be interpreted as evidence that the tool captured the same information clinicians use throughout the stay.
The authors flag a limitation on their own side of the comparison. Case manager estimates were visible to care teams and may be partly self fulfilling, while the AI estimates sat in the record without being switched on automatically.
Case manager estimates were part of usual care and visible to care teams, so they may be partly self-fulfilling and, thus, advantaged nearer discharge; AI estimations were accessible in the EHR but not enabled automatically.
The authors call for prospective outcome based evaluation before anyone decides the tool's extra accuracy in some subgroups is worth what it takes to run it.
Whether the AI's incremental accuracy in selected subgroups justifies the complexity of implementation warrants prospective, outcome-based evaluation.
The number of encounters compared is reported separately at each of the three time points.
admission, n = 21 710; 48 hours, n = 9236; 24 hours, n = 12 172
The comparison carries no patient demographics at all, which the authors state outright as a limitation.
No baseline patient demographic data were available.
The study reports itself against the SQUIRE reporting guideline, the standard for quality improvement work.
The study followed the SQUIRE reporting guideline.
Three named people are on the study's byline, two with an MD and one with a BS. The byline is the whole of what the source states about them.
Havish S Kantheti, MD, Connor Dolan, BS, Jordan Dale, MD
The study appeared in JAMA Network Open on September 3rd, 2026.
Published: September 3, 2026. doi:10.1001/jamanetworkopen.2026.32033








