AI in health care delivery works only when trust, explainability and implementation science come before the model. At an MIT HEALS session on 27 April 2026, Hartford HealthCare's Barry Stein, MIT's Dimitris Bertsimas, anthropologist Amy Moran-Thomas and Nobel laureate Simon Johnson each described a different blocker, from pulse oximeter calibration to governance and education.
What the MIT HEALS panel actually argued about AI in health care delivery
AI in health care delivery is a translation problem before it is a modelling problem, according to four speakers at the MIT HEALS session held on 27 April 2026. The panel, chaired by MIT economist Joe Doyle, treated deployment, monitoring and trust as the hard part, and model accuracy as necessary but insufficient. MIT HEALS is a cross-institute collaboration at the Massachusetts Institute of Technology that spans bioengineering, clinical medicine, anthropology and economics.
Barry Stein, chief clinical innovation officer and chief medical informatics officer at Hartford HealthCare, framed the session around one risk that outranks the rest. He described patient harm, physical and emotional, as fundamental to nearly every other concern about clinical AI, and said the concerns have to be understood and mitigated before deployment accelerates.
Dimitris Bertsimas, MIT's Vice Provost for Open Learning and Associate Dean for AI and Online Education at MIT Sloan, presented holistic AI for medicine, which he abbreviates HAIM, the Hebrew word for life. Amy Moran-Thomas, the Andrew W. Mellon Associate Professor of Anthropology at MIT, presented ethnographic research on device design. Simon Johnson, the Ronald A. Kurtz Professor of Entrepreneurship at MIT Sloan and a 2024 Nobel laureate in economic sciences, presented an institutional plan for public health.
Each speaker identified a different binding constraint: governance and monitoring, explainability, instrument calibration, and funding incentives. All four agreed that education and open conversation are prerequisites rather than afterthoughts.
How Hartford HealthCare runs AI in health care delivery, from governance to simulation
Hartford HealthCare governs clinical AI through a multidimensional review and continuous monitoring rather than a one-time procurement, Stein said. The health system's Center for AI Innovation organises the work around six pillars, with an ecosystem of academic, corporate, startup, venture, government and MIT partners described as the most important of them.
Stein's central design argument is that AI differs from a drug or device because it changes. A clinician prescribing penicillin is assumed to know what it does, how to spot a complication and how to treat one; a model in production can behave differently after the version changes, so the system needs clinical accountability, transparency and continuous monitoring to hold trust.
The governance model distributes risk ownership across clinical, operational, technical, legal and ethics reviewers before anything reaches patients. Stein said Hartford HealthCare has worked with Bertsimas for roughly 10 years and with more than 20 of his doctoral students, a relationship he presents as the mechanism behind its bench-to-bedside speed.
Education and simulation carry equal weight in Stein's account. Hartford HealthCare runs a long-established simulation environment for robotic surgery, IV placement and trauma, used by the US Navy and local police as well as clinicians. Stein argues that increasingly autonomous clinical AI should be trained and tested the same way pilots are, because autonomy fails in situations that only rehearsal prepares people for.
What HAIM and xHAIM add to multimodal clinical models
HAIM is a multimodal modelling framework that combines images, clinical language, tabular records, time series and genomics by placing them in one shared representation, Bertsimas said. The premise is that humans perceive the world through sight, hearing, smell, taste and touch at once, so clinical models that combine modalities should outperform single-mode models.
The first version converted every modality into numbers. Images were compressed into embeddings, text was summarised with large language models such as BERT, and the concatenated vector was fed to conventional machine learning. Bertsimas said this worked well but was not explainable, a limitation he traced to a paper his group published in 2022 in Nature Digital Medicine.
xHAIM, the explainable variant, changes the shared language from numbers to words. Tabular values are verbalised as sentences, such as a body mass index of 25 or a systolic pressure of 142, and a large language model reasons over that text. The output includes a written explanation with references, which Bertsimas said physicians require alongside accuracy.
Subdural hematoma, length of stay and cardiac detection: what the numbers show
Hartford HealthCare's clearest published-style result is a roughly half-day reduction in average length of stay, from about 5.4 days to about 4.9 days, without an increase in the probability of readmission. Bertsimas presented that figure at the 27 April 2026 session as an implemented system rather than a pilot, and tied it to throughput at Hartford Hospital, which he described as an 860-bed facility.
The arithmetic Bertsimas gave was directional rather than audited. He noted that freeing beds raises daily admissions and then multiplied the effect across roughly 800 beds, without publishing a controlled comparison. Treat the 5.4-to-4.9-day change as a vendor-site report from a single health system, not as a general effect of clinical AI.
For subdural hematoma, bleeding between the brain and its outer covering, Bertsimas described a pipeline that estimates the probability of a bleed, localises it on the scan and predicts whether it will worsen over the following days. Combining methods reached what he called a practical 94% to 95%, while the worsening prediction was weaker though still useful; he did not name the dataset or the evaluation split, so the figure cannot be independently checked from the talk.
Other applications he listed include detecting multiple cardiac diseases from an electrocardiogram, deciding between surgical and transcatheter aortic valve replacement, psychiatric diagnosis and treatment where explanation matters, and identifying intimate partner violence from radiology reports and clinical notes. The subdural hematoma work was described as a pilot; the length-of-stay programme was described as implemented.
Why pulse oximetry bias became a test case for clinical AI inputs
Pulse oximeters, which estimate blood oxygen by shining light through a finger, were calibrated in ways that made them less accurate for patients with darker skin, and the resulting errors feed directly into downstream algorithms. Amy Moran-Thomas described tracing that problem through ethnographic fieldwork and engineering conversations at MIT rather than through a clinical trial of her own.
The documented history is that light- and colour-sensing technologies were developed with predominantly white reference groups. Moran-Thomas noted that oximeters work by transmittance, shining an LED through the finger to a sensor, which makes skin and tissue pigmentation more consequential than in a reflectance-based system.
After she and MIT colleagues wrote about the gap, clinicians who examined their own hospital data found that oximeter discrepancies large enough to change clinical management were about three times more common for patients with darker skin tones. In one analysis, roughly 16% of Black patients were placed in a no-treatment category when treatment was in fact needed. One case she cited involved an arterial blood gas reading of 83 while the oximeter showed 97.
The downstream risk is what Moran-Thomas calls inputs rather than data. Oxygen values are binned inside tools such as the Rothman Index, and AI-guided ventilators may consume the same signals, so a small measurement error can propagate through systems that treat vital signs as unmediated truth.
What a public health engineering lab would need to work
The MIT Public Health and Resilience Lab is an initiative in formation, not a launched institution, and its premise is that public health funding and responsibility in the United States are fragmented, fragile and underdeveloped. Simon Johnson described the effort, jointly developed with Michael Mina, as targeting reactive funding, fragmented responsibility and weak market incentives.
Johnson's comparison is with clinical care, which he said has strong systems in place, while public health lags in resilience. During the COVID-19 pandemic, vaccines and diagnostic tests existed but distribution, allocation and authorisation moved slowly, and cheap surveillance testing was available early without being approved and rolled out.
The proposed mechanism is a prevention economy in which individuals, employers, insurers and industry invest in reducing risk before disease occurs, covering early detection, sensing, population-scale monitoring and pandemic preparedness. Johnson was explicit that relying on a single federal funding stream is a single point of failure, and that self-sustaining revenue matters as much as the technology.
He coined public health engineering, credited to Mina, as a discipline that revives the early MIT civil engineering tradition of sewer systems and infectious disease control while adding market design. The lab's stated horizon is roughly five to 10 years, and Johnson described it as a policy and social incubator rather than a narrow commercial one.
Where the four speakers disagree on the binding constraint
The speakers agreed on trust as a prerequisite and disagreed on what to fix first. The table below compares their positions as stated on 27 April 2026; each entry is one speaker's own framing, not an independent finding.
| Speaker | Role | Named binding constraint | Evidence offered |
|---|---|---|---|
| Barry Stein | Chief clinical innovation officer, Hartford HealthCare | Implementation and execution science, after trust | Deployment experience across the health system |
| Dimitris Bertsimas | Vice Provost for Open Learning, MIT | Education of clinicians and nurses | Teaching sessions with roughly 300 Hartford HealthCare staff at a time |
| Amy Moran-Thomas | Associate Professor of Anthropology, MIT | Social and design gaps in instruments and research scoping | Ethnographic fieldwork and clinician data on oximeter error |
| Simon Johnson | Professor of Entrepreneurship, MIT Sloan | Fragmented, fragile, underdeveloped public health institutions | Institutional and funding analysis |
Stein said implementation is extraordinarily difficult for three recurring reasons: many sites lack the infrastructure to run models at scale, patient data must be protected, and a model that ignores clinician workflow can create more risk than it removes. He summarised the problem as the knowledge gap and the trust gap.
Bertsimas answered the same question with education, describing the artificial intelligence knowledge of physicians, residents and nurses he works with as extremely limited. He noted that medical education still follows a curriculum model from the 1920s and has not absorbed the last decade of AI progress, which is why his Open Learning effort has begun teaching health care professionals.
Moran-Thomas warned against framing the choice as patients versus profits, calling that a false binary, and argued that when social knowledge has no channel into system design, patients and families eventually refuse or withdraw trust. Johnson's answer was that people must understand what they need and why before political change becomes possible, and that shortening the distance between expert conversation and public conversation is the work.
How trust is rebuilt in practice
Trust is rebuilt through disclosure, accountability and transparency, according to Stein, who said patients should know when AI is involved in their care and clinicians should be able to demonstrate competence with the tool. He described trust in health care as something that takes years to build and minutes to break.
Bertsimas gave a concrete account of how trust formed on a cardiac surgery project involving six physicians. The group met every two weeks for 30 to 60 minutes, and the physicians who had initially withheld trust eventually insisted on presenting the work themselves rather than letting him present it, because they had become owners of the ideas.
The practical sequence that emerged across the session was: involve clinicians early and repeatedly, educate patients and staff about what the tool does, publish and monitor what happens after deployment, and accept open criticism rather than defending expertise. Johnson said his podcast with former SEC chair Gary Gensler exists for that reason, and that MIT's reputation for telling people what is and is not established technology is an asset in Washington.
FAQ: AI in health care delivery and the MIT HEALS session
- What is AI in health care delivery? It is the use of machine learning, multimodal models and automation inside the clinical and operational workflow, rather than in a research notebook. At the 27 April 2026 MIT HEALS panel, it covered radiology detection, ambient documentation, 24/7 patient chatbots with clinician handoff, subdural hematoma triage and hospital length-of-stay management.
- What is the MIT HEALS session that these remarks come from? It is a panel titled Delivering on the Promise, chaired by MIT economist Joe Doyle on 27 April 2026, with speakers from Hartford HealthCare, MIT Sloan and MIT SHASS. The session focused on translating research into policy and practice.
- What is xHAIM, the explainable version of HAIM? HAIM is a multimodal clinical modelling framework that unifies images, text, tabular data, time series and genomics into one representation. xHAIM keeps that structure but expresses inputs as language instead of embeddings, so a large language model can produce an explanation alongside the prediction.
- Did the pulse oximetry finding come from the MIT researchers? No. Moran-Thomas and MIT colleagues raised the design-history concern and published it, and clinicians who then examined their own hospital data reported that clinically significant discrepancies were about three times more common for patients with darker skin tones. The data finding belongs to those clinical teams.
- Does the session describe a way to make clinical AI safe for regulated settings? Not directly. The speakers described governance review, continuous monitoring, simulation-based training and education, which are controls a health system can build. They did not present compliance certification or regulatory approval for any tool.
Fork this article
Start a new branch from the same video, shaped your way. You keep the credit; the original keeps the attribution.
A fork in another language is filed as a translation of this article, so the two pages point at each other. You can unlink it later from the editor.
0/240
You are creating
- Format
- For
- Language
- Source
- Your angle
You will be asked to sign in before it is generated.
Buy credits