Podcast
Getting Control of AI in Healthcare
Sep 25, 2026

“These are the open questions for this generation to figure out. We absolutely know the tools are here. We absolutely have to use them because I think they will be better than the current state. And we have to understand how to train people to use them in the right ways, and to open their eyes to where things could go astray. I’m not sure we’ve had that kind of core training yet.”
Tejal Gandhi, MD
Chief Safety and Transformation Officer at Press Ganey
CRICO: Welcome to Safety Net. I’m your host Tom Augello. And today we have two very special guests to talk about how to manage AI at your healthcare organization. All across healthcare today, artificial intelligence is showing up in documentation and decision support, patient messages, operations, vendor tools. The real question isn’t whether AI will be used, but how can healthcare organizations use it safely and responsibly?
The good news is today we have two people who have been working and thinking with their colleagues very hard about the governance around AI and healthcare. First, we have Dr. Kate Humphrey. She is the Associate Medical Director for the Academic Medical Center AMC Patient Safety Organization, and she practices clinically as a pediatric hospitalist at Boston Children’s Hospital, where she is the Medical Director for Patient Safety.
Kate, thank you for joining us today and leading the discussion.
Kate: Great. Thank you. Tom.
CRICO: Now, I know that AI is moving so quickly. Most organizations are still figuring out how to control it in a way that keeps pace. To help us unpack the issues, I’m going to let you introduce our very special guest, Dr. Gandhi, and then lead that discussion around what AI governance looks like in practice. Kate?
Kate: Thank you so much. As we get going. I want to provide a little bit of an introduction. AI governance can sound broad or technical, but for hospitals it’s really quickly becoming practical. So we want to dig into who reviews these tools. How do we know what’s being used and how do we monitor safety issues to help us think through these questions, I’m pleased to introduce my good friend Dr. Gandhi who is the Chief Safety and Transformation Officer at Press Ganey, where she’s responsible for advancing the Zero Harm movement, improving patient and workforce safety, and developing innovative health care transformation strategies throughout her career. Dr. Gandhi is focused on how technology can improve quality and safety, and the implementation of these systems. Thank you so much for being on the program to talk about AI governance with us today.
AI is really such a hot topic, and there isn’t anyone better positioned to talk about the intersection of safety and technology than yourself.
Tejal: Well, thanks so much for having me here.
Kate: Great. Well Tejal, as we get started, I thought first we could talk a little bit about governance and governance structures. So when you think about strong AI governance, governance processes, who needs to be at the table and how can structures like steering committees, AI Centers of Excellence or similar groups help to clarify decision making, accountability and prioritization across these numerous stakeholders?
Tejal: Yeah, and there’s been a lot of examples published about how organizations have been setting up these structures. And I think they all have some common characteristics. And the first is it really does have to be a multidisciplinary committee or structure that is overseeing AI. And when I say that, I mean it’s not something that can be owned solely by your information technology group, for example.
It absolutely needs to have senior leadership from the organization, from the clinical areas, especially depending on the technology, whichever clinical area might be being affected. It might have operations folks because it impacts workflow, etc. and actually very critically, in one area that I see a lot of gaps is it should have that structure, should have quality and safety leaders from the organization as part of it, because often that lens can potentially get missed if that expertise isn’t at the table as well.
And I’ve talked to quality and safety leaders around the country to ask, are you part of that governance structure? And many times the answer is no. And so I just think that as organizations are setting it up, they need to make sure they’re pulling in all of the relevant expertise. And then when they have these structures, it is really looking at a wide array of dimensions as technologies are considered.
And you definitely have to look at the characteristics of the technology. How is it designed, what were the characteristics of the algorithm that was created, etc. And also think about safety implications. Think about equity implications, think about effectiveness. How effective is it truly in impacting whatever outcome you’re trying to impact, etc. and having very clear guidelines or algorithms yourself around how to walk through these various components and rate them and think about does it meet your criteria?
The last thing I would say is if something is directly impacting patients, it may have a different set of cutoffs or criteria than something that say is meant more for something operational or behind the scenes. And so one size doesn’t necessarily fit all in the selection criteria.
And I said that was the last thing. But there is one more thing. You know, you may remember from the era when we were implementing electronic health records, there were a lot of organizations that had various technologies coming in through maybe departments or divisions or a pet project from a few physicians who bring in a technology. And I think what I’ve seen is health systems really trying to ensure that doesn’t happen this time around.
And there is more of a consistent, centralized oversight. So we don’t have kind of I don’t want to say—I’ll use the term rogue—applications out there that we don’t have oversight of, because hopefully we’ve learned our lesson from some of those past challenges with EHR adoption.
Kate: And I think that’s just a wonderful kind of overview. And I want what I’d like to do is dig into a couple of those pieces in a little bit more detail. So from our PSO convening sessions, many of our organizations have initiated much of what you’re describing, kind of a risk-based review of their AI applications and tools.
Can you help us to think about that a little bit more deeply?
And what we’ve heard, Tejal, from our various teams and groups is that some of them are leaning into some of the CHAI-based frameworks around risk assessment of their AI tools. Other teams have created their own internal review processes. And I think this is a nice segue as well into thinking about patient-facing versus non-patient-facing generative AI versus kind of large language model type reviews.
And how can they help to make this a part of their standard process of evaluation?
Tejal: Yeah, I mean I think first of all there are a lot of frameworks out there that organizations are using, some that they develop themselves, some that have come out from organizations like CHAI. And it is we are definitely in a rapid learning phase around which framework might work best for an organization and how to use that framework.
I also do want to make the caveat that you know what an academic medical center can do, maybe very different than what a small community hospital can do. And that’s the challenge because we don’t want small community hospitals to get behind. We want AI to be implemented just as safely there as the big AMCs. And I think there are challenges with having expertise, resources, etc. to do this well in smaller, less resourced places.
And then as you think about assessing risk, there are differences between, say, an algorithm that’s a predictive model based on discrete data elements or what we in the past might have called decision support, where it’s pretty much going to give you the same consistent answer every time compared to more generative AI types of tools where you may get a different response every time, and there’s a lot more variation in what comes out of it.
And so and just based on the prompt, things can change entirely, etc. So I think understanding how to assess risk of those more discrete types of algorithms versus the more generative is going to be really important, because I think it’s much more controlled when you’re talking about a very specific prediction algorithm versus, again, using large language models to do chart summarization in clinic or whatever it is.
So again, we are just learning some of the potential risks of these tools, especially the generative AI tools, because they don’t behave like your typical decision support. So I think those are the ones that that we have to maybe be much more cautious about, because there’s more risk that we maybe don’t understand versus with the algorithms that are very discrete, we know, oh, if the data set gets impacted, that’ll change the way the algorithm works, etc.
We have to understand that, but it’s the generative AI space is a is a whole other ball game. So I think there’s going to need to be different criteria based on the type of AI we’re talking about.
Kate: And Tejal, one of the pieces that I wanted to get some of your thoughts on are the is the use of AI tools that are patient facing? So for example, say you have your organization’s webpage where a patient can go in and put in a symptom and then can get an AI-generated output or guidance for medical advice. Those are some of the types of tools that we’ve heard organizations have paws around. And I’m curious for your thoughts or your reflections.
Tejal: Well, I mean, we have seen patients are having access to these tools, maybe on your website, but also just out in the AI world with various tools where they can input their medical record, they can ask questions, etc. there were studies that show that from a triage perspective, some of these tools actually don’t work that well. And for symptoms that require an emergency room visit are not telling people to go to the emergency room, for example.
So there is a lot of risk with these kinds of tools. I think as organizations maybe want to implement those on their website, etc., to improve access. And I’ll come back to that. There has to be really rigorous testing about what the performance is of the tool and where it fails, and making sure, especially that it’s not failing around those really critical types of issues.
I mean, if it’s failing around symptoms that are not concerning, maybe it’s that’s acceptable, but maybe if it’s failing around symptoms that like chest pain that needs to go to an emergency room kind of thing, you don’t want it failing around those.
But I will say we have to make sure also that we aren’t throwing out these tools or being too cautious with these tools because we’re looking for perfection. Because we know that access is such a problem and we want patients to be able to get seen at appropriate places, be able to share their concerns and know what they should do next.
And in some ways, I’d rather they were doing it with something that’s been curated by a health system rather than out in the wild west of the internet. So it’s a big concern for people that they can’t get primary care access, for example. So if these tools can help improve access to primary care from a safety perspective, that is actually a big win.
And so we just have to make sure we have some of those guardrails, but not necessarily feel like it has to be perfect, because not having access at all is certainly not perfect either. So I just think we have to weigh that every time we are looking at one of these tools. What’s the current state and is this going to actually advance it even though it might bring some additional risk, but it’s probably still better than the current state.
Kate: Absolutely. Agreed. And I feel like some of our stakeholders as a part of the convening sessions, we’re really talking about the human in the loop and trying to think through where the right places, as you’re talking about guardrails, where are the right places for the human to be a part of that guardrail-related system? You know, whether it’s checks and balances, whether that the human is actually inputted within the system at certain decision points. I think there was a lot of conversation that was incredibly interesting there.
Tejal: Yeah. And I think I like the idea of at certain critical junctures, having a human being able to oversee or double check. Et cetera. And I also worry a lot about humans in the loop, because we as humans are not the best at overseeing these kinds of tools and finding a very subtle, potentially error that’s occurring.
I think in certain scenarios, like maybe in the triage type of scenario, it’s probably a better option. Where I worry about it is and we think about documentation, for example, an ambient AI and asking a human to review all these AI-generated notes and find perhaps the one thing that isn’t exactly right.
We know humans are not that great at this, and we’ve seen this from the era of dictation, when we were supposed to review our dictations and no one ever did. And there all these errors and dictations and are we going to kind of come back to that again? So I think we also need to really understand where can humans be effective when they’re in the loop, and where is it really not going to be effective?
And maybe we need other safeguards for those types of situations. It’s not kind of a one size fits all type of thing.
Kate: And Tejal, I think you’re taking us in a direction that were some of the other pieces that were on my mind I wanted to talk a little bit about during this conversation. As we’re thinking about say the use of ambient technologies and being able to validate and verify that the information is actually accurate and then is putting out the type of a report that we find to be valuable and meaningful, both in clinical care but also for our patients.
It raises for me the thoughts around monitoring of AI-based tools. And can you help to dig in a little bit and thinking for organizations, what are things that they should be watching for in their use of AI tools once they are approved? And can you help to describe anything that you’ve seen done well around making it easy for clinicians or staff or for patients to help them raise up concerns? If a tool isn’t acting in the way that they anticipate or intended to.
Tejal: Yeah, well, for that latter part, I think we do need to encourage cultures where people are willing and able to speak up easily when something’s not working as intended to. Often, I think we sort of fix it in the moment and don’t think to tell anyone that something’s broken. So that is a lot of the safety culture.
And we’ve been talking about this since for a long time, but especially I remember in my days as a safety leader if something in the EHR was broken, like, you have to report it, you have to tell someone, otherwise it’ll never get fixed. That messaging has to be super clear to people that don’t just assume somebody knows it’s not working. You should be reporting it.
The other thing, though, that you mentioned is how do we… So monitoring is a really important piece to this because with more traditional decision support, generally speaking. I mean, yes, if somehow a data feed is broken or something changes, it can stop working. But generally once it’s implemented, it works pretty much the same way going forward over time. I’m thinking about like a drug interaction or even a sepsis prediction model or something like that. But with the generative AI tools, performance can definitely shift over time. And so I think there may have been initial characteristics when implemented, but it needs to be monitored first. You know, when it’s working on your own data set because it might work very differently, differently on your own data versus maybe what where it has been tested before, but also your data sets going to change over time. Is the performance going to change over time? So that’s going to be a really key piece as well.
And so building in the time and resources to do that constant monitoring is a big change and lift that we’re going to need organizations to be doing.
And then the other thing you mentioned is how do we make it easier for clinicians to identify when there might be mistakes happening? And that is actually, I think, a really important question. And I don’t know if we’re there yet, but even thinking about, ambient, could the AI highlight areas where it’s less certain? And those are the areas that we are sort of pointing the clinicians towards, hey, double check this because the AI put this, but it’s not as certain about it.
Or there’s been conversation around even having errors in notes periodically put in to see how often physicians catch them, because it’s just sort of it’s like a continuous feedback like, hey you need to be better at X, Y, and Z to catch these. So like having basically it be a test that goes on occasionally to make sure people aren’t just saying, yeah, it’s fine, yeah, it’s fine and doing things like that.
So there’s a lot of different strategies that I think we’re still very early in on understanding. How do we most effectively implement and then monitor and, and make sure these things are having the impacts we want over time?
Kate: No. That’s great. And through the convening sessions, what we did hear from a number of our stakeholder groups is building in and kind of baking in that monitoring strategy up front. So as they’re establishing their AI governance committees, really creating validation checkpoints to say, are we seeing what we intended to see? But what I think you’re also really describing is taking the responses from a black and white, yes, 100 percent “we are certain of the information,” and really doing that quality control and doing a more thorough review and assessment of information. And what I was also hearing you say is lean into some of those traditional systems whether it’s a safety event reporting system, a voluntary reporting system for reporting errors or concerns that folks are seeing, but also trying to think about strategies for just-in-time, in-the-moment, recognition and identification of challenges they may be seeing in tools.
Tejal: Exactly. Yeah. No, I think that’s a great summary.
Kate: And I want to shift gears a little bit because I think we’ve talked a little bit about implementation of tools. We’ve talked a little bit about monitoring and reporting. And I want to shift us a little bit into thinking about our workforce and our clinical staff and how they’re actually engaging with these AI-related tools. And what can clinicians and staff, what do they need to know to use AI responsibly? And how can organizations keep ethical principles, things like fairness, appropriateness, abuse, validity, effectiveness, safety at the center of their work, especially as they’re navigating and using these AI-related tools? Do you have any thoughts there?
Tejal: Yeah. I mean, to me, this is I mean, this is $1 million question because we know the tools are out there. We know people are using them. I don’t know that we have really figured out how best to train people to use these tools safely and effectively. And I mean, I think of even example I mean, it seems so easy, like you just put a prompt in and you get a summary of my patients’ diabetes history and it seems great because, again, it might be better than the current state. Which the current state is I have to go read notes from the last year, and it takes a long time that I don’t have. There was a study that said, if you are in clinic and you’re reviewing notes from the past year, it’s like reading Fahrenheit 451. So, I mean we don’t have that kind of time.
So it does seem great. But then I also have seen studies that say just based on the way you ask the prompt, you might get very different answers back on what that summary of that condition is. And I don’t know that our clinicians understand that or realize that, because I don’t know that we’ve really trained on the implications of how different ways of prompting, for example, can give you entirely different answers, because I didn’t know that until I saw that study.
And I was kind of shocked that it could be so variable based on just how you word the question and there’s all kinds of implications around that. And so we have to understand how best to teach and train people on this.
And then there’s the whole concern about de-skilling as well. As we start relying on these tools, are we going to be losing some really important skills? And there’s a lot of debate about de-skilling. You know, everyone says, Oh, well this generation now doesn’t know how to read a map. And it’s fine because they all have GPS on their phones. And so is that a skill that, is it terrible that we’ve lost? Maybe not, but there are some skills that we don’t want to lose. And I think about that clinical judgment, that something doesn’t feel right. I should escalate. I really do worry about overreliance on technology, losing some of that key clinical intelligence that that we’ve had and that we may just miss things because we are relying on technology that isn’t perfect. So, so but we don’t have answers yet about how to train our physicians and nurses, etc. now that these tools are here, how to train them so they will optimize it.
So I just think these are the open questions for this generation to figure out is we absolutely know the tools are here. We absolutely have to use them because I think they will be better than the current state. And we have to understand how to train people to use them in the right ways, and to open their eyes to where things could go astray. Which, again, I’m not sure we’ve had that kind of core training yet, though I would love to hear at Boston Children’s, for example, maybe I’m wrong, maybe there’s great training that’s already in place, but I know people are talking about it. I’m not sure I’ve seen it implemented.
Kate: And I feel like so many of your comments resonate with some of the next step convening sessions that we’ve started to have. So our next group related to AI that we’ve really been digging into have been with our educators and really thinking about upskilling de-skilling, thinking about is static education really the right way to go? Probably not, for many of the reasons that you’re talking about, and really ensuring that clinicians are having experiential learning opportunities where they do have to interact with the system to really understand how the system works and how they receive the information.
I think what has been incredibly interesting, too, is this dichotomy of learners, right, where you have these incredibly seasoned clinicians who really never had interfaces that involved so much technology with AI use embedded in their EHR and other tools. And then you have your younger clinicians who may still be in training or coming out from training, where they may be very, very savvy with some of the tools. But as you’re describing, they may not have some of the clinical acumen or that experience learning over time from engaging with patients. And I think we’re going to have a lot more to come and hopefully really dig into from some of those convening sessions.
Tejal: So, and I’ll give you an example of a case that that just crystallized for me how important this is, which is this was a true story from a health system where they had an AI tool that was reading radiology imaging, and it told the team that the patient had bilateral pneumothoraces. So air in the lungs on both sides.
So the team called the surgical team that puts in chest tubes. And that surgical team came. And yep the reading was bilateral pneumothoraces. They were kind of at the bedside already to go and put in these chest tubes. But somebody looked at the patient and was like, this patient does not look sick enough to have this reading.
And so they went back and re-reviewed the images. And it turned out the AI was wrong and it wasn’t bilateral pneumothoraces. And what had happened was there had been an update to some algorithm and they didn’t know about it at their organization. Which is another thing as you think about governance, it’s not just the initial, but when there’s updates to the system, how do you check on it? Make sure it’s not starting to perform in a different way. But that person on the team basically halted and said, wait. And then they went and looked and realized. And so they did not insert these chest tubes.
But that just shows the importance still of that clinical judgment. And I do think we need to make sure we don’t lose that in this whole process. And so that was just a clear example to me of where you could potentially over rely on the technology, but you have to still feel like you can question it, escalate, stop the line, all of those kinds of things.
Kate: Gosh, and Tejal, I’ll just say, I think that is such an incredible example because if we bring it back, as you’re suggesting to the AI governance pieces: the monitoring, the evaluation of the system, and then how did those clinicians, how were they able to provide feedback to say, gosh, this is not functioning as we intended. But then in the end, this is real-world clinical application of these tools and how it’s actually affecting our care at the bedside and our clinical decision making.
So I just think that is a tremendous example that you just highlighted.
And Tejal, I want to thank you first for all of your time and conversation today. And as we start to close up our conversation, I was wondering if you had one thing that you could say to our listeners that they could bring back to their organizations about AI governance around the implementation of these AI tools and their health care systems, what would that one thing be?
Tejal: I would say, well, it might be two things. One is definitely please, please, please make sure that as you do this governance, you include people with the expertise from quality and safety because they can definitely be helpful both with that initial piece, but also that monitoring plan that’s needed and thinking about metrics and all those kinds of things.
And then I love what you said, and I’m just going to double down on it, building in the monitoring from the beginning, because I just think that we’ve done a lot of work on governance in technology for the last 20 years or so, and so I think we’re pretty good at that initial selection.
But this is a new day about having to monitor over time that I don’t know that a lot of organizations have built in. And so I think really doubling down on having that monitoring piece thought through from the very beginning and building in the time and resources to do it is going to be absolutely critical.
Kate: Well Tejal, again, thank you. This has just been great. You know, as I reflect back on AI governance it really ultimately is about patient safety, trust and organizational accountability. And the challenge that we know is that AI tools are being used. We have to focus our oversight on the highest-risk applications, including those patient-facing ones. We have to monitor for those unintended consequences and make sure that people really do understand and know how to raise up those concerns.
Thank you.
Tom?
CRICO: Thank you. Kate, and thank you for joining us today, Tejal. I hope you come back. Dr. Gandhi is the Chief Safety and Transformation Officer at Press Ganey, where she is responsible for advancing the Zero Harm movement, improving patient and workforce safety, and developing innovative healthcare transformation strategies. And a very special thanks to our facilitator, Dr. Kate Humphrey. Kate is the associate medical director for the Academic Medical Center, AMC Patient Safety Organization for CRICO, and she practices clinically as a pediatric hospitalist at Boston Children’s Hospital, where she is the Medical Director for Patient Safety.
And I’m Tom Augello for Safety Net.
Commentators
- Kate Humphreys, MD, MPH
- Tejal Gandhi, MD
About the Series
We’ve got you.
Our Safety Net podcast features clinical and patient safety leaders from Harvard and around the world, bringing you the knowledge you need for safer patient care.
Episodes
Safe Births are a Team Effort
Taking the Mystery Out of Being a Clinician Defendant
Defending Providers is Different Today, Says Legal Expert After 45 Years
How Application Forms and Burnout Threaten MD Mental Health and Patient Safety