The AI Is in the Scope. The Evidence Isn’t.
- By Winston Thomas
- September 29, 2026

Mark Lambrecht has a friend who finds polyps for a living. The friend’s colonoscopy system has an AI detection function. He keeps it off.
“He says, ‘If I would turn on the function, it would be faster than I,’” Lambrecht says. Accuracy is not the issue. His friend did not study medicine for 12 years to be replaced by an AI.
The story carries extra weight coming from Lambrecht. In 2019, he told Becker’s ASC Review that deep learning could help gastroenterologists evaluate polyps found during colonoscopy. Gastroenterology has since become the most tested field in clinical AI: of the 86 randomized trials reviewed in The Lancet Digital Health, 37 were in gastroenterology. Even the best-tested use case in medicine can lose to one doctor’s reluctance.
Lambrecht is the senior director and global head of health care and life sciences at SAS. He joined the company in 2005 after a PhD at KU Leuven and postdoctoral research at Stanford. When we spoke in Singapore during SAS Innovate 2026, he kept returning to one idea: the gap between what AI can do and what doctors actually use.
Approvals are easy. Evidence is hard.
The numbers show how wide that gap is. As of March 2026, the FDA’s list showed 1,524 authorized AI-enabled medical devices, 76% of them in radiology. Very few have been tested the way drugs are, in trials that compare patients who use a tool with similar patients who don’t.
In August, a study in PLOS Digital Health examined the 1,357 AI devices the FDA had cleared through Dec. 5, 2025. Only 34 were linked to registered clinical trials, and only three had been tested on patient-centered outcomes, such as death rates or hospital readmissions.
“The big problem is prospective evidence,” Lambrecht says. Prospective means the study is planned before patients are treated, not pieced together from old records. Doctors, he says, are not the obstacle. “There’s not a resistance, but doctors need evidence that the technology will truly help their patients.”
His answer: “We need to start treating AI as if it’s a medicine.” In other words, test it the way you would test a drug: Does it improve patient outcomes? “That is not happening consistently.”
Lambrecht estimates that only 2% of algorithms reach the clinical workflow. Asia-Pacific data is less bleak but points the same way. Bain & Co. surveyed 600 doctors in Australia and the Philippines for its 2026 regional healthcare report. The firm says only 30% of proof-of-concept projects reach production. About one in three doctors said their organizations are not ready to deploy AI at scale.
“We need to start treating AI as if it’s a medicine.” — Mark Lambrecht, SAS
Projects that make it through do pay off. Bain reports that more than half of organizations already deploying AI see a meaningful return within 12 months. Drugmakers show a similar split. In an April 2026 Deloitte survey of 150 life sciences executives, 45% said their AI initiatives had produced measurable improvement, but only 13% saw it at scale.
The returns are real, but they go to the few projects that survive the move from lab to ward. Lambrecht says most fail at that move, when models trained on clean, narrow data meet real hospitals. “You come into a messy population with different backgrounds,” he says.
Some health systems make that move easier than others. Lambrecht points to the one where we met.
Singapore, the test bed
Singapore has three public healthcare clusters, about 6 million people and a small land area. Its data is still fragmented, Lambrecht says, but less than in larger systems.
It also shows what working clinical AI looks like: useful, supervised and far less glamorous than a demo. Singapore’s national diabetic retinopathy program screens patients with SELENA+, a deep learning system approved as a medical device.
SELENA+ does not work alone, and its pilot data shows why. It caught 94.7% of patients who needed referral and correctly cleared 82.2% of those who did not. Human graders scored 98.9% and 97.2%. The AI adds speed and scale, and trained humans make the final call. This is the partnership Lambrecht argues for.
He wants retinal scans to do more. Citing cardiologist Eric Topol, he says a single retinal image can already signal diabetes or Alzheimer’s risk, yet almost no clinic uses it that way. Singapore reads retinas for eye disease at national scale. Reading them for disease elsewhere in the body is the next step.
Generative AI is moving in too. By June 2025, SingHealth’s Note Buddy, which turns consultations into draft clinical notes, had helped more than 2,100 healthcare workers create over 16,000 notes.
Regulators are carefully clearing a path. In February 2026, the Health Sciences Authority confirmed a regulatory sandbox, a controlled testing space, that lets public healthcare institutions share in-house AI tools without product registration. It covers only lower-risk software, and a senior clinician must oversee each tool.
The reason for the rush is age. In May 2026, Health Minister Ong Ye Kung said more than 21% of citizens are 65 or older. Japan is at 30%. Hiring more doctors and nurses will not close the gap, Lambrecht says, because there are not enough of either. “We can mitigate the impact at least if we do investments wisely.”
For Lambrecht, investing wisely starts with a design question: which kind of AI gets to make which kind of call?
Keep the chatbot out of the decision
Lambrecht’s answer is architectural. A deterministic model always gives the same output for the same input. Large language models (LLMs), which generate answers based on probability, do not. So SAS keeps them out of clinical decisions.
Models that don't give the same answer every time, he says, are “not core to the decision that is being made to those patients.” In his setup, AI agents act as a harness. They take a doctor’s question in plain language, run it through SAS’s deterministic models and return the answer. “We can show the code, we can reproduce it.”
Ask a chatbot like ChatGPT the same question twice, Lambrecht says, and you may get two different answers. “That is not acceptable in a medical environment.”
It’s not about perfection. Doctors vary too, he notes. One who had a heavy weekend may decide differently on Monday than a well-rested one on Friday. The real test is whether the risk to the patient can be controlled.
“There’s not a resistance, but doctors need evidence that the technology will truly help their patients.” — Mark Lambrecht, SAS
Regulators are moving the same way. On Jan. 14, 2026, the FDA and the European Medicines Agency jointly released 10 principles for good AI practice in drug development. South Korea’s AI Basic Act, in force since Jan. 22, 2026, covers “high-impact” AI, a category that includes healthcare. It requires operators to explain AI outcomes to users, provide human oversight and document safety measures.
For CDOs, the design doubles as a compliance strategy. Doctors get a conversational front end, and auditors get math underneath that they can check and rerun.
Both the principles and the law lean on the same safeguard: a human in the loop. That safeguard is weaker than it looks.
The loop has a human. Is anyone reading?
Asked where the human-in-the-loop line should sit, Lambrecht says the question has not really arrived yet. “It’s always the doctor that will make the decision.” AI agents beat doctors in simulations, he says, but not in real practice.
The line wears thin in paperwork. AI scribes listen to consultations and draft the notes. A Texas malpractice insurer warns that doctors may rubber-stamp those notes without careful review as they come to rely on them. Singapore planned to offer Note Buddy across all public healthcare institutions by the end of 2025.
Consent is a second risk. Lawsuits in California and Illinois allege that health systems used ambient scribes without patients’ informed consent.
The doctor still signs every note. The CDO’s question is how long the doctor spent reading it first. A rushed note is one input that can go wrong without anyone noticing. The data that trains the model is another.
Bias is not poisoning
I asked Lambrecht how SAS stops poisoned real-world data, such as readings from wearables and consumer apps, from corrupting clinical models. He answered as a statistician.
SAS has “hundreds of PhDs in statistics” working to remove bias, Lambrecht says. Causal inference, a set of methods that separate cause from coincidence, can find cause and effect in data that was never collected for that purpose. Wearable data is consumer grade, not medical grade. And whatever the source, “you have to use it in provenance,” meaning you must know where each piece of data came from.
That is a strong answer, but to a different question. Bias is accidental skew. Poisoning is deliberate: an attacker plants false data on purpose. A Nature Medicine study found that replacing just 0.001% of training tokens, the word fragments a model learns from, with medical misinformation produced harmful models. Those models scored as well as clean ones on standard medical tests. A 2025 study by Anthropic, the UK AI Security Institute and the Alan Turing Institute found that as few as 250 malicious documents could plant a hidden trigger, known as a backdoor, in models of every size tested. The number of poisoned documents needed did not grow with the model.
Statistics can correct skew, but it cannot stop an attacker. That makes poisoning a job for the CISO, not only the data scientist.
Provenance matters more as data moves between organizations. In healthcare, the biggest movement runs between hospitals and drugmakers.
Two industries, one patient
Lambrecht oversees both SAS’s healthcare and life sciences communities. He thinks hospitals underestimate how much drugmakers want to work with them. They meet at real-world data: records of how medicines and devices perform in ordinary patients, outside clinical trials.
Twenty years ago, he recalls, a hospital might keep 20 computers for 20 drug companies, with nurses typing the same patient data into each one. Today many drugmakers invest in pulling that data directly from electronic health records instead.
For rare diseases, where there are too few patients for a full trial, SAS uses real patient data to generate synthetic data: artificial records that mimic real ones. The method only works with real data underneath. “You need real data to make sure that your statistical signal is in fact in the large data,” Lambrecht says.
His 2026 prediction for SAS follows the same logic: significant investment in joining discovery and clinical analytical data. Joining the data also means joining the people who own it.
No more islands
Will the healthcare CDO absorb the clinical authority of the CMIO, the chief medical information officer who links clinical staff with health IT? Lambrecht says no. He still sees the CMIO as the bridge to clinicians, including training them to use AI. He expects the two roles to work more closely and partly blend, as long as leaders stop “sitting on islands.”
He has warned about islands before. In 2020, he told Pharmaceutical Executive that the biggest pitfall was building analytics platforms on an “island,” set up without a clear purpose. Six years later, the island is the org chart.
Agentic AI is coming to hospitals, he says, “whether they like it or not.” His friend with the colonoscope will have to adapt, “but it will take time.” The AI is already inside the scope. What it still needs is enough evidence to win over the doctor holding it.
Image credit: iStockphoto/aurielaki
Winston Thomas
Winston Thomas is the editor-in-chief of CDOTrends. He likes to piece together the weird and wondering tech puzzle for readers and identify groundbreaking business models led by tech while waiting for the singularity.