Why Medical AI Needs a Referee | Protege's Engy Ziedan
- 01The Training Data Ceiling Will Define AI's Limits
- 02Benchmark Performance Is Dangerously Disconnected From Real-World Clinical Utility
- 03Subtle Misalignment Is More Dangerous Than Catastrophic Failure
- 04AI in Healthcare Has No Independent Referee
- 05Benchmark Contamination Is a Systemic and Underappreciated Problem
- 06Static, Retrospective Quality Measurement Cannot Keep Pace With Rapidly Evolving AI
1. Key Themes
The Training Data Ceiling Will Define AI's Limits
Protege's founding insight, written in a 2024 Google Doc by co-founder Bobby, was that models would be constrained by the quality and reality of their training data — long before the broader market arrived at that conclusion.
"Models are going to be inhibited in their usefulness by the training data available for them. And that goes beyond healthcare, in any domain. And so it became our mission to provide safe and aligned data, essentially the teachings that would make AI useful for humans." [00:00:11]
"He also had I think like the eye for the humans who annotate and generate data will reach the saturation of their imagination, and that there is nothing like reality. Today you kind of see the market now saying, oh, it's actually real world data that we want, and like synthetic data maybe is insufficient." [00:04:33]
Benchmark Performance Is Dangerously Disconnected From Real-World Clinical Utility
A model can ace thousands of standardized test questions and still fail at the actual clinical tasks it is deployed to perform. The gap between credentialed performance and operational performance is the central problem Protege is solving.
"Models, you know, can score up to like 92% on a licensing exam but 45% on a real-world clinical task. And so, like, acing a benchmark doesn't make an AI ready for the hospital just like a perfect MCAT score... doesn't make for a great surgeon." [00:20:13] (Daisy Wolf summarizing Engy Ziedan's prior statements)
"If you are a physician and maybe I ask you like as a patient, right, you're about to get spinal surgery. They tell you this model has passed like a 5,000 question, 10,000 question Q&A exam. Would you trust it versus would you trust like a physician that has done 4,000 of these surgeries, right? What you care about actually isn't the general knowledge of the model... I care, has it been in this scenario before? And has it been helpful?" [00:21:28]
Subtle Misalignment Is More Dangerous Than Catastrophic Failure — And Harder to Catch
The industry focuses on preventing obvious, dramatic AI failures. But the more pervasive and harder-to-detect risk is subtle bias and misalignment embedded in everyday clinical workflows, often invisible until damage is done.
"When people talk about kind of like safety and alignment, right, there is the catastrophic failure, which is something goes wrong. Can you prevent that? Actually, I do think that that problem is easier because catastrophic failure is like narrow... The harder one is misalignment broadly and subtle bias. First, it's very hard to define. It's even harder sometimes to detect." [00:09:55]
"Is it fair to say that if you were to be prescribed a medication, it's okay for the model to first check if you owe the hospital any money? Is it fair to say that if the hospital said, don't refer people out of network, right, that that would be okay, even though your preference might be I'd like to go out of network sometimes." [00:11:15]
AI in Healthcare Has No Independent Referee — Everyone Grades Their Own Homework
Every vertical AI vendor publishes their own evals claiming to be best, creating a market with severe information asymmetry. There is currently no arm's-length, independent body continuously evaluating clinical AI performance.
"Every vendor of an AI tool, every vertical AI builder, would have, like, a little brochure that says, I am best. And so the no one is that if everyone says they are best, no one knows who's best. Also, there is no incentive to kind of pinpoint to where your product fails. And so really, on behalf of, like, healthcare systems or patients, there is no one that is doing this independently, that's on arm's length removed, that's looking beyond the iceberg of, like, catastrophic failures and misalignment." [00:13:57]
Benchmark Contamination Is a Systemic and Underappreciated Problem
Models may be memorizing benchmark answers rather than genuinely reasoning, because the vast majority of available healthcare data has already been used in training. Creating truly uncontaminated evals requires pulling data that has never been seen by any model.
"80% of what we have has been given for training... if your definition of basically independent is a patient that the model hasn't ever seen before, you may have a challenge in front of you." [00:30:31]
"For like the oncopathology example, we literally had to pull patients whose whole slide images had never been scanned before. We contacted hospitals and we're like, in your pathology drawer, who hasn't been scanned before? And they're like, net new slides." [00:30:59]
Static, Retrospective Quality Measurement Cannot Keep Pace With Rapidly Evolving AI
The existing healthcare quality infrastructure (Value-Based Purchasing, retrospective reporting) was built for slow-moving interventions. AI changes so fast that annual or even quarterly evaluation cycles create dangerous blind spots.
"You risk, not to fearmonger, but you risk a situation where we enter, like, the opioid epidemic when it peaked in 2010. That's only when people started saying, okay, National Task Force, we're calling it an epidemic, we're acting on it, let's move. We don't want to get there." [00:17:07]
"In AI evaluations, you can actually ask a thousand questions about how an agentic system is behaving inside the hospital." [00:29:12]
Physician Preference Variation Breaks Real-World AI Evaluation
Because physicians have sticky, idiosyncratic clinical preferences, using physician decisions as the ground truth for AI evaluation introduces systematic error — the physician being used as a reference may themselves be wrong.
"There are physicians who never do a full knee replacement. He's only partial. And then physicians who, if he touches a knee, it's a full knee replacement. Never does partial. And that's his beat. You take this medical record, this real-world case, and you put it in front of a model, and you say, what would you recommend? And it would recommend the opposite of what the physician did. And you say, oh, model's wrong. But actually, it could be that the physician is wrong, right? He has a sticky preference, a hysteresis almost." [00:22:54]
Ambient AI Documentation May Inadvertently Reduce Clinical Bias
One underappreciated upside of ambient documentation tools: they strip out the subjective, often biased language physicians embed in notes, potentially producing more clinically accurate records and fairer downstream care.
"A tool that records a conversation between a patient and a physician kind of like cuts that out completely. What was said is now documented into a soap note, right? None of that is in there. And now we have a very accurate clinical description of the conversation. A third-party physician that's looking at the conversation would be like, oh, I should probably order a clinical exam to investigate the cause of this pain. I don't think it's just in her head." [00:12:11]
2. Contrarian Perspectives
The Government Is Not Going to Solve AI Safety in Healthcare — And Waiting for It Is the Unsafe Choice
The conventional instinct is to defer to regulators. Engy argues the opposite: waiting for government action is itself a safety failure, and the right move is for industry to act now and invite challenge.
"A lot of people just say like, why isn't the government doing this? Because the government is not going to do it. And like basically what's more unsafe, right, is like sitting there and saying, well, I'm going to wait for the government to do it. No, we think what's more safe is just do it. Right? And then let people challenge you on it... The government can't even get you to file taxes online. Like people had to invent like a marketplace online for that to file for the government." [00:28:15]
Catastrophic AI Failures Are the Easy Problem; Subtle Misalignment Is the Real Threat
Most AI safety discourse focuses on dramatic failures — a model that kills someone. Engy argues these are actually the easier problems to solve because they're narrowly definable. The harder, more insidious threat is misalignment that benefits institutional actors at the expense of patients, which is nearly invisible.
"Catastrophic failure is like narrow. Like you could just define it as like mortality, right? The harder one is misalignment broadly and subtle bias. First, it's very hard to define. It's even harder sometimes to detect." [00:09:55]
"There are two indifference curves that are coming together on a contract curve. One is trying to maximize total revenue and basically minimize payouts. And one is trying to maximize revenue recovered, right? They are both using agentic tools. And the patient is in the middle. The patient has no agency." [00:10:25]
AI Performing Worse Than Random on a Task May Signal Misalignment, Not Irrelevance
The standard research community reaction when a model performs below chance is to dismiss the task or tolerate slight contamination. Engy argues the opposite — below-random performance is a red flag for serious misalignment that demands investigation.
"Worse than random could be a detection that there is serious misalignment too. Like, why is it doing worse than random? We should just care that this patient has not entered training previously." [00:31:54]
Real-World Healthcare AI Evaluation Requires Testing at the Point of Care, Not Just Retrospectively
The standard approach is to evaluate models offline against historical data. Engy argues that for high-stakes, expert-level tasks, this will eventually become insufficient — testing must happen live, at the point of care.
"At some point, a task will be so expert, right, so high risk, that testing at point of care will become really warranted." [00:23:21]
The Data Supplier Is Best Positioned to Be the Independent Evaluator — Because of Network Effects, Not Conflict of Interest
The intuitive reaction is that a company supplying training data to models cannot objectively evaluate those models. Engy argues the opposite: because Protege knows exactly which data each model has seen, they are uniquely positioned to create uncontaminated evals, and network effects force strict impartiality.
"We know which patients have entered which models. We've given this data for pre and mid training." [00:31:29]
"There's also an incentive to completely remain impartial, right? Because like any network effect idea, lack of trust destroys the entire network almost immediately. It's like explosive." [00:27:15]
3. Companies Identified
Protege
Healthcare AI data and evaluation company. Co-founded by Engy Ziedan (Chief Scientific Officer) and Bobby. Supplies pre- and mid-training data across healthcare (medical notes, endoscopy video, pathology images, audio) as well as other verticals including robotics and veterinary care. Mentioned as having data relationships with nearly all major foundation models, and now building an independent evaluation and benchmarking function for clinical AI — acting as an arbiter between competing vertical AI builders.
"Almost all the models have been pre- and mid-trained on our healthcare data. We provide data in multiple verticals, audio, video, robotics." [00:03:19]
"Once the evaluation is done and we hold this test set that never leaves. If you have a deficiency in an attribute that's now been revealed, we can tell you exactly what data would improve the model's performance on that deficiency." [00:26:50]
Open Evidence
Clinical AI tool used by physicians to query medical knowledge in the course of practice. Mentioned as a real-world example of AI being actively used at the point of care today.
"Doctors are using tools like Open Evidence to ask medical questions to AI in the course of their practice every day." [00:13:05] (Daisy Wolf)
ChatGPT (OpenAI)
Mentioned as the most visible consumer-facing AI tool through which hundreds of millions of people seek health information, with no independent verification of answer safety or accuracy.
"Hundreds of millions of people ask ChatGPT questions about their health." [00:00:00] (Daisy Wolf)
4. People Identified
Engy Ziedan
Co-founder and Chief Scientific Officer of Protege. Healthcare economist, assistant professor at Indiana University (previously at Tulane). Her work has been featured in the New York Times and cited by the CDC. Wrote her dissertation on hospital readmission penalty policy under Value-Based Purchasing. Brings a rigorous economic framework (hedonics, comparative advantage, information asymmetry) to AI evaluation problems in healthcare.
"I often talk to the team about the effect of having a doctor in the family. There's research from Sweden that shows like just a random assignment in who goes to med school based on your grade in the med school entry exam. Randomly assigns a doctor in the family and people live X years longer, the entire family." [00:08:22]
Bobby (Protege Co-Founder)
Co-founder of Protege (last name not stated). Previously worked for a data facilitation company. Had early, prescient insight that models would be constrained by training data quality and that real-world data would ultimately prove superior to synthetic data — documented in a February 2024 Google Doc that Engy re-reads every few months as a founding document.
"He had this ambition... and in typical Bobby style, it was human-written, very simple, and it said something like, what do we hold to be true about the future? And in it, it basically lays out this hypothesis... that models are going to be inhibited in their usefulness by the training data available for them." [00:02:53]
"He also had I think like the eye for the humans who annotate and generate data will reach the saturation of their imagination, and that there is nothing like reality." [00:04:33]
5. Operating Insights
Flywheel of Eval-to-Data Improvement as a Proprietary Moat
Protege has constructed a closed loop that competitors cannot easily replicate: run an independent evaluation, hold the test set privately, identify the model's specific deficiencies, then immediately prescribe and supply the exact training data needed to fix those deficiencies. This turns evaluation from a one-time report into a recurring revenue and stickiness engine.
"Once the evaluation is done and we hold this test set that never leaves. If you have a deficiency in an attribute that's now been revealed, we can tell you exactly what data would improve the model's performance on that deficiency. And because of the size of the data set we have and the various sources that we have, it's a very fast way for you to evaluate and then reinforce, right? Where the model is failing with like new teachings." [00:26:50]
Trust Architecture for a Multi-Party Network: Blind Methodology, Silent Exit Rights
When running head-to-head evaluations between competitors, Protege preserves impartiality through structural design: methodology is shared with all participants upfront, competitor identities are concealed until the end, and participants retain the right to exit silently without their results being disclosed. This removes the incentive to refuse participation and preserves the integrity of the network.
"All the methodology is shared broadly with everyone in the subnode. We don't reveal the names of the people, the vertical AI builders in the subnode until the very end. They have the right to a silent exit and things like that." [00:27:46]
Use Economic Theory (Hedonics, Comparative Advantage) to Justify Product Positioning
Engy explicitly applies academic economic frameworks — hedonics for pricing AI value, information asymmetry for healthcare market structure, comparative advantage for why Protege belongs in evals — to make product and strategic decisions legible to sophisticated buyers and investors. Translating domain expertise into established theoretical language accelerates trust-building with enterprise customers.
"If you think about it, economists back out kind of value from things using an approach called hedonics, which is like someone's willingness to pay for the thing would give you the value of the thing... No one ever asked, what is the value of Uber? Like show me the eval. People just paid and gone on the ride. But in today's AI market, there is a need for the pricing to be accurate." [00:07:03]
6. Overlooked Insights
Knowing Which Data Entered Which Model Is Itself a Rare and Enormously Valuable Asset
This was mentioned almost in passing, but the strategic implication is massive. Protege knows, with specificity, which patients and data sets have been fed into which foundation models for pre- and mid-training. This is not just useful for contamination-free benchmarking — it is a form of a proprietary map of the AI training landscape that no other entity outside the model labs themselves possesses. This knowledge could underpin regulatory compliance infrastructure, audit trails for liability, and entirely new data licensing models as AI liability frameworks develop.
"We know which patients have entered which models. We've given this data for pre and mid training." [00:31:29]
"This moratorium on any data set or any patient that has entered training. They are not allowed to be in a benchmark or an eval." [00:31:54]
Medical Notes Contain Embedded Social and Racial Bias That AI Will Systematically Amplify — and Ambient Tools May Accidentally Fix It
Mentioned as a brief illustrative example, this is actually a deeply significant finding with regulatory and commercial implications. Protege has access to hundreds of billions of medical notes and has identified that clinician-written notes routinely encode social judgments that drive diagnostic bias downstream. This means every model trained on EHR data inherits that bias. But ambient documentation tools — which transcribe the actual clinical conversation rather than the clinician's interpretation — may structurally eliminate this bias pathway. This creates a strong, evidence-based clinical argument for ambient AI adoption that goes far beyond productivity gains, and implies that models trained on ambient-sourced notes would be measurably less biased.
"Some medical notes will say things like, she looks really disheveled. Her husband is asking good questions. What they're actually saying in the note is that she's not trustworthy. And suddenly you see a diagnosis that's like, oh, that pain she has is potentially mental illness, right? We're not going to clinically look into it. Now, a tool that records a conversation between a patient and a physician kind of like cuts that out completely... And now we have a very accurate clinical description of the conversation." [00:11:44]