Diadia Transparency Lab · Methodology
Checking medical claims without trusting the model
The problem with frontier LLM models is that they can produce unreliable, opaque results that are hard to verify without special tools and domain knowledge of a clinician. What might sound plausible to a regular user may not be trustworthy to a clinician with deep expertise in a particular area.
It creates obvious risks for regular users who trust the model output for medical guidance in their most vulnarable state. But it also creates risks for clinicians and makes their work more challenging, all of which combined makes it harder to deploy these models in real medical settings.
We believe that AI is supposed to empower clinicians to do more with less effort, and with higher quality, accuracy and better outcomes for patients.
Similarly, we believe that patients should be able to trust AI to help them navigate and root cause their medical problems.
In order to address this problem we have developed a system that reasons across thousands of sources of evidence and verifies each claim through a constructed mechanisms graph (shown below). In this study we show how it works and how such system can improve the accuracy and trustworthiness of AI diagnosis.
The problem with confident answers
There is a decent amount of research on this topic. One 2025 study found that 30-50% of LLM medical statements were lacking proper evidence.11 Medical Hallucinations in Foundation Models and Their Impact on Healthcare (Kim et al., arXiv, 2025). In a survey of healthcare professionals, 91.8% of clinicians said they'd run into confident but wrong medical statements from AI systems.
A plausible sounding claim cites a credible journal, uses correct terminology, and even gets the tone right. These are so called - faithful hallucinations.
A simple fact-checking or LLM search alone do not solve this problem, both provide true or false result on the whole claim, but it does not work when the claim is only partially correct. In medicine, the complexity is hidden in these exact borderline cases which only an experienced clinician can reason about.
The other technique that's commonly used - is asking a second model to review the result of the first one. But in practice, models often disagree. Depending on how a given model was post-trained, what kind of data was used and what the model was optimized for determines its problem solving capabilities and its final verdict. Even running the same model twice on the same input may produce subtly different results.
Mechanistic reasoning
Here is a claim one model actually produced:
“Chronic stress disrupts cortisol signalling, leading to HPA axis dysregulation and fatigue.”
Let's break it down into its three separate causal steps:
The first two have decades of direct well researched evidence behind them.22 Chronic stress and the hypothalamic-pituitary-adrenocortical axis in humans (Miller et al., Psychological Bulletin, 2007). The third is a bit weaker, a real mechanism but with much weaker direct evidence. Any single verdict on the entire sentence has to average over that, and averaging is exactly the wrong operation.
Instead of asking for verdicts, we decompose the claim into a mechanisms graph. Nodes are the biological entities involved (a biomarker, a process, a symptom), while edges are the causal links the claim asserts between the biological entities.
We then use our specifically tuned model to evaluate each edge at a time in the overall patient's context, with the evidence gathered and ranked for that edge, and answer a question: does this evidence support this link? That's a task models are actually good at.
The claim-level conclusion is then computed from the labelled graph by deterministic rules. No model touches the aggregation, so the same inputs give the same result on every run.
The animation below walks through the whole pipeline on that this claim.
Three kinds of reasoning
It became clear early that simply searching for a respected trial for the claim was not the right approach for a lot of legitimate medicine and leads to a lot of false positives or false negatives. Simply put, the absence of a trial does not mean the claim is false — it simply means no one has tested it yet.
Sometimes studies demonstrate the link directly. Chronic stress raises cortisol; people have measured it in controlled settings many times.33 Acute stressors and cortisol responses: a synthesis of laboratory research (Dickerson & Kemeny, Psychological Bulletin, 2004). Most textbook medicine is like this and it's the easy case.
Sometimes nobody has tested the claim end to end, but every intermediate step is correct. GLP-1 drugs delay gastric emptying. Delayed emptying reduces appetite. Reduced appetite produces weight loss, and weight loss improves insulin sensitivity. Four well-studied links, but no single trial covering the full chain.44 Mechanisms of Action and Therapeutic Application of Glucagon-like Peptide-1 (Drucker, Cell Metabolism, 2018).
Sometimes the connection runs through a shared pathway. A gene variant drives appetite problems through a particular brain circuit, and a drug is known to act on the same circuit. No trial connects the gene to the drug, and there may never be one — the population is too narrow. But the shared circuit is a real mechanistic reason to consider the connection, and each half of it can be checked against evidence. This pattern shows up constantly in precision medicine, and it's the one LLMs struggle with and binary tools are not able to handle.55 Network pharmacology: the next paradigm in drug discovery (Hopkins, Nature Chemical Biology, 2008).
What "plausible" means
Our system tags each edge in the mechanisms graph with one of three labels.
Supported by science: direct, consistent evidence demonstrates the link. In most cases, it will be backed by recent high-quality SR/MA, large RCTs, meta-analyses, etc.
Unsupported: the evidence directly contradicts the link, or no coherent mechanism can be built at all. Note what's missing from that definition. An absence of studies does not make an edge unsupported. Only contradiction does.
Plausible: is everything in between. The label means a sound reasoning path exists but the evidence for it is indirect. Animal work, small trials, a chain of evidenced steps never tested end to end. When the system assigns it, it has to store the reasoning path in the graph itself, so anyone can inspect why the connection was considered plausible instead of taking our word for it.
In our evaluation, 47% of the 145 claims came out plausible. More than supported (30%) and unsupported (23%). We'd argue this tier is the whole point of the system. A binary tool either throws these claims away for lacking trials or waves them through. Either way the clinician loses the one thing they needed to know: which parts of the reasoning are direct and which are inferred.
The conclusion is deterministic
After all edges are labelled the system traverses every path through the graph from a root cause to a final outcome. A path containing a contradicted critical edge is broken. A path where every edge is supported is fully supported. Anything else counts as plausible. Then the claim: all paths supported means supported, any broken path means unsupported, otherwise plausible. Confidence follows the weakest link.
When a clinician asks why a claim was rejected, there's always a specific edge with specific evidence against it.
The demo below is the real roll-up running on a two-path example. Change any edge's label and watch the claim-level result recompute.
- Path 1: e1 → e2 → e3 — plausible
- Path 2: e1 → e4 → e5 — broken
Frontier model performance
We evaluated the system on four different frontier models: GPT-5.2, Claude Sonnet, Claude Opus, and Gemini, using data drawn from 20 real medical reports, containing 145 claims of various medical severity.66 Claim-Level Transparency Analysis of LLM-Generated Diagnostic Reports (bioRxiv, 2026 — full evaluation protocol).
Every model produced claims that failed verification, but the rates spread widely: 11% unsupported for GPT-5.2, 15% for Claude Sonnet, 24% for Claude Opus, 34% for Gemini. The dispositions were stable across patients. Gemini speculated the most. Sonnet hedged the most — 65% of its claims came out plausible. GPT-5.2 stayed nearest to established medicine, which is not the same thing as being most useful; its reports simply took fewer risks. Report length predicted nothing: Gemini's reports were half the length of everyone else's with the same claim density.
Two cases stuck with us.
A 32-year-old woman's panel showed fasting insulin of 1.4 uIU/mL, below the reference range, glucose normal at 72. One model called this "the most clinically significant finding in your entire panel" and the likely cause of her fatigue. All three steps of that reasoning failed evaluation; there's no clinical evidence tying isolated low insulin with normal glucose to fatigue. And the evidence search surfaced two corrections on its own. Low fasting insulin with normal glucose usually indicates good insulin sensitivity, a healthy finding. And the standard workup for hypoglycemia requires documented symptoms with the low value, not one number on one draw. The pipeline didn't just reject the claim, it found the correct interpretation that the model missed.
The other case: a mycotoxin panel with gliotoxin at 78.68 ng/g, inside the lab's reference range (upper limit 207.87). A model described this value as "extremely high," blamed Candida overgrowth, and called it a severe mitochondrial toxin. Per edge, the picture split cleanly. "Extremely high" contradicts the lab's own range. The Candida attribution contradicts the microbiology. But Aspergillus really is the primary source of gliotoxin, and immunosuppression really is a documented effect. Half wrong, half right, and a true/false answer would have flattened it either way.
The matrix below is the 19% agreement figure, per patient and per model.
| Sonnet | Opus | GPT-5.2 | Gemini | |
|---|---|---|---|---|
| Patient 1 | 33% | 10% | — | 40% |
| Patient 2 | 20% | 44% | — | 44% |
| Patient 3 | 10% | 25% | 11% | 40% |
| Patient 4 | 0% | 22% | 10% | 40% |
| Patient 5 | — | 22% | — | 0% |
The spread isn't noise. Each model has a stable personality, visible only after decomposition. That, more than any single number, is the argument for doing verification this way. Models will keep getting stronger and their claims will keep sounding right. But sounding right, as we have shown, was never the problem.
Conclusion
Medical AI has a trust problem. Models state claims with confidence but without reliability, and the usual ways of checking them either flatten the answer to true or false or handoff the judgment to another model with the similar weaknesses and its own biases.
The approach on this page is a structural answer. Take the claim apart into the biological steps it actually asserts. Check each step against published evidence. Compute the conclusion from the checked steps with fixed rules. The model does the one narrow thing it is good at, while the arithmetic does the rest. Every label traces back to a specific link and a specific source, and the same inputs always produce the same result.
Every claim in every Diadia report goes through the pipeline described here, and the catalogue on this site is the output: each published claim comes with its graph, its evidence, and its reasoning open for inspection. The evaluation protocol and all 145 claims from the experiment are in the paper below.
We think this division of labor, models for narrow evaluation and fixed rules for aggregation, is a general pattern for building AI systems in fields where being wrong is expensive.