Does a Report Generation Model Know When an Image Is Bad?
Emory HITI Lab Health AI Datathon 2026 · Team 3, "The Hallucinators" · Third Place
July 2026 - August 2026
In April 2026, I sent an email to Dr. Saptarshi Purkayastha at IU Indianapolis with the subject line "high schooler in Indy built bias tool your Lancet paper inspired." I’d been following his group's work on racial encoding in medical imaging models, and DermEquity, the fairness framework I built for skin cancer screening, came partly out of reading it. He answered, and on a subsequent Zoom, I walked through the project and limitations. By the end, Dr. Purkayastha invited me on the lab's team and to join their group at the 2026 Health AI Datathon at Emory University.
My registration and stay at the datathon were graciously covered through Dr. Judy Gichoya's grant, who runs the HITI Lab at Emory and is Dr. Purkayastha's academic partner-in-crime of sorts. I flew to Atlanta in July for nine days of lectures, a symposium, and a three-day datathon, as the youngest researcher in the building by about a decade.
My team first drew the prompt: characterize how report quality degrades under graded image perturbations. However, upon consideration, I thought the framing was wrong and voiced this opinion to my team. A good report, I said, should change when the image degrades, since a radiologist for a rotated film would write "rotated, evaluation limited". Measuring how much a model's report changed collapses two different problems into one number: a brittle model flips its diagnosis under trivial perturbation, which is bad. A confabulating model returns a nearly identical report since it stopped reading the image and fell back on the population prior, which is worse and scores as maximally robust on any similarity metric.
What I proposed was to measure two things separately: does the finding survive, and does the model say the image is bad? Logically, the first should hold steady under degradation and the second should rise. Held apart, the confabulating case becomes visible since it is the one where findings quietly disappear while the hedging rate remains constant. This proposed reframe was voted on and passed unanimously. Thus, The Hallucinators ran on this framing for the Datathon (and placed 3rd!).
My contributions to the project included the reframe, the dataset and quality-awareness layers, and co-presenting the final talk. The perturbation grids, embedding analysis, and natural-variability track were built by my incredible teammates and are credited as such.
With much love to The Hallucinators (Rohan, Ryan, Kanika, Sameer, Myeisha, & Dileep),
Angie X.
A radiologist learns to read chest X-rays by looking at thousands of them, until they can glance at a film and tell which shadow is pneumonia and which is normal anatomy. AI systems can now learn the same way, from millions of images paired with the reports doctors wrote about them. If you feed one a new X-ray, it can produce a written diagnostic report on its own. These systems are already being deployed, particularly in settings where radiologists are scarce, with models writing a first draft that a physician then reviews.
But x-ray films aren't always optimal. A patient coughs and the image might blur. If they turn slightly, it comes out crooked and if machine settings drift, it comes out too dark. Patients too sick to stand are photographed with a portable machine wheeled to their bedside, and those images are worse still. When a radiologist gets a low-quality film like that, they say so in the report: this study is limited by motion, I cannot fully evaluate the left base. That sentence tells whoever reads the report that the conclusion is provisional.
So, we asked whether AI models write that sentence. To find out, we took real chest X-rays from Emory University's dataset and perturbed them deliberately, in ways that occur in actual hospitals: rotated 10-15 degrees, darkened, blurred, cropped, compressed. We generated 5,000+ reports from these perturbed images and compared each against the report the same model produced from the clean, non-perturbed version.
The first thing we found was that on a film with pneumonia, rotating the image by 10-15 degrees (which is a normal variation nobody orders retakes for) made the finding completely disappear from the model's report. And the report offered no explanation, no hedging, and no mention that the image was crooked. It read just as confidently as the model's report on the clean image, which is exactly what makes it dangerous. A wrong answer can be caught, but a missing answer delivered with total confidence looks identical to a correct one.
Our second finding was that the model's failures are not uniform. Different models are blind in different places: one model noticed underexposed images and said so explicitly. Another noticed compression artifacts but went the wrong way on motion blur (the blurrier the image, the more confident its report became). And motion blur, which is the most common quality problem in real radiology, erased the most findings while producing the least warning.
Report generation models are increasingly deployed as a first-pass layer, particularly in resource-limited settings, producing a draft that a radiologist reviews. However, the position in the workflow is potentially problematic: an report generator model sits upstream of the radiologist and cannot request a reshoot, meaning it simply produces a report on whatever it is handed.
If a model receives a low-quality film and produces a confident, clean-sounding report, that report becomes a severe automation-bias risk. A radiologist who may have already felt uneasy about the image might override their own doubt when the model reads it as normal. This is the specific harm the project was designed to detect: silent, confident errors on an input the model should have flagged.
With regard to prior work, there are 3 main literatures:
1. Synthetic augmentation studies (Ogawa et al. 2019; Zhou et al. 2021; Lin et al. 2024) apply rotation, contrast, brightness, and blur to medical images, almost always in the context of training-data augmentation (not deployment robustness evaluation)
2. Clinical image-quality studies (Walz-Flannigan et al. 2018) document that real acquisition artifacts cause diagnostic error, but do not connect specific artifacts to specific model behaviors.
3. AI report evaluation studies (Guan et al. 2021; Nousiainen et al. 2021) compare generated reports against reference reports using surface metrics (did not vary image quality as an experimental factor).
4. Robustness benchmarks (Koh et al. 2021; Phillips et al. 2020) test models under distribution shift, WILDS on naturally occurring shift and CheXphoto on synthetic corruption, but neither connects the shift to what a generated report says about the image it received.
The gap here is what we worked on: real-world image unusualness hadn't yet been connected to specific report failures.
Before degrading anything, Ryan ran three synthetic images through all six report models: a cartoon chest, an image of the moon, and a field of vertical stripes (shown below). From these gimmick images, five of six models produced a confident, structured radiology report regardless: the moon returned as "Mild cardiomegaly. Left rib fractures. External electronic device obscuring the left apex" from model-f. The stripe field returned as "There is a linear lucency through the posterior aspect of the left scapula. No definite fracture is identified" from model-a. The cartoon chest returned a clean normal report from model-c, and only model-d flagged the inputs as non-diagnostic.
This is the failure mode demonstrated at its most extreme, but isn’t quite evidence on its own. Reza Chavoshi (HITI Lab mentor) made a point directly during feedback rounds, stating that these images would never reach a radiologist, so a model's behavior on them says little about deployment. So what our pilot establishes is narrower, but still useful: a fluent report is not evidence that the model looked at a chest X-ray. Everything following asks whether this same thing happens at severities that might actually occur.
The datathon provided nine chest radiograph datasets. I surveyed all of them and produced a data inventory documenting images, reports, and metadata availability for each.
DS1 and DS2, both MIMIC-derived, have human reports but almost no demographics beside ssex and age. DS3 from VinDr and DS6 from SIIM have neither reports nor demographics. DS5, the TAIX-Ray subset, has reports only in Spanish. DS4, the Emory subset, is the only one with images, human-written reports, and full patient demographics including sex, age, BMI, race, and ethnicity (15,858 image rows across 9,998 unique studies, one study per patient). The six subsets discussed here are compared in Table 1 below.
Before using the report column, I verified first that it was real text. Median report length is identified as 463 characters and 95.8% contain a FINDINGS or IMPRESSION section. Words are sometimes concatenated without spaces (e.g. strings like ofpneumonia and Negativechest). COMPARISON fields carry redaction asterisks and report text is also duplicated across views within a study, which requires deduplication before sampling.
Seven anonymized foundation models, model-a through model-g (names were withheld by the organizers to prevent researcher bias). Six generate reports, whereas model-b is embeddings-only. All report generation ran at temperature 0, so any difference between a clean and a perturbed report is attributable to the image rather than to sampling variation.
The initial 250-study sample was stratified by race. Rohan flagged that findings-stability metrics only work on studies where the clean report asserts a finding, which would have left 88 of 250 studies contributing to the primary outcome. The sample was rebuilt stratified 125/125 on no_finding, and demographic balance fell out naturally: 115 Black, 112 White, 10 Asian, and 13 across the remaining categories; 128 female, 122 male; median age 59.
Dileep's image-quality grid used a separate 150-study cohort at 71% normal to preserve the natural class balance of chest radiography and the dataset itself. This is a deliberate design difference and means that two cohorts aren't directly comparable, which is why every result reported below is within-model, same studies, clean versus degraded.
Mentor feedback from Reza Chavoshi and Hari Trivedi reshaped our perturbation design partway through. The point they made was that perturbations have to be clinically possible: in reality, a 90-degree rotation gets corrected in the PACS viewer and a completely black image gets reshot. But a 10-degree rotation is difficult for the human eye to catch and nobody reshoots for it. Underexposure and overexposure, patient rotation, and motion blur are also real world occurences.
In total, 8 perturbation families were built, 5 of them at 3 ordered severity levels. Rotation runs at 5-10, 10-15, and 15-20 degrees, standing in for patient positioning error. Underexposure runs at dose fractions of 0.25, 0.10, and 0.04, standing in for acquisition parameter error. Motion blur runs at 3, 8, and 15 millimeters for movement during exposure. Anatomical clipping removes 5, 15, and 30% of the field for collimation error. JPEG compression runs at quality factors 20, 8, and 3 for storage and transmission loss. Three further families run at a single level: grayscale inversion and laterality flip, both processing or display errors, and quantization to 8 grey levels. However, quantization is labelled and functions only as a stress test throughout since real radiographs have 12-16 bits of depth, so 8 gray levels isn't a real-world defect. Further, not every family ran in every grid. The image-quality grid that includes the hedging analysis covers 15 conditions, namely clean plus 3 levels each of underexposure, motion blur, anatomical clipping and JPEG compression, together with grayscale inversion and laterality flip. Rotation and quantization ran separately on model-d in the earlier report-side grid, which is where the erasure results in Phase 4 are derived.
Two families run backwards from what the parameter name suggests, and both orderings were verified before any trend testing. Underexposure is parameterised as a dose fraction, so 0.04 is darker and worse than 0.25. JPEG uses the standard quality factor, so q03 is more compressed and worse than q20. Getting either ordering wrong would have inverted the direction of every trend test downstream. The full library with plausibility tiers is in Table 2 (above section), and the quality-control previews confirming each defect applied as specified are in Figure 2 below.
Motion blur is specified in millimeters, which does require an assumption since PNG exports carry no pixel spacing. Image width is taken to span roughly 35 centimeters of chest.
Every model in the evaluation reads multiple views, which was a point of contention regarding lateral views. Leaving it pristine would let a model recover a finding from the intact view, and the amount recovered would vary with how heavily each model weights its two inputs, so attenuation would differ across models for reasons having nothing to do with robustness. Thus, both views were degraded at matched severity throughout.
As aforementioned, based on mentor feedback, our team also decided to look for perturbations in our existing data. Instead of asserting that a given perturbation severity is realistic, a part of the team (led by Rohan) measured how far each one moves an image from the natural distribution of real no-finding scans, in embedding space.
The result is a calibration curve placing every synthetic defect on the same axis as real clinical variation. Rotation, contrast adjustment, and mild motion sit inside the range of ordinary variation in the dataset. Quantization and heavy JPEG sit outside it.
Total scored: 5,356 generated reports, compared against 15,858 human radiologist reports. The bulk of that is the image-quality grid at 4,500 reports, 2 models across 15 conditions at 150 per cell, with the remainder coming from the report-side grid on model-d at 174 per condition, a 100-report pipeline test, and a 60-report clean run across all 6 models. The archive has 6,182 generated reports in total, including earlier grid runs that were never scored.
In Study DS4_patient_01078, the human radiologist reported "bilateral pleural effusions." On the clean image, model-d reported "…small left pleural effusion…" On the same study rotated 10-15 degrees, model-d reported "The left lung is clear. No pleural effusion."
The finding is completely gone, with no hedging language or comment about image quality. Nothing in the perturbed report indicates that anything had changed.
At scale, the erasure rate is defined as: the clean report named a finding, and the perturbed report dropped it. Across 4 models, namely model-c, model-e, model-f and model-g, and 5 conditions, erasure ran between roughly 30%-70%. Motion blur was the most destructive condition, erasing 53% of findings on average and up to 71% for a single model. Approximately 15% of surviving findings returned with laterality swapped, i.e. a left-sided finding was reported as right-sided or vice versa.
Erasure exceeded hallucination under every condition we tested, meaning these models delete findings rather than invent them. This direction matters clinically: a fabricated finding gets investigated and disproven, while an erased finding gives nothing to investigate. Model-a was excluded from this analysis for insufficient baseline-finding denominator (n < 10), and model-d appears separately in Figure 4 at per-finding resolution rather than in the cross-model comparison.
An earlier analysis on model-d, run on the report-side perturbation grid rather than the image-quality grid, illustrates why single-metric evaluation would have missed this entirely. Overall accuracy against ground-truth labels stayed at 0.82 across every condition. Underneath that flat number, pneumonia recall fell from 0.35 to 0.13. The model shifted toward asserting "clear," which mechanically improved its false-positive rate and kept accuracy steady while recall collapsed.
I decided to score every report on 2 binary questions. The first was whether the model's report contains uncertainty language (hedging), including words like possible, probable, may represent, cannot exclude, suggestive of, suspicious for, concerning for, worrisome for, questionable, likely, most likely, could represent, differential, or versus. The second asks whether the report mentions the quality of image itself, including phrases like limited by, suboptimal, technically limited, rotated, underpenetrated, overpenetrated, motion, degraded, obscured by, poor inspiration, or difficult to evaluate. The second list explicitly excludes "no prior films available for comparison," which concerns missing priors and appears often enough that counting this would inflate the measure substantially. Both are also computed at report level, meaning the fraction of reports containing at least one match rather than a count of matches within a report, which is what makes the numbers directly comparable to published prevalence results.
I validated this three ways: List 1 is our original fourteen terms. List 2 is the same list with "most likely" excluded but "likely" retained, on the reasoning that the Brigham and Women's Hospital Diagnostic Certainty Scale (Shinagare et al. 2020) maps "most likely" to greater than 90% probability, which is assertion, and "likely" to 75-90%, which is hedging. In practice, this distinction proved almost irrelevant, since across all 5,356 reports the two lists differ by at most 2.7 percentage points and the human rate moves only from 15.2% to 14.9%- "most likely" is simply rare in radiology reports. List 3 is Callen et al. 2020, a published 44-term uncertainty lexicon validated on 642,569 report impressions from 171 radiologists.
In regard to validation results, Callen's lexicon applied to our 15,858 human DS4 reports returns 24.7%, against Callen's reported 29.5% for general radiography. This gap is likely consistent with DS4 being exclusively chest radiography and not the full exam-type mix Callen covered. A separate external check points the same direction, as our human quality-statement rate of 1.0% sits in the same range as Mabotuwana et al. (2018), who found technical quality concerns in 2.4% of 1.2 million radiology exams, with motion the single most common category at 31.6% of flagged studies. This second figure is what makes our motion result clinically relevant, since the defect these models are worst at are also the ones documented to occur most often. Relative ordering between models and conditions were consistent across all three lists, despite absolute numbers shifting by roughly 10 percentage points depending on lexicon breadth.
Keyword matching cannot handle negation, despite negation encompassing much of the clinical content in a radiology report. The phrase "no suspicious nodule" contains suspicious and would be incorrectly flagged as hedging. To bound this error, Gemma 4 31B Instruct at temperature 0 re-read 600 reports drawn from the 4 cells most relevant to our project's claims, with explicit rules in the prompt: negation is not hedging, "most likely" is assertion under the BWH scale, "likely" alone is hedging, and recommendations such as "may benefit from additional views" are not diagnostic uncertainty.
Both methods agree on direction and magnitude in every cell, as shown below in Table 3. This correction also ran both ways; the model caught contextual hedging no word list would find, including "questionable small left pleural effusion" and "minimal linear scar or atelectasis," where the uncertainty sits in the framing. It also threw out recommendation phrasing that the broad Callen list had over-counted in model-a's underexposure cell, where matches on the bare word may were coming from sentences about repeating the study rather than concerning the diagnosis.
Cochran-Armitage trend tests were run across ordered severity at n=150 per cell, giving 8 tests across 2 models and 4 defect families, of which 3 came back significant. The full matrix is below in Figure 8.
Model-a responds to underexposure and to nothing else. Under the Callen lexicon, its hedging climbs from 1.3% on clean images to 16.0% at 4% dose, a total increase of 14.7 percentage points, with the Cochran-Armitage trend across the 3 severity levels significant at p = 0.0002. The independent LLM parse of the same two cells moves the same way, 0.0% to 10.0% at p = 0.0001. Model-a also mentions the problem explicitly, writing that the lung bases are not included in the image, and quality statements in that parse rise from 0.7% to 8.7% at p = 0.001. Motion, clipping and JPEG all return as flat. Model-e reverses under motion blur, with hedging falling from 24.7% on clean images to 12.0% at 15 millimeters, a drop of 12.7 points at p = 0.0003, monotonic across all 3 severity levels and confirmed by the clean-vs-severe comparison at p = 0.005 (the more perturbed the image, the more confident the report). That same model responds correctly to JPEG compression, increasing 8.6 points from 24.7% to 33.3% at p = 0.014, and returns flat on underexposure and clipping.
The flat cells are themselves a finding. Neither model is calibrated across defect types and there is no overlap in what they catch, with model-e detecting a digital artifact and missing a physical one, which is close to the opposite of what clinical usefulness would require. Joining this to the erasure results gives our headline claim that motion blur erases the most findings and triggers the least warning. One plausible mechanism for the model-e reversal, though not one we tested, is that severe blur collapses the visual signal so completely that the model falls back on a confident normal-template report, and distinguishing that from the alternatives would require attention analysis our project did not run.
Everything above imposes degradation. The complementary question, raised by our mentors, was whether these defects already exist in real data, because if naturally degraded images can be found then no argument about the realism of synthetic perturbation is needed at all. Rohan led this track, along with Kanika and Myeisha.
Starting from 15,858 DS4 image rows, a homogeneous cohort was built on adult age, exact PA view, No Finding = 1, image file present, non-portable evidence, and upright status confirmed or presumed. Portability was inferred conservatively from the current-exam report heading only, examining text before COMPARISON or FINDINGS so that a mention of a prior portable exam would not be misread as describing the current one. 41 studies were excluded for portable, mobile or bedside wording and 3 for missing, conflicting or unclear evidence, leaving a final cohort of 4,732 unique patients with one image and one study each. The point of all that filtering is that if every patient is a healthy adult imaged the same way, then any variation left between images has to be fro the image itself, and not the pathology or acquisition mode.
For each image, we took the global embedding, L2-normalized, reduced to 50 principal components, then computed the mean Euclidean distance to its q nearest neighbors within the cohort (written as d_q(i) = (1/q) Σ ||z_i − z_j||₂ over the q nearest neighbors of i). Higher mean neighbor distance means the image sits in a sparser region of embedding space, which is just saying that fewer similar images exist. Studies were ranked on that score and split into 5 percentile bands running from very typical to extremely unusual, with 50 random samples from each.
The clean reports for the sampled studies were scored against the human reports using BERTScore F1, RadGraph-F1, CheXbert positive F1, and a Gemma-scored clinical discrepancy metric, across 5 cohorts: no finding, pleural effusion, cardiomegaly, atelectasis, and pneumothorax. Report quality degrades as images become more unusual, with BERTScore and RadGraph-F1 both decreasing across the bands while the Gemma discrepancy score increases.
One exception was for pneumothorax, where BERTScore F1 went up as images became more anomalous, opposite to every other cohort and metric. On manual inspection, BERTScore was scoring both the explicit presence and the explicit absence of pneumothorax as a match, since the surrounding language is nearly identical and the word appears either way. RadGraph and the Gemma judge caught the discrepancy that BERTScore could not. This is a good example of why surface metrics fail on negation, and it is the same weakness the LLM cross-check meant to address in the hedging analysis.
Kanika and Myeisha then measured surface image properties directly on the sampled images, specifically mean brightness, contrast, and pixel variance. All three correlate with the unusualness band, which is what breaks the circularity here: the unusualness score derives from model embeddings and the reports come from the same models, so on its own "the model does worse on images it finds unusual" is pretty close to tautological. Brightness, contrast and pixel variance are computed from raw pixels with no model involved anywhere, so the chain becomes physical property, then embedding position, then report degradation, with the first link grounded outside the model. Ryan noted during the presentation that the brightness differences, while more variant in anomalous images, remain within normal bounds, so the models are not simply reading brightness.
To join the two tracks then, the synthetic track imposed underexposure and motion blur whereas the natural track found that deviant images differ in brightness and pixel variance. Underexposure is a brightness reduction, and motion blur collapses pixel variance because blurring averages neighboring pixels toward one another. The natural track identifies which axes have the variation and not which direction along them. Naturally unusual images run darker and higher in variance, while the imposed defects move along those same two axes in the directions a technologist would recognize.
WILDS (Koh et al. 2021) captures natural distribution shift without synthetic corruption, while CheXphoto (Phillips et al. 2020) applies synthetic corruption to chest radiographs without grounding in real-world deviation. This project does both, calibrating corruption severities against a learned anomaly metric derived from real atypical images, which quantifies how realistic each perturbation actually is.
On the hedging side, the closest published comparison is CheXthought (arXiv 2604.26288), which reports radiologists expressing uncertainty in 44.2% of chains-of-thought on optimal studies versus 84.1% on suboptimal ones. That figure is confounded for our purposes, since the same annotator rated the image quality and expressed the uncertainty within a structured framework that required commenting on technical adequacy. Our design assigns quality independently of whoever writes the report.
Dileep wrote a formal pre-registration before any perturbed report existed, committing to the cohort, the primary outcome of CREST risk-weighted error per study measured within-model clean versus perturbed, 5 hypotheses, the analysis plan, and 11 stated limitations. Two of its hypotheses were predictive: H1 predicted that models with a small vision encoder relative to their language model would shift toward fabrication under degradation as the language prior fills the gap, noting that model-e has a 1.3% vision share against model-c's 16.4%. H4 predicted that photometric inversion would separate the models more sharply than any graded defect, on the reasoning that a model returning a confident normal report on an inverted radiograph is not checking image quality at all. We saw that H4 held up, where grayscale inversion produced the widest separation between the two models of any condition in the grid, with model-a hedging on 0.7% of inverted images against model-e's 31.3%, a wider gap than any graded defect produced. H1 couldn't be tested, since its contrast is model-e against model-c and compute limited the severity grid to model-a and model-e. The document also contains a yardstick measured before any perturbation: different models disagree with model-a by 6.3-10.7 CREST points per study on identical clean images, which frames the scale of any defect effect against the effect of simply swapping models.
Compute limited the full severity grid to two models, so the hedging trend analysis covers model-a and model-e only. The design generalizes to any report generation model; the coverage does not yet, and no claim here should be read as a statement about report generation models in general.
DICOM view labels are unreliable, and laterals mislabelled as PA are present in the cohort. Hari identified this during mentor rounds, noting that the standard correction is a learned view classifier instead of trusting the header field. Kanika produced a PA-only version of the contrast analysis to control for it in one place, but it is not controlled everywhere.
Keyword detection is imperfect and has no negation handling. It is applied identically to clean and degraded reports on the same studies, so it cannot manufacture a difference between them, but it is LLM-verified on only 600 of 5,356 reports and the rest carry this uncorrected error. Quantization to 8 grey levels is not clinically possible and is thus labelled a stress test throughout; the pre-registration notes that the active-learning decision boundary in an exploratory analysis was driven by that cell and not by realistic degradation.
The unusualness score is model-derived and partially circular by construction, with the brightness and contrast result being what grounds it externally. Model-c returned cosine similarity of exactly 1.0000 under every perturbation condition in the embedding analysis, flagged as possible embedding degeneracy rather than true invariance and never resolved. The two-patient sanity check, comparing two different patients' clean images under model-c, would settle it in minutes and remains the clearest open loose end in our project.
The images are presentation data (not raw detector output), so the display LUT is already applied and pixel values are not linear in dose, which is why severity is reported by measured SNR in place of nominal dose. No radiologist adjudicated any output, with CREST risk weighting substituting for a clinical consequence grade. And the cohort skews toward normal, since roughly 70%-80% of DS4 has no positive finding, which shaped both sampling strategies and constrains every erasure denominator.
What this produces is a reusable audit protocol for black-box report generation models: a small labelled defect benchmark with severities calibrated against real image variation, two metrics kept deliberately separate asking whether the finding holds and whether the model flags the image, and validated hedging and quality-statement detection with published-lexicon cross-reference. Any report generation model can be run against it without access to weights, training data, or internals.
The next steps are to validate the anomaly metric across imaging modalities beyond chest radiography, and to generate perturbations that mimic real anomalies automatically by learning the direction of real deviation in embedding space and synthesizing along it, which would move the field from hand-picked perturbations toward ones mined from data.
The 2026 Emory Health AI Datathon was built mostly around evaluation, and the resurfacing theme was that the metric you choose decides what you can see. A commercial model reports 96% sensitivity and returns 82% on Emory data. Accuracy holds flat while recall collapses underneath it. In each case the numbers are correct, but answer a question that nobody had checked was the right one.
This project is a small instance of that. The specific percentages will not survive a larger study unchanged, since they come from two models on one dataset over nine days. What holds is the reframe. Splitting the question into whether the finding survived and whether the model said the image was bad makes two opposite failures distinguishable, and the second one turns out to be where the danger sits. Every model we tested was blind somewhere, none was blind everywhere, and no two were blind in the same place. If quality awareness is a set of narrow architecture-specific detectors and not a general capability, then scoring robustness on a single axis is measuring something that does not exist.
The part that sticks with me moving forward is the direction of the failure. These models delete rather than fabricate. Most of the concern about AI in medicine is about what these systems will say that is wrong, yet the harder problem may very well be what they quietly decline to say at all, which is invisible by construction and shows up in no metric built to catch errors.
Cheers,
Angie X.