A bounded visual probe
Normalize the clean–perturbed probability contrast to obtain a token-level score between −1 and +1.
GenMI-LabMedical vision–language research
VGS–Decoding
* Equal first authors. Govinda Kolli and Adinath Madhavrao Dukre contributed equally.
Probe how an image changes the next token.
Use that signal to
guide decoding—without retraining.
· Preprint available on arXiv
01 / ABSTRACT
An answer can be fluent without being supported by the image. VGS-Decoding asks a focused question: how does a candidate token’s probability change when the visual input is perturbed?
Medical vision–language models can generate plausible answers that rely on learned language patterns rather than the image in front of them. We introduce VGS-Decoding, a token-level mechanism for probing this visual dependence and using it during generation. The method compares the same model’s predictions for an original image and one perturbed version, turning their normalized probability difference into a bounded Visual Grounding Score. This score supplies a separate multiplicative adjustment for every candidate token. The model remains frozen: no retraining, additional expert model, or architectural modification is required.
A seeded Gaussian–Poisson perturbation supplies the comparison image. Both branches share the same question and generated answer prefix, while maintaining independent visual contexts and decoding caches.
We study this approach across LLaVA-Med, CheXagent, and MedGemma on VQA-RAD, SLAKE, and MIMIC-Diff-VQA. Our approach connects token-level probing with a lightweight decoding policy, making visual sensitivity directly inspectable during generation.
WHY VGS-DECODING?
A token-level intervention.
The same model, a different
decoding signal.
A medical answer can change meaning with a single organ name, side, or number. VGS compares clean and perturbed image-conditioned probabilities to expose that dependence.
The resulting score is used directly in the next-token decision, creating a connection between visual probing and answer generation.
Explore all six appendix examples ↗Normalize the clean–perturbed probability contrast to obtain a token-level score between −1 and +1.
Use each token’s own VGS to scale the original distribution before choosing the next token.
Keep the backbone frozen and reuse one perturbed image throughout the answer.
02 / METHOD
Keep the model and text fixed.
Change the image, then inspect
the difference.
Create one seeded Gaussian-plus-Poisson image variant. Keep that variant fixed throughout the answer.
original image → perturbed imageCompare clean and perturbed next-token probabilities under identical text histories and independent caches.
Scale the clean probabilities, normalize, and select the highest-probability token. Append it to both paths.
p′ ∝ p · max(1 + α · VGS, δ)One original-image and one perturbed-image forward pass per step. The distorted image is fixed for the entire answer; model weights remain frozen.
Move the slider to see how VGS redistributes probability in a synthetic four-token example.
Illustrative token probabilities; α = 0 recovers the clean distribution.
Selected next token Bafter reweighting
03 / THE PAPER
Three medical VLMs. Three benchmarks.
A shared
visual-dependency perspective.
We evaluate the same probing and decoding principle across three medical vision–language model families.
We study medical visual question answering across these datasets, with comparisons to greedy, VCD, DoLA, and OPERA decoding.
We analyze VGS across laterality, numeral, anatomy, and function tokens. This view reveals how perturbation sensitivity varies across token categories and model families.
A positive VGS means that a token is more probable with the original image than with its perturbed counterpart.
04 / EXPERIMENTAL RESULTS
Greedy · VCD · DoLA · OPERA · VGS
Three backbones, three
dataset settings.
Open answersToken-level answer recall
Closed answersClosed-answer accuracy
OverallQuestion-count-weighted mixed score
451 questions · 200 open / 251 closed
| Decoding methodFive methods per backbone | OpenRecall (%) | ClosedAccuracy (%) | OverallMixed score (%) | Δ vs greedyPercentage points |
|---|---|---|---|---|
| Backbone 01LLaVA-Med | ||||
| Greedy | 34.45 | 68.92 | 53.64 | — |
| VCD | 30.85 | 61.20 | 47.71 | -5.93 |
| DoLA | 32.76 | 58.96 | 47.34 | -6.30 |
| OPERA | 33.22 | 61.69 | 49.05 | -4.59 |
| VGS-Decoding | 38.90 | 72.91 | 57.75 | +4.11 |
| Backbone 02CheXagent | ||||
| Greedy | 22.02 | 70.92 | 49.24 | — |
| VCD | 21.73 | 68.53 | 47.78 | -1.46 |
| DoLA | 20.73 | 68.92 | 47.55 | -1.69 |
| OPERA | 20.50 | 69.32 | 47.67 | -1.57 |
| VGS-Decoding | 23.48 | 71.31 | 50.10 | +0.86 |
| Backbone 03MedGemma | ||||
| Greedy | 49.50 | 61.75 | 56.32 | — |
| VCD | 50.29 | 57.77 | 54.45 | -1.87 |
| DoLA | 51.91 | 72.51 | 63.38 | +7.06 |
| OPERA | 48.90 | 65.74 | 58.27 | +1.95 |
| VGS-Decoding | 53.05 | 73.58 | 64.48 | +8.16 |
Swipe or scroll sideways to view all metrics. The method column stays visible.
13,121 questions · 8,380 open / 4,741 closed
| Decoding methodFive methods per backbone | OpenRecall (%) | ClosedAccuracy (%) | OverallMixed score (%) | Δ vs greedyPercentage points |
|---|---|---|---|---|
| Backbone 01LLaVA-Med | ||||
| Greedy | 28.04 | 48.39 | 35.39 | — |
| VCD | 25.98 | 46.42 | 33.33 | -2.06 |
| DoLA | 28.75 | 47.94 | 35.68 | +0.29 |
| OPERA | 21.18 | 46.02 | 30.14 | -5.25 |
| VGS-Decoding | 30.84 | 55.79 | 39.86 | +4.47 |
| Backbone 02CheXagent | ||||
| Greedy | 44.06 | 82.07 | 57.79 | — |
| VCD | 38.88 | 79.14 | 53.42 | -4.37 |
| DoLA | 39.78 | 81.94 | 55.02 | -2.77 |
| OPERA | 36.19 | 82.03 | 52.75 | -5.04 |
| VGS-Decoding | 43.99 | 82.53 | 57.91 | +0.12 |
| Backbone 03MedGemma | ||||
| Greedy | 25.97 | 73.55 | 43.16 | — |
| VCD | 29.38 | 68.76 | 43.61 | +0.45 |
| DoLA | 32.56 | 76.82 | 48.55 | +5.39 |
| OPERA | 28.35 | 75.53 | 45.40 | +2.24 |
| VGS-Decoding | 34.95 | 82.91 | 52.28 | +9.12 |
Swipe or scroll sideways to view all metrics. The method column stays visible.
1,061 questions · 645 open / 416 closed
| Decoding methodFive methods per backbone | OpenRecall (%) | ClosedAccuracy (%) | OverallMixed score (%) | Δ vs greedyPercentage points |
|---|---|---|---|---|
| Backbone 01LLaVA-Med | ||||
| Greedy | 40.81 | 62.25 | 49.22 | — |
| VCD | 39.50 | 60.56 | 47.76 | -1.46 |
| DoLA | 42.54 | 61.97 | 50.16 | +0.94 |
| OPERA | 31.25 | 58.59 | 41.97 | -7.25 |
| VGS-Decoding | 41.11 | 75.21 | 54.48 | +5.26 |
| Backbone 02CheXagent | ||||
| Greedy | 44.14 | 69.30 | 54.00 | — |
| VCD | 43.01 | 66.20 | 52.10 | -1.90 |
| DoLA | 42.95 | 69.01 | 53.17 | -0.83 |
| OPERA | 38.19 | 69.30 | 50.39 | -3.61 |
| VGS-Decoding | 43.75 | 70.14 | 54.10 | +0.10 |
| Backbone 03MedGemma | ||||
| Greedy | 54.74 | 73.56 | 62.12 | — |
| VCD | 54.00 | 66.35 | 58.84 | -3.28 |
| DoLA | 58.71 | 79.57 | 66.89 | +4.77 |
| OPERA | 58.61 | 76.92 | 65.79 | +3.67 |
| VGS-Decoding | 54.25 | 85.82 | 66.63 | +4.51 |
Swipe or scroll sideways to view all metrics. The method column stays visible.
Bold marks the highest displayed score for each model and metric. Δ is the overall-score change from greedy, in percentage points. On MedGemma–SLAKE, DoLA has the highest overall score (66.89 versus 66.63). Open-answer recall is verbosity-sensitive and is not, by itself, evidence of reduced hallucination or clinical correctness.
05 / TOKEN-LEVEL ANALYSIS
Laterality. Numerals. Anatomy. Function words.
Visual
dependence varies by token and backbone.
We investigate the score itself as well as the decoding policy. The mean-score chart above summarizes each category; the distribution view below shows the variation within categories.
Computing VGS is a diagnostic step. It affects generation only when its values are fed into the probability-scaling rule.
We measure score magnitude and argmax changes across Gaussian–Poisson settings with the token sequence held fixed.
An additional study replaces the image and measures probability drops, examining image dependence without using reference answers.
06 / ABLATIONS & EXTENDED STUDIES
Seventeen detailed tables.
Explore each configuration.
We study component choices, guidance strength, perturbations, decoding strategy, tuned comparators, and transfer settings. Open each table to explore the configuration and results.
VQA-RAD; first three rows: mean VGS. Last three: standard deviation / token count.
| Model / statistic | Laterality | Numeral | Anatomy | Function |
|---|---|---|---|---|
| LLaVA-Med · mean | 0.061 | 0.094 | -0.107 | 0.047 |
| CheXagent · mean | 0.103 | 0.268 | 0.078 | 0.058 |
| MedGemma · mean | 0.233 | 0.051 | 0.182 | 0.219 |
| LLaVA-Med · std / n | 0.226/152 | 0.174/37 | 0.259/86 | 0.218/20006 |
| CheXagent · std / n | 0.225/39 | 0.311/10 | 0.181/17 | 0.196/1222 |
| MedGemma · std / n | 0.384/134 | 0.181/27 | 0.370/302 | 0.402/19022 |
VQA-RAD. Probe-only is corrected to greedy equality: measuring an unused score cannot change the selected tokens. Other results are unchanged.
| Configuration | Open recall | Closed accuracy | Overall | Δ vs greedy |
|---|---|---|---|---|
| Baseline (Greedy) | 34.45 | 68.92 | 53.64 | — |
| + Fixed weight (VCD-style) | 30.85 | 61.20 | 47.71 | -5.93 |
| VGS probe only (= Greedy) | 34.45 | 68.92 | 53.64 | 0.00 |
| + VGS + Prob. scaling | 38.90 | 72.91 | 57.75 | +4.11 |
VQA-RAD. Probe-only is corrected to greedy equality, supported by the saved-output identity check. Other results are unchanged.
| Configuration | Open recall | Closed accuracy | Overall | Δ vs greedy |
|---|---|---|---|---|
| Baseline (Greedy) | 49.50 | 61.75 | 56.32 | — |
| + Fixed weight (VCD-style) | 50.29 | 57.77 | 54.45 | -1.87 |
| VGS probe only (= Greedy) | 49.50 | 61.75 | 56.32 | 0.00 |
| + VGS + Prob. scaling | 53.05 | 73.58 | 64.48 | +8.16 |
VQA-RAD, unified harness; α_cd controls the VCD contrast strength.
| Setting | LLaVA-Med | MedGemma | ||||
|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | |
| Greedy | 34.45 | 68.92 | 53.64 | 49.50 | 61.75 | 56.32 |
| VCD α_cd=0.5 | 32.91 | 66.14 | 51.40 | 49.74 | 65.34 | 58.42 |
| VCD α_cd=1.0 | 33.65 | 66.14 | 51.73 | 50.43 | 65.34 | 58.73 |
| VCD α_cd=1.5 | 34.65 | 65.74 | 51.95 | 49.02 | 65.34 | 58.10 |
| VCD α_cd=2.0 | 35.44 | 66.53 | 52.75 | 48.57 | 64.54 | 57.46 |
| VCD α_cd=2.5 | 36.28 | 65.74 | 52.68 | 50.61 | 64.94 | 58.59 |
| VCD α_cd=3.0 | 35.16 | 67.33 | 53.06 | 51.86 | 64.94 | 59.14 |
| VCD α_cd=5.0 | 32.73 | 67.73 | 52.21 | 49.27 | 64.54 | 57.77 |
| VGS (Ours, α=1) | 38.90 | 72.91 | 57.75 | 53.05 | 73.58 | 64.48 |
VQA-RAD. The released-default row is included alongside low, high, and spread layer buckets.
| Setting | LLaVA-Med | MedGemma | ||||
|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | |
| low | 34.47 | 68.13 | 53.20 | 32.46 | 78.88 | 58.30 |
| high | 32.70 | 63.35 | 49.76 | 28.58 | 74.10 | 53.91 |
| spread | 33.22 | 67.33 | 52.20 | 32.46 | 78.88 | 58.30 |
| released default‡ | 32.76 | 58.96 | 47.34 | 51.91 | 72.51 | 63.38 |
| VGS (Ours, α=1) | 38.90 | 72.91 | 57.75 | 53.05 | 73.58 | 64.48 |
VQA-RAD, α = 1.0. The release default is δ = 0.01.
| Setting | LLaVA-Med | MedGemma | ||||
|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | |
| 0.0 | 38.78 | 71.71 | 57.11 | 49.72 | 76.49 | 64.62 |
| 0.001 | 39.02 | 71.31 | 57.00 | 50.47 | 75.30 | 64.29 |
| 0.01 | 38.73 | 72.51 | 57.53 | 50.47 | 72.91 | 62.96 |
| 0.05 | 38.51 | 72.11 | 57.21 | 51.15 | 75.70 | 64.81 |
| 0.1 | 38.38 | 72.11 | 57.15 | 49.47 | 74.90 | 63.62 |
LLaVA-Med / VQA-RAD. The release uses one perturbed view, M = 1.
| M | Open recall | Closed accuracy | Overall | Approx. cost |
|---|---|---|---|---|
| 1 | 38.73 | 72.51 | 57.53 | 2× |
| 3 | 38.98 | 72.91 | 57.86 | 4× |
| 5 | 38.98 | 72.51 | 57.64 | 6× |
VQA-RAD with VGS. Our greedy-only baselines for this study are 51.93 (LLaVA-Med) and 56.74 (MedGemma).
| Setting | LLaVA-Med | MedGemma | ||||
|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | |
| Greedy | 38.73 | 72.51 | 57.53 | 50.47 | 72.91 | 62.96 |
| Temp. T=0.7 | 33.82 | 71.71 | 54.91 | 54.40 | 78.49 | 67.80 |
| Top-p p=0.9 | 34.87 | 72.91 | 56.04 | 53.41 | 78.49 | 67.36 |
VQA-RAD training subset: first 250 questions. This table is separate from the test-set sweep.
| Setting | LLaVA-Med | MedGemma | ||||
|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | |
| Greedy | 30.41 | 77.10 | 54.87 | 30.25 | 72.52 | 52.40 |
| 0.5 | 31.75 | 81.68 | 57.91 | 35.38 | 75.57 | 56.44 |
| 1.0 | 28.21 | 81.68 | 56.23 | 35.42 | 79.39 | 58.46 |
| 1.5 | 30.52 | 80.92 | 56.93 | 35.66 | 75.57 | 56.57 |
| 2.0 | 33.28 | 79.39 | 57.44 | 38.67 | 75.57 | 58.01 |
Guidance-strength sensitivity on the VQA-RAD test set, with scores and p-values for each setting.
| α | LLaVA-Med overall | Reported p | MedGemma overall | Reported p |
|---|---|---|---|---|
| 0.0 | 53.64 | — | 56.32 | — |
| 0.5 | 56.95 | 0.046 | 59.00 | 0.002 |
| 1.0 | 57.75 | 0.025 | 64.48 | 0.001 |
| 1.5 | 58.79 | 0.001 | 64.37 | <0.001 |
| 2.0 | 57.50 | 0.014 | 65.90 | <0.001 |
VQA-RAD, α = 1.0. σ controls Gaussian noise; λ is the Poisson scale.
| Setting | LLaVA-Med | MedGemma | ||||
|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | |
| mild (0.03,200) | 38.51 | 72.51 | 57.43 | 50.43 | 75.30 | 64.27 |
| (0.05,70) | 38.41 | 72.51 | 57.39 | 49.43 | 75.30 | 63.83 |
| (0.07,70) def. | 38.73 | 72.51 | 57.53 | 50.47 | 72.91 | 62.96 |
| (0.10,70) | 38.15 | 72.11 | 57.05 | 50.59 | 74.50 | 63.90 |
| (0.15,70) | 38.77 | 72.11 | 57.33 | 51.47 | 74.10 | 64.07 |
| (0.07,50) | 38.53 | 72.11 | 57.22 | 50.97 | 76.10 | 64.95 |
| (0.07,100) | 38.00 | 72.11 | 56.98 | 50.51 | 75.70 | 64.53 |
| strong (0.15,30) | 39.22 | 71.31 | 57.08 | 51.42 | 75.30 | 64.71 |
| Poisson (0,70) | 37.91 | 72.11 | 56.94 | 53.62 | 72.41 | 64.63 |
LLaVA-Med / VQA-RAD. Individual noise types, blur, and brightness are compared with combined noise.
| Perturbation | Configuration | Overall |
|---|---|---|
| Greedy (no perturbation) | — | 53.64 |
| Gaussian noise only | σ = 0.07 | 56.18 |
| Gaussian noise only | σ = 0.15 | 54.03 |
| Poisson noise only | λ = 70 | 55.91 |
| Poisson noise only | λ = 30 | 54.47 |
| Gaussian blur | k = 3 | 54.82 |
| Brightness jitter | ± 0.1 | 54.34 |
| Gaussian + Poisson (Ours) | σ=0.07, λ=70 | 57.75 |
LLaVA-Med on the PathVQA test set: 6,719 questions.
| Method | Open recall | Closed accuracy | Overall |
|---|---|---|---|
| Greedy | 10.05 | 67.49 | 38.79 |
| VCD (α_cd=1) | 10.66 | 67.01 | 38.86 |
| VGS (Ours) | 13.46 | 79.12 | 46.31 |
SLAKE, 1,061 questions. The region comparator is an ARCD-style port, not the native ARCD system. The mask column describes the configuration used in this comparison.
| Setting | LLaVA-Med | MedGemma | Mask? | ||||
|---|---|---|---|---|---|---|---|
| LLaVA-Med Open recall | LLaVA-Med Closed accuracy | LLaVA-Med Overall | MedGemma Open recall | MedGemma Closed accuracy | MedGemma Overall | ||
| Greedy | 38.47 | 62.74 | 47.99 | 54.67 | 73.32 | 61.98 | — |
| Instruction-CD | 37.80 | 65.38 | 48.61 | 54.60 | 75.48 | 62.79 | no |
| VCD (global noise) | 37.92 | 66.83 | 49.26 | 53.72 | 74.52 | 61.88 | no |
| ARCD-style port (region) | 38.37 | 65.87 | 49.15 | 52.83 | 64.90 | 57.57 | yes |
| VGS (Ours) | 41.00 | 74.28 | 54.05 | 54.25 | 85.82 | 66.63 | no |
LLaVA-Med / VQA-RAD, 150 questions with fixed token sequences. These are sensitivity measurements, not answer-quality scores.
| (σ, λ) | Mean |VGS| | Content VGS | |VGS| > 0.1 | Argmax flips |
|---|---|---|---|---|
| mild (0.03,200) | 0.0067 | 0.0014 | 0.4% | 0.3% |
| (0.05,70) | 0.0095 | 0.0030 | 1.4% | 0.7% |
| (0.07,70) def. | 0.0106 | 0.0043 | 1.6% | 0.7% |
| (0.10,70) | 0.0125 | 0.0063 | 2.4% | 0.9% |
| (0.15,70) | 0.0169 | 0.0109 | 3.6% | 1.2% |
| strong (0.15,30) | 0.0194 | 0.0136 | 4.1% | 1.2% |
| Gaussian-only 0.07 | 0.0078 | 0.0015 | 0.6% | 0.6% |
| Gaussian-only 0.15 | 0.0153 | 0.0093 | 2.9% | 1.3% |
| Poisson-only 70 | 0.0085 | 0.0027 | 1.1% | 0.6% |
LLaVA-Med / SLAKE. Image-swap probability drops measure image dependence, not clinical correctness.
| Token category | Mean VGS | Mean swap drop |
|---|---|---|
| Anatomy | 0.035 | 0.184 |
| Numeral | 0.018 | 0.097 |
| Other | 0.004 | 0.037 |
| Laterality | 0.002 | 0.020 |
| Function | 0.002 | 0.013 |
LLaVA-Med / VQA-RAD. Values use this appendix study's own scoring configuration.
| Decoding | Reported score |
|---|---|
| Greedy | 54.05 |
| VGS (full) | 57.12 |
| VGS, function-word reweighting neutralised | 57.12 |
07 / APPENDIX CASE STUDIES
Comparing answers across decoding methods.
LLaVA-Med on
VQA-RAD.
We compare VGS-Decoding with four baselines on six illustrative VQA-RAD cases. Each card shows the question, reference answer, and all five model responses. Selected examples illustrate behavior; they do not estimate an overall success rate or verify every additional clinical statement in a generated answer.
Reference Lung
The main organs are the lungs, specifically focusing on a right-sided lung mass.
Reference No
No, it appears to be an AP chest X-ray focused on the thoracic region, not the neck.
Reference Lower Right Lung
The pneumonia appears to be located at the right lower lobe of the lung.
Reference Head
The image shows a CT scan of the head, focusing on the paranasal sinuses.
Reference Liver
The gray part on the left represents a portion of the liver.
Reference Small Bowel
The gray part on the right represents a portion of the small bowel.
INTERACTIVE RESEARCH DEMO
MedGemma · LLaVA-Med
Public VQA-RAD examples or research
images.
The prepared demo lets you choose a backbone, adjust VGS strength, compare against an α = 0 greedy control, and inspect selected-token scores. It uses the same shared decoder as the code release.
The live Space link will be added after the hosting account and GPU are confirmed. This is a research tool, not a clinical system; use only public or de-identified research images.
08 / CODE RELEASE
LLaVA-Med and MedGemma on VQA-RAD.
Two runners. One shared VGS
decoder.
The Hugging Face-compatible Mistral-7B conversion, with the original-image and perturbed-image branches sharing the same question prompt.
vgs_llavamed_vqarad.py Community HF-format checkpointMedGemma-4B-IT, using its native processor chat template and explicit end-of-turn handling for concise medical-image question answering.
vgs_medgemma_vqarad.py MedGemma checkpoint · accept model terms for accessSeparate visual contexts; identical answer history.
Stable per-example noise across dataset partitions.
JSONL answers, run metadata, and optional token traces.
After installing the package in the appropriate model environment, run:
python vgs_llavamed_vqarad.py \
--device cuda:0 --limit 3 \
--output-dir outputs/llava-med-smoke
Use a GPU allocation. Authenticate with Hugging Face for gated model access. Each output directory must be new; existing runs are never overwritten.
Developed and shared through GenMI-Lab.
09 / CITATION
If you build on the method or code, please cite the linked arXiv preprint.
@misc{kolli2026vgs,
title = {VGS-Decoding: Visual Grounding Score Guided Decoding for
Hallucination Mitigation in Medical VLMs},
author = {Govinda Kolli and Adinath Madhavrao Dukre and
Behzad Bozorgtabar and Dwarikanath Mahapatra and Imran Razzak},
year = {2026},
eprint = {2603.20314},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2603.20314},
url = {https://arxiv.org/abs/2603.20314}
}
RESEARCH CONTEXT
VGS is a perturbation-sensitivity signal, not a clinical confidence score. The experiments concern research benchmarks and short answers. Lexical metrics and selected examples do not establish diagnostic safety; the method is not intended to replace expert image interpretation.