GenMI-Lab logo GenMI-LabMedical vision–language research

VGS–Decoding

VGS-Decoding: Visual Grounding Score
Guided Decoding for
Hallucination Mitigation in Medical VLMs

Govinda Kolli*,1Adinath Madhavrao Dukre*,1Yifan Lu1Ziyun Zou1Dwarikanath Mahapatra3Behzad Bozorgtabar2Imran Razzak1

* Equal first authors. Govinda Kolli and Adinath Madhavrao Dukre contributed equally.

1 MBZUAI, Abu Dhabi, UAE 2 Aarhus University, Aarhus, Denmark 3 Khalifa University, Abu Dhabi, UAE

Probe how an image changes the next token.
Use that signal to guide decoding—without retraining.

· Preprint available on arXiv

Training-freeToken-levelVisual sensitivity
VGS-DECODING ARCHITECTUREVisual sensitivity guides each decoding step.
The same question and answer prefix are processed with an original image and a perturbed image. Two independent model caches produce probability distributions. Their normalized difference gives VGS, which reweights the clean distribution to select the next token.
Our VGS-Decoding framework. Original and perturbed images pass through the same frozen medical VLM. Their token distributions produce VGS, which guides probability reweighting. View full-resolution figure ↗

01 / ABSTRACT

A closer look at visual dependence.

An answer can be fluent without being supported by the image. VGS-Decoding asks a focused question: how does a candidate token’s probability change when the visual input is perturbed?

Medical vision–language models can generate plausible answers that rely on learned language patterns rather than the image in front of them. We introduce VGS-Decoding, a token-level mechanism for probing this visual dependence and using it during generation. The method compares the same model’s predictions for an original image and one perturbed version, turning their normalized probability difference into a bounded Visual Grounding Score. This score supplies a separate multiplicative adjustment for every candidate token. The model remains frozen: no retraining, additional expert model, or architectural modification is required.

A seeded Gaussian–Poisson perturbation supplies the comparison image. Both branches share the same question and generated answer prefix, while maintaining independent visual contexts and decoding caches.

We study this approach across LLaVA-Med, CheXagent, and MedGemma on VQA-RAD, SLAKE, and MIMIC-Diff-VQA. Our approach connects token-level probing with a lightweight decoding policy, making visual sensitivity directly inspectable during generation.

WHY VGS-DECODING?

Fluent answers need
visual evidence.

A token-level intervention.
The same model, a different decoding signal.

Comparison of greedy, VCD, DoLA, OPERA, and VGS answers to an abdominal CT question; the reference answer is small bowel.
Comparing greedy, VCD, DoLA, OPERA, and VGS-Decoding on an abdominal CT question.

Ask which tokens depend on the image.

A medical answer can change meaning with a single organ name, side, or number. VGS compares clean and perturbed image-conditioned probabilities to expose that dependence.

The resulting score is used directly in the next-token decision, creating a connection between visual probing and answer generation.

Explore all six appendix examples ↗
01

A bounded visual probe

Normalize the clean–perturbed probability contrast to obtain a token-level score between −1 and +1.

02

Token-adaptive guidance

Use each token’s own VGS to scale the original distribution before choosing the next token.

03

Training-free integration

Keep the backbone frozen and reuse one perturbed image throughout the answer.

02 / METHOD

Two distributions.
One decoding decision.

Keep the model and text fixed.
Change the image, then inspect the difference.

01

Perturb the image

Create one seeded Gaussian-plus-Poisson image variant. Keep that variant fixed throughout the answer.

original image → perturbed image
02

Measure sensitivity

Compare clean and perturbed next-token probabilities under identical text histories and independent caches.

VGS=p − qp + q + ε
03

Reweight and decode

Scale the clean probabilities, normalize, and select the highest-probability token. Append it to both paths.

p′ ∝ p · max(1 + α · VGS, δ)
Guidance α
1.0
Gaussian σ
0.07
Poisson scale λ
70
Positive floor δ
0.01
Perturbed views M
1

One original-image and one perturbed-image forward pass per step. The distorted image is fixed for the entire answer; model weights remain frozen.

EXPLORE THE RULE

What changes when
guidance gets stronger?

Move the slider to see how VGS redistributes probability in a synthetic four-token example.

0 · clean distribution2 · stronger guidance

Illustrative token probabilities; α = 0 recovers the clean distribution.

Original probabilityVGS-reweighted

Selected next token Bafter reweighting

03 / THE PAPER

From token-level signals
to medical VQA.

Three medical VLMs. Three benchmarks.
A shared visual-dependency perspective.

MEDICAL VISION–LANGUAGE MODELS

Three complementary backbones

  • LLaVA-Med
  • CheXagent
  • MedGemma

We evaluate the same probing and decoding principle across three medical vision–language model families.

MEDICAL QUESTION ANSWERING

Three benchmark settings

  • VQA-RAD
  • SLAKE
  • MIMIC-Diff-VQA

We study medical visual question answering across these datasets, with comparisons to greedy, VCD, DoLA, and OPERA decoding.

A TOKEN-LEVEL VIEW

One answer.
Different visual dependencies.

We analyze VGS across laterality, numeral, anatomy, and function tokens. This view reveals how perturbation sensitivity varies across token categories and model families.

A positive VGS means that a token is more probable with the original image than with its perturbed counterpart.

Mean VGS with standard-deviation bars for laterality, numeral, anatomy, and function tokens across LLaVA-Med, MedGemma, and CheXagent on VQA-RAD.
Mean VGS ± standard deviation across token categories on VQA-RAD.

04 / EXPERIMENTAL RESULTS

Our benchmark
comparison.

Greedy · VCD · DoLA · OPERA · VGS
Three backbones, three dataset settings.

Open answersToken-level answer recall

Closed answersClosed-answer accuracy

OverallQuestion-count-weighted mixed score

Teal rows: VGS-Decoding Bold: best within each backbone and metric Δ: overall change from the same backbone’s greedy baseline

VQA-RAD

451 questions · 200 open / 251 closed

Scores (%) · higher is better
VQA-RAD results
Decoding methodFive methods per backbone OpenRecall (%) ClosedAccuracy (%) OverallMixed score (%) Δ vs greedyPercentage points
Backbone 01LLaVA-Med
Greedy 34.45 68.92 53.64 —
VCD 30.85 61.20 47.71 -5.93
DoLA 32.76 58.96 47.34 -6.30
OPERA 33.22 61.69 49.05 -4.59
VGS-Decoding 38.90 72.91 57.75 +4.11
Backbone 02CheXagent
Greedy 22.02 70.92 49.24 —
VCD 21.73 68.53 47.78 -1.46
DoLA 20.73 68.92 47.55 -1.69
OPERA 20.50 69.32 47.67 -1.57
VGS-Decoding 23.48 71.31 50.10 +0.86
Backbone 03MedGemma
Greedy 49.50 61.75 56.32 —
VCD 50.29 57.77 54.45 -1.87
DoLA 51.91 72.51 63.38 +7.06
OPERA 48.90 65.74 58.27 +1.95
VGS-Decoding 53.05 73.58 64.48 +8.16

Swipe or scroll sideways to view all metrics. The method column stays visible.

MIMIC-Diff-VQA

13,121 questions · 8,380 open / 4,741 closed

Scores (%) · higher is better
MIMIC-Diff-VQA results
Decoding methodFive methods per backbone OpenRecall (%) ClosedAccuracy (%) OverallMixed score (%) Δ vs greedyPercentage points
Backbone 01LLaVA-Med
Greedy 28.04 48.39 35.39 —
VCD 25.98 46.42 33.33 -2.06
DoLA 28.75 47.94 35.68 +0.29
OPERA 21.18 46.02 30.14 -5.25
VGS-Decoding 30.84 55.79 39.86 +4.47
Backbone 02CheXagent
Greedy 44.06 82.07 57.79 —
VCD 38.88 79.14 53.42 -4.37
DoLA 39.78 81.94 55.02 -2.77
OPERA 36.19 82.03 52.75 -5.04
VGS-Decoding 43.99 82.53 57.91 +0.12
Backbone 03MedGemma
Greedy 25.97 73.55 43.16 —
VCD 29.38 68.76 43.61 +0.45
DoLA 32.56 76.82 48.55 +5.39
OPERA 28.35 75.53 45.40 +2.24
VGS-Decoding 34.95 82.91 52.28 +9.12

Swipe or scroll sideways to view all metrics. The method column stays visible.

SLAKE

1,061 questions · 645 open / 416 closed

Scores (%) · higher is better
SLAKE results
Decoding methodFive methods per backbone OpenRecall (%) ClosedAccuracy (%) OverallMixed score (%) Δ vs greedyPercentage points
Backbone 01LLaVA-Med
Greedy 40.81 62.25 49.22 —
VCD 39.50 60.56 47.76 -1.46
DoLA 42.54 61.97 50.16 +0.94
OPERA 31.25 58.59 41.97 -7.25
VGS-Decoding 41.11 75.21 54.48 +5.26
Backbone 02CheXagent
Greedy 44.14 69.30 54.00 —
VCD 43.01 66.20 52.10 -1.90
DoLA 42.95 69.01 53.17 -0.83
OPERA 38.19 69.30 50.39 -3.61
VGS-Decoding 43.75 70.14 54.10 +0.10
Backbone 03MedGemma
Greedy 54.74 73.56 62.12 —
VCD 54.00 66.35 58.84 -3.28
DoLA 58.71 79.57 66.89 +4.77
OPERA 58.61 76.92 65.79 +3.67
VGS-Decoding 54.25 85.82 66.63 +4.51

Swipe or scroll sideways to view all metrics. The method column stays visible.

Bold marks the highest displayed score for each model and metric. Δ is the overall-score change from greedy, in percentage points. On MedGemma–SLAKE, DoLA has the highest overall score (66.89 versus 66.63). Open-answer recall is verbosity-sensitive and is not, by itself, evidence of reduced hallucination or clinical correctness.

View the results overview
Results comparing open, closed, and overall scores across five decoding methods for LLaVA-Med, CheXagent, and MedGemma.
Performance across models and decoding methods. See the tables above for exact scores. Full-resolution PDF ↗

05 / TOKEN-LEVEL ANALYSIS

Look beyond the
final answer.

Laterality. Numerals. Anatomy. Function words.
Visual dependence varies by token and backbone.

We investigate the score itself as well as the decoding policy. The mean-score chart above summarizes each category; the distribution view below shows the variation within categories.

VGS distributions across four token categories for LLaVA-Med, MedGemma, and CheXagent.
VGS distributions across token categories and model families. Positive scores indicate higher probability under the original image than under the chosen perturbation. They do not certify correctness.
View the additional mean-score figure
Mean VGS and standard-deviation bars by token category and model.
Comparing mean visual sensitivity across token categories and model families. Error bars show standard deviation.

Score versus policy

Computing VGS is a diagnostic step. It affects generation only when its values are fed into the probability-scaling rule.

Perturbation sensitivity

We measure score magnitude and argmax changes across Gaussian–Poisson settings with the token sequence held fixed.

Independent image-swap probe

An additional study replaces the image and measures probability drops, examining image dependence without using reference answers.

06 / ABLATIONS & EXTENDED STUDIES

Explore the settings.
Inspect the evidence.

Seventeen detailed tables.
Explore each configuration.

We study component choices, guidance strength, perturbations, decoding strategy, tuned comparators, and transfer settings. Open each table to explore the configuration and results.

O / C / Ov mean open-answer recall / closed-answer accuracy / the weighted mixed score, in percent unless noted. The appendix uses several configurations, so compare methods within the same table. Probe-only entries are corrected to equal greedy.
Token-category statistics

VQA-RAD; first three rows: mean VGS. Last three: standard deviation / token count.

Token-category statistics
Model / statistic Laterality Numeral Anatomy Function
LLaVA-Med · mean 0.061 0.094 -0.107 0.047
CheXagent · mean 0.103 0.268 0.078 0.058
MedGemma · mean 0.233 0.051 0.182 0.219
LLaVA-Med · std / n 0.226/152 0.174/37 0.259/86 0.218/20006
CheXagent · std / n 0.225/39 0.311/10 0.181/17 0.196/1222
MedGemma · std / n 0.384/134 0.181/27 0.370/302 0.402/19022
Component ablation · LLaVA-Med

VQA-RAD. Probe-only is corrected to greedy equality: measuring an unused score cannot change the selected tokens. Other results are unchanged.

Component ablation · LLaVA-Med
Configuration Open recall Closed accuracy Overall Δ vs greedy
Baseline (Greedy) 34.45 68.92 53.64 —
+ Fixed weight (VCD-style) 30.85 61.20 47.71 -5.93
VGS probe only (= Greedy) 34.45 68.92 53.64 0.00
+ VGS + Prob. scaling 38.90 72.91 57.75 +4.11
Component ablation · MedGemma

VQA-RAD. Probe-only is corrected to greedy equality, supported by the saved-output identity check. Other results are unchanged.

Component ablation · MedGemma
Configuration Open recall Closed accuracy Overall Δ vs greedy
Baseline (Greedy) 49.50 61.75 56.32 —
+ Fixed weight (VCD-style) 50.29 57.77 54.45 -1.87
VGS probe only (= Greedy) 49.50 61.75 56.32 0.00
+ VGS + Prob. scaling 53.05 73.58 64.48 +8.16
VCD contrast-strength sweep

VQA-RAD, unified harness; α_cd controls the VCD contrast strength.

VCD contrast-strength sweep
Setting LLaVA-Med MedGemma
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
Greedy 34.45 68.92 53.64 49.50 61.75 56.32
VCD α_cd=0.5 32.91 66.14 51.40 49.74 65.34 58.42
VCD α_cd=1.0 33.65 66.14 51.73 50.43 65.34 58.73
VCD α_cd=1.5 34.65 65.74 51.95 49.02 65.34 58.10
VCD α_cd=2.0 35.44 66.53 52.75 48.57 64.54 57.46
VCD α_cd=2.5 36.28 65.74 52.68 50.61 64.94 58.59
VCD α_cd=3.0 35.16 67.33 53.06 51.86 64.94 59.14
VCD α_cd=5.0 32.73 67.73 52.21 49.27 64.54 57.77
VGS (Ours, α=1) 38.90 72.91 57.75 53.05 73.58 64.48
DoLA layer-selection sweep

VQA-RAD. The released-default row is included alongside low, high, and spread layer buckets.

DoLA layer-selection sweep
Setting LLaVA-Med MedGemma
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
low 34.47 68.13 53.20 32.46 78.88 58.30
high 32.70 63.35 49.76 28.58 74.10 53.91
spread 33.22 67.33 52.20 32.46 78.88 58.30
released default‡ 32.76 58.96 47.34 51.91 72.51 63.38
VGS (Ours, α=1) 38.90 72.91 57.75 53.05 73.58 64.48
Probability-floor sensitivity

VQA-RAD, α = 1.0. The release default is δ = 0.01.

Probability-floor sensitivity
Setting LLaVA-Med MedGemma
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
0.0 38.78 71.71 57.11 49.72 76.49 64.62
0.001 39.02 71.31 57.00 50.47 75.30 64.29
0.01 38.73 72.51 57.53 50.47 72.91 62.96
0.05 38.51 72.11 57.21 51.15 75.70 64.81
0.1 38.38 72.11 57.15 49.47 74.90 63.62
Number of perturbed views

LLaVA-Med / VQA-RAD. The release uses one perturbed view, M = 1.

Number of perturbed views
M Open recall Closed accuracy Overall Approx. cost
1 38.73 72.51 57.53 2×
3 38.98 72.91 57.86 4×
5 38.98 72.51 57.64 6×
Decoding-strategy study

VQA-RAD with VGS. Our greedy-only baselines for this study are 51.93 (LLaVA-Med) and 56.74 (MedGemma).

Decoding-strategy study
Setting LLaVA-Med MedGemma
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
Greedy 38.73 72.51 57.53 50.47 72.91 62.96
Temp. T=0.7 33.82 71.71 54.91 54.40 78.49 67.80
Top-p p=0.9 34.87 72.91 56.04 53.41 78.49 67.36
Held-out guidance-strength selection

VQA-RAD training subset: first 250 questions. This table is separate from the test-set sweep.

Held-out guidance-strength selection
Setting LLaVA-Med MedGemma
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
Greedy 30.41 77.10 54.87 30.25 72.52 52.40
0.5 31.75 81.68 57.91 35.38 75.57 56.44
1.0 28.21 81.68 56.23 35.42 79.39 58.46
1.5 30.52 80.92 56.93 35.66 75.57 56.57
2.0 33.28 79.39 57.44 38.67 75.57 58.01
Test-set guidance-strength sensitivity

Guidance-strength sensitivity on the VQA-RAD test set, with scores and p-values for each setting.

Test-set guidance-strength sensitivity
α LLaVA-Med overall Reported p MedGemma overall Reported p
0.0 53.64 — 56.32 —
0.5 56.95 0.046 59.00 0.002
1.0 57.75 0.025 64.48 0.001
1.5 58.79 0.001 64.37 <0.001
2.0 57.50 0.014 65.90 <0.001
Gaussian–Poisson perturbation grid

VQA-RAD, α = 1.0. σ controls Gaussian noise; λ is the Poisson scale.

Gaussian–Poisson perturbation grid
Setting LLaVA-Med MedGemma
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
mild (0.03,200) 38.51 72.51 57.43 50.43 75.30 64.27
(0.05,70) 38.41 72.51 57.39 49.43 75.30 63.83
(0.07,70) def. 38.73 72.51 57.53 50.47 72.91 62.96
(0.10,70) 38.15 72.11 57.05 50.59 74.50 63.90
(0.15,70) 38.77 72.11 57.33 51.47 74.10 64.07
(0.07,50) 38.53 72.11 57.22 50.97 76.10 64.95
(0.07,100) 38.00 72.11 56.98 50.51 75.70 64.53
strong (0.15,30) 39.22 71.31 57.08 51.42 75.30 64.71
Poisson (0,70) 37.91 72.11 56.94 53.62 72.41 64.63
Perturbation types and intensities

LLaVA-Med / VQA-RAD. Individual noise types, blur, and brightness are compared with combined noise.

Perturbation types and intensities
Perturbation Configuration Overall
Greedy (no perturbation) — 53.64
Gaussian noise only σ = 0.07 56.18
Gaussian noise only σ = 0.15 54.03
Poisson noise only λ = 70 55.91
Poisson noise only λ = 30 54.47
Gaussian blur k = 3 54.82
Brightness jitter ± 0.1 54.34
Gaussian + Poisson (Ours) σ=0.07, λ=70 57.75
PathVQA transfer study

LLaVA-Med on the PathVQA test set: 6,719 questions.

PathVQA transfer study
Method Open recall Closed accuracy Overall
Greedy 10.05 67.49 38.79
VCD (α_cd=1) 10.66 67.01 38.86
VGS (Ours) 13.46 79.12 46.31
Same-backbone comparisons on SLAKE

SLAKE, 1,061 questions. The region comparator is an ARCD-style port, not the native ARCD system. The mask column describes the configuration used in this comparison.

Same-backbone comparisons on SLAKE
Setting LLaVA-Med MedGemma Mask?
LLaVA-Med Open recall LLaVA-Med Closed accuracy LLaVA-Med Overall MedGemma Open recall MedGemma Closed accuracy MedGemma Overall
Greedy 38.47 62.74 47.99 54.67 73.32 61.98 —
Instruction-CD 37.80 65.38 48.61 54.60 75.48 62.79 no
VCD (global noise) 37.92 66.83 49.26 53.72 74.52 61.88 no
ARCD-style port (region) 38.37 65.87 49.15 52.83 64.90 57.57 yes
VGS (Ours) 41.00 74.28 54.05 54.25 85.82 66.63 no
Perturbation strength at the signal level

LLaVA-Med / VQA-RAD, 150 questions with fixed token sequences. These are sensitivity measurements, not answer-quality scores.

Perturbation strength at the signal level
(σ, λ) Mean |VGS| Content VGS |VGS| > 0.1 Argmax flips
mild (0.03,200) 0.0067 0.0014 0.4% 0.3%
(0.05,70) 0.0095 0.0030 1.4% 0.7%
(0.07,70) def. 0.0106 0.0043 1.6% 0.7%
(0.10,70) 0.0125 0.0063 2.4% 0.9%
(0.15,70) 0.0169 0.0109 3.6% 1.2%
strong (0.15,30) 0.0194 0.0136 4.1% 1.2%
Gaussian-only 0.07 0.0078 0.0015 0.6% 0.6%
Gaussian-only 0.15 0.0153 0.0093 2.9% 1.3%
Poisson-only 70 0.0085 0.0027 1.1% 0.6%
Image-swap dependency analysis

LLaVA-Med / SLAKE. Image-swap probability drops measure image dependence, not clinical correctness.

Image-swap dependency analysis
Token category Mean VGS Mean swap drop
Anatomy 0.035 0.184
Numeral 0.018 0.097
Other 0.004 0.037
Laterality 0.002 0.020
Function 0.002 0.013
Function-word ablation

LLaVA-Med / VQA-RAD. Values use this appendix study's own scoring configuration.

Function-word ablation
Decoding Reported score
Greedy 54.05
VGS (full) 57.12
VGS, function-word reweighting neutralised 57.12

07 / APPENDIX CASE STUDIES

Six questions.
A closer look at the answers.

Comparing answers across decoding methods.
LLaVA-Med on VQA-RAD.

We compare VGS-Decoding with four baselines on six illustrative VQA-RAD cases. Each card shows the question, reference answer, and all five model responses. Selected examples illustrate behavior; they do not estimate an overall success rate or verify every additional clinical statement in a generated answer.

VQA-RAD research image for case A
LLAVA-MED · VQA-RAD · CASE A

What is the main organ in the image?

Reference Lung

VGS-Decoding

The main organs are the lungs, specifically focusing on a right-sided lung mass.

Compare all four baselines
Greedy
The main organ is the esophagus, a muscular tube connecting the mouth to the stomach.
VCD
The main organ is the mediastinum, the central thoracic compartment between the lungs.
DoLA
The main organ is the thoracic aorta, carrying oxygen-rich blood from the heart.
OPERA
The main organ is the thymus gland.
VQA-RAD research image for case B
LLAVA-MED · VQA-RAD · CASE B

Is this a study of the neck?

Reference No

VGS-Decoding

No, it appears to be an AP chest X-ray focused on the thoracic region, not the neck.

Compare all four baselines
Greedy
Yes, the image is a chest X-ray which includes the neck region.
VCD
Yes, the image is an AP view of the neck, focusing on the right lateral side.
DoLA
Yes, it appears to be a study of the neck in an anterior view, showing bones, blood vessels, and soft tissues.
OPERA
Yes, it appears to be a study of the neck.
VQA-RAD research image for case C
LLAVA-MED · VQA-RAD · CASE C

Where is the pneumonia in the lung?

Reference Lower Right Lung

VGS-Decoding

The pneumonia appears to be located at the right lower lobe of the lung.

Compare all four baselines
Greedy
The pneumonia is located in the right upper lobe of the lung.
VCD
The pneumonia is located in the right upper lung.
DoLA
The pneumonia is in the left lower lobe of the lung.
OPERA
The pneumonia is located in the right lung (lobe unspecified).
VQA-RAD research image for case D
LLAVA-MED · VQA-RAD · CASE D

What is the organ system visualized?

Reference Head

VGS-Decoding

The image shows a CT scan of the head, focusing on the paranasal sinuses.

Compare all four baselines
Greedy
The organ system is the paranasal sinuses (substructure, not system).
VCD
The organ system is the temporomandibular joint (TMJ).
DoLA
The organ system is the parotid gland, a salivary gland near the ear.
OPERA
The organ system visualized is the paranasal sinuses (substructure, not system).
VQA-RAD research image for case E
LLAVA-MED · VQA-RAD · CASE E

What organ is the gray part on the left?

Reference Liver

VGS-Decoding

The gray part on the left represents a portion of the liver.

Compare all four baselines
Greedy
The gray part on the left is the spleen.
VCD
The gray part on the left is the spleen.
DoLA
The gray part on the left is the spleen, an organ in the abdomen.
OPERA
Cannot provide specific organ information; likely an abdominal structure (unspecified).
VQA-RAD research image for case F
LLAVA-MED · VQA-RAD · CASE F

What organ is the gray part on the right?

Reference Small Bowel

VGS-Decoding

The gray part on the right represents a portion of the small bowel.

Compare all four baselines
Greedy
The gray part on the right is the right kidney.
VCD
The gray part on the right is the right kidney.
DoLA
The gray part on the right is the liver (left lateral segment).
OPERA
Cannot provide specific organ; gray part on the right could be related to the bladder.

INTERACTIVE RESEARCH DEMO

Explore the released decoder.

MedGemma · LLaVA-Med
Public VQA-RAD examples or research images.

HUGGING FACE SPACE · HOSTING PENDING

The prepared demo lets you choose a backbone, adjust VGS strength, compare against an α = 0 greedy control, and inspect selected-token scores. It uses the same shared decoder as the code release.

The live Space link will be added after the hosting account and GPU are confirmed. This is a research tool, not a clinical system; use only public or de-identified research images.

Demo code & setup

08 / CODE RELEASE

Run VGS-Decoding.
Explore the implementation.

LLaVA-Med and MedGemma on VQA-RAD.
Two runners. One shared VGS decoder.

SHARED IMPLEMENTATION

Seeded perturbations, visual grounding scores, probability reweighting, and independent clean/perturbed decoding caches.

vgs_decoding.py ↗

Independent caches

Separate visual contexts; identical answer history.

Seeded perturbations

Stable per-example noise across dataset partitions.

Inspectable outputs

JSONL answers, run metadata, and optional token traces.

QUICK START

Generate your first answers.

Full setup guide ↗

After installing the package in the appropriate model environment, run:

python vgs_llavamed_vqarad.py \
  --device cuda:0 --limit 3 \
  --output-dir outputs/llava-med-smoke

Use a GPU allocation. Authenticate with Hugging Face for gated model access. Each output directory must be new; existing runs are never overwritten.

CODE CONTRIBUTORS

Developed and shared through GenMI-Lab.

09 / CITATION

Cite this work.

If you build on the method or code, please cite the linked arXiv preprint.

@misc{kolli2026vgs,
  title = {VGS-Decoding: Visual Grounding Score Guided Decoding for
           Hallucination Mitigation in Medical VLMs},
  author = {Govinda Kolli and Adinath Madhavrao Dukre and
            Behzad Bozorgtabar and Dwarikanath Mahapatra and Imran Razzak},
  year = {2026},
  eprint = {2603.20314},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  doi = {10.48550/arXiv.2603.20314},
  url = {https://arxiv.org/abs/2603.20314}
}

RESEARCH CONTEXT

Interpreting VGS responsibly.

VGS is a perturbation-sensitivity signal, not a clinical confidence score. The experiments concern research benchmarks and short answers. Lexical metrics and selected examples do not establish diagnostic safety; the method is not intended to replace expert image interpretation.