Skip to main navigation Skip to search Skip to main content

Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models

  • Bidur Khanal
  • , Sandesh Pokhrel
  • , Sanjay Bhandari
  • , Ramesh Rana
  • , Nikesh Shrestha
  • , Ram Bahadur Gurung
  • , Cristian A. Linte
  • , Angus J M Watson
  • , Yash Raj Shrestha
  • , Binod Bhattarai* (Corresponding Author)
  • *Corresponding author for this work
  • Rochester Institute of Technology
  • Nepal Applied Mathematics and Informatics Institute for Research
  • Kathmandu University
  • University of Lausanne

Research output: Chapter in Book/Report/Conference proceedingPublished conference contribution

9 Downloads (Pure)

Abstract

Vision-Language Models (VLMs) are becoming increasingly popular in the medical domain, bridging the gap between medical images and clinical language. Existing VLMs demonstrate an impressive ability to comprehend medical images and text queries to generate detailed, descriptive diagnostic medical reports. However, hallucination--the tendency to generate descriptions that are inconsistent with the visual content--remains a significant issue in VLMs, with particularly severe implications in the medical field. To facilitate VLM research on gastrointestinal (GI) image analysis and study hallucination, we curate a multimodal image-text GI dataset: Gut-VLM. This dataset is created using a two-stage pipeline: first, descriptive medical reports of Kvasir-v2 images are generated using ChatGPT, which introduces some hallucinated or incorrect texts. In the second stage, medical experts systematically review these reports, and identify and correct potential inaccuracies to ensure high-quality, clinically reliable annotations. Unlike traditional datasets that contain only descriptive texts, our dataset also features tags identifying hallucinated sentences and their corresponding corrections. A common approach to reducing hallucination in VLM is to finetune the model on a small-scale, problem-specific dataset. However, we take a different strategy using our dataset. Instead of finetuning the VLM solely for generating textual reports, we finetune it to detect and correct hallucinations, an approach we call hallucination-aware finetuning. Our results show that this approach is better than simply finetuning for descriptive report generation. Additionally, we conduct an extensive evaluation of state-of-the-art VLMs across several metrics, establishing a benchmark.
Original languageEnglish
Title of host publicationMedical Image Computing and Computer Assisted Intervention – MICCAI 2025
Subtitle of host publication28th International Conference, Daejeon, South Korea, September 23–27, 2025, Proceedings, Part X
EditorsJames C. Gee, Daniel C. Alexander, Jaesung Hong, Juan Eugenio Iglesias, Carole H. Sudre, Archana Venkataraman, Jinah Park
Place of PublicationCham, Switzerland
PublisherSpringer
Pages235–245
Number of pages11
ISBN (Electronic)978-3-032-05127-1
ISBN (Print)978-3-032-05126-4
DOIs
Publication statusPublished - 20 Sept 2025
Event28th International Conference On Medical Image Computing And Computer Assisted Intervention: MICCAI 25 - Daejon Convention Center, Daejon, Korea, Republic of
Duration: 23 Sept 202527 Sept 2025
https://conferences.miccai.org/2025/en/default.asp

Publication series

NameLecture Notes in Computer Science
PublisherSpringer
Number15969
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference28th International Conference On Medical Image Computing And Computer Assisted Intervention
Country/TerritoryKorea, Republic of
CityDaejon
Period23/09/2527/09/25
Internet address

Bibliographical note

Submission history
[v1] Sun, 11 May 2025 14:54:11 UTC (5,502 KB)
[v2] Sun, 22 Jun 2025 20:58:41 UTC (5,502 KB)

All accepted papers will be made available by Springer's Lecture Notes in Computer Science no earlier than two weeks prior to the conference.

Keywords

  • multimodal data
  • Gastrointestinal image analysis
  • Vision Language Model
  • Hallucination
  • Hallucination-aware finetuning

Fingerprint

Dive into the research topics of 'Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models'. Together they form a unique fingerprint.

Cite this