Hyderabad:Artificial intelligence (AI) tools that help doctors interpret chest X-rays may identify a disease correctly but highlight the wrong area, a study by researchers at the International Institute of Information Technology Hyderabad (IIIT-H) has found. The study has been accepted for the iMIMIC satellite event at MICCAI 2026 in Strasbourg.

Researchers at the Language Technologies Research Centre (LTRC), led by Prof. Parameswari Krishnamurthy, examined four AI models — MAIRA-2, MedGemma-4B, LLaVA-Med-1.5 and LLaVA-1.5 — using thousands of publicly available chest X-rays. Two radiologists also assessed the areas highlighted by the models.

AI tools can mark suspected abnormalities on X-rays using boxes, outlines or heatmaps, helping radiologists review areas that may require attention. “We essentially wanted to examine whether the heatmaps created by vision language models actually correspond to where radiologists who look at the image would say the disease lies,” said principal investigator Dr Syed Faizan.

The researchers found that even when a model highlighted the correct area, it did not necessarily identify the disease in the same way a radiologist would. When information about the diagnosis was removed, the models became less effective at locating the abnormality.

“An AI model may appear to highlight the correct part of an image, but that does not necessarily mean it has identified the disease in the same way a radiologist would,” Dr Faizan said. “The model may first arrive at a diagnosis and then use that diagnosis to determine where to place its heat map. In other words, it may be working backwards.”

The study also found differences between computer-based assessments and the radiologists’ evaluations. MAIRA-2 performed better when its highlighted areas were compared with doctors’ markings, but the two radiologists rated MedGemma higher.

Dr Faizan said radiologists might prefer a broader view around a suspected abnormality to assess its extent, rather than a narrowly marked spot. “A radiologist might want to look at not only areas with the disease, but also a broader surrounding area,” he said.

The LTRC is also examining whether AI medical tools provide consistent answers when doctors phrase the same question differently. A separate study, accepted at EMNLP and conducted with a radiologist from CMC Vellore, tested different ways of asking medical questions.

“Our lab’s efforts are focused on how we can use NLP for healthcare,” Prof. Krishnamurthy said. The broader aim is to reduce routine tasks such as documentation and report writing while keeping human decision-making central to healthcare.

Tags: