Skip to main content Skip to main navigation

Publication

MedHyperGraph: EHR-Integrated Multimodal Hyperedges for Clinical VQA

Mina Heinein; Mina Youssef; Kirellos Nashed; Tamer Basha; Hasan Md Tusfiqur Alam; Abdulrahman Mohamed Selim; Omair Shahzad Bhatti; Daniel Sonntag
In: Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges. Medical Image Computing and Computer Assisted Intervention (MICCAI-2026), September 27 - October 1, Vol. LNCS 17266, Springer Nature Switzerland, 2026.

Abstract

Many clinical questions about chest X-rays cannot be answered from images alone, as they depend on longitudinal context, including prior studies, earlier reports, diagnoses, and the timing of clinical events. Most medical vision-language models still process isolated imagequestion pairs. We present MedHyperGraph, a per-patient multimodal hypergraph for electronic health record (EHR)-grounded medical visual question answering (VQA). It integrates structured EHR events, prior report entities, imaging studies, ontology concepts, and image regions into a single temporally indexed structure; evidence is then retrieved under explicit temporal constraints. On EHRXQA boolean questions (n = 1,900), the retrieval baselines remain close to a majority-class predictor on the MedGemma backbones; MedHyperGraph improves on all three backbones, reaching 73.95% accuracy (95% CI [71.9, 75.9]), without task-specific training. On open-ended questions, it raises weighted F1 by 10.0 points at 27B, driven by multi-study questions (5.2 → 43.4). Ablations attribute the gain primarily to the structured EHR (−9.1), indicating that EHR-grounded VQA benefits from retrieval that preserves the temporal and n-ary structure of clinical evidence

Projects

More links