Skip to main content Skip to main navigation

Publication

Structural Bottleneck Reasoning: Efficient Medical VQA via Concept Alignment

Hoang Nguyen; Dang Le; Ta Duc Huy; Han Nguyen; Vy Tuong Dang; Duy Duong-Tran; Khoa D Doan; Anji Liu; Li Shen; Pengtao Xie; others
In: ICML 2026 3rd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences. International Conference on Machine Learning (ICML), ICML 2026 3rd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences, 2026.

Abstract

The adoption of a medical visual Chain-of-Thought (CoT) is fundamental to establishing trustworthy predictions; by externalizing the model's internal logic, it allows clinicians to verify that a diagnosis is derived from valid clinical indicators rather than spurious correlations. However, the expert-driven CoT paradigm is often bottlenecked by the prohibitive cost of dense manual annotations. To resolve this, we introduce \textsc{StructAlign}, a framework that enables the seamless integration of expert knowledge via a predefined list of clinical concepts. Rather than requiring expensive, sentence-level supervision, our method leverages these concepts as \textit{latent causal mediators} to anchor the model’s reasoning. This prevents \textit{reasoning collapse} and ensures diagnostic consistency, providing a scalable, cost-effective path toward interpretable AI in high-stakes medicine. Our results demonstrate that StructAlign maintains high performance even in data-scarce regimes (10% and 40%), achieving large-margin improvements of up to +24.64%, demonstrating that conceptual grounding can effectively restore high-fidelity diagnostic reasoning, thereby reducing reliance on exhaustive expert supervision.