Publikation
LocSAM3: Box-Supervised Adaptation of SAM3 for Text-Only Chest X-Ray Segmentation
Kirellos Nashed; Mina Heinein; Mina Youssef; Tamer Basha; Hasan Md Tusfiqur Alam; Abdulrahman Mohamed Selim; Omair Shahzad Bhatti; Daniel Sonntag
In: Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges. Medical Image Computing and Computer Assisted Intervention (MICCAI-2026), September 26 - October 1, Vol. LNCS 17266, Springer Nature Switzerland, 2026.
Zusammenfassung
Pixel-level segmentation masks are expensive to annotate, particularly in chest X-rays, where anatomical structures overlap and pathological boundaries are often ill-defined. Bounding boxes (BBoxes) provide a lower-burden form of localization supervision and are available in several chest X-ray datasets. We present LocSAM3, a mask-free adaptation of SAM3 for text-only anatomy and pathology segmentation that uses only BBoxes and concept names during training. Rather than learning radiograph-specific masks directly, LocSAM3 adapts SAM3’s concept-grounding and localization modules while keeping its pretrained mask decoder and more than $95\%$ of its vision backbone frozen. The fixed decoder then converts the learned in-domain localization into a pixel-level prediction. During inference, the model receives only an image and a text concept. On a 10-class MIMIC-CXR benchmark, LocSAM3 improves text-only segmentation over Medical-SAM3 from $0.48$ to $0.69$ mIoU for anatomical structures and from $0.09$ to $0.53$ for pathological findings. These results demonstrate that coarse localization annotations can support effective text-to-pixel adaptation in chest X-rays without requiring dense mask annotations for training or spatial prompts at inference.
Projekte
- No-IDLE - Interactive Deep Learning Enterprise
- NoIDLEChatGPT - No-IDLE meets ChatGPT
- TrustML - Trustworthy Machine Learning
