Skip to main content Skip to main navigation

Publication

Synthesizing Visual Concepts as Vision-Language Programs

Antonia Wüst; Wolfgang Stammer; Hikaru Shindo; Lukas Henrik Helff; Devendra Singh Dhami; Kristian Kersting
In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2026, Denver, CO, USA, June 3-7, 2026. International Conference on Computer Vision and Pattern Recognition (CVPR), Pages 17346-17356, Computer Vision Foundation, 2026.

Abstract

Vision-Language models (VLMs) achieve strong perfor- mance on multimodal tasks but often fail at systematic vi- sual reasoning tasks, leading to inconsistent or illogical outputs. Neuro-symbolic methods promise to address this by inducing interpretable logical rules, though they ex- ploit rigid, domain-specific perception modules. We pro- pose Vision-Language Programs (VLP), which combine the perceptual flexibility of VLMs with systematic reasoning of program synthesis. Rather than embedding reasoning in- side the VLM, VLP leverages the model to produce struc- tured visual descriptions that are compiled into neuro- symbolic programs. The resulting programs execute di- rectly on images, remain consistent with task constraints, and provide human-interpretable explanations that enable easy shortcut mitigation. Experiments on synthetic and real-world datasets demonstrate that VLPs outperform di- rect and structured prompting, particularly on tasks requir- ing complex logical reasoning.

More links