Publication
Synthesizing Visual Concepts as Vision-Language Programs
Antonia Wüst; Wolfgang Stammer; Hikaru Shindo; Lukas Henrik Helff; Devendra Singh Dhami; Kristian Kersting
In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2026, Denver, CO, USA, June 3-7, 2026. International Conference on Computer Vision and Pattern Recognition (CVPR), Pages 17346-17356, Computer Vision Foundation, 2026.
Abstract
Vision-Language models (VLMs) achieve strong perfor-
mance on multimodal tasks but often fail at systematic vi-
sual reasoning tasks, leading to inconsistent or illogical
outputs. Neuro-symbolic methods promise to address this
by inducing interpretable logical rules, though they ex-
ploit rigid, domain-specific perception modules. We pro-
pose Vision-Language Programs (VLP), which combine the
perceptual flexibility of VLMs with systematic reasoning of
program synthesis. Rather than embedding reasoning in-
side the VLM, VLP leverages the model to produce struc-
tured visual descriptions that are compiled into neuro-
symbolic programs. The resulting programs execute di-
rectly on images, remain consistent with task constraints,
and provide human-interpretable explanations that enable
easy shortcut mitigation. Experiments on synthetic and
real-world datasets demonstrate that VLPs outperform di-
rect and structured prompting, particularly on tasks requir-
ing complex logical reasoning.
