Functional
Compare what objects can do
Which vehicle is more suitable for carrying larger cargo?
We introduce PRIMA, a multi-image vision-language model that couples cross-image reasoning with pixel-grounded explanations. Its SQuARE module uses learnable relational queries to gather shared context across images, then injects that context into compact visual tokens before language fusion. PRIMA generates answers and segmentation masks together, trained and evaluated on M⁴Seg’s approximately 744K multi-image question-answer pairs with object and part masks.

PRIMA compares image sets and associates the objects and parts referenced in its answers with segmentation masks.
Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning segmentation alongside PRIMA an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4Seg a new multi-image reasoning segmentation benchmark consisting of ~744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with 7.83% and 11.25% improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships.
PRIMA compares multiple images while grounding the answer in object and part masks. SQuARE builds a shared relational representation across the image set, conditions compact visual tokens on that context and the instruction, and passes them to a language model that generates the explanation and segmentation tokens.


The visual encoder first extracts features for each image. Learnable relational queries attend over the concatenated features, gathering relationships shared across the image set.
Relational attention builds a shared representation from all input images.
The shared context is added to the text-conditioned query pathway before global attention extracts compact visual features.
The language model uses these enriched visual tokens to reason across images and identify the objects or parts that its segmentation tokens should mask.
M⁴Seg trains and evaluates multi-image pixel-grounded reasoning. Each example links an image set and a reasoning question to a natural-language answer with object and part segmentation annotations. Success requires comparing the images and identifying the pixels that support the answer.
Multiple views or scenes supply evidence that must be compared.
The instruction asks about relations, counts, functions, or broader context across the images.
The answer explains the comparison and links referenced objects and parts to segmentation masks.
Image pairs and triplets are sampled from ADE20K-Part-234, Pascal-Part, and PACO-LVIS. Questions compare similar scenes containing related objects. Filtering removes answers without grounding across images, references to absent objects or parts, and examples with more than 16 segmentation masks.
Compare what objects can do
Which vehicle is more suitable for carrying larger cargo?
Compare positions and orientations
Which bird faces toward the left across the images?
Compare quantities across scenes
Which image has fewer suitcases, and by how many?
Explain relationships in context
How do the tables’ functions relate to the surrounding objects?
Experimental Results on M4Seg. PRIMA significantly outperforms both general-purpose and pixel-grounding LVLM baselines. The general-purpose LVLMs struggle with this task, a limitation we attribute to their lack of task-specific training and cascading errors from using SAM on the text-based grounding information. While Gemini-2.5 Pro is the top performer in this category, its performance remains limited with 30.21% mIoU and 68.51% SS. The reasoning segmentation baselines perform better, as they leverage pixel grounding capabilities and are finetuned on M4Seg. GLaMM, for instance, outperforms the general-purpose LVLMs with 38.12% mIoU and 74.05% SS. Compared to all baselines, PRIMA sets a new benchmark, surpassing the next best baseline by 8.11% and 12.33% in terms of mIoU and I-SS, respectively. PRIMA also achieves a gain of up to 6.45% and 11.25% on the text metrics SS and SIoU.
| Model | mIoU ↑ | Recall ↑ | I-SS ↑ | I-SIoU ↑ | SS ↑ | SIoU ↑ | METEOR ↑ |
|---|---|---|---|---|---|---|---|
| Open-source LVLMs · post-hoc grounding | |||||||
| Qwen3-VL-8B | 21.52 | 12.03 | 20.76 | 16.05 | 68.44 | 53.57 | 20.47 |
| Nemotron-Nano-12B-v2 | 10.45 | 4.48 | 7.76 | 4.29 | 56.45 | 41.83 | 18.69 |
| Qwen3-VL-235B-A22B | 27.52 | 15.2 | 25.05 | 18.00 | 71.29 | 52.85 | 22.24 |
| InternVL3-78B | 19.65 | 10.4 | 19.07 | 15.35 | 65.02 | 49.23 | 20.03 |
| Pixtral-Large-124B | 28.54 | 14.20 | 26.90 | 20.72 | 70.88 | 54.29 | 21.44 |
| Closed-source LVLMs · post-hoc grounding | |||||||
| GPT-4o | 25.47 | 12.19 | 21.59 | 15.76 | 68.62 | 50.55 | 21.56 |
| Gemini 2.5 Pro | 30.21 | 15.25 | 24.69 | 15.51 | 68.51 | 45.21 | 18.89 |
| Pixel-grounding LVLMs | |||||||
| LISA | 37.21 | 21.36 | 39.34 | 33.31 | 73.32 | 63.86 | 25.47 |
| GLaMM | 38.12 | 22.78 | 41.91 | 35.76 | 74.05 | 64.59 | 25.33 |
| SQuARE ablations | |||||||
| Pooling instead of SQuARE | 44.34 | 26.59 | 45.79 | 37.11 | 72.37 | 60.05 | 26.24 |
| Projection instead of SQuARE | 45.08 | 28.12 | 50.41 | 45.91 | 78.14 | 71.33 | 25.88 |
| PRIMA (Ours) | 46.23 | 30.61 | 54.24 | 51.62 | 80.50 | 75.84 | 25.70 |
| Δ PRIMA − GLaMM | +8.11 | +7.83 | +12.33 | +15.86 | +6.45 | +11.25 | +0.37 |
@article{wahed2024prima,
title={PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation},
author={Wahed, Muntasir and Nguyen, Kiet A and Juvekar, Adheesh Sunil and Li, Xinzhuo and Zhou, Xiaona and Shah, Vedant and Yu, Tianjiao and Yanardag, Pinar and Lourentzou, Ismini},
journal={arXiv preprint arXiv:2412.15209},
year={2024}
}