PRIMAMulti-Image Vision-Language Models for Reasoning Segmentation

♦University of Illinois Urbana-Champaign●Virginia Tech
* Equal Contribution

TL;DR

We introduce PRIMA, a multi-image vision-language model that couples cross-image reasoning with pixel-grounded explanations. Its SQuARE module uses learnable relational queries to gather shared context across images, then injects that context into compact visual tokens before language fusion. PRIMA generates answers and segmentation masks together, trained and evaluated on M⁴Seg’s approximately 744K multi-image question-answer pairs with object and part masks.

M4Seg functional, spatial, numerical, and open-ended pixel-grounded reasoning examples

Reason across images, ground the answer in pixels

PRIMA compares image sets and associates the objects and parts referenced in its answers with segmentation masks.

Abstract

Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning segmentation alongside PRIMA an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4Seg a new multi-image reasoning segmentation benchmark consisting of ~744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with 7.83% and 11.25% improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships.

Contributions

  • Novel Task. We propose the novel task of multi-image pixel-grounded reasoning segmentation which requires fine-grained comparison and contextual understanding across multiple images at pixel level and natural language reasoning.
  • New Benchmark. We curate M4Seg a new challenging benchmark with ~744K multi-image QA pairs, annotated with multiple object and part segmentation masks to train and evaluate multi-image pixel-grounding models.
  • New Multi-Image Pixel-Grounded LVLM. We propose PRIMA an LVLM designed to perform instruction-guided cross-image alignment of relevant visual features via a novel SQuARE module, enabling reasoning with contextually grounded segmentation masks across multiple images. Experiments demonstrate PRIMA's impressive performance compared to strong baselines across segmentation (+↑8.11% mIoU and +↑7.83% Recall) and text-based metrics (+↑6.45% Semantic Similarity and +↑11.25% S-IoU).

PRIMA Method

PRIMA compares multiple images while grounding the answer in object and part masks. SQuARE builds a shared relational representation across the image set, conditions compact visual tokens on that context and the instruction, and passes them to a language model that generates the explanation and segmentation tokens.

PRIMA architecture: SQuARE, language model with LoRA, segmentation token projection, and SAM mask decoder
SQuARE supplies cross-image visual context to the language model. Segmentation tokens associated with objects in the answer are projected into a SAM-based decoder to produce their masks.
SQuARE module: relational queries attend across image features and enrich the global query pathway

How SQuARE connects the images

The visual encoder first extracts features for each image. Learnable relational queries attend over the concatenated features, gathering relationships shared across the image set.

  1. Gather relational context

    Relational attention builds a shared representation from all input images.

  2. Condition the global queries

    The shared context is added to the text-conditioned query pathway before global attention extracts compact visual features.

  3. Ground the explanation

    The language model uses these enriched visual tokens to reason across images and identify the objects or parts that its segmentation tokens should mask.

M⁴Seg Dataset

M⁴Seg trains and evaluates multi-image pixel-grounded reasoning. Each example links an image set and a reasoning question to a natural-language answer with object and part segmentation annotations. Success requires comparing the images and identifying the pixels that support the answer.

743,864
Question-answer pairs
207
Object categories
181
Part categories
3.49
Masks per answer

What an example contains

  1. Image set

    Multiple views or scenes supply evidence that must be compared.

  2. Reasoning question

    The instruction asks about relations, counts, functions, or broader context across the images.

  3. Grounded answer

    The answer explains the comparison and links referenced objects and parts to segmentation masks.

Image pairs and triplets are sampled from ADE20K-Part-234, Pascal-Part, and PACO-LVIS. Questions compare similar scenes containing related objects. Filtering removes answers without grounding across images, references to absent objects or parts, and examples with more than 16 segmentation masks.

Four kinds of cross-image reasoning

Functional

Compare what objects can do

Which vehicle is more suitable for carrying larger cargo?

Spatial

Compare positions and orientations

Which bird faces toward the left across the images?

Numerical

Compare quantities across scenes

Which image has fewer suitcases, and by how many?

Open-ended

Explain relationships in context

How do the tables’ functions relate to the surrounding objects?

Quantitative Results

Multi-image reasoning segmentation

Experimental Results on M4Seg. PRIMA significantly outperforms both general-purpose and pixel-grounding LVLM baselines. The general-purpose LVLMs struggle with this task, a limitation we attribute to their lack of task-specific training and cascading errors from using SAM on the text-based grounding information. While Gemini-2.5 Pro is the top performer in this category, its performance remains limited with 30.21% mIoU and 68.51% SS. The reasoning segmentation baselines perform better, as they leverage pixel grounding capabilities and are finetuned on M4Seg. GLaMM, for instance, outperforms the general-purpose LVLMs with 38.12% mIoU and 74.05% SS. Compared to all baselines, PRIMA sets a new benchmark, surpassing the next best baseline by 8.11% and 12.33% in terms of mIoU and I-SS, respectively. PRIMA also achieves a gain of up to 6.45% and 11.25% on the text metrics SS and SIoU.

Multi-image reasoning segmentation
Model mIoU ↑ Recall ↑ I-SS ↑ I-SIoU ↑ SS ↑ SIoU ↑ METEOR ↑
Open-source LVLMs · post-hoc grounding
Qwen3-VL-8B 21.52 12.03 20.76 16.05 68.44 53.57 20.47
Nemotron-Nano-12B-v2 10.45 4.48 7.76 4.29 56.45 41.83 18.69
Qwen3-VL-235B-A22B 27.52 15.2 25.05 18.00 71.29 52.85 22.24
InternVL3-78B 19.65 10.4 19.07 15.35 65.02 49.23 20.03
Pixtral-Large-124B 28.54 14.20 26.90 20.72 70.88 54.29 21.44
Closed-source LVLMs · post-hoc grounding
GPT-4o 25.47 12.19 21.59 15.76 68.62 50.55 21.56
Gemini 2.5 Pro 30.21 15.25 24.69 15.51 68.51 45.21 18.89
Pixel-grounding LVLMs
LISA 37.21 21.36 39.34 33.31 73.32 63.86 25.47
GLaMM 38.12 22.78 41.91 35.76 74.05 64.59 25.33
SQuARE ablations
Pooling instead of SQuARE 44.34 26.59 45.79 37.11 72.37 60.05 26.24
Projection instead of SQuARE 45.08 28.12 50.41 45.91 78.14 71.33 25.88
PRIMA (Ours) 46.23 30.61 54.24 51.62 80.50 75.84 25.70
Δ PRIMA − GLaMM +8.11 +7.83 +12.33 +15.86 +6.45 +11.25 +0.37

Qualitative Results

Qualitative Examples.
Qualitative Results. PRIMA exhibits strong qualitative performance in both segmentation and reasoning. Each example shows the posted question, the textual responses from each model, and the corresponding segmentation masks alongside the Ground Truth reference. For clarity, we use consistent colors between the segmentation masks and the highlighted text spans referring to each object or part. The results indicate that PRIMA produces high-quality textual responses together with crisp, well-localized segmentation masks, whereas GLaMM often yields noisy masks, attends to incorrect objects (e.g., treating the cabinet as a surface for placing small items), and hallucinates categories (e.g., misidentifying the horse as a truck or a dog). We also include a failure case in which PRIMA predicts the correct number of objects but fails to precisely distinguish individual instances, instead producing overlapping masks.

PRIMA in the Wild

PRIMA in the wild results.
PRIMA's performance on unseen images found on the web. Notably, PRIMA generates superior segmentation masks compared to baselines, as exemplified by the fine-grained details of the hammers in Example 3 and the dining chairs in Example 4. Moreover, the baselines are more prone to hallucinations, e.g., "pizza is shown with multiple slices" in the first example (GLaMM), "watch" in the second example (both GLaMM and LISA). Beyond visual recognition, these examples underscore PRIMA's capability to retain compositional reasoning capabilities when transferring to diverse and noisy real-world images.

BibTeX

@article{wahed2024prima,
  title={PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation},
  author={Wahed, Muntasir and Nguyen, Kiet A and Juvekar, Adheesh Sunil and Li, Xinzhuo and Zhou, Xiaona and Shah, Vedant and Yu, Tianjiao and Yanardag, Pinar and Lourentzou, Ismini},
  journal={arXiv preprint arXiv:2412.15209},
  year={2024}
}
Template acknowledgements

This site is built upon the work of Nerfies, made available under the Creative Commons Attribution-ShareAlike 4.0 International License. We gratefully acknowledge LLaVA, GLaMM, DINOv2, Q-Former, and Cheetah for open-sourcing their models and code.