ELSA3D is a unified 3D foundation model that structures language and geometric reasoning along matched abstraction scales. Scale-aware octree tokenization retains coarse structural cues and fine geometric detail. Sparse Anchor Tokens provide an explicit cross-modal interface: selected semantic tokens retrieve geometric evidence at a routed scale and fuse it back into the unified representation. An elastic router jointly controls block execution, MLP width, anchor selection, and geometric-scale assignment. Evaluations cover image-to-3D generation, text-to-3D generation, and object captioning, with ablations examining grounding and computational efficiency.
ELSA3DElastic Semantic Anchoring for Unified 3D Understanding and Generation
NeurIPS 2026
TL;DR
We introduce ELSA3D, a unified 3D model built around elastic semantic anchoring: sparse Anchor Tokens link selected language tokens to geometric evidence at the relevant scale. One router jointly controls anchor selection, geometric scale, block execution, and MLP width, improving language–geometry grounding while roughly halving FLOPs and inference latency compared with the non-elastic variant.

ELSA3D generates 3D objects from an image while preserving overall shape and local appearance.
Abstract
Contributions
- Elastic semantic anchoring structures language and geometric reasoning jointly. ELSA3D makes text–3D interaction explicit across matched abstraction scales, rather than collapsing coarse structural cues and fine geometric detail into an undifferentiated sequence. Scale-aware octree tokenization supplies structural bits, scale-specific content codes, positional embeddings, and fixed scale tags for multiscale geometry.
- Anchor Tokens provide a sparse, dynamic cross-modal interface. Each anchor selects a semantic token, routes it to the most relevant 3D scale, retrieves scale-specific geometric evidence, and writes the fused language–geometry signal back into the unified sequence, keeping interaction sparse yet precise.
- An elastic router jointly controls computation and grounding. One routing scheme controls block execution, MLP width, anchor selection, and geometric-scale assignment. The elastic router concentrates cross-modal capacity where semantic–geometric alignment is needed, adapting both computation and grounding to the input.
- ELSA3D improves unified 3D performance while reducing computation. The model achieves state-of-the-art results on image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline on every reported metric while roughly halving FLOPs and inference latency relative to a non-elastic variant of the same model.
ELSA3D Method
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, where one model can reconstruct an object from an image, generate one from text, describe its structure, and support downstream reasoning over geometry. Yet their text–3D interaction remains largely implicit: existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention to discover cross-modal correspondences, collapsing coarse structural cues and fine geometric detail into a single undifferentiated representation. What is missing is not merely a stronger hierarchical 3D representation, but a unified design that structures language and geometric reasoning jointly. ELSA3D addresses this with elastic semantic anchoring, enabling structured interaction between semantic cues and geometric content.
Scale-aware octree tokenization
Each 3D object is represented as a multiscale octree constructed from a canonicalized 128³ voxel grid, recursively subdivided to a maximum depth. The first three octree levels are fully populated to provide a stable global scaffold; beyond this coarse scaffold the octree is sparse, with nodes added only where they intersect surface geometry. Each node carries a structural bit indicating whether it is subdivided and a content token encoding its local geometry, quantized against a scale-specific codebook so that each vocabulary specializes in geometric primitives at its resolution. Every content token is further augmented with a learned positional embedding and a fixed, non-trainable scale tag, making geometric resolution recoverable from a token's embedding and available to the transformer. Within each depth, nodes are serialized in Morton (Z-order) to preserve spatial locality in the token sequence.
Scale-aware octree tokenization. ELSA3D's octree VQ-VAE encodes a voxelized 3D shape into multiscale structural bits and scale-specific content codes, then decodes them to reconstruct the shape. Nodes are organized by octree depth and serialized with Morton/Z-order to preserve spatial locality within each scale.
Anchor Tokens
Dense interaction between all semantic tokens and all 3D tokens is computationally wasteful and semantically noisy, since many words provide only global or contextual constraints while only a subset requires precise geometric grounding. Anchor Tokens form a sparse semantic–geometric interface that creates cross-modal interaction only where it is useful. At each block, every selected semantic token is used as a query to cross-attend over the 3D tokens at its routed scale, and the retrieved scale-specific evidence is fused with the semantic state to form a transient, block-local anchor.
Dynamic routing and elastic reasoning
Because input difficulty varies across examples and only a subset of semantic tokens requires precise geometric grounding, a lightweight per-block router makes reasoning elastic in both computation and grounding. From a mean-pooled block context it decides whether to execute the block and how much MLP width to allocate, discretized into four levels. At the token level, it predicts for each semantic token an anchor gate — whether the token instantiates an anchor — and a soft distribution over geometric scales. By first selecting a scale and then attending within it, each token performs a coarse-to-fine search for its most relevant geometric correspondence.
Quantitative Results
Text-to-3D generation
Prompt alignment and generated-shape quality, measured with CLIP, FD, KD, and Q-Align.
| Method | CLIP↑ | FD↓ | KD↓ | Q-Align↑ |
|---|---|---|---|---|
| Text-/image-conditioned generators | ||||
| Shap-E | 24.94 | 53.24 | 1.13 | 1.45 |
| LN3Diff | 18.79 | 68.09 | 2.24 | 2.14 |
| XCube | 26.37 | 31.82 | 0.42 | 1.68 |
| SAR3D | 23.21 | 22.43 | 0.23 | 2.91 |
| 3DTopia-XL | 25.89 | 43.46 | 1.18 | 1.47 |
| GaussianAnything | 24.76 | 28.94 | 0.51 | 2.27 |
| Trellis | 29.43 | 21.61 | 0.11 | 3.42 |
| Unified 3D models | ||||
| ShapeLLM-Omni | 27.98 | 24.40 | 0.15 | 3.21 |
| CoRe3D | 37.66 | 20.55 | 0.11 | 3.68 |
| ELSA3D | 39.01 | 16.50 | 0.10 | 3.82 |
3D object captioning
Lexical and semantic similarity between generated object descriptions and reference captions. Higher scores are better.
| Model | BLEU-1↑ | ROUGE-L↑ | METEOR↑ | Sentence-BERT↑ | SimCSE↑ |
|---|---|---|---|---|---|
| General vision-language models | |||||
| LLaVA-13B | 4.01 | 8.18 | 13.18 | 46.97 | 48.86 |
| Qwen2.5-VL-7B | 4.05 | 7.85 | 14.23 | 48.90 | 50.86 |
| 3D-specialized understanding models | |||||
| 3D-LLM | 15.11 | 17.84 | 19.22 | 42.36 | 43.58 |
| LEO | 16.98 | 20.12 | 20.91 | 48.01 | 47.25 |
| PointLLM-13B | 3.18 | 7.54 | 12.24 | 47.89 | 49.01 |
| Unified 3D understanding and generation models | |||||
| ShapeLLM-Omni | 18.92 | 21.46 | 22.12 | 49.43 | 50.72 |
| CoRe3D | 24.02 | 26.45 | 24.98 | 51.17 | 52.79 |
| ELSA3D | 26.22 | 28.15 | 25.78 | 54.73 | 54.11 |
Ablation Studies
Anchor design
| Variant | Text-to-3D | Image-to-3D | Captioning | Cost | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CLIP↑ | FD↓ | KD↓ | CLIP↑ | FD↓ | KD↓ | MET.↑ | SimCSE↑ | FLOPs (G)↓ | Lat. (s)↓ | |
| No Anchors | 36.72 | 21.84 | 0.17 | 85.94 | 14.78 | 0.12 | 23.42 | 50.87 | 568 | 15.4 |
| Direct Cross-Attn | 38.32 | 17.91 | 0.12 | 88.41 | 10.72 | 0.08 | 24.96 | 53.34 | 1081 | 29.8 |
| Dense Anchors | 38.57 | 18.34 | 0.11 | 88.63 | 11.16 | 0.07 | 25.21 | 53.62 | 865 | 23.6 |
| ELSA3D | 39.01 | 16.50 | 0.10 | 89.21 | 9.23 | 0.06 | 25.78 | 54.11 | 632 | 17.2 |
Scale-aware routing
| Variant | Text-to-3D | Image-to-3D | Captioning | Lat. (s)↓ | |||||
|---|---|---|---|---|---|---|---|---|---|
| CLIP↑ | FD↓ | KD↓ | CLIP↑ | FD↓ | KD↓ | MET.↑ | SimCSE↑ | ||
| All-Scale Attn. | 38.72 | 17.21 | 0.11 | 88.71 | 10.18 | 0.07 | 25.36 | 53.76 | 22.4 |
| Coarse-Only | 37.84 | 19.76 | 0.14 | 86.92 | 13.40 | 0.10 | 24.46 | 52.41 | 16.5 |
| Fine-Only | 38.09 | 18.88 | 0.13 | 87.63 | 12.46 | 0.09 | 24.73 | 52.86 | 18.6 |
| Random Scale | 35.31 | 21.24 | 0.19 | 81.05 | 14.82 | 0.14 | 22.02 | 49.21 | 17.6 |
| ELSA3D | 39.01 | 16.50 | 0.10 | 89.21 | 9.23 | 0.06 | 25.78 | 54.11 | 17.2 |
Random scale assignment performs worst, so scale diversity alone is insufficient. Constraining every cue to a single resolution (coarse-only or fine-only) underperforms, since category-level cues favor coarse geometry while part and appearance cues require finer scales. Attending to all scales is competitive but slower (22.4s), whereas learned routing attains the best quality at near coarse-only latency.
Elastic computation — adapt depth and width
| Variant | Text-to-3D | Image-to-3D | Captioning | Compute | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CLIP↑ | FD↓ | KD↓ | CLIP↑ | FD↓ | KD↓ | MET.↑ | SimCSE↑ | FLOPs (G)↓ | Lat. (s)↓ | |
| Full-Compute (reference) | 39.18 | 16.22 | 0.09 | 89.34 | 8.92 | 0.06 | 25.89 | 54.28 | 1284 | 34.6 |
| Depth-Only Elastic | 38.52 | 17.61 | 0.12 | 88.46 | 10.42 | 0.08 | 25.23 | 53.64 | 811 | 24.1 |
| Width-Only Elastic | 38.21 | 18.04 | 0.12 | 88.12 | 10.91 | 0.08 | 25.05 | 53.37 | 985 | 27.0 |
| ELSA3D | 39.01 | 16.50 | 0.10 | 89.21 | 9.23 | 0.06 | 25.78 | 54.11 | 632 | 17.2 |
The full-compute model — every block executed at full width — is a quality upper bound but costs 1284 GFLOPs and 34.6s. ELSA3D retains near-full-compute quality at 632 GFLOPs and 17.2s, roughly a 2× reduction. Adapting depth or width alone is weaker than adapting both, indicating that efficient unified 3D modeling benefits from jointly choosing which blocks execute and how much capacity each uses.
Qualitative Results
Image-to-3D generation. From a single input image, ELSA3D more faithfully preserves both the global shape and the local appearance cues of the input than strong generative and unified baselines.
Text-to-3D generation. ELSA3D more faithfully follows category-level intent and fine-grained prompt constraints, including object parts, support structures, materials, and surface appearance.
3D object captioning. Given a 3D shape, ELSA3D produces richer, better-grounded descriptions, capturing both structure and appearance detail that terser baselines omit.
BibTeX
@inproceedings{yu2026elsa3d,
title={ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation},
author={Yu, Tianjiao and Li, Xinzhuo and Shen, Yifan and Susladkar, Onkar and Liu, Yuanzhe and Zhou, Xiaona and Lourentzou, Ismini},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}
Template acknowledgements
This site is built upon the work of Nerfies, made available under the Creative Commons Attribution-ShareAlike 4.0 International License.