Pith. sign in

Paper Citation Record · LEDGER

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency

As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2512.01008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.01008 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:22:04.169020Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee625a1d-d930-4b74-9f6f-9b4ed33aef11 · outbound

This paper cites ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.068354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.068354Z digest=sha256:a23967a11c04a91b9c7997447a5cfb6a9820e22086ec9a071862d8b94a0e6787

Observation 59d774a7-62da-4978-9f45-16185531320c · outbound

This paper cites Qwen2.5-VL Technical Report, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Qwen2.5-VL Technical Report, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.073150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.073150Z digest=sha256:aa9254bba4191d36d0cebbc9fb9a62d398e58278d89f0a2b6d116a21c3f3f424

Observation e9d708dc-b5f9-4f7c-867a-d211eaae0c0b · outbound

This paper cites From thousands to billions: 3d visual language ground- ing via render-supervised distillation from 2d VLMs.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency From thousands to billions: 3d visual language ground- ing via render-supervised distillation from 2d VLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.076635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.076635Z digest=sha256:8d188cc16f362b851e7e5ae7fa168732af99abceae847feb53a4489eb42de816

Observation 1f688f06-5710-4471-828a-053e172b4c09 · outbound

This paper cites SAM 3: Segment Anything with Concepts, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SAM 3: Segment Anything with Concepts, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.080543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.080543Z digest=sha256:648f9ee127f24489d5b9709b1cffb2057d4d7491154406b4038333b68fe95ee5

Observation 6ed007e5-8c81-49ac-adc2-f42b51050920 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.084353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.084353Z digest=sha256:70e517f74e84783fcbbd00137288a2205f8410cb39b345c49c9aa29aac2b425b

Observation 9bb644bd-eed4-4f30-a43a-53a1277142c0 · outbound

This paper cites Sdfusion: Multimodal 3d shape completion, reconstruction, and generation.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Sdfusion: Multimodal 3d shape completion, reconstruction, and generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.088236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.088236Z digest=sha256:f7bc77a98f6c0642a1343a2908d0ddb1ebaa82deae9a5a345d09d6346e8f0ae8

Observation cc2bb181-470e-4cfa-8d68-c55480adc5a8 · outbound

This paper cites Lam3d: Large image-point clouds align- ment model for 3d reconstruction from single image.Ad- vances in Neural Information Processing Systems, 37:4454– 4480, 2024.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Lam3d: Large image-point clouds align- ment model for 3d reconstruction from single image.Ad- vances in Neural Information Processing Systems, 37:4454– 4480, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.092329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.092329Z digest=sha256:c04b178fd05383f4a775255a5aaf7793124f2446b0ef6381c0ed53c8dd479830

Observation f447d6a5-bab7-42b5-b350-ff14773c9387 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.095972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.095972Z digest=sha256:d208c150b0ca81699c35daade7bd95dc200c25648bd4bb355315692d314b2159

Observation 99c161c3-e72b-4eaf-a6b1-ce0fddbfbc87 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.099573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.099573Z digest=sha256:87e90d15e61638832c95615fe09357048c4984350a5fd3b9698e53ecb4309deb

Observation 5ec97523-7de9-4617-9329-5c491c15786a · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.103070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.103070Z digest=sha256:ebc00b087a1409fb275b448f92881bcf333513bac0ff64b122133e495ca26862

Observation d2565d1d-0aab-4bf0-a4b1-d374143ef46d · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.107066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.107066Z digest=sha256:2bcbdd3e20cb315e57282565da858d2ebdea3c750caabc09489e9fb8dd0ef483

Observation 143fcfd2-9e46-45a1-92cc-8e347d6e4ef9 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.111126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.111126Z digest=sha256:eeef465497149b48f8dc868f0c4d08e9c26d60af3ae0f6de88cc8968247b240f

Observation 0c6be40f-b30c-41f9-9bb0-a6a6d6a64354 · outbound

This paper cites LISA: Reasoning Segmen- tation via Large Language Model, 2024.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency LISA: Reasoning Segmen- tation via Large Language Model, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.114579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.114579Z digest=sha256:c826ca34e8e1529229b4a5dae16613115c05990854d01a4a0415ebf8e524c680

Observation 99dca848-b166-4097-a0a1-c8e8aa0820c1 · outbound

This paper cites Part123: part-aware 3d reconstruction from a single-view image.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Part123: part-aware 3d reconstruction from a single-view image

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.117895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.117895Z digest=sha256:87bfdf673cbad93023c0d2087b776d1f4ca05aeeafa94be3c91f970c8c24366b

Observation 88db3e1a-84d6-4151-a57e-b0d6ae85998b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.121349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.121349Z digest=sha256:9ed5ef08fa38430390719a3ad4779661e4be21c8a75c631670c41fd9046ef3e2

Observation fd11c686-cc90-433f-a5e6-a0a8125de3d9 · outbound

This paper cites Decoupled Weight Decay Regularization.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Decoupled Weight Decay Regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.125440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.125440Z digest=sha256:a266eed91ce90239ae50fe820497dcfb76c066789791c5ec477365bbd414dd47

Observation 98f6dc14-9a7c-40ce-88c4-044f2359a9cf · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Learning transferable visual models from natural language supervi- sion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.129035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.129035Z digest=sha256:a8271f2c349dacc687fb1b7ba678050aa5b14bb1309fbd7b2d9a028347b4a6a5

Observation 59a19e38-e36d-4cc6-9bbe-5ded861c5608 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos,.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SAM 2: Segment Anything in Images and Videos,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.132279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.132279Z digest=sha256:63a5f4300d1d6ecf44f99af9f6105992ddf71516f10e9170b9e607b0dfa18575

Observation 1bcf6dc4-b341-4eeb-9ffe-29b565129a99 · outbound

This paper cites A survey of language-grounded mul- timodal 3d scene understanding.Knowledge-Based Systems, 321:113650, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency A survey of language-grounded mul- timodal 3d scene understanding.Knowledge-Based Systems, 321:113650, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.135636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.135636Z digest=sha256:1f86a182e2dbea143cb20d2452fcd5d15c326e49afa064246a73075ed54e62ab

Observation 1f29bdc8-0905-4165-94ac-ee4b37488acc · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.138576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.138576Z digest=sha256:984acd8e8d3a29d7a0500f950e50aa96a163e97e3cb5bcfd3043764214d79e24

Observation 2241c7dc-e393-4d51-b749-64b7fcaa0f37 · outbound

This paper cites Anything-3D: Towards Single-view Anything Reconstruction in the Wild.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Anything-3D: Towards Single-view Anything Reconstruction in the Wild

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.142399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.142399Z digest=sha256:af4c52a769e42b9bac011521c410cac9f4a2eae334e64f26fb90d4a30406f83f

Observation addd49e9-abbb-4e9a-9fe3-b4f0933fc794 · outbound

This paper cites Point-gnn: Graph neural net- work for 3d object detection in a point cloud.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Point-gnn: Graph neural net- work for 3d object detection in a point cloud

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.145796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.145796Z digest=sha256:bffb20df71f43e89c0b2dd0f8dd1c9711dd6c9003260f19cb07c8e1880d5d85c

Observation e89dc7f5-0491-427d-a054-7ffc7845727a · outbound

This paper cites Splat-MOVER: Multi-Stage, Open- V ocabulary Robotic Manipulation via Editable Gaussian Splatting.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Splat-MOVER: Multi-Stage, Open- V ocabulary Robotic Manipulation via Editable Gaussian Splatting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.148641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.148641Z digest=sha256:71dcf616dd565f96fa95dfeb0972d2cef2414e576baa47746f75dd62034ad88a

Observation 021f976c-f0b5-47ba-948b-2c75e928d469 · outbound

This paper cites Natural Language Guided Goals for Robotic Manipulation.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Natural Language Guided Goals for Robotic Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.151842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.151842Z digest=sha256:64f09ce769142d95bbe69d959f0447c6a6f913192611f8047356545874d3508e

Observation e388a48c-70e8-48a6-94dd-43d8215c9c91 · outbound

This paper cites SAM 3D: 3Dfy Anything in Images.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SAM 3D: 3Dfy Anything in Images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.155120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.155120Z digest=sha256:f835487c99970b69ba54999271b1e0cff1b62a5a991acec1e4a58dce912cba4b

Observation 01114c86-0576-4c11-9c4d-97d0a888b6b4 · outbound

This paper cites Sa2V A: Marrying SAM2 with LLaV A for Dense Grounded Understanding of Images and Videos, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Sa2V A: Marrying SAM2 with LLaV A for Dense Grounded Understanding of Images and Videos, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.159081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.159081Z digest=sha256:1782c53d57ae07d630e93c0e53cc11c81f74c50ee65c6784626930d32aae5bf4

Observation 4022fd66-06cf-4a6e-9034-bcf8553acea4 · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.162258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.162258Z digest=sha256:4c3e8fd1cd60bb385d120771a487c88447606c1fd48d733409e96f1f87d460e8

Observation 91e3baa7-afcb-44f5-84ed-f9500e5b0f16 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.165316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.165316Z digest=sha256:6540ca8812eb58207a1c62f3e7c893cd4e19a5e14c0c84ed2e26ec12e50f0bfa

Observation 6d03deb5-8ba9-48af-8377-2da9c989cb31 · outbound

This paper cites EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing,.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.169020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.169020Z digest=sha256:c7b576e212bb23d01dbfe3e9d63283ebeba6cffefe5af015be60e0d28d7192a7

Pith citing papers

No inbound Pith citation observations are available.