Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:26.911809Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.22817.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:26.911809Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d862c8fb-0622-4778-bee2-d0b2969728d3 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5c8f2f-b8cc-417e-b8bc-0075048536cb · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Matterport3d: Learning from rgb-d data in indoor environments
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 817ea282-ef21-49ab-90b6-975c251dccff · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Clip2scene: Towards label-efficient 3d scene understanding by clip
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221a3ce0-75cf-48e6-acc7-849d3156425a · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Yolo-world: Real-time open-vocabulary object detection
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aade28f8-34cc-4f69-8720-4b9de7222f8d · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e325d129-4696-43da-a644-0f31ed7de852 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c98312e-0c7c-4ea7-ab69-f040c085cf36 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pla: Language-driven open- vocabulary 3d scene understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1c40434-2502-4382-a8c3-10198e2ea9ce · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A density-based algorithm for discovering clusters in large spatial databases with noise
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da4eb8b9-77c3-4f66-a972-72f2d89e7aac · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Efficient graph-based image segmentation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f8aed7b-1828-4be3-91f1-360252aefda2 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scal- ing open-vocabulary image segmentation with image-level labels
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85ed6f15-1795-40cc-a14d-7ed5193835f3 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sam-guided graph cut for 3d instance segmentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c56b82a9-481e-4cc5-b0d2-88f168114bda · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2ae8e89-bc58-444b-9583-e40b705df8d1 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-Set Image Tagging with Multi-Grained Text Supervision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c22408fb-30cc-454f-ba23-2975b517077a · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Odin: A single model for 2d and 3d segmentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9ab477e-a527-40b0-9369-b7991f928d63 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Con- ceptfusion: Open-set multimodal 3d mapping
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02e28432-4ef9-4e2a-9420-90ca782c20ab · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0aedb603-85e6-4dfb-8e34-285348d27db4 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointgroup: Dual-set point group- ing for 3d instance segmentation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1f5d4af-1562-41c8-9058-30f798b0e6a8 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 481c95b7-e3f8-4e1a-8427-60d192cda968 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment any- thing
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d55d3f15-4868-472e-90ca-007d5cf55a39 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment Any 3D Object with Language
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f496db-a38c-4b33-be66-2d9be5002250 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language-driven semantic seg- mentation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c277b8fd-9aff-4d86-a887-ac667d61faeb · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed9a6c5e-9f19-464d-81d0-5ecd83f37dae · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounded language-image pre-training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d752310-2e3d-42d9-9dcc-5ae62c6f966b · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Decap: Decoding clip latents for zero-shot captioning via text-only 2 training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef56cedc-8d59-4cab-be92-20c9776cffaf · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary semantic segmentation with mask-adapted clip
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a463a6-d0dd-44dc-a926-ee60422fd869 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bef844e-217b-4ee7-89b5-d1e83e7972a0 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78e18653-9ae1-4f41-ba5d-378c4fac1c9e · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding An end-to- end transformer model for 3d object detection
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ded09a3-37f7-4e31-a942-01732b4775f8 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff8a22de-058e-4002-811f-b1881c0ff054 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ea5db78-d815-4c22-a769-3b417885b848 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding V oxel cloud connectivity segmentation- supervoxels for point clouds
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d8875fd-fabc-4f8b-b423-6a146be19250 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41261cd9-0f92-471a-be01-3c8e93769ee4 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a9ab7c-b2b5-4246-b961-4ba0d4178294 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1245972d-d8c0-4a64-99c6-c2ec725bae19 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnext: Revisiting pointnet++ with improved training and scaling strategies
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e783c82f-3083-4223-ae58-2b2dabd020fd · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Learn- ing transferable visual models from natural language super- vision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce3af4e4-8fb1-45ef-843c-4c957316bc50 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87199b5b-123c-4058-af14-2459ac33dc25 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Dense multimodal alignment for open-vocabulary 3d scene understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation caba26fd-c95e-4145-a951-3415bfd46ee5 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 144baabf-9704-493e-af1d-e4065f70590c · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointr- cnn: 3d object proposal generation and detection from point cloud
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1fc3a7d-c0f7-4299-946c-047ecd3e4457 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding The Replica Dataset: A Digital Replica of Indoor Spaces
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29eb27ca-f877-4a1e-afe7-a95ac6682dbb · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open- mask3d: open-vocabulary 3d instance segmentation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06d30a65-8d4b-412b-b94d-c319ffb436ab · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e55810b-e7e8-4d45-a634-fb19c99fbd84 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open vocabulary 3d scene under- standing via geometry guided self-distillation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2c0fcc5-b4fe-4ca7-943f-3b736b1ba0d6 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detr3d: 3d object detection from multi-view images via 3d-to-2d queries
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f98a4414-c302-4e2e-bc9c-b6af28403af1 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Uni3detr: Unified 3d detection trans- former
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d53afc04-3ee3-4b33-aaa1-191d221b7f27 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer v3: Simpler faster stronger
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b13d381d-c0aa-4ba0-933c-e9943c23fb9b · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc4587d7-fd21-485f-be34-8dd8e54e0d60 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1465d24-4171-442b-b92f-ff475a070587 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25cb3969-22a2-40b6-ad59-feb63a1b2f00 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A unified framework for 3d scene understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bad85a9-7d29-4045-b462-e393dfbb24f5 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86e4464a-03c7-4444-9c8e-6d82710e2d65 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sa3dip: Segment any 3d instance with potential 3d priors
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57d72f2d-c14b-4814-b926-729e614e05bd · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c503099-ae0c-4611-8aa9-392a2abfb99c · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point deformable network with enhanced nor- mal embedding for point cloud analysis
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4edc5ae-8f26-490b-8637-de3c652bf272 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sai3d: Segment any instance in 3d scenes
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87f8f197-2039-4b13-a046-9793e04d82f7 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9e1f414-50e4-40cc-afdc-03647a7ff35b · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Recognize anything: A strong image tagging model
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4399e95b-2d94-4dc2-a7c2-d401a625b62f · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87774ca9-b572-4089-94c6-198c6dd3e27e · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Se- ssd: Self-ensembling single-stage object detector from point cloud
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dd8fe42-6988-4e8c-aafe-49a356da40d7 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detecting twenty-thousand classes using image-level supervision
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0eefbddc-203e-45da-a043-65961e46fda9 · outbound
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with text-to-image diffusion models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.