Pith. sign in

Paper Citation Record · LEDGER

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.22817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22817 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:26.911809Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d862c8fb-0622-4778-bee2-d0b2969728d3 · outbound

This paper cites Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:22.010410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:22.010410Z digest=sha256:6611620d1f52fc810b5f4f74cdbe3fc72b56a74dcb3d0a82213732922b8de9cd

Observation 0a5c8f2f-b8cc-417e-b8bc-0075048536cb · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Matterport3d: Learning from rgb-d data in indoor environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.605437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.071577Z digest=sha256:d7732e2d923546b0f955df661fa585d7094e10a17703b02ba8284b89eedae359

Observation 817ea282-ef21-49ab-90b6-975c251dccff · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:22.157362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:22.157362Z digest=sha256:35d6e6c668392f4e31151f479f00933dd0a9d430c2b87fa7b2ea61ba6ae55409

Observation 221a3ce0-75cf-48e6-acc7-849d3156425a · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.427717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.249266Z digest=sha256:355b5b05156b2ba1c68bf754b8a37ad6b2690994daecda9e2590cfb4ec034808

Observation aade28f8-34cc-4f69-8720-4b9de7222f8d · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.252491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.355844Z digest=sha256:c6a3597537daf0eb3a219ec0ed10f1af95bd01fa9c8c7d0d816675a41314df9a

Observation e325d129-4696-43da-a644-0f31ed7de852 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.057704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.437904Z digest=sha256:ad620ef2c69956907593335bfe9ff7f9009f5a87a73c80a5fcc6b3d9d5a68fc6

Observation 2c98312e-0c7c-4ea7-ab69-f040c085cf36 · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pla: Language-driven open- vocabulary 3d scene understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.820884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.513194Z digest=sha256:7ff749d8797dd7fd60ba1c66a7c5118a3ee0c96c81b6e2003aa6a3268246b192

Observation d1c40434-2502-4382-a8c3-10198e2ea9ce · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.597980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.618688Z digest=sha256:5c7b87fe35629fb949a2a4d341f365290dd3fe2a2b98cc9ad9a561816defe481

Observation da4eb8b9-77c3-4f66-a972-72f2d89e7aac · outbound

This paper cites Efficient graph-based image segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Efficient graph-based image segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.336113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.699936Z digest=sha256:4d7641997ef13c79f45d4847f725a753b9cccce0d805ddcea56450368f183a8a

Observation 1f8aed7b-1828-4be3-91f1-360252aefda2 · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scal- ing open-vocabulary image segmentation with image-level labels

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.102671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.777482Z digest=sha256:b1794c8e734841589e5f629c01db697713c1a534a96f82eff03315f516b1b539

Observation 85ed6f15-1795-40cc-a14d-7ed5193835f3 · outbound

This paper cites Sam-guided graph cut for 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sam-guided graph cut for 3d instance segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.842643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.821605Z digest=sha256:72139b0f511db016e824ed83dc98fbb940ec68aab3d081d09e405aa5a8228dc7

Observation c56b82a9-481e-4cc5-b0d2-88f168114bda · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.620721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.909272Z digest=sha256:2a7015898d88d7743d2e4ecbcaf26a6cfc41079f61a86cef247ce0d92e774f6e

Observation e2ae8e89-bc58-444b-9583-e40b705df8d1 · outbound

This paper cites Open-Set Image Tagging with Multi-Grained Text Supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-Set Image Tagging with Multi-Grained Text Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:04:27.162531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:22.991770Z digest=sha256:2ae852c2d6c2d522ceee734cffda01afd02df0bb8ab6113c67305739b93bb45b

Observation c22408fb-30cc-454f-ba23-2975b517077a · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Odin: A single model for 2d and 3d segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.404545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.080658Z digest=sha256:0aa12fce060e267cc4686e8262b5f00a2cfc9e165888dde9ebe30e01aab65153

Observation e9ab477e-a527-40b0-9369-b7991f928d63 · outbound

This paper cites Con- ceptfusion: Open-set multimodal 3d mapping.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Con- ceptfusion: Open-set multimodal 3d mapping

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.160072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.155746Z digest=sha256:6f038f59fb577409e659869e088a7d8f74153b2e57ae1be9050e05a121dc6dd5

Observation 02e28432-4ef9-4e2a-9420-90ca782c20ab · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.833358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.203242Z digest=sha256:4c5bbefef49192d2cc76a4776706006d159982dacb196fff6ab7a20cbb12c039

Observation 0aedb603-85e6-4dfb-8e34-285348d27db4 · outbound

This paper cites Pointgroup: Dual-set point group- ing for 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointgroup: Dual-set point group- ing for 3d instance segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.536514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.293731Z digest=sha256:4a6ef46ad4799714a290932d9c7651a501e0a3abfa5977dde419610739389321

Observation c1f5d4af-1562-41c8-9058-30f798b0e6a8 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.256176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.352475Z digest=sha256:261c761f04523f0d44d744fdc2de3824a7ff784d5f3a158b553700fdbc56ea8e

Observation 481c95b7-e3f8-4e1a-8427-60d192cda968 · outbound

This paper cites Segment any- thing.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment any- thing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.977129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.440915Z digest=sha256:8a0d02874bcb3143d2f7f771607050ad16bf5f5c8a631edbe116048eaed15b17

Observation d55d3f15-4868-472e-90ca-007d5cf55a39 · outbound

This paper cites Segment Any 3D Object with Language.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment Any 3D Object with Language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:23.521003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:23.521003Z digest=sha256:82fe4438be290676a95277c1ea5f3111ee093142d748b5c83e31bd1b59bf732d

Observation 87f496db-a38c-4b33-be66-2d9be5002250 · outbound

This paper cites Language-driven semantic seg- mentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language-driven semantic seg- mentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.711837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.650481Z digest=sha256:60ee2ff612e353fbab1dd99604e4c2c1377778944ebe315965cd4da5675ffca9

Observation c277b8fd-9aff-4d86-a887-ac667d61faeb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.445243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.704502Z digest=sha256:30f0d37a273d312d0473714b4715de8cf549c73205de53cb0ed412d7fe28d443

Observation ed9a6c5e-9f19-464d-81d0-5ecd83f37dae · outbound

This paper cites Grounded language-image pre-training.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounded language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.237398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.809703Z digest=sha256:2fbcd4c044c4154fa6afe961a843f3103a08a70a2e4a45c4a87f8fd2fe43e6fe

Observation 2d752310-2e3d-42d9-9dcc-5ae62c6f966b · outbound

This paper cites Decap: Decoding clip latents for zero-shot captioning via text-only 2 training.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Decap: Decoding clip latents for zero-shot captioning via text-only 2 training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.035731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:23.876577Z digest=sha256:4a9261022096d77e109960e3ad4999b5e8bb3ec307ea5fa99b2fed08134f8b61

Observation ef56cedc-8d59-4cab-be92-20c9776cffaf · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary semantic segmentation with mask-adapted clip

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:23.977911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:23.977911Z digest=sha256:35bbe24e82dc5c31175d668a2528b5f9bcb3b4333f949e4eaed1526b6baede4a

Observation 49a463a6-d0dd-44dc-a926-ee60422fd869 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.039336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.039336Z digest=sha256:684beabe0fc51b6881e7933be73792d44331c53083ce8e6df3275d78578e942e

Observation 2bef844e-217b-4ee7-89b5-d1e83e7972a0 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.758417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.101366Z digest=sha256:2113dcf5f07962eced41effb1947a7e481fe1e97ea42ac8839315c952fdb14ed

Observation 78e18653-9ae1-4f41-ba5d-378c4fac1c9e · outbound

This paper cites An end-to- end transformer model for 3d object detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding An end-to- end transformer model for 3d object detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.177736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.177736Z digest=sha256:206f8c48fb7b7e93f07ad4f2eb1c9c58c5319114fc06f8f185930a78256a34c0

Observation 1ded09a3-37f7-4e31-a942-01732b4775f8 · outbound

This paper cites Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.504554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.257808Z digest=sha256:41c3dca384680c24dba07fe85c20293055946d2636b8bbabc7da65948d40a85b

Observation ff8a22de-058e-4002-811f-b1881c0ff054 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.288532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.338008Z digest=sha256:f17aac35dc9ccfa38196ffdbca4c0b1420ed6b67923b6fd377570a84a7b02238

Observation 6ea5db78-d815-4c22-a769-3b417885b848 · outbound

This paper cites V oxel cloud connectivity segmentation- supervoxels for point clouds.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding V oxel cloud connectivity segmentation- supervoxels for point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.020530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.406111Z digest=sha256:ebc8fd539db089885e89905d5c390c6fec872195a4472dacbd70f6d56addcf89

Observation 6d8875fd-fabc-4f8b-b423-6a146be19250 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.783522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.455441Z digest=sha256:5c247e8a76b7127a76eaac137c4addd9702e77090c9375b03cb030d202a6b128

Observation 41261cd9-0f92-471a-be01-3c8e93769ee4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.595455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.595455Z digest=sha256:d69938951222ccb7af00484a99d0a649f84ef8ed17fa4dfb153c11e33924f91a

Observation f8a9ab7c-b2b5-4246-b961-4ba0d4178294 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.530755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.712330Z digest=sha256:b5ae46c29fc41696851dfdee7f91d922ff53cad055bdeeaf7d4303fbe822e94b

Observation 1245972d-d8c0-4a64-99c6-c2ec725bae19 · outbound

This paper cites Pointnext: Revisiting pointnet++ with improved training and scaling strategies.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnext: Revisiting pointnet++ with improved training and scaling strategies

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.233921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.813661Z digest=sha256:4ab160163ff42caa8a29dd1e757395972a03dcd4b6fda6223477892e9745fb87

Observation e783c82f-3083-4223-ae58-2b2dabd020fd · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Learn- ing transferable visual models from natural language super- vision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.018565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.906952Z digest=sha256:8114a90a62a47545dce5c561a0bce61026950d5c054cc0426b0a331ef68bef13

Observation ce3af4e4-8fb1-45ef-843c-4c957316bc50 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.777998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:24.978354Z digest=sha256:c852d6337f0e01e9b11701aac4cd5705eaeaa5d63fc08baf610e7760267a7a27

Observation 87199b5b-123c-4058-af14-2459ac33dc25 · outbound

This paper cites Dense multimodal alignment for open-vocabulary 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Dense multimodal alignment for open-vocabulary 3d scene understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.587167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.074672Z digest=sha256:4df30b136234d1a359c0899cdaf279e8d2e67149b55d2df307056849ca0abff2

Observation caba26fd-c95e-4145-a951-3415bfd46ee5 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.417674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.178981Z digest=sha256:1b5ab62dab61f9cd41503cc55a36783383bfdf758d715f67cfc8702e4804dd39

Observation 144baabf-9704-493e-af1d-e4065f70590c · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.277779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.277779Z digest=sha256:e2ef131e14a8bca9cd71412224fc04cffbce0fdea8c5369a225676e4d6c77bc7

Observation e1fc3a7d-c0f7-4299-946c-047ecd3e4457 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.352234Z digest=sha256:96f82782b0b797807325b6dbfb08e6f836e561d50d7aa88f43efbc443d339b02

Observation 29eb27ca-f877-4a1e-afe7-a95ac6682dbb · outbound

This paper cites Open- mask3d: open-vocabulary 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open- mask3d: open-vocabulary 3d instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.244663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.418913Z digest=sha256:b4dbdcdb85152bc66255ba782cd81ebfee395de35bed4d0fcb1400cdd6edbdd2

Observation 06d30a65-8d4b-412b-b94d-c319ffb436ab · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.152314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.484739Z digest=sha256:baf5e293d28a9038f434cf8d0fc5dc1cdd2c45fbbe80480fa232c571a0c98994

Observation 6e55810b-e7e8-4d45-a634-fb19c99fbd84 · outbound

This paper cites Open vocabulary 3d scene under- standing via geometry guided self-distillation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open vocabulary 3d scene under- standing via geometry guided self-distillation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.046440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.564953Z digest=sha256:1d0626c999baeb1d5dcd97c60c07dacf78a0207f6d9ce0a76643431f88ef23f4

Observation c2c0fcc5-b4fe-4ca7-943f-3b736b1ba0d6 · outbound

This paper cites Detr3d: 3d object detection from multi-view images via 3d-to-2d queries.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detr3d: 3d object detection from multi-view images via 3d-to-2d queries

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.937754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.629129Z digest=sha256:5fd95586d30e150fd24988566ae832eff2f1c17e2ef43819a43f1540f0795c21

Observation f98a4414-c302-4e2e-bc9c-b6af28403af1 · outbound

This paper cites Uni3detr: Unified 3d detection trans- former.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Uni3detr: Unified 3d detection trans- former

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.837411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.674003Z digest=sha256:ff15136874af977c50c1635eed95382999e45c079f1f77cbb11ba71ac5c116ad

Observation d53afc04-3ee3-4b33-aaa1-191d221b7f27 · outbound

This paper cites Point transformer v3: Simpler faster stronger.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer v3: Simpler faster stronger

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.715963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.735877Z digest=sha256:b5ca1b243abb68b0939624c6d1be856621063f0c857f1a44209449b79228a3d5

Observation b13d381d-c0aa-4ba0-933c-e9943c23fb9b · outbound

This paper cites Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.633199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.810601Z digest=sha256:7a86eb6dc1b281df5d757cfabcc7f4bcac10321779979901ba0ec21d44ad1a7d

Observation fc4587d7-fd21-485f-be34-8dd8e54e0d60 · outbound

This paper cites Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.543457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:25.853406Z digest=sha256:8a15e84609249f38f62c0b73c0d1923cd1bfba298b68abb472506833a332eaac

Observation a1465d24-4171-442b-b92f-ff475a070587 · outbound

This paper cites SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.946778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.946778Z digest=sha256:0795db0ca682cec4d357d9e43536523b959eed4cfceec916243ba1d88aeff4a2

Observation 25cb3969-22a2-40b6-ad59-feb63a1b2f00 · outbound

This paper cites A unified framework for 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A unified framework for 3d scene understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.453467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.007884Z digest=sha256:4af24956afda2c4ca2a5ecec29650957e851a4752e68ccc8183065546927bf95

Observation 1bad85a9-7d29-4045-b462-e393dfbb24f5 · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.365669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.049587Z digest=sha256:240b0a2c52e4b58d007ac559debc41aeda5b7ab19708f550b11f52d4e0cfb9d7

Observation 86e4464a-03c7-4444-9c8e-6d82710e2d65 · outbound

This paper cites Sa3dip: Segment any 3d instance with potential 3d priors.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sa3dip: Segment any 3d instance with potential 3d priors

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.252504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.130338Z digest=sha256:36a71003c22b0301444784385c7ee46b5f5f34ac81ee6f51b39c21e1ae37ea7c

Observation 57d72f2d-c14b-4814-b926-729e614e05bd · outbound

This paper cites SAM3D: Segment Anything in 3D Scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:26.211071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:26.211071Z digest=sha256:1bdc2c4c75ac5dfc83fd3c891d0397132b57dbda52f49e37abd3dedbef85bba1

Observation 4c503099-ae0c-4611-8aa9-392a2abfb99c · outbound

This paper cites Point deformable network with enhanced nor- mal embedding for point cloud analysis.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point deformable network with enhanced nor- mal embedding for point cloud analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.148885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.321851Z digest=sha256:adede1e23b9e63373a2df07da9b291712c63e9b0d05326957e96cf37623b6db4

Observation a4edc5ae-8f26-490b-8637-de3c652bf272 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sai3d: Segment any instance in 3d scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.062390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.399770Z digest=sha256:51da907245d0fda67b41ad60611be5ae220d92973793e92c837650caf8801639

Observation 87f8f197-2039-4b13-a046-9793e04d82f7 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.968280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.485535Z digest=sha256:14ea5a3337b0a45d02cca542fc51c7ea14fd6e70a0b748e3f24cabbb7fae1573

Observation d9e1f414-50e4-40cc-afdc-03647a7ff35b · outbound

This paper cites Recognize anything: A strong image tagging model.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Recognize anything: A strong image tagging model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.875800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.585203Z digest=sha256:8cb10b53f6d1bb1de7f930628a204403fd650463e233176d448a443386a46bb8

Observation 4399e95b-2d94-4dc2-a7c2-d401a625b62f · outbound

This paper cites Point transformer.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.789413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.664878Z digest=sha256:09fdec12bed7e273b35386af9a107b5c66d916c85a88f40e2d9804299a6cac63

Observation 87774ca9-b572-4089-94c6-198c6dd3e27e · outbound

This paper cites Se- ssd: Self-ensembling single-stage object detector from point cloud.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Se- ssd: Self-ensembling single-stage object detector from point cloud

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.655081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.765134Z digest=sha256:3d7d2f625ec427fdfaca3d3fa0f0fbda1224ab2e9ee6953df30467b409e1966b

Observation 9dd8fe42-6988-4e8c-aafe-49a356da40d7 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detecting twenty-thousand classes using image-level supervision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.511718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.846855Z digest=sha256:b33525759a4205292ecf1f8bfae925ed648a62bc84b912218e0993ca36156ce1

Observation 0eefbddc-203e-45da-a043-65961e46fda9 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with text-to-image diffusion models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with text-to-image diffusion models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.346606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:04:26.911809Z digest=sha256:e5046ba5d77dcbe415a6273a15c5a1af19d655111180b5fba1b7ea8b21a41ac0

Pith citing papers

No inbound Pith citation observations are available.