Pith. sign in

Paper Citation Record · LEDGER

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.02980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02980 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:29.029373Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aeb6dbff-f449-4e49-b91a-9442bef094cc · outbound

This paper cites ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.719317Z digest=sha256:0683c4de797f3d2b77d124ebba72dc76b9b6ddd1d5e3b29e5f94391699d54303

Observation 1d6a1d8e-9086-4208-94a4-ffcdb6f39801 · outbound

This paper cites Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.165375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.724810Z digest=sha256:c7214910fc8ab9582fc988a01f598ce18ebe5e4f65035ced086c9defbf4b0cac

Observation 2705b699-df6a-4def-ba9e-90ca31a5b5a5 · outbound

This paper cites Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.148181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.729999Z digest=sha256:d3b036ce63a22a1f47597e659a0565195e963df7ab7e8886911d6fdba2f31216

Observation bf16414f-d807-45ae-885d-41f2b0be43cc · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.129808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.734941Z digest=sha256:b41d47467f0f6cccccdb4fc8090e95240c5815c888601a7957d7fb06014b7092

Observation a53fa677-6872-40d5-97c4-0fd172862f17 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:30.112824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.740493Z digest=sha256:b7e490ccac8a6f8ee57620e3cf6e94b54747eda75e3a4e0e81b5e8ce841b72bb

Observation 07d077cc-e702-40ff-b6d3-9e1d824d8e99 · outbound

This paper cites Token merging: Your ViT but faster.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Token merging: Your ViT but faster

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.095012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.745801Z digest=sha256:996878ed98e0f6e83494648ceb2d7915a317cc0cc375035799943c74253904af

Observation 0c2c2a72-6b94-49ed-823a-92146925c91d · outbound

This paper cites From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.078294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.751420Z digest=sha256:f0dad26ef5f5ce0f1f06e8d58d0754479578cbf97fd64bde8fe827b2311e4cc0

Observation 558c4219-0243-4c7c-a0e5-5490ec30a8e7 · outbound

This paper cites End- to-End Object Detection with Transformers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding End- to-End Object Detection with Transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.059219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.756349Z digest=sha256:e3f365b6eb1159cfbc08360b91e31eb58a3cdcdc4871d6f5db3727c9b12e94f7

Observation ec62c190-df38-4cbf-8493-7fc304c59ed1 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.761505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.761505Z digest=sha256:2ab19952cb1637cc1f997dfca07dd339375167fe252955d8675e2cb8294a57da

Observation 9bb4ba43-8281-40f9-b5dc-512086bd13d1 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.039282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.766679Z digest=sha256:8328e6c480fd50cfe4eb75327db618f8d349dc74220f61ba27d2800e5ae8ba02

Observation ca779a4b-4832-461b-8f06-d871b5697ef7 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.017462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.771739Z digest=sha256:4a8c7e982f6443dc6b32e75dd840090e25fd67624024bed475692b298c6dd21a

Observation 619f4d38-021e-4b34-a0e4-bcd3f7a9c974 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.776677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.776677Z digest=sha256:698107c00a1d0bc797ac4d7e1bd76b0c8e0c178635ce098d04f76bf3d0593295

Observation 7196c203-57d3-4550-bb01-9805a5c02bfb · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.996076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.782179Z digest=sha256:60be472d57916b2c642a64dd28a00a37ba9573ae3098a49d9dce0f735c9eb087

Observation ff2b045f-ce34-4731-8b73-0618518d85e3 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.979160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.787025Z digest=sha256:e4b16f75a9ff13f56db5a31a996e2727e47899bd9dfcd12986af0e68f55a49f0

Observation e479c723-70a6-4685-a861-2e2752513adf · outbound

This paper cites Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.962201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.791806Z digest=sha256:550e2aa1aa1ed6def38fd43180ad265e52f5942554f7c6c89c710298153be818

Observation 335a98b9-e5cf-4641-8d30-f44a2c5d3c9b · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.796534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.796534Z digest=sha256:5d8848bbd252e415b876087d2c255fe0dc25e14e0aad32f243aacb741c2e3a9b

Observation 9d43c921-b730-4e1a-a0e7-c6ff566238d8 · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.802169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.802169Z digest=sha256:8324f09b1720802ed43d77e1b34e7879e1b127d3ea9dcd999d862e8d5ecdae2b

Observation 4bcbb5c7-fc2b-41e5-a0ee-43312f09a4b7 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.807289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.807289Z digest=sha256:2e4692817263b7ab77d5e328865823eff118c1b946832da7f3a26b4902e56670

Observation 0639deb9-e224-41af-832e-f8b83e7fa876 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.921345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.812043Z digest=sha256:66de48e18d21547acc464849b79fe65de16708266ae26ec76ea6343c7caafc8b

Observation 826b1c76-0114-4b09-a29b-1337a1ccb579 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.816998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.816998Z digest=sha256:6e06115a367ef721868d4e7c02960289377d74728e9ef91c9b4b9d34f5f12ac6

Observation da6b087d-9517-436b-b3b5-7d670b149057 · outbound

This paper cites Revisiting multimodal positional encoding in vision-language models, 2026.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Revisiting multimodal positional encoding in vision-language models, 2026

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.904945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.822679Z digest=sha256:fb286f20d9defe88fdc31a8335d444a5705b8efd9a64dad4197bf0e35165d20b

Observation dc93f922-558f-492f-95c5-56de9aac65c1 · outbound

This paper cites Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.887836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.827621Z digest=sha256:2c4ac6d33d922e19bde329cfd90fd14b40de77b16395a4bfbe42333b70386e5a

Observation 6d18d69e-bbb6-403e-a26e-21a3aeeb3ec8 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.869456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.832603Z digest=sha256:5a153c0e6645f92da9cfe396d61d0786595a71485aee5bcb373f019a663357e1

Observation 00114ae5-a656-48af-af38-86faf6e5ad96 · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Odin: A single model for 2d and 3d segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.850869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.837664Z digest=sha256:82a3c7e93e904cbe969560a03ef17921be395ceead9613ae6b3ecbff1b42ffc3

Observation 1f15f975-3b82-4750-8445-9ef82ff9df72 · outbound

This paper cites Unifying 2d and 3d vision-language un- derstanding, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 2d and 3d vision-language un- derstanding, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.832652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.842550Z digest=sha256:8a2d48bd1aff4a55c8188d21bc945c5ab360514cce25681c07e6762404efcffe

Observation 6594ca3a-8222-4c0b-9389-eafbf8bf2361 · outbound

This paper cites MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.814536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.847319Z digest=sha256:d28f3231842e7c8a40b77caeb80271955afc1af165613b6053753b480aeeceac

Observation 54add548-3d69-4b75-b0ed-14a7a4c44850 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.797281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.852077Z digest=sha256:73838ec539726745557bdf7635baf37bedcd38e7e72e7afc3f982a82caaddcf6

Observation 4df9d2ee-444b-4b06-8fcf-80a493fc90c8 · outbound

This paper cites Restr: Convolution-free referring image segmentation using transformers, 2022.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Restr: Convolution-free referring image segmentation using transformers, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.780428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.856930Z digest=sha256:dfb002a4f95d8b15362d53e6bb91fc376ea4401fa7d756bd8e462c6b3931d463

Observation 7a97f72d-02bf-433a-9bf6-2406accc3ba8 · outbound

This paper cites Mask-attention-free transformer for 3d in- stance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask-attention-free transformer for 3d in- stance segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.862157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.862157Z digest=sha256:9228b26a163e2e168ca07270248b7197ce89f2d15a274e4f99e64593e3938d19

Observation 1fbc27d2-3930-4d67-b19e-c3dcb341710d · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Lisa: Reasoning segmenta- tion via large language model, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.752413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.867247Z digest=sha256:905a05c973dbc3dbb77af23e7b1746eed69e28396ea0c665aec3a4580e3df268

Observation 2bc1063e-20fe-48e6-a533-28a0ba3899a5 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.872639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.872639Z digest=sha256:1e3fefaca27a823ed18c86f36eda551cda965efe0656fed2f14e66e3600f10ea

Observation 990060e3-85a9-4452-b100-14f7a6892a78 · outbound

This paper cites Grounded language-image pre-training.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.723467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.877707Z digest=sha256:2c2d0c3ef23acbcd9eb2d1fa6885e7cc2eeb3db33d2a767bc1a616362b5280af

Observation d6cdb1fe-4044-4bc4-84f5-41148dfb81d4 · outbound

This paper cites 3eed: Ground everything everywhere in 3d.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3eed: Ground everything everywhere in 3d

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.702798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.882564Z digest=sha256:bb834377842d3ca3cf6796250dc45879273ab568422c3a68cf60c8eb459e3898

Observation fff96ced-d235-4551-8a5b-736904880d32 · outbound

This paper cites Microsoft coco: Common objects in context.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Microsoft coco: Common objects in context

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.682245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.887595Z digest=sha256:ca389c9a2a2b44048e5d95e1d99a9cc7d5587f61e37a804c67cd2d528f3ba24a

Observation 679bd081-9eab-4124-b646-e150e493fdae · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.662921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.893221Z digest=sha256:f3e72b3eab52215ba97d61a03d79b66efee714642abc1b20395d0803c24d3ef6

Observation 86f5f433-2743-4c66-9336-3161c0ddaef9 · outbound

This paper cites View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.898167Z digest=sha256:b1e6eb18c87746af323a3f50b8e199830db80e09d30f14e772ae388ffd2b2e85

Observation e7e3f48c-5888-47b3-99cf-02e104416982 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.627699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.903472Z digest=sha256:497f7164f07790d198b94265b796c3334d6049be074e16e6afc75b8b67fedd2b

Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.908340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.908340Z digest=sha256:1ea34f2762df1d68d9f1cd228ac7698a8a207792497b4e48c4e7c398f0f24e9c

Observation 9f5db08d-fb08-4d91-97e2-a62a87f4d3cf · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.610276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.913803Z digest=sha256:0f3d69401aefb268d094d5fb57aa62c37c3ad3d7217166e98c2fd39d052613bb

Observation 89d94f3b-0fd8-4117-84e4-7eb4187182f3 · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Languagerefer: Spatial-language model for 3d visual grounding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.918804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.918804Z digest=sha256:f8bc02615a267f529ca492de62c11a4b62b6d41b38fe35a8097f9ac2158781ff

Observation 804b1176-f87a-4488-82aa-f34154d5fdb3 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.578334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.923771Z digest=sha256:d36ae759061ad073650293795d414e7bc0432ec7ad61fb84813aab6ce5921efb

Observation e8181c31-468e-47a5-9aaa-621d58963501 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.558961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.928579Z digest=sha256:684726f3e59db002c04717139e75b3c482a230a76e63aa967d64cdf0d3af1888

Observation e85d7520-83ca-4c62-a919-c53c0a1894e0 · outbound

This paper cites Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.540015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.933509Z digest=sha256:f81b76dfc7916c169273a133562378adbf5ef7ebf91b87c485838df219912572

Observation fb31858b-2724-4037-a0f0-14fc4a0e0baa · outbound

This paper cites Hashimoto.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hashimoto

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.521967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.938741Z digest=sha256:88d99c4936f8014f70b269d906f3547a176be3fa1be0a1e722ba07f6a6c87786

Observation e7588881-b2a1-403d-981b-8ab81995c8c4 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Gemini: A family of highly capable multimodal models, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.502906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.944625Z digest=sha256:7730536f8d2f9b1a56f2e693b97acbd33828b3c7032c8aec0edbc9c50ddeed33

Observation 00d5c209-b687-46ef-930b-6564ef3fe0c7 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.486307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.949602Z digest=sha256:70a5a54ed56321d1d5e366973515dea8498f9d2dff2966b68a0ff79488ddde2f

Observation 40bfa89e-34f3-4e62-89db-af3f72fd4069 · outbound

This paper cites MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.954483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.954483Z digest=sha256:b72fe1620d5107ed35c838f330fdc5bf6b00d4e23b03a5c3bacdc5b82611158e

Observation 9dfa92bb-95b8-4e8f-a1ee-e45f1e9792f5 · outbound

This paper cites Realworldqa.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Realworldqa

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.469572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.959712Z digest=sha256:0c2c23203ecdb5ff811ae08315fb65c5910d1187a3ec0e230ff0624d5b2e6719

Observation 6889bd48-7bd7-473d-ad12-834ff21fc8c2 · outbound

This paper cites Qwen2.5 Technical Report.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.964720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.964720Z digest=sha256:ae39ffdb87d00094eface480d3b399a04ebc1a27f2972428849127cf8cbe4306

Observation 73561748-6326-40c7-8411-bbce61381ead · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Sat: 2d semantics assisted training for 3d visual grounding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.452483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.969841Z digest=sha256:42636d4c4fc02c818606b5d683a18eff217d1948ffef3e690327bb4c300de725

Observation 08c79efa-76bc-4065-9f52-83fed13b6322 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.435741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.974617Z digest=sha256:6a77086e17a2df83075082bc4eda2f732c227b8f84ffb6fe72f69889ee7cab31

Observation 8e754b19-6361-4677-b4d2-9b35de80d58c · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.979475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.979475Z digest=sha256:2176d8a8580fbe4237c6926a1aab58abc0795eab665278b8834f31ff90640644

Observation 72dadb02-5187-48be-8e89-950c586ccd5d · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.408883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.984880Z digest=sha256:819d807ed11d6b20edae0864795c6f6915bde2d44a6cd5c10e11e4a8f29ad0b4

Observation a4be2bb5-2751-4d74-a4b4-c5f4ac0c264a · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.391572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.989830Z digest=sha256:77735b2e3d673eca64e075b2b4f1daa8350712cfde72233585464c5c2974de15

Observation 7ba9a399-a68b-4a16-b07c-e8d9b19d776e · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Towards learning a generalist model for embod- ied navigation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.371586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:28.994481Z digest=sha256:b1caac6e469124d1faba600b19b53efcff2a5375e4d9762d125b147a5a2a6a6f

Observation 473c6c45-a62f-48ee-946f-aeb34c1cb832 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.999121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.999121Z digest=sha256:21238a1908efa98f6291237d0f0f31b1928457a62003c2f00ac57f28e0fb1db5

Observation 1d49da31-7603-4865-8b38-9e7ca9fa0cb9 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.353145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:29.004503Z digest=sha256:e3345fdc734635112b3cdfcc52cb28fb7972b7cf4754388b2a7b1632d141e4b4

Observation bc1e7350-31e4-4b27-9f50-6a1eb06190e7 · outbound

This paper cites Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.336725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:29.009579Z digest=sha256:1ced98b1ea76862806f2785dea35ff66a5bca919fbb5ba3bb18bfa9b860d09a9

Observation aede57e3-b847-485b-8cae-c13d3ab279a2 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.320100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:29.014371Z digest=sha256:c980d926562d0bc72a4c45470cbf9fa06cff6aaecd7a0d78914532c1dfb5adf1

Observation 87b2bc23-7840-435c-aa55-b96ad2a15573 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:27:29.079215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:29.019048Z digest=sha256:d383bc5983b759338092c2d21f43dd2746337c7a5d9e04656a51cf1616f3867d

Observation 5d76e898-d0c7-45e5-a8e8-862ac983013b · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Generalized decoding for pixel, image, and lan- guage

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.303684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:29.024405Z digest=sha256:da413d019e57e0c465c0141a9b623b4b9824b84dedc44a4cb6d872ce2baf33c3

Observation 78a29522-3b1a-456e-9cc1-91dc6ac60d28 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T04:27:29.286999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T04:27:29.029373Z digest=sha256:9aa191f3d410dabf5f2b234438527bc9f7a280738ff61e98c723fabdc49ec4e2

Pith citing papers

No inbound Pith citation observations are available.