Pith. sign in

Paper Citation Record · LEDGER

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

As of 16 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.02980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02980 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:29.029373Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aeb6dbff-f449-4e49-b91a-9442bef094cc · outbound

This paper cites ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.719317Z digest=sha256:1339d1b701bec15a9d1a99fa54bbd95af88e9a647309a8d1bc59847838fc3e07

Observation 1d6a1d8e-9086-4208-94a4-ffcdb6f39801 · outbound

This paper cites Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.165375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.724810Z digest=sha256:5fe7151636f589c609ece4fe32187b4c6460d6c42ccf4482565279d0880eadf9

Observation 2705b699-df6a-4def-ba9e-90ca31a5b5a5 · outbound

This paper cites Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.148181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.729999Z digest=sha256:7fa2295011c5df1115bf4b240170cd6411b0142c0082711a832eca46548da143

Observation bf16414f-d807-45ae-885d-41f2b0be43cc · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.129808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.734941Z digest=sha256:ddf2540acb734a8f4fdc65b4656faeb5c771a57a31fda3f54cd52a710661460d

Observation a53fa677-6872-40d5-97c4-0fd172862f17 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:30.112824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.740493Z digest=sha256:35db8fe9f3089ae1cd837bf1a443f647930dcca185815bcc33d318da52db922e

Observation 07d077cc-e702-40ff-b6d3-9e1d824d8e99 · outbound

This paper cites Token merging: Your ViT but faster.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Token merging: Your ViT but faster

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.095012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.745801Z digest=sha256:01a0a7fc1f069a2c5651abf1bbba09f572ba54a683a6fe82792eec174cce13ba

Observation 0c2c2a72-6b94-49ed-823a-92146925c91d · outbound

This paper cites From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.078294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.751420Z digest=sha256:814334c934e2279471466b7b1b1d4d7341c097eb803e8d685966b0d96b16da1c

Observation 558c4219-0243-4c7c-a0e5-5490ec30a8e7 · outbound

This paper cites End- to-End Object Detection with Transformers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding End- to-End Object Detection with Transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.059219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.756349Z digest=sha256:ace70986356e49a3bcd834ced09d62bf7c8f79102b80d0e612d740cd18c35068

Observation ec62c190-df38-4cbf-8493-7fc304c59ed1 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.761505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.761505Z digest=sha256:e96687360717be0f4bcab23e16544ab118821fb75b2b492a7d2e9143999798bf

Observation 9bb4ba43-8281-40f9-b5dc-512086bd13d1 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.039282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.766679Z digest=sha256:f4dad568efc34c1d2489d3c737b397c3642bb1f096223bc2864621d67e96160e

Observation ca779a4b-4832-461b-8f06-d871b5697ef7 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.017462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.771739Z digest=sha256:2cb552405581fece5bd71602e7b249b6a49c7c0827a025e9a34f21db3221289b

Observation 619f4d38-021e-4b34-a0e4-bcd3f7a9c974 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.776677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.776677Z digest=sha256:d6aebb1e4fddc1050ae650c2c0b4236b10708a50c5da6d67d5927b84f031f183

Observation 7196c203-57d3-4550-bb01-9805a5c02bfb · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.996076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.782179Z digest=sha256:0a0e9663ee04c598f51a05166da5da4c02e40dcabf54321f92d049f80a89ca21

Observation ff2b045f-ce34-4731-8b73-0618518d85e3 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.979160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.787025Z digest=sha256:a116922115fffa5e40a30b9fda1dfe1c5dec50fafa9ddb972042625b2ed0f776

Observation e479c723-70a6-4685-a861-2e2752513adf · outbound

This paper cites Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.962201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.791806Z digest=sha256:6e078b1b60e1cb416cd2c9566cd800b115e4975c3ae0658aaef5f72e443c84ab

Observation 335a98b9-e5cf-4641-8d30-f44a2c5d3c9b · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.796534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.796534Z digest=sha256:286a4ff049fb8c680f1017448b9f0b4869510d3c843503453f8713f1c8172a91

Observation 9d43c921-b730-4e1a-a0e7-c6ff566238d8 · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.802169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.802169Z digest=sha256:444cd25def9e8c9f700b15b719836d33ab7b5c181515defa91c5e55335595f49

Observation 4bcbb5c7-fc2b-41e5-a0ee-43312f09a4b7 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.807289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.807289Z digest=sha256:2ca0b2bc42962d788568c3b9290ed22f38e50905f0b0cd1f7c4c0f222156d8a9

Observation 0639deb9-e224-41af-832e-f8b83e7fa876 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.921345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.812043Z digest=sha256:86afc713d6cf134e0b710e6173c7b5408432cecef1fcf5948f783ee0314d192d

Observation 826b1c76-0114-4b09-a29b-1337a1ccb579 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.816998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.816998Z digest=sha256:f743a37231c3e8eebe963024543e3d21464bd12edd0ef91b0af78cd888f6ea0e

Observation da6b087d-9517-436b-b3b5-7d670b149057 · outbound

This paper cites Revisiting multimodal positional encoding in vision-language models, 2026.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Revisiting multimodal positional encoding in vision-language models, 2026

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.904945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.822679Z digest=sha256:50927485122447d08a76de9dd4ccf0212b18d3f1093140a119ee11eb7197154c

Observation dc93f922-558f-492f-95c5-56de9aac65c1 · outbound

This paper cites Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.887836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.827621Z digest=sha256:8a0f8ec0c91150d9f6c2e8a93ed1b5af3bcc187d83b4ff71877934af24f0c509

Observation 6d18d69e-bbb6-403e-a26e-21a3aeeb3ec8 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.869456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.832603Z digest=sha256:bd7ac7258df1740dd168590e16b38ead332fd46fd1dd9ee4657a5ad4f8a7647a

Observation 00114ae5-a656-48af-af38-86faf6e5ad96 · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Odin: A single model for 2d and 3d segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.850869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.837664Z digest=sha256:c0f56b3318d83b0b45b682aa19587bdcfe4e58f136600019a30de968c68d51dc

Observation 1f15f975-3b82-4750-8445-9ef82ff9df72 · outbound

This paper cites Unifying 2d and 3d vision-language un- derstanding, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 2d and 3d vision-language un- derstanding, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.832652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.842550Z digest=sha256:e048983d67406dcdbaa6fe6df0b5b9aa4b833001983e782415cfcd2f14d43af3

Observation 6594ca3a-8222-4c0b-9389-eafbf8bf2361 · outbound

This paper cites MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.814536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.847319Z digest=sha256:63f0082dea3d0f7229c1c98e93bc080778128577936246e364d95085e0bfe7c6

Observation 54add548-3d69-4b75-b0ed-14a7a4c44850 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.797281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.852077Z digest=sha256:18c47f1ffa90d824bdfad880f0525a42ec6fdf1e39ba1137e7c8dd8738982156

Observation 4df9d2ee-444b-4b06-8fcf-80a493fc90c8 · outbound

This paper cites Restr: Convolution-free referring image segmentation using transformers, 2022.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Restr: Convolution-free referring image segmentation using transformers, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.780428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.856930Z digest=sha256:f4952e414b3be4c6aa17b6acf3adba7593d44feac1079a3a691bef9a9ec6f327

Observation 7a97f72d-02bf-433a-9bf6-2406accc3ba8 · outbound

This paper cites Mask-attention-free transformer for 3d in- stance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask-attention-free transformer for 3d in- stance segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.862157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.862157Z digest=sha256:4a3bb88e34b8dee1d2d96702be691ed7d63a4ef222d5f279efb6f7457ca837aa

Observation 1fbc27d2-3930-4d67-b19e-c3dcb341710d · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Lisa: Reasoning segmenta- tion via large language model, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.752413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.867247Z digest=sha256:8d54a57752dfceba0ddddb0ab9f43a5b072fe1fdac136cf5a6c05ae38723cac3

Observation 2bc1063e-20fe-48e6-a533-28a0ba3899a5 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.872639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.872639Z digest=sha256:774e69dac448bcb786eec7d3e9a8eedb0a5bdaefe36375b35dfefb0534d55f98

Observation 990060e3-85a9-4452-b100-14f7a6892a78 · outbound

This paper cites Grounded language-image pre-training.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.723467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.877707Z digest=sha256:dbc15f85124357becf2fc8f3dcce870d2c6720e881358b7f018a3ac90524774c

Observation d6cdb1fe-4044-4bc4-84f5-41148dfb81d4 · outbound

This paper cites 3eed: Ground everything everywhere in 3d.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3eed: Ground everything everywhere in 3d

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.702798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.882564Z digest=sha256:55c54df2c21654e21b67c9b98bd88835a615e29865fea1469a83838acb86f945

Observation fff96ced-d235-4551-8a5b-736904880d32 · outbound

This paper cites Microsoft coco: Common objects in context.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Microsoft coco: Common objects in context

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.682245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.887595Z digest=sha256:4306b31cb5d7297487bc48bdb4d519d020adb72b79171eafc190e5c811e43908

Observation 679bd081-9eab-4124-b646-e150e493fdae · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.662921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.893221Z digest=sha256:cb6ce8cefdb71c3a3be2ef3b19d438b07b31cb7a55fcf17f5c548e48a29ab92f

Observation 86f5f433-2743-4c66-9336-3161c0ddaef9 · outbound

This paper cites View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.898167Z digest=sha256:c048c25def351591458cf34b3f01356a501b3c82bfa421925a541b137b4553ac

Observation e7e3f48c-5888-47b3-99cf-02e104416982 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.627699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.903472Z digest=sha256:b3921470dca21198e83ebbb1f0b77443b6adfeec9cf2ba28960de8b0809936e1

Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.908340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.908340Z digest=sha256:0a8d13a85d88ce8620823aef1557f2442bde7af2419fffcb634cf124215a2592

Observation 9f5db08d-fb08-4d91-97e2-a62a87f4d3cf · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.610276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.913803Z digest=sha256:281b2dc4fdde17aa69c4dc92ece8182d08f0d2b65ad9d18f7dd481e695cf03c2

Observation 89d94f3b-0fd8-4117-84e4-7eb4187182f3 · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Languagerefer: Spatial-language model for 3d visual grounding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.918804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.918804Z digest=sha256:f0e9f1514b82a3197817e936565c094c3dbf6afde3ab34c0671a42ea0ddf2eb3

Observation 804b1176-f87a-4488-82aa-f34154d5fdb3 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.578334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.923771Z digest=sha256:aea2628b566267c6f8a5f1e44d550454f296ad7c57b0bbb379dd2b190db184b0

Observation e8181c31-468e-47a5-9aaa-621d58963501 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.558961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.928579Z digest=sha256:a8ad311a1cedcc0a286a88372436458a8b2a4964e678d8af62c908d174a7d550

Observation e85d7520-83ca-4c62-a919-c53c0a1894e0 · outbound

This paper cites Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.540015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.933509Z digest=sha256:2576b3c19b8506e3ca919061dd33a8491273a124d88b9b5067fe0215c09986a8

Observation fb31858b-2724-4037-a0f0-14fc4a0e0baa · outbound

This paper cites Hashimoto.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hashimoto

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.521967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.938741Z digest=sha256:0e0556a61dbd4049518232f3a8200ca5cae16e424309740a762a7ab08aa135db

Observation e7588881-b2a1-403d-981b-8ab81995c8c4 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Gemini: A family of highly capable multimodal models, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.502906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.944625Z digest=sha256:0023a540e87c4b9c1c1a45f29c7b6f42c70a8f7108cb4e2edb158cbdb077d50b

Observation 00d5c209-b687-46ef-930b-6564ef3fe0c7 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.486307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.949602Z digest=sha256:fcc8ef5bc6f83e60e140c9826a8244fcc809ec26e77bbee783a728b6c26d6ad9

Observation 40bfa89e-34f3-4e62-89db-af3f72fd4069 · outbound

This paper cites MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.954483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.954483Z digest=sha256:fdb6a1f0e651d0fea871dbef085309c36a1ed8dd54c96b3e2115d4eaaac9d71f

Observation 9dfa92bb-95b8-4e8f-a1ee-e45f1e9792f5 · outbound

This paper cites Realworldqa.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Realworldqa

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.469572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.959712Z digest=sha256:ff3b95523a588404c8752e3749a425f23a8147ebf8bb06fc4d3ee5e0c8b3a00a

Observation 6889bd48-7bd7-473d-ad12-834ff21fc8c2 · outbound

This paper cites Qwen2.5 Technical Report.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.964720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.964720Z digest=sha256:f55e6a1636a6039db24d96085c9ea5b94d4ac06ac89bf7f422c622eebc0c2448

Observation 73561748-6326-40c7-8411-bbce61381ead · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Sat: 2d semantics assisted training for 3d visual grounding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.452483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.969841Z digest=sha256:79f499d15aac9c70e739cb3bc043defded58196d73ddb60318b2512f0db72707

Observation 08c79efa-76bc-4065-9f52-83fed13b6322 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.435741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.974617Z digest=sha256:ab40d8cfaa208dfd849f14d33e687b6f8f405f901de75bda805395f53925c28a

Observation 8e754b19-6361-4677-b4d2-9b35de80d58c · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.979475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.979475Z digest=sha256:8691339a0d3f9dc46b1760c9d42fefb043d97f63f8a25f4a53548a7c264abeb9

Observation 72dadb02-5187-48be-8e89-950c586ccd5d · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.408883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.984880Z digest=sha256:a4bc96d5099809d8112e08ac1ea7955cf12ce038bc85fb45eecc0ab66ca9fafa

Observation a4be2bb5-2751-4d74-a4b4-c5f4ac0c264a · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.391572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.989830Z digest=sha256:0a21908662c6f8691f6f323de78ef2790ffec5abc817ef8219d677a90993565c

Observation 7ba9a399-a68b-4a16-b07c-e8d9b19d776e · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Towards learning a generalist model for embod- ied navigation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.371586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:28.994481Z digest=sha256:8f3625961162ac7491c99120c663163b6f16a02705b526d1ff195cc012cc9410

Observation 473c6c45-a62f-48ee-946f-aeb34c1cb832 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.999121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.999121Z digest=sha256:333f2c28e2da5a2dc73f75406022a0b06936207febccc260af5d2171924f4c95

Observation 1d49da31-7603-4865-8b38-9e7ca9fa0cb9 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.353145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:29.004503Z digest=sha256:a77e13a1ca9b90378cfa0904f67c76af71d68f385aa517f9e6948894c02cd169

Observation bc1e7350-31e4-4b27-9f50-6a1eb06190e7 · outbound

This paper cites Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.336725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:29.009579Z digest=sha256:1f7f11036278ff957b34180e4f690c8cfe9949d4f19a7c643ac5c16f73dc2799

Observation aede57e3-b847-485b-8cae-c13d3ab279a2 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.320100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:29.014371Z digest=sha256:a6201796b93df9c77a6b30643f98d5b223e88b9ed4c24abbbd29cbe0d730247b

Observation 87b2bc23-7840-435c-aa55-b96ad2a15573 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:27:29.079215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:29.019048Z digest=sha256:7dff37223ddb7f82b171bbb1534d5acc2bbc2313d1c1e0218d8cfdd7ac71e35f

Observation 5d76e898-d0c7-45e5-a8e8-862ac983013b · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Generalized decoding for pixel, image, and lan- guage

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.303684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:29.024405Z digest=sha256:2e610578e48e8ced8ffab28c9289b01709dedb082843d6efb9b149477450a918

Observation 78a29522-3b1a-456e-9cc1-91dc6ac60d28 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T04:27:29.286999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T04:27:29.029373Z digest=sha256:a2fb09948493f43f68e92350b2d6dea912c89c4a8320be2e17880c997ba635be

Pith citing papers

No inbound Pith citation observations are available.