Pith. sign in

Paper Citation Record · LEDGER

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

As of 4 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2603.27507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.27507 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T21:31:00.764078Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T07:16:36.444118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact35
  • verified fuzzy41
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16aa3dc8-112a-4a2d-a7fa-25e2553c55ab · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.413375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:e6caa0fe7c328058c703a42d15e0c51fcfafd3b6091d30a2e3eb19c52e3e89e6

Observation 6aadeb4d-7842-407c-950b-957abc2fe343 · outbound

This paper cites GPT-4 Technical Report.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.412498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:3214af85500a5c04709beecf9f90966c6fccd08e4619db55b5b06c7127b427a2

Observation 0fee2bf1-61eb-4e95-93e1-847603cf5e41 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.474291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:b00ae689124dbb704cbb994b375e68aaabee8130c0a15a0324fa0cc77fcee556

Observation 7a18ae23-810a-4cc8-9b7f-fd473b22b536 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM PaLM: Scaling Language Modeling with Pathways

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.440255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f3c6be7815e9b1448ff4a818082bed978f17c17da90ac44f7609c3d75413488d

Observation 603e4b0d-d928-43ff-96ac-d27b9c253815 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.463336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:56e734a94fdfc199625b99b7fc587f5200b09f6dab72ed2585d62b42ba2d72d9

Observation 6ae8429c-9bda-42ff-aaee-d4356946c85d · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:48.053680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:b202c175bf11a911378a54daabe7965c399b7f787f8be63294e6a39aab54951d

Observation 11d5b250-1ed7-4655-9210-aecb585e6306 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM VideoChat: Chat-Centric Video Understanding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.466784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:dd769d907864212b0c2e644cb54a84bd50c5a72cf75fbe211f50a178e62da3ea

Observation 84620b2b-d6f2-4a23-99b7-08c196405e01 · outbound

This paper cites Visual Instruction Tuning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Visual Instruction Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.437554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f12dbbe7171e4b0fea7ad139813e9538873e788c91d54286de53847b617abf8d

Observation 56aa4678-3825-4853-b98c-4ce24168722d · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:32:59.446971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:b155392061a59a603a73e738dd18c00d513ef6c233fdcaec18a9b65740d029b9

Observation 1062295b-d3f2-4bc0-ab56-8a9b9227c363 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.360234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:c50b915462c1b3ef8ce49e3592658e244aeee05f5a03d744e9ab6884a0ec76b4

Observation f517c01a-f9c3-44e2-be25-65b4ac68f3db · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Improved Baselines with Visual Instruction Tuning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.423703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:e16c6b65650c1b2a0e749242acc2aca4248b2fa6d8da4466c971f0d68186ab10

Observation 708a6f30-9b32-4808-ba10-c61796725e5d · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM LISA: Reasoning Segmentation via Large Language Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.377974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:112848543fd74b490192973a2d18256373590ea9b3619e2374703cd060be46c8

Observation e78b1977-d0e1-4623-80a2-a7a49a5740f0 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:29:05.422888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:a970f72cbbb6ba5c929e6ff31a42ad63a3dfb84375699e8ca6f073a0deaeb05f

Observation 6f3e2468-dfa2-49f4-ab46-e0270013755c · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.388670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:584f43b0df43586281cfe457d066aaa81a85f77ff755bb6ffdacbff8915ed9fb

Observation 58a9b2d0-0451-46aa-958f-05af1ee84c92 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Multi3drefer: Grounding text description to multiple 3d objects

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.384644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:8d0c184b1d8409184501d8870b17b6402d440c2e6e0095e99254324d61856daa

Observation 5dbac903-47f7-42b9-9393-5331806da6ed · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real- world scenes.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Referit3d: Neural listeners for fine-grained 3d object identification in real- world scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.352199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:adf8f897a182eed96b21791d781d82c856f98c9821d3f9052d64cef5f1bb1b9b

Observation b807cb88-8b5c-4fd8-92fd-b6ef966219a4 · outbound

This paper cites Scan2cap: Context- aware dense captioning in rgb-d scans.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Scan2cap: Context- aware dense captioning in rgb-d scans

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.430111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:b45e8e1c3008ec520df45eb5536e26df1649871e3cc403ee27a8a3270dab1e45

Observation ec5de08d-4130-4acc-bbfb-c499d5a8cb20 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Scanqa: 3d question answering for spatial scene understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.325082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:0b8b6fc3a4df4e5b97f64ffd49c6000b03b08ea837ea110747a973797de0eab1

Observation 6692311b-c692-414b-9e12-bbd1b51aa19b · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM SQA3D: Situated Question Answering in 3D Scenes

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.431171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f4ea5d56f9ed4ce5db0e5a1afc42b2e305705b04df63db9ed29c9f476dfb2025

Observation ff6de004-3de8-42cc-926a-e86596b02e19 · outbound

This paper cites LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.427774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:9e0272145cbfb691557893f48f7ed2b7d08563361ecb4fe190c145eb0dcfee6f

Observation d7218bc5-1ec5-48ce-8ea3-691b7b68ae46 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM An Embodied Generalist Agent in 3D World

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:22:18.773673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:a1f39e680eecbc31ecf6e0e557afe63bab3506c0050c23052a789be15995fc7d

Observation 109c2f19-c992-43a6-a4fc-f20609c1acac · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.420455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:43f9687c181f2eb076c8bc8b1a3a8c14768f4417eefd48d0ebfa0eb1d1f061f0

Observation f8549aaf-f536-48b2-85a0-2b1c00ebc3ec · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM 3D-LLM: Injecting the 3D World into Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.373925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f8d1752ee6f684436a5665416e3dcc34415dc23c0af13f7ab1d4b483b218b562

Observation ca0ba364-1786-4756-aea1-a7c1e37193e8 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.425578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:bd451741c7699ad638cffbcdc74a7120f9126c47ca99c494018dd5d1b22cd5cf

Observation ca7453c6-b447-4d92-9b99-6c2e28ae05ae · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.373125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:c5bde9f50e05d528664a3e1c173a88831dbd61ce41142674d70add644ba3d5ee

Observation 20484126-c8cd-431b-b1e4-db100ee229a6 · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Uni3D: Exploring Unified 3D Representation at Scale

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.406470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:fd3bb82c5bc0fa309848d02a9f21576e2e56a5adce9699b20084f898c97e6398

Observation 67bc82bf-57ec-431f-ae29-86e6cfb0e795 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Chain-of-thought prompting elicits reasoning in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.393027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:0b9f797f0fce10d9f5e7005ab188afe8f439955178d459e76779fd8112cc5e8f

Observation 86a452fc-3963-4cd9-b041-be127654e63b · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:26.109012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:79261b6e2d6ad8a487cc49ea64bec88ca00f9cc44fbef44448d7d8eedb4eae7d

Observation 845b5c78-7297-445e-8109-9ba35228f073 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.478593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f7d2df41db7660ae1bcec10451f9ca52d738075cba01f0c1607a8d586b84a0dc

Observation 5982d1e5-e30b-468e-9eca-f9a903a4799a · outbound

This paper cites Openai o1.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Openai o1

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.473606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:828efbc3650522ab64f056c85f6ef90663a2a45a17e6498c2b636c613b496042

Observation 8e4c536d-6e8f-448e-b4af-ee974f3581d2 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.348045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:2b79e488235ce29c87ad18d0a83ba5b49d992f583476fe69236bb0e17a66449f

Observation c18351e8-3af1-4581-9f24-e252d6df0b67 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.289031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:b81abb4d77ad2910191da649af083c7849860adc1249020721ad9a1f83e75aa1

Observation f4530d82-d8c6-42ca-bcd8-5db962f1e896 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.343726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:673d0e96c9c5110341daae0624dec0d19c1a38b0f7d579724d1ca29dbe194717

Observation c22aba20-c486-4ef2-8ed5-1514b3794cfe · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Text-guided graph neural networks for referring 3d instance segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.448513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:1ef3f1e957147575e8c24dd7ac0699e260946cd14c53b523128f9d807c04dd83

Observation 644acb1b-7296-4740-a89a-e1362384f5b7 · outbound

This paper cites Multi-view transformer for 3d visual grounding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Multi-view transformer for 3d visual grounding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.434036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:da658aacba706f0daff9edf054b423ca06357568c415903958ea68e479f1f12c

Observation 7b67c4af-04b1-4dbd-a7cb-395d1d25dafe · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Language conditioned spatial relation reasoning for 3d object grounding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.421533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:48ed3a643e8f02042742fc8aecb289d3198daf56b95f4d9e012ba43e02fa2056

Observation c1583646-69a3-439e-bed7-f2e5cef40f0e · outbound

This paper cites 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.457187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:04aa0b48e87c4fd16c2fd532aa3bd3e31a20173b39d7bef781edb5a6fc38f60f

Observation d649e7b5-2d42-4444-a584-ca9082302440 · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.452520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:901df5a4298ce762ce5dac5e56a38988fe5a6c4d897a0eb4f1f281c5b02a99c4

Observation 6613a71e-cf1c-4dca-85fb-d45ccd811202 · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly super- vised 3d visual grounding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Distilling coarse-to-fine semantic matching knowledge for weakly super- vised 3d visual grounding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.481599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:9d8e06e132f10b3725025c431ab5bd9b203a85c30b5bf11603bba2dcf1ac345f

Observation 182fbc81-06c2-4ee0-8c8c-dd153e32baff · outbound

This paper cites Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.400054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:82e3b9396c5285f08569e7d11ea727615a878d90ab3d610eb1a8aa908b6f912a

Observation 5fe9ebf6-8b2a-4974-8950-1f1f536636f7 · outbound

This paper cites X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.329454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:68559111a50e926abc3615039356ba59b64143cea65fcfaaee68caad3ae6deab

Observation 7776047d-b9b6-419f-b3f8-426a38769b15 · outbound

This paper cites More: Multi-order relation mining for dense captioning in 3d scenes.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM More: Multi-order relation mining for dense captioning in 3d scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.456600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f6cb7e217cbc15171b48f99395f916e88bd848ec00792aa04f4a37043bf2c62c

Observation 239814d5-b26a-45c2-a53a-496472ff87ab · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM End-to-end 3d dense captioning with vote2cap-detr

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.401559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:1741842a9471c1342c430f871eea1103e49f46652d1e7d16d7799b8dcffc1099

Observation 35fb9b4b-74d8-416f-8b6a-aa077934eb19 · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.364680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:79c54284b710286f4d7a490d8b2c76d36d5dba7d82228b85f8368261a609dc65

Observation 389ba71c-af69-4f43-a89e-bf5b28d5862c · outbound

This paper cites Clip-guided vision-language pre-training for question answering in 3d scenes.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Clip-guided vision-language pre-training for question answering in 3d scenes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.460679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:c32d9675a60249f3de62520ac427ca1af4cd8c2664e699de5c646380127e0650

Observation 3b1af93f-4963-41a6-a8c7-967be0e39941 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.320017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:3ee249e75971d980770911eec306e4a7f7c3a654e08bd5d9734ea7119d6f480c

Observation 44aad0be-7ca0-4686-a864-00d99a5befe1 · outbound

This paper cites D3net: A unified speaker-listener architecture for 3d dense captioning and visual grounding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM D3net: A unified speaker-listener architecture for 3d dense captioning and visual grounding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.360377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:c7ec432a8b5fde1c7f55d82726850df3f583afa360a65aed9ebec3db4cf1c9d4

Observation f70a4011-4e10-4031-946e-dfcc52af0f34 · outbound

This paper cites 3d-vista: Pre- trained transformer for 3d vision and text alignment.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM 3d-vista: Pre- trained transformer for 3d vision and text alignment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.469417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:32f29eeaeafa30d8dc85b616427601f058f20a612da1afdc2629fd779699ca0f

Observation f18c41f6-ea42-4ad2-a03f-2e6a1aaa39fd · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Context-aware alignment and mutual masking for 3d-language pre-training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.405458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:a6e07761e40f46e5f12573e14488d52104f7afe70e0c91991d8fd88ec9c3b27f

Observation 27354c12-1c47-4ecf-87e5-bcbd9b6ad2e1 · outbound

This paper cites A survey on fine-grained multimodal large language models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM A survey on fine-grained multimodal large language models

Reference 50

Resolution
verified exact
doi, observed 2026-05-14T21:32:58.940800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:7c8c0bc4f5d0f21b9c4c9cf47fc88801f7a5ebf425b2ea81d722b96b9246134d

Observation ba682909-3eb1-441a-8ff6-940dd1e335e9 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM ImageBind-LLM: Multi-modality Instruction Tuning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.382205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:1b7cd7cb8a6b4edd33a46cfd2ea202459c6c4c478fce202a7ab7109d687e1bc1

Observation 2a95df81-557a-4628-91a3-784c68deacc3 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.409437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:a81027fa8b6bc4b92c3175f4f9282f6d3cf6d992e125cba8759e970f24800b38

Observation fb7f232f-1d86-428d-a350-aee2585f70f7 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:32:59.393660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:737b861256d47f7555e4bca9543bb136e3ecf34df2a4c62d55ffc01d7ed355de

Observation e5f1089f-3bcc-48ec-a7cd-0fb2861ab47e · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Grounded 3D-LLM with Referent Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.390113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:5e644a0a8a90fc7529dc6c732e9e5ff06ad8eabf169a8bc0ca3e8caadc81e113

Observation 56f6f3e7-a376-44f7-8870-4ee557420637 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.454076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:d9228c9ded4fc4a6bab251629edef2ed224d94b5a78c0ceda559a84801847860

Observation 5d1b91f0-8416-40d0-a92f-658e7cb217bb · outbound

This paper cites Towards reliable multimodal intelligence via uncertainty-aware inference.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Towards reliable multimodal intelligence via uncertainty-aware inference

Reference 56

Resolution
verified exact
doi, observed 2026-05-14T21:32:58.937969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:4779467e05849232de32484d01f9735dd32f51465e9a2c8bcec555a4c57a5abe

Observation 5b59aea3-f7a6-4afc-a370-7ffbe29cc824 · outbound

This paper cites Formation-guided multimodal representation learning for group action quality assessment.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Formation-guided multimodal representation learning for group action quality assessment

Reference 57

Resolution
verified exact
doi, observed 2026-05-14T21:32:58.934628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:98f3fbaa837e828930c061b33239ea7b5e9d4b352d8af240e132c44fefe5603a

Observation 677ceca4-5977-436a-a1b4-4f77eabf31dd · outbound

This paper cites Point-bert: Pre-training 3d point cloud transformers with masked point modeling.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Point-bert: Pre-training 3d point cloud transformers with masked point modeling

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.409289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:8ed4c52fd47b40125f735b6c93aa31426ccd58dc02e4e0aa84aa00faba86ba72

Observation bff9e9b1-9ea3-4aed-a536-4f4be48a5d14 · outbound

This paper cites Masked autoen- coders for point cloud self-supervised learning.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Masked autoen- coders for point cloud self-supervised learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.356288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:4c55d64d5eb822b12dbca19c85c07271aac71b031b3a4fb183bda00f42c681f3

Observation 9763cc4d-e3d2-4c5b-a665-d9fc7edbcf29 · outbound

This paper cites Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.477561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:5bcf8f519e6b5922e14eda6a80b9911b7906924de30c5140e388e818e5732777

Observation 0d8e6480-efec-4036-a62e-b3ccd57c8991 · outbound

This paper cites Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.438269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:be7a67846ae98a28a7b229b24b390ace7416e79546899e2ca7df8334fbf400cd

Observation c9c1b719-29d0-4b9b-bbf8-7262063ba30e · outbound

This paper cites OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.369561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:880b21f33ad2e243f71cde1ad4720aed67d8e001e561d15ec1d03be6c2b0d46c

Observation 31a4d617-3c6a-4192-8ec7-b5355f6e3aee · outbound

This paper cites Learning transferable visual models from natural language supervision.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Learning transferable visual models from natural language supervision

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.417738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:a1648bb56bf318c54b1b4c2c11be26d3fec119e2f85a83cc3b3d00515a9fb5fe

Observation 9c1a45b5-07d4-4ece-9dc7-dddcea3e4cfc · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.397407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:05667fd6ba11b32f8c2c2475e98722874c2cd16feca6c8ea199f2e734e469fe2

Observation fef9480e-27ea-4279-bd78-437b308e4736 · outbound

This paper cites Grounded language-image pre-training.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Grounded language-image pre-training

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.381011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:fda781d68260dce9a0d9e8ff632556c27252fecd594fc0c244f06c7e0a568d0a

Observation 0b1777ac-44ed-46da-87e6-15c09d6787e5 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM DINOv2: Learning Robust Visual Features without Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.416888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:217e64bc31bb71a3dd5e6c0bce9c0fb4e683134e2363af07c78b795b7b224d11

Observation f6eebdc6-848c-4e0d-baa4-be7f4bfadea9 · outbound

This paper cites Cider: Consensus-based image description evaluation.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Cider: Consensus-based image description evaluation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.369123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:099341021a9aa2bdaab697b819c4f452c3429fdb1351931cf85cb5d93c137da8

Observation 5b27c794-257f-498e-8a0d-9968aceba22d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Bleu: a method for automatic evaluation of machine translation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.339405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:154c45836a429d1a9c40e7db83a1f216e7edeee6e917cb8d6d2ac984f58ce192

Observation 0b99d65a-7151-424f-b512-32f7ca26b936 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM OPT: Open Pre-trained Transformer Language Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.470401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:cd6210f18262b3e3aa878c59110922f203bc98f0b6ac43a8779ff3751a1e08c4

Observation 601c0069-d347-4286-8446-25c651a7387e · outbound

This paper cites The Llama 3 Herd of Models.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM The Llama 3 Herd of Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.364986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:b296eed17f984ee9a4fba60a5a9efd077e873cb3fd68b70a117f1a3caa894f65

Observation cba2dd93-57a0-4401-bb22-4ab1420a4258 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Rio: 3d object instance re-localization in changing indoor environments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.442627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:8171f79e02c143963513c5803c7e2724127f8ae27226f8db109194ce7391f0c5

Observation ebde41d2-be6f-4bec-9919-75cee9bb9b16 · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.333876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:ec1ee73c828d4063d9b8b20de26c044fed4a24574128dd231b18640701338870

Observation 50275376-df9f-4088-903e-8e9ed5ccf014 · outbound

This paper cites Lidar-llm: Exploring the potential of large language models for 3d lidar understanding.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Lidar-llm: Exploring the potential of large language models for 3d lidar understanding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.464639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:548522ff0ca43bc74e90a18b5d364389873ea696f988688d38b830a356b6d404

Observation 3579223c-cd8d-4521-be4a-69f92a597ae6 · outbound

This paper cites OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.482555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:0e4ebd4ee1c0a97f989308bedc11c36bccfb799d1786f8de851202a52098f0db

Observation fb05a9f4-3725-4a25-9033-f69062f51098 · outbound

This paper cites Lscenellm: Enhancing large 3d scene understanding using adaptive visual preferences.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Lscenellm: Enhancing large 3d scene understanding using adaptive visual preferences

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.313425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:85329c37cdc21c2cbd79859d101353fed4f678e734127848ef9318512322adf0

Observation a7b28b3f-c286-42bf-baa0-0a9d7ef22b0a · outbound

This paper cites Opendrivevla: Towards end-to-end au- tonomous driving with large vision language action model.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Opendrivevla: Towards end-to-end au- tonomous driving with large vision language action model

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.386757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:1918f42d48e2311ab4134c9959cd066738c3d91eaaddf2aa5ef60f84d2b127d0

Observation 8f745aad-cdc2-4d94-bca7-845173a14791 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:41:05.025329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:23694670ced4d6d7f73220f1b0a1f79bb2f269fadf0d8029af2e8352ed0f983c

Observation 8c9b2c79-9979-482d-99bf-a2bbfe1ccac1 · outbound

This paper cites Tracking anything with decoupled video segmentation.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Tracking anything with decoupled video segmentation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:38:18.377281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:1cde0953156d2c808bfee665ae38d50da92ae09a29798b763497089273785d15

Pith citing papers

Observation 0cf9400b-0a9b-4d0e-b5c3-a7af00965c2e · inbound

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA cites this paper.

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T07:16:36.444118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:16:36.444118Z digest=sha256:aa8e4e8d0bb509bd9d9c7d902102c9c1b5632221613666cd9fdbd31aac8dbc65