Pith. sign in

Paper Citation Record · LEDGER

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2412.01292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01292 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:37.029873Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.376214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T10:49:47.240788Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3fe63a2a-e12e-4997-9012-8109844ba71e · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.070144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.758229Z digest=sha256:290d8c4d11cb48f3f951455574aa2e097d6dbd182296264e50b829fafaa7c634

Observation cce39ea2-b811-48b5-afb0-58f06112dc11 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.053855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.764332Z digest=sha256:5d982cdf97601f948932042b5240daf3f4a3ba036d80b98b487327a35ef46110

Observation 89625eb6-3a0a-47af-b56b-0c07ee467f36 · outbound

This paper cites Qwen Technical Report.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.770448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.770448Z digest=sha256:993376e2eb26734b11005710f337fde651e4092e8a54d8cee67450b378ee8ec9

Observation 0f011e68-62db-40b9-9a65-35802a41da18 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.034213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.775993Z digest=sha256:5dbebef2e965d47c33de08b40a37c0befc4ce67f0a6a87e60c428c532a1ec037

Observation 8314fb60-969c-494a-8404-e6af4f4ace65 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.781754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.781754Z digest=sha256:73b75c9ca9d2efeaf1b75f405f92aa5cd0511cba1ac015116d8958b26b34d9ef

Observation 766b527c-8945-45cc-85b4-3158c4abfe4d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.787399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.787399Z digest=sha256:dfbf9f700c7000376480ef459e2306ca3ac933a0b06b3c87fbb5b01754804dbe

Observation 99481ef7-81e6-43aa-8a80-6e604f803d84 · outbound

This paper cites $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.793861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.793861Z digest=sha256:3ee2a6c83ad8e9d90ed96d10652df2670d10e73bf409e6428557cd9980f1c5ea

Observation 0ed7d7af-6a74-46ef-a9c0-42411919e95f · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.007324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.799342Z digest=sha256:8e8c62d1f69fbcd7d440335db62e19600a84e03627864465d8f5e2c36c68bc52

Observation 55b59e00-9332-4f1b-85b8-fea3868d9847 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap- detr.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences End-to-end 3d dense captioning with vote2cap- detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.988183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.805160Z digest=sha256:3f5ed02d5b5de11bbeccdc15e02e1e75585a066408a35061323f5a2d398706a0

Observation 0084ae39-f534-471a-b858-6b40fbd4d9d1 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.810807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.810807Z digest=sha256:6277569b76d27b03aeabca183a11e561d4350b2edfa9df1de784e09c601b5197

Observation 95bbf4d7-0bfb-4bee-9130-8cbbc9b3daf7 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.816633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.816633Z digest=sha256:bd39d1b33c50deaf486bf856f66d28bc42c8975e589d4999da499673d2234f3e

Observation b7b8680e-eed0-4471-94fc-ec625400368c · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.823134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.823134Z digest=sha256:9e7c0d7bda0bea31e28ceaca1024c85be122cbf605d6a4305275409adaeea9d6

Observation 831ae781-da79-4ceb-86bc-b97f7354b054 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.830027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.830027Z digest=sha256:7aa8481799444862a8b702bdac619ac8e848a09516967c70300613c9332b7434

Observation a4448104-fa27-41da-bf9f-d48299e08c4d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.836304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.836304Z digest=sha256:62942fe66296e87af9330bbbecd4fe0dfff98d9eaa1eb7ff55f9d8da9ae1644f

Observation 9598c877-e40e-4dca-a05f-1c8b005f0efc · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3d-llm: Injecting the 3d world into large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.844087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.844087Z digest=sha256:d52a655dae16dae8bc5f1dfef701664406203d90c01cc57a6d1492d6821b119f

Observation 1a5b9862-ebd9-4560-9457-c18768df34dd · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.933127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.852883Z digest=sha256:bd707cdcf619cf71a2cf3fbb6717a0d5a87b7a2e772908fc09faf901f05efd82

Observation 27f7e11d-5cba-4ba7-9b8b-aa69f5a921d3 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.858454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.858454Z digest=sha256:0d96b6f124722d1d61e384d95ee81303ddb2a02c4e2aba1b18ccc7928d1e6283

Observation b8e95ab9-4675-4a26-8643-a5b8ab002919 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences An Embodied Generalist Agent in 3D World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.865168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.865168Z digest=sha256:ac0acd956049f914b00ab2e1b1faa68d3219dbe2ed650547e25ad99acc9eef00

Observation dc713f0b-4911-4c64-bd7e-847baf796a85 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.871188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.871188Z digest=sha256:cd8d30f3813e1597239d0be7348b080b30706d7402fd077f1d581692c913fd9d

Observation 16476d8f-cf5d-460d-8cbd-6c8c3c118703 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.876736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.876736Z digest=sha256:e7ed4f6d5b9d4c3c9fe7da82cb1196bf1d016cdb0bdaa90e8b5628ab26fea19d

Observation ac1239c6-4c5f-47d1-a80f-5c66a1dc739c · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Context-aware alignment and mutual masking for 3d-language pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.916311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.882736Z digest=sha256:7b88acdcf41140adf79e93ee1ff5cb69cbee115f3ff434979adac7a0deb6688a

Observation 91ac9420-3e8a-4900-b510-e51b59bc384a · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Beyond the nav-graph: Vision-and-language navigation in continuous environments

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.897796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.887832Z digest=sha256:82cb5a7f541a4014b65ac3b9c88ca89641939916e46c27d7c7b0e163fcf43d76

Observation f90ef089-cb8b-405a-a583-f9d18e022e63 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.894102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.894102Z digest=sha256:db1ccc5d7bfc95bfcc0b2ec26e9066dff69743b02ee70daacb3a63d5fb9aed25

Observation 28576c4a-b2b0-42e9-9c68-fe7aef66df65 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rouge: A package for automatic evaluation of summaries

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.900360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.900360Z digest=sha256:3ffd410cb40a0ba81eb0c492272c9037d8a42fe8c19a653f7d11e8260bc7101f

Observation e4249674-4ce6-4727-a50f-667fa6c79017 · outbound

This paper cites Visual instruction tuning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Visual instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.905491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.905491Z digest=sha256:cbd05aeaf194f7515a9912be97f5275a1934a4b38346592c62393c98cb18af40

Observation 7b1d5d94-7ef6-46d9-b3b1-145ecc5d3287 · outbound

This paper cites Decoupled Weight Decay Regularization.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.911294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.911294Z digest=sha256:e4ea8e8950b5e4d30508be8e2fc3b141066792ebce4d1da9d4fba7edc61aaa2d

Observation 03f47964-be3c-47ca-a0ed-e811e786e0ab · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Openscene: 3d scene understanding with open vocabularies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.828298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.916675Z digest=sha256:051e1a54573c0377ad123f31bbeffc482ff15d7541086240b3a31a30453dfbc2

Observation aa0ba61a-c73e-45dd-8cca-f1cafe004031 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.921651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.921651Z digest=sha256:764d7515c8ef0ff543582f6decc0c227e957f7d3a9aa6a2ab5b5c64009bd61a7

Observation 2cf8fe14-0928-4d78-bda2-90b5bf9784aa · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.926180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.926180Z digest=sha256:611e4f54b4c6747c69399c8cef155f382b0e1d5933b53f25c3bec423f252926f

Observation d32fbdfe-714f-49aa-8c0e-604f83201fc7 · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.931272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.931272Z digest=sha256:2b10c5933d944630b3ae374531e986dfb7dc6d5096bc243d47c56805353731b8

Observation 0b67ca78-061a-4d9c-bb91-f968e9452ec5 · outbound

This paper cites Eye movements in iconic visual search.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Eye movements in iconic visual search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.781151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.937120Z digest=sha256:9878828f195af7adf945d24cc07e4aa1aa45ef461c4c7975f1b18333e3656a28

Observation 4a93a63b-c80b-4b5e-9944-eeee7bd933a5 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.942835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.942835Z digest=sha256:2769708f9659c808c838e52897e5868bb25d8a113052c0decf41bfa08be92bad

Observation 49ba5686-d19b-4bfd-9813-167bdebcc65a · outbound

This paper cites Indoor scene segmen- tation using a structured light sensor.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Indoor scene segmen- tation using a structured light sensor

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.743530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.947620Z digest=sha256:d10aa39d5cea634fa170e664ef51d6067a7f42c362f88f2d4a55d9db246899c1

Observation 5eec47bb-6fef-436a-9893-525d27d01bd3 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.953258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.953258Z digest=sha256:7c449a957dafdee1155c141e6369c88e3c219608ff60b10f9e26439fde384e27

Observation e3ae5547-944b-4072-b5ab-31fefc9fe72a · outbound

This paper cites Fgprompt: fine-grained goal prompting for image-goal navigation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Fgprompt: fine-grained goal prompting for image-goal navigation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.704457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.958878Z digest=sha256:68b2f8d264fda8702ab56f22057d27cb7e77489a33c74d975293c1cdebe21c09

Observation 2abc01b9-451d-48fa-9b09-c1b2d4935959 · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.964112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.964112Z digest=sha256:2ddac09034ac021e1147a52639cb7c46a01674458361588adddc8ba8fe622704

Observation d38280e9-3764-4b09-9095-8f114c10ff3e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.969395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.969395Z digest=sha256:c09120ea201f62794d53efdd356da1d570f018200e1f49d311b4d32d7e60a3ed

Observation 0fd286fb-d35a-4bb2-86ca-7f27d0c4447f · outbound

This paper cites A feature-integration theory of attention.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences A feature-integration theory of attention

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.684125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.974290Z digest=sha256:e83421b3044f1be4121d9439d73c7b5fc4579ec54eb8fddfc0d11e6c8e673a18

Observation 44949439-52ab-4c0d-89a1-1e7d694004e8 · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Cider: Consensus-based image description evalu- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.666338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.979864Z digest=sha256:873ab1cc51a741a935dd736f6f1a940aa916b4783d59712ba16f094e10a7e996

Observation 6070df36-5bfd-4529-b8ac-c1d978427a70 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rio: 3d object instance re-localization in changing indoor environments

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.649299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:36.985024Z digest=sha256:4bcb6bdb29f9a777bec6556ed0d0a37bb75fb3da5907396a9c2c42a3461c0685

Observation 470c7e11-cbf2-4bdb-b133-de04097cc0c4 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.991429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.991429Z digest=sha256:2cf0d744a88bc5bd2d1536bea20d0a9b821a20b8048d9b70116a8ec30672b8c5

Observation 4a82ddc5-e9f0-4d88-8fdc-c8425cae31b4 · outbound

This paper cites OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.997620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.997620Z digest=sha256:47376e36b9963942cb118d67ada4d641875dec9bfb1b919c56046a38f4c30151

Observation bb40708d-d3b5-4e09-aa2b-cb2dba6761f4 · outbound

This paper cites What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.629469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:37.004468Z digest=sha256:0fd4004a1e3910996433ea3d09f36e70bc87fe14d30616c32d490824e151e6ea

Observation 6e6ffa06-16de-4985-b0fd-5779beaf94e6 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.009769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.009769Z digest=sha256:59f1eb872fb9aacae38e6928055af6dfbb48a8ebb7b3c6510d3c1714a914c523

Observation 8aca1262-4a18-4809-adfd-c15946878935 · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.014890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.014890Z digest=sha256:a7f67f0546912549fbe13f6abe883a376f3d5ebcb3f3a6fcd773fcd2e117da92

Observation 6ba662bc-0477-4522-9029-b31ae0f3edb1 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.020259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.020259Z digest=sha256:d0e35cd45f212c603b02af0caf549d493c00d3e73ee1f3b8b9e0a10a675fd89a

Observation 11dde869-e4c9-4999-a23d-7b200624bf83 · outbound

This paper cites Uni3d: A unified baseline for multi-dataset 3d object detection.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Uni3d: A unified baseline for multi-dataset 3d object detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.606717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:35:37.025291Z digest=sha256:a9e0ac6bd11daf5dadc4536b2a6bbf4074b22d3583b6e3ce747545a2cdff6eab

Observation 83f99ab1-5a8a-4afa-a952-dbb517c0ee36 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.029873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.029873Z digest=sha256:ce734f0cb8161a37f62d044430f0300e4ac96e48bef97c4360bf42e0fbfb6fc7

Pith citing papers

Observation 7e8b58bb-e232-48c0-9cc4-1cc51e9b19ae · inbound

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models cites this paper.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.376214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.376214Z digest=sha256:a745fd76b3f142db0ef43ea7a4798e7a02b85e1eb927b8a342f51e7376157a18

Observation 9faf74b7-9151-42b8-a42e-61024e8edca4 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.765433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.765433Z digest=sha256:c603701c67ffabd3f7c99f429b9ce605b9714b24f6f408a3feef44f28e664974

Observation 91e291e6-f4df-43a5-88c0-2cb0ff3dd88f · inbound

A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark cites this paper.

A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:49.176786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:49.176786Z digest=sha256:cec001430461bd6b943cec1ba5c025dce6591c2309894ea21930a22e38275d47

Observation 537570de-39ae-44c7-b1c1-77720a093e87 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:48.086502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:48.086502Z digest=sha256:558331950de01353083c2cb32b1157c7ea44280f2c1f03ea744451ebce33499a

Observation 5587a349-bc45-4921-9800-e501f3c2a3f7 · inbound

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond cites this paper.

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:41.554989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:41.554989Z digest=sha256:5c8c577628417b53b0afcd9ded9830b7498b74fa426ea7e53af81d93aa8fddfb

Observation 966be097-444d-475d-91d4-6aabacd26c1e · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:49:47.249706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T10:49:47.135209Z digest=sha256:126136070c8cef3f9aaf75c499315a6f5176df213c7204d199e6daff7ae38c36

Observation c02b4ead-fde1-4b93-92b0-417ce4d8e435 · inbound

Nav-R1: Reasoning and Navigation in Embodied Scenes cites this paper.

Nav-R1: Reasoning and Navigation in Embodied Scenes LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T17:31:43.089996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:31:43.089996Z digest=sha256:a56aca8a5f6aadea2168784c6deefb7bc0480a1bd5a65ad2b05c9dec5910fa9e