Pith. sign in

Paper Citation Record · LEDGER

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2412.01292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01292 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:37.029873Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.376214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T10:49:47.240788Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3fe63a2a-e12e-4997-9012-8109844ba71e · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.070144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.758229Z digest=sha256:d42300d786d705f1e6cbe72f27369d56ea32b0db7bae2be824c7193cc8663c64

Observation cce39ea2-b811-48b5-afb0-58f06112dc11 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.053855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.764332Z digest=sha256:361d5b65fd3ef1878c595c2f7a343ac0b0dcd1ff7369163e422979cb08445bce

Observation 89625eb6-3a0a-47af-b56b-0c07ee467f36 · outbound

This paper cites Qwen Technical Report.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.770448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.770448Z digest=sha256:7280352ece8432d8f13768edcc95c0c46a5b749ebecc6a9cee1f0eeb20a3083e

Observation 0f011e68-62db-40b9-9a65-35802a41da18 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.034213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.775993Z digest=sha256:e65214ccac794226a3adb11e37bb0af8239619c318ab5c1639d8de52f9c55ecc

Observation 8314fb60-969c-494a-8404-e6af4f4ace65 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.781754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.781754Z digest=sha256:7a44637ccb40edfedfd7e5e15627f4ea4f1e94ddde5730680bf6e9fd3a69a013

Observation 766b527c-8945-45cc-85b4-3158c4abfe4d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.787399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.787399Z digest=sha256:0e08c9cc04bc787fa6ad711a1356c1bbb67f49a551dbcbfd670bf82cd984f495

Observation 99481ef7-81e6-43aa-8a80-6e604f803d84 · outbound

This paper cites $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.793861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.793861Z digest=sha256:0e72b458472d0a1ef5fefd9a41b7889e6df550cf61285075bd637f9b5e3ccd7b

Observation 0ed7d7af-6a74-46ef-a9c0-42411919e95f · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.007324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.799342Z digest=sha256:60f2d3146d87922d97548a0c4d40119db25051dec3680a3369e4a6acedeb63c8

Observation 55b59e00-9332-4f1b-85b8-fea3868d9847 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap- detr.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences End-to-end 3d dense captioning with vote2cap- detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.988183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.805160Z digest=sha256:0f4489dc74028951e0b3b8ce430ff9107e2f7e10973929ea6376fa0ba1cdd53b

Observation 0084ae39-f534-471a-b858-6b40fbd4d9d1 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.810807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.810807Z digest=sha256:74cf7972e369f0b2703feb7580c5659322f2c71516b20908143de58d49a97ca5

Observation 95bbf4d7-0bfb-4bee-9130-8cbbc9b3daf7 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.816633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.816633Z digest=sha256:c1f3c9053f529376dd1a6497bb83ee6509fa89be7f5c84dcb12bc212b437d510

Observation b7b8680e-eed0-4471-94fc-ec625400368c · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.823134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.823134Z digest=sha256:02fdb896c1997d6ec38a837290e875f4f62cf2e6aecd1bbe142059d93a1255e4

Observation 831ae781-da79-4ceb-86bc-b97f7354b054 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.830027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.830027Z digest=sha256:f436472d70cba8af74d034806c1a83cd859b45ea3c09afa9833401ceb9ccfa3c

Observation a4448104-fa27-41da-bf9f-d48299e08c4d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.836304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.836304Z digest=sha256:d893f54350c236591cb3f60d7643ef6ebcc85f1332654bee9604bc14e1d65b4f

Observation 9598c877-e40e-4dca-a05f-1c8b005f0efc · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3d-llm: Injecting the 3d world into large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.844087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.844087Z digest=sha256:a5e8a52de7b90a29ee0b9bf8950f04b4cce27011715ed7132b149f7c0131d258

Observation 1a5b9862-ebd9-4560-9457-c18768df34dd · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.933127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.852883Z digest=sha256:63509f3405202857a86a9b6773d0db1c1b59b47f28d97760147766c19272e0c2

Observation 27f7e11d-5cba-4ba7-9b8b-aa69f5a921d3 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.858454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.858454Z digest=sha256:649c9af3b429c95b6e313ed287908ddcc61445ea037656eb837b064487f09deb

Observation b8e95ab9-4675-4a26-8643-a5b8ab002919 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences An Embodied Generalist Agent in 3D World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.865168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.865168Z digest=sha256:b727eaea91069e11ccca54ceeee6a0393f502b5bb0c84c132ae05d1fd46b4b80

Observation dc713f0b-4911-4c64-bd7e-847baf796a85 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.871188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.871188Z digest=sha256:ea57dfc755b437bf6e5537f9f395192c24e87565807bb53c86845e0596bc4bf1

Observation 16476d8f-cf5d-460d-8cbd-6c8c3c118703 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.876736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.876736Z digest=sha256:35c7a4155fc26abba55aa59c5eaa0560a30ab591141f68f20ee0039e65303d47

Observation ac1239c6-4c5f-47d1-a80f-5c66a1dc739c · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Context-aware alignment and mutual masking for 3d-language pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.916311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.882736Z digest=sha256:7acb20d763f78044bbfd75a4a9d7aabd3c9e40a67621714987a60335ac8bcc01

Observation 91ac9420-3e8a-4900-b510-e51b59bc384a · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Beyond the nav-graph: Vision-and-language navigation in continuous environments

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.897796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.887832Z digest=sha256:00eef134b437397d887f612c6cd7f1d44dcb82da8c0923eff6fb7d95b419af20

Observation f90ef089-cb8b-405a-a583-f9d18e022e63 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.894102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.894102Z digest=sha256:4a8666653485e172097c75082b3770790dd32a096368e88814bf04dbe76a729f

Observation 28576c4a-b2b0-42e9-9c68-fe7aef66df65 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rouge: A package for automatic evaluation of summaries

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.900360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.900360Z digest=sha256:2b163fade97558588e7f4bf3ce6420864e0105af2e543d1c024095736518cb9e

Observation e4249674-4ce6-4727-a50f-667fa6c79017 · outbound

This paper cites Visual instruction tuning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Visual instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.905491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.905491Z digest=sha256:747c9def0da54a0ba827a66f7b52f398453ef37cf6f6172585a1a0d303eaeb27

Observation 7b1d5d94-7ef6-46d9-b3b1-145ecc5d3287 · outbound

This paper cites Decoupled Weight Decay Regularization.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.911294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.911294Z digest=sha256:452acd36f30a5a17bba6a624f3b621bac65698bafaa273535c459464c6c2f4e6

Observation 03f47964-be3c-47ca-a0ed-e811e786e0ab · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Openscene: 3d scene understanding with open vocabularies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.828298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.916675Z digest=sha256:eb4fd6eedb7a75198d0e96e4d79679fee39f93c0ad94b0ff3393df43bc75db76

Observation aa0ba61a-c73e-45dd-8cca-f1cafe004031 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.921651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.921651Z digest=sha256:7b56f1d5b3e0d57e7ffd7dfe25206926e13f44aa77ae4b44a132345b268cafe9

Observation 2cf8fe14-0928-4d78-bda2-90b5bf9784aa · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.926180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.926180Z digest=sha256:50c11cd53c4735a81cfe87a7042dfa5ba198bdd6cd198d1f1b868073e73c868f

Observation d32fbdfe-714f-49aa-8c0e-604f83201fc7 · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.931272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.931272Z digest=sha256:344403175ab4495cd51c294edf6955e4177f177442a02bf065f7ba08eebeeb8d

Observation 0b67ca78-061a-4d9c-bb91-f968e9452ec5 · outbound

This paper cites Eye movements in iconic visual search.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Eye movements in iconic visual search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.781151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.937120Z digest=sha256:2134cfa41a7d4e4d3b35ece0854f6e1a379bece57280623eff3dd2c4926565f9

Observation 4a93a63b-c80b-4b5e-9944-eeee7bd933a5 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.942835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.942835Z digest=sha256:f6f3c4f077b04518b7972e3ac5193e289ef1986f98f355f6d1d2da104fa394e7

Observation 49ba5686-d19b-4bfd-9813-167bdebcc65a · outbound

This paper cites Indoor scene segmen- tation using a structured light sensor.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Indoor scene segmen- tation using a structured light sensor

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.743530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.947620Z digest=sha256:694fa7d241985dd619325b89d92452cf50084c1e8d453b060b2d55506e14094e

Observation 5eec47bb-6fef-436a-9893-525d27d01bd3 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.953258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.953258Z digest=sha256:d84032dac41f5b868c4a890acc8b1f362441710312bd94105a8144a3d93cbd4b

Observation e3ae5547-944b-4072-b5ab-31fefc9fe72a · outbound

This paper cites Fgprompt: fine-grained goal prompting for image-goal navigation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Fgprompt: fine-grained goal prompting for image-goal navigation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.704457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.958878Z digest=sha256:aaae20b6166d6a9c5e2d92324205273704d2b2c7560337105bd0ece7964847e5

Observation 2abc01b9-451d-48fa-9b09-c1b2d4935959 · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.964112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.964112Z digest=sha256:04292420d9f7170002958321c2d987badb0ba90cffed3731625427de980e5c25

Observation d38280e9-3764-4b09-9095-8f114c10ff3e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.969395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.969395Z digest=sha256:9746d8cdc06ca7f5361837693b6b2691642c9327cee1fa30d2de4e4d55fe0cbc

Observation 0fd286fb-d35a-4bb2-86ca-7f27d0c4447f · outbound

This paper cites A feature-integration theory of attention.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences A feature-integration theory of attention

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.684125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.974290Z digest=sha256:b341892366aebcc9a29f3ea3a25ac7d59d026db3d45b66d7591e83a7b23f5327

Observation 44949439-52ab-4c0d-89a1-1e7d694004e8 · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Cider: Consensus-based image description evalu- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.666338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.979864Z digest=sha256:d91c2e0a659623d94b3e18956902764248e094cc8fa23a4847c2f624488412f8

Observation 6070df36-5bfd-4529-b8ac-c1d978427a70 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rio: 3d object instance re-localization in changing indoor environments

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.649299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:36.985024Z digest=sha256:baa514200ebcbbaa27bcdec7c20f8450e095320e60e86de09131bd3225d6a1ce

Observation 470c7e11-cbf2-4bdb-b133-de04097cc0c4 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.991429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.991429Z digest=sha256:5dc1e73cedcb7b38f0593e69f1d202aedc7ae8c71d40e53c444b07ec3ac46a49

Observation 4a82ddc5-e9f0-4d88-8fdc-c8425cae31b4 · outbound

This paper cites OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.997620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.997620Z digest=sha256:fd089e540e1cc3f8c5a92b43ac01e90d208f86dd6005539b34345d2aeaa52371

Observation bb40708d-d3b5-4e09-aa2b-cb2dba6761f4 · outbound

This paper cites What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.629469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:37.004468Z digest=sha256:bd989603c1f9daf18b84f5d9e3386ed5d3effc3b36c0bb99f34ecda37abd4287

Observation 6e6ffa06-16de-4985-b0fd-5779beaf94e6 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.009769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.009769Z digest=sha256:d7a03372416002f592cb69ed47eabcab08d981e8b9ea3a44ebfb7efdea996105

Observation 8aca1262-4a18-4809-adfd-c15946878935 · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.014890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.014890Z digest=sha256:85a896f052814fcedeeacbccc981703a078a379a426cb04578a647251fb344ac

Observation 6ba662bc-0477-4522-9029-b31ae0f3edb1 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.020259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.020259Z digest=sha256:9beb3a6ad628bd29944f0e9806163435d7e6d16a2b69245fda0b6216f9740d99

Observation 11dde869-e4c9-4999-a23d-7b200624bf83 · outbound

This paper cites Uni3d: A unified baseline for multi-dataset 3d object detection.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Uni3d: A unified baseline for multi-dataset 3d object detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.606717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:35:37.025291Z digest=sha256:c1c6ed747a0b04c54042ea7b26726c391e11276868d3c3acd75873714b5f3be8

Observation 83f99ab1-5a8a-4afa-a952-dbb517c0ee36 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.029873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.029873Z digest=sha256:c335e84db7a4c0200d3f4b1fbb8ee1526394f25ef7616e7b11546b1e53737a10

Pith citing papers

Observation 7e8b58bb-e232-48c0-9cc4-1cc51e9b19ae · inbound

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models cites this paper.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.376214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.376214Z digest=sha256:c6fe0933d03b227ac99e8ae99e83226112c4d0ce0c936b1561f4622680b8046e

Observation 9faf74b7-9151-42b8-a42e-61024e8edca4 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.765433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.765433Z digest=sha256:320c8577eff3829dda5aa2cb26d179675a0198a0f5058a9da1a8dcfd215fdca3

Observation 91e291e6-f4df-43a5-88c0-2cb0ff3dd88f · inbound

A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark cites this paper.

A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:49.176786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:49.176786Z digest=sha256:05415d94fd8b41488234908793ca4a0d26bbf11dc6e7bc6e738ab5e492f13074

Observation 537570de-39ae-44c7-b1c1-77720a093e87 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:48.086502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:48.086502Z digest=sha256:c1d00d00abf7d5abb9cfc979469b5db482397b30c6521e8dc8a3eb9607964a12

Observation 5587a349-bc45-4921-9800-e501f3c2a3f7 · inbound

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond cites this paper.

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:41.554989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:41.554989Z digest=sha256:7585665b54c06648240920d08709574170211812fff3bce0eba4860609fbf4bc

Observation 966be097-444d-475d-91d4-6aabacd26c1e · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:49:47.249706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T10:49:47.135209Z digest=sha256:9945165c7dcbf578b230e75535777ade24df6d463782967cd06c27c04e09c7d0

Observation c02b4ead-fde1-4b93-92b0-417ce4d8e435 · inbound

Nav-R1: Reasoning and Navigation in Embodied Scenes cites this paper.

Nav-R1: Reasoning and Navigation in Embodied Scenes LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T17:31:43.089996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:31:43.089996Z digest=sha256:f0610ab25fda1cd6d3fc365ae0f8a1c81539605b32fa0cabfeadc871ee058042