Pith. sign in

Paper Citation Record · LEDGER

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2507.09334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09334 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:03.566081Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f48ccb24-04c1-4af9-ab85-115adccb60b5 · outbound

This paper cites GPT-4 Technical Report.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.124609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.124609Z digest=sha256:690e7d8eb5b3d2fc9cbba0d177b39bdeae1d6a05de0d2d4dd34c24608fd9ff7e

Observation c97d3428-4d4b-4ca9-bfc9-53a764938a49 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.157375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.157375Z digest=sha256:d21268edb9b39d0ba4c9e0d232dc36c71689629ab31cc4e15050a76b660626f5

Observation 451f4ec4-0adb-4693-b0e8-cf0c196b7a00 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.223606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.223606Z digest=sha256:7c77dcf8dd4b455f822ab4bccc243eebabddf664a808807e386d4675ea505efb

Observation 5ed70302-e3d8-48dd-b8b9-60d04bc948aa · outbound

This paper cites Token Merging: Your ViT But Faster.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.302019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.302019Z digest=sha256:ab1f7ec5fa6fec70cc9fda39325f733a4f8bcb1f9d043b2f9dd63c8b4e391983

Observation 3b25b53c-588f-4af3-9966-29ef4776d0ad · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.364155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.364155Z digest=sha256:6b129756099c261d29108287d79c201a89b785811042fa154b825ff244df1ef7

Observation e4cbb8fe-1c56-4b2f-8079-d6820a086803 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.422074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.422074Z digest=sha256:25541110fb6bd8ee514566dfce5330ed29b8c5e5900c89dc90e6b3a22f1b1821

Observation 772f2383-4452-4096-8cc5-fa90cab7567e · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.504730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.504730Z digest=sha256:6d4a3e08cf82255da7b6a3869cf44851ee514f68fb202ea8ee8de7c98da4f3d7

Observation 0e65442a-8ef8-44e2-86c1-fa56cf9150ee · outbound

This paper cites Efficient Large Multi-modal Models via Visual Context Compression.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.561451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.561451Z digest=sha256:dcbd2938278cc8d92baa648919416bdbf14fe101d812707ef5585e25c312faa4

Observation b05cc5d5-aee2-4d4d-885f-6acb86921161 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.638753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.638753Z digest=sha256:112846a896112ceb0919d1d21be3a4eccd7c6ebb7410d02593140f63b501af2c

Observation 015c0e41-aed1-4605-bda8-030c0e18aaa1 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.697380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.697380Z digest=sha256:699120d3c83a5c0c31af66a36cd0127404f10e95a2692b67edca8fb1a1548d1b

Observation 07554a33-1893-4957-9184-ecb0773e65cf · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.759815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.759815Z digest=sha256:04338a5b059083294091b03a314acb61820dbeb6a67b56fdd421d8e81a8ed0f4

Observation 3c7f5ec5-5aab-427a-bf75-90febf63009d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.817271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.817271Z digest=sha256:ded01ade7bb594815fc4dde1198ff25bd30b5b15601e1963075737a2a08921b1

Observation 428ab502-aaaf-46c8-9e4d-2c368183afa6 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Grounded 3D-LLM with Referent Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.856085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.856085Z digest=sha256:951af8c1fc0a8290aa013965c036e24550291735d54f5b43f7fb2b875f3021c2

Observation 58363efd-d264-40bf-8056-4fe131cdcc46 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.914716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.914716Z digest=sha256:c6c54b41f9646775c802a39e25234a4660cfcbaea0eb9e6b2128336d72f9c12e

Observation 7854f2f5-0054-4eab-884b-f903fa972108 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.972389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.972389Z digest=sha256:8faeb6d556863b5855ff764fb885399f2a713a1744c3f9d7686fdf6ec7356608

Observation b81c9a9c-0bd5-4dd9-aa3b-85195db93ebe · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.036403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.036403Z digest=sha256:99b945a35be9804d486730cf16981649051d15f769b3cdfa15a531a3a5c0951f

Observation fa1e8111-6433-4fed-b898-564526b8756d · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.100588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.100588Z digest=sha256:d58f51800b6e75f2010aebd7b538464a6176d2c196aaa48635145497d6bf916e

Observation fc7c0d4e-de84-4322-8dfb-1d57500b0b96 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.151702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.151702Z digest=sha256:9ae79aea026247fcda7a44207e535612cbea6f5037a453d83ffab388a206cb59

Observation c344f0a6-47bd-4900-b813-e89346807cdc · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.210209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.210209Z digest=sha256:ebb78a8690ac5d38e633e971fd5dd966a85efb2e7d3196fa31c259bd295f2637

Observation 9261afb5-5d0a-4769-93d5-a5c5a9b0cb68 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.260129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.260129Z digest=sha256:51f1d9d77a202a45ba6402978a2c14fbc9a3ae88111a77eb0b9e143eb0c0e6be

Observation 68ae865f-2237-4e64-b034-6bb324ef0489 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.310302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.310302Z digest=sha256:23a2755364c61f5c43215bdd94e5e232907fc0016a8140848e449560eed9a481

Observation 18c0dbd6-b05f-4c15-a4df-c38a6f616566 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.359288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.359288Z digest=sha256:76b575f6f27f18c17a85852562c2d72622cbfd665a1853b8ca4ca3e7b2c9c848

Observation 11b29fae-b0ba-4098-9893-4791cfeeef98 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding ImageBind-LLM: Multi-modality Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.427373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.427373Z digest=sha256:f095878e81f84495413458506ce3869d2e52fe910ad568908477b8c28981f7da

Observation dac39ce8-5bc7-49d5-bca1-aad175bc498d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.538187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.538187Z digest=sha256:8f6951bc9c5055e6440de929fa481b850854ed247060f8b7533d8a78d93610c7

Observation 12dbf14c-b7ad-45d6-a642-dda8bf02b171 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.785935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.785935Z digest=sha256:e1ab2eaf375169bfda1508a7fb19c510c4be7a35b53ef9c87d91e485f67a23c7

Observation d3b9fa8c-0657-4b68-ac8d-f777b35998c1 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.896886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.896886Z digest=sha256:924b47f819705ba5962a9a1a0dcd8141533a1b2d75ef2ef9708077d264c90f4c

Observation ee98cb67-0a3a-478b-9812-a7ad333c6f60 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.333211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:02:59.961573Z digest=sha256:bc6dd697f763dbed1cca3515627d7b929d5405ef543eac0c11e42a42eec27296

Observation 3eb360f5-3190-4106-8bdc-1133acb30210 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.981076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.981076Z digest=sha256:e6f21866259524a4471080bf7849e7cdb59e3573d7af48fcdce7d1d55736b387

Observation d929c75e-a43d-4e5b-a955-1e35386283a2 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.320759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:00.100464Z digest=sha256:20f707c69e427765f21555032e59318c0b4bf37b5f42568dcfd799276f8dd784

Observation 14f17569-c587-46af-bf64-25ab994e7398 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.188535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.188535Z digest=sha256:cedfdac009feb31d90d452b4439d943b7149ef76fe011e4fef5c8a6860267f57

Observation 6d6927c8-a06d-4f62-82f4-2feec4d9382f · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding An Embodied Generalist Agent in 3D World

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.275263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.275263Z digest=sha256:82a5040c5d6c9272827496144c26a51205f232d516d1b18f51fb571dc29e57b1

Observation 31965077-7e67-48fe-aaa2-ced25d9f7d75 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.312278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:00.390648Z digest=sha256:1f0410147d76d6671165f7746ce06a6f2694eb0e4b3622bf0b9137a856505128

Observation e7bd4f78-18c4-45ac-8ee2-59ab9ed91133 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.515248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.515248Z digest=sha256:a0243eaa150b81cb1f5b6ccd822eb56aad90eb30c41b15efa16f8077effed16f

Observation 943b7901-d968-4229-9a6e-ea74284ac571 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.299584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:00.648257Z digest=sha256:bd0956b6dca790c7ae4ca4e7f4da3cc792e68dcc2d1dfe37030d9e5530110f8b

Observation 4d279c18-dd29-4371-a6dd-b9afe731d86f · outbound

This paper cites What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.712937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.712937Z digest=sha256:47a4f5d09df7493065bbf07e7842243c95025310e4e20057f785ef1cba0fe66d

Observation 263f6650-c2bf-406b-bdf2-e5da196521b9 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.291211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:00.796668Z digest=sha256:2f66ccb4d3f4e68a6dd3a9575b5c40d7fc37af45bafeb5fdf53b1aa2250cf273

Observation a944f717-4782-49d0-8b83-340b38a70338 · outbound

This paper cites Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.870564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.870564Z digest=sha256:7e85947381164fdb2a53bd8ae5838ea5b5a6186ab52b7c4d2ffb9048a9571ac6

Observation 7346c2f8-6b98-4f16-8a59-a2848d296f0d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.965145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.965145Z digest=sha256:4dc3ec9597f6cfd30129450649ed34ab181e0fe2c8b066fb213e0a01a2bbe409

Observation 58936284-d05d-4d50-9943-35cc6971894e · outbound

This paper cites RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.186306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.186306Z digest=sha256:2169c3014df1725f3798799f7a7084bf6940cc01748dbd3ea393c24d0b6d2df6

Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.298449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.298449Z digest=sha256:4b12491123ffcdedec6ae79361500bae520fc5274dee593e0ba515b38d7e4b2a

Observation 7d46be39-c1eb-43b0-a664-74208764c257 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.384970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.384970Z digest=sha256:7d8a8a799bfc3137e2eb8bfe450ee755f518772cab0ba88753f53104e5c0830d

Observation ba871690-ba73-4293-b9cb-ef2c964ae505 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.459589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.459589Z digest=sha256:59100f06bccef1a001c8bebe4b3ad4756e18b6b6b4b90b78ac08f2402b571864

Observation 8b21d098-658d-4f82-8300-bd55089c689f · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.514560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.514560Z digest=sha256:0a6f736544e371e7dec151fb3648cefcad418e955c72472859d1af5d2dd3c5f4

Observation d5ab953b-5a41-4c58-abc4-31cbd716dbe0 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.268643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:01.598313Z digest=sha256:f701496d3934b7e605d3feb75960bfe243a7dbe8aae0acac02d48e96b201b419

Observation 4209fa98-41a5-4b90-8146-f2305a0afdd0 · outbound

This paper cites A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.736901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.736901Z digest=sha256:71bfc65e1612ee8267f72f22a84511723ed9726adf482b006cd9021a12c1b0fb

Observation 63d6651f-e8da-4f2f-b415-e3cce06f15ba · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.862985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.862985Z digest=sha256:6166c3603480e9e4331545278a2d1dd123551c29d73e168247e555ef050f3093

Observation aaad7d22-b811-41a8-bde0-57f758a96a4a · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.260407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:01.958274Z digest=sha256:d6b426d104992299c76eca0138117710447c50c6eaeee293ec18fc459cddc857

Observation c4dfc515-b507-46d5-bb4a-5c2bee356090 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.096446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.096446Z digest=sha256:a4cf8e4eb411059d4b4e16bedeff8f95c55a00de5188f537a72b7d511062857a

Observation 8e495855-3c2f-4e86-871d-c49d21bf92f7 · outbound

This paper cites Decoupled Weight Decay Regularization.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Decoupled Weight Decay Regularization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.177529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.177529Z digest=sha256:3684ccb5247a209d132a1bf283749610dac824c269b0c9d510839c7ef2f4f4c4

Observation c05b8ad3-7462-4d43-a348-e4e2897b91a0 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.275383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.275383Z digest=sha256:9722b942282799e617d567cfa30ba10b91f1afede147e11e713a15e89dbf20e8

Observation cc8d66a7-de7c-48f2-ab20-092c5b8e0252 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.247401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:02.387663Z digest=sha256:1126ab8abed36f7d57594dfa11b757c27d9e29e340ab6e7abda474d52043fcad

Observation c6747194-ffa5-4fad-a468-8e298098a2ef · outbound

This paper cites DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.529236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.529236Z digest=sha256:b2c1b714c1aa98cb491a7fd0fed1e23b6bfacfd22d2215bc572fe71a94ae0962

Observation b4fb18f8-b5c6-409b-bd09-a7820b23a8e8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.638809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.638809Z digest=sha256:f1245f6e316565fec841ceebeacdcebcb54b46d35d7a5582275f4cc4b64f853f

Observation aa459827-2401-489d-9602-c3cc74856acb · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.775854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.775854Z digest=sha256:4a72bee1f1e8e19087b109c6b47b69578b06f8285380f4722a3015222f20577b

Observation f00de20a-bd16-44f9-a5d5-e95b21d78b99 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.234376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:02.909116Z digest=sha256:ab414ad91e2db773d421fb13bfe09b798ec46f331a441f7d3565a52a85df820d

Observation beb01344-99fc-4f67-8d51-9643a7b40d2b · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.225873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:02.982828Z digest=sha256:64d38525a44fa51c5c1cdcb23214a0cdaca61ca1db67b07fcf86f1985d842044

Observation 10c1911f-5be5-43f9-a8e1-23036f06c7d8 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.159380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.159380Z digest=sha256:4b36427d4fc31bf7967a4bf7fa89219ca32dff6a7f2a7fc791e7471f15859eff

Observation 85fc4b70-9b34-4241-9f9d-26410c2934b0 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.419032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.419032Z digest=sha256:2cd87ae88b02d7bd1a0e394ff1930f2c942b35f62d507b98a2f0628fa8738fff

Observation d3d5e37a-1175-423d-9e07-11cfb0186ccc · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.196313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.476300Z digest=sha256:991d3562690f284876858c41f95170dfbf9d5b05d6fd157b413cc5a908397d78

Observation 1d2efe2b-79bc-4dd7-8171-3abf4037646f · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.479044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.479044Z digest=sha256:9b0eb9b611ffc0e3cf993d0ab14561411652a571453f555074b748ccf52a4221

Observation 826bb000-010a-4fe9-af72-5ff6515c331f · outbound

This paper cites CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.482002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.482002Z digest=sha256:e44562a7a3fce631b358b56617de644e64cbf5b2511ac15408a5b44938a05626

Observation f625af16-3a32-43a3-b587-62a26cd8d262 · outbound

This paper cites Advances in neural information processing systems 34 (2021), 13937–13949.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Advances in neural information processing systems 34 (2021), 13937–13949

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.207795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.473632Z digest=sha256:b7a3d6ee021351f177f646435b7ba11827b6ab43b7a9cbe09c77d5a75fe6eacb

Observation 8e66426c-04bb-4b6e-ba8d-17fa51174131 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.488481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.488481Z digest=sha256:5ef859c442c5b3e9b17321a61087313e28fc8c7d687f893d45cdc551df370c5c

Observation fd9a1140-07e0-493f-9aba-660ef3bda262 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.491561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.491561Z digest=sha256:b5b7e951655eb2ff83d7955bfe5144d5aa8b5e825fefca169c02e96e5d5ba003

Observation 8cd510ac-b256-40bb-8011-9c7fe44dc19e · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.494966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.494966Z digest=sha256:05959a22cb1df672c5360aa836b819e7ea66140edf60ca4d218f28dcf8e6f774

Observation a7e98b14-bf4d-4ba0-b598-48ecc44be977 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.485061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.485061Z digest=sha256:bcdcb8ed6dd030af40008e9872d52f22011a3d7b7e0bafc74261bee965f888a3

Observation 12439054-5c02-4e68-95df-1195896b9604 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.163793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.500422Z digest=sha256:f99216064911c161c8f9e2a91374a7bd3f9af8e0e2150b067c49d48aa997e3f9

Observation 31a64dc9-6dae-47d2-8735-d4d256d72e78 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.506443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.506443Z digest=sha256:e99615caa3cd118e636d3528a3b51c19ff59149ad53b62b3263c69630114419d

Observation 1c0768b6-6793-4731-a6e9-a2dbe0c99a8a · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.509381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.509381Z digest=sha256:e609dab9d97f0eca593391327624c3bc7fa01dd94d2fea041ab3d15ad538cc79

Observation 6db3fe3b-5dfe-4b00-85e9-0b3513d74123 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.171716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.497751Z digest=sha256:4aeb38c667f23cad6d186ba8dd008fa0caca45e4817bc2516d585813cf8a3d2d

Observation b402ff32-e7ce-435a-9dfc-b27fa2980217 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.135222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.517025Z digest=sha256:0c83eb05f7353fc09dc3b41b35be8ff4e4ad4a02a9ae8f7ca83ead61134f8d05

Observation bf08dcb9-4566-433e-b52b-f72d9f749cab · outbound

This paper cites [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.503424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.503424Z digest=sha256:5795e673aaceaf76878bec3bf01e777a382aca74098a1fa804b22f2900aa2874

Observation 710f3385-21c6-4bd4-b355-9f7a54843c1a · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.521866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.521866Z digest=sha256:8e6e18ed680fbd3590df85b5eaef0675c8bd43528139df3daecd7c48da75a567

Observation 9d18b4bf-9bd0-4340-9b6c-470f3b9efc82 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.524477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.524477Z digest=sha256:e3c04528c63a39766e7f3f084468b26e4b3c03f6f75df440992aff2e0aef98ec

Observation 093a4a1a-8679-4475-b912-ffb8858f64e3 · outbound

This paper cites In Findings of the Association for Computational Linguistics: NAACL.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding In Findings of the Association for Computational Linguistics: NAACL

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.144483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.511790Z digest=sha256:22501195a06c1f0248b2b741ab65223b63c5fcbdcab901e9faa8493b6b0ef5d5

Observation 787ddb93-6028-4bca-8479-89bbd6852216 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.514224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.514224Z digest=sha256:7920bc9bf2a8e4dbc2f6b9ba97c649c4b7209d33413536b69e990a33f6f002f2

Observation 1643fde7-2768-4b46-bc1c-8a6a60bbdb1c · outbound

This paper cites 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.534576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.534576Z digest=sha256:fae08a182c66e1fdb685c9390d4bde1ba70a5a8aaa65370177477e6973438d11

Observation 69cbbbb6-fb80-4d49-9b2c-da94686edd61 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.126834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.519321Z digest=sha256:e6b663b1b62c7ef0eb210ddb020284591f89325819c233c7872c4f5b235255ca

Observation 4583d938-0a81-440b-aa6e-02f54282bb68 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.097685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.542291Z digest=sha256:f02b1bd106a9dac213ba4c3132f1eadf5f566694f9b1f023a4fafb82e1c4eb8c

Observation 35ef04ae-baf7-481c-b4bf-7ee6ea162144 · outbound

This paper cites Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.544777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.544777Z digest=sha256:6c64ed581eb2d08177a183682ad0e214a5804d9ea789d34d19d7a2b5aaebf970

Observation c5403498-287f-4bc9-8c20-0ef179aba5b5 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.118907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.527088Z digest=sha256:ff80c9aaf68524cec21268e96e7c7f57af4a36c35e25fc4c8eda620c1513dbaf

Observation 84c13542-3f1b-49c4-b010-e4b1694cf24c · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.529576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.529576Z digest=sha256:6771a752f50bc0d58f814695dc193d12303922a5544929dec724f909c58c193e

Observation 3b558d08-a2d2-449c-80fc-1505b8cbf49d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.111097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.532227Z digest=sha256:ea444fefe8bfa6787d0001f4f8e7524cdd07a114c9eb59b5d98a2954ffa96034

Observation b344fd13-b7d1-46a2-af72-80b526c10748 · outbound

This paper cites AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.555597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.555597Z digest=sha256:87fcaefa6604621439c41b076aac973bd9d4685b622b026b8cfe8bd6fbbd60f7

Observation ebf472a7-2538-4504-9b01-4bad846c8f82 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.537492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.537492Z digest=sha256:b4c5eaa90a5bfdd0d8a5d794430361d7df6c0d5d7b32a30fcd5500aae539f6b9

Observation 098af109-893c-4b64-a642-5475658e6011 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.539721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.539721Z digest=sha256:adf05c527aafdf606759ebc8147f9eb1f44eaa93163ac3e5b3ec7f9b334c3a7f

Observation 688744cc-7d1a-49a3-a462-de351c6c7125 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.088401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.563680Z digest=sha256:f6fbd2e31b7206203909b32e52e7464f41e077d00cca0d1e9ca3da3c89b3393c

Observation bded14ac-604b-402d-8a64-1dfa3143d598 · outbound

This paper cites A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.547382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.547382Z digest=sha256:3f9eb7491e6c82a17c9f73c84ee056479ba54b619f2334c857e46e6bc7eaf006

Observation 87499f39-b287-4cd0-b263-0a6eadf54e95 · outbound

This paper cites Dynamic Diffusion Transformer.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Dynamic Diffusion Transformer

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.550076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.550076Z digest=sha256:d2240edeed78e9c9e5571d1894515ddf39b13ee1ac17660f5767e17b50bb3f32

Observation 6a42d758-339c-4ce3-842c-429e342f8015 · outbound

This paper cites Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.553163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.553163Z digest=sha256:f3be72a2ff48aad564191b536a0eb5609cac27f427ed902d786301cfd64717f4

Observation 5276eb9e-4006-4c13-852f-11657fbe3d35 · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Uni3D: Exploring Unified 3D Representation at Scale

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.558444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.558444Z digest=sha256:4c1c1a6fcd44c82af39b52c5ea27238d91d26d17f927aa4b824ce22b128056e4

Observation 88aba4f5-63ef-40eb-bcbc-6f0cab7a4d4c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.561057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.561057Z digest=sha256:507c8c7f2943395b4bd3ee92fc7913b89984bff73490dc6ca140e98a43b6e7d8

Observation b6237571-f163-467e-89ff-16f506e8ad34 · outbound

This paper cites ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.566081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.566081Z digest=sha256:05aba8429a5432fc88e70c4e516ed9010cd1dca91e8bcdf817797bf7a0ba1baf

Observation ae7110bb-e2d5-42b2-8559-7285a281dae2 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 11 (2021), 7436–7456.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 11 (2021), 7436–7456

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.680523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.680523Z digest=sha256:64f06225e724653dba37b3bc58e3889bbf744dce5c15abbd09492c5357ae01bb

Observation 65fb1d5d-d32b-4f50-9781-e7455cb29b16 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.029960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.029960Z digest=sha256:e97d6e045ed59522eda6c34eec9371618bdecaa29669cabb895dfed36086f9c3

Observation 79870942-2f3b-4997-919b-9adaf1f622a7 · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.297222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.297222Z digest=sha256:a4f8f821ccbf5ce66a8a671c294689c20b8f44d3275636756b66ae18129b3ebc

Pith citing papers

No inbound Pith citation observations are available.