Pith. sign in

Paper Citation Record · LEDGER

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

As of 14 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 4 inbound Pith citation observations for arXiv:2505.21079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21079 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:18.874616Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:07.208379Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:15:22.050578Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abe1037d-31f3-4496-9bcd-17cc3458f866 · outbound

This paper cites Hierar- chical open-vocabulary 3d scene graphs for language-grounded robot navigation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Hierar- chical open-vocabulary 3d scene graphs for language-grounded robot navigation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.590741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.590741Z digest=sha256:1182beb8596a30c3e5a6c55d390a86c2fdbd52b9004b3ab80f25e94a6d95aae7

Observation 0e020615-500f-4fae-9edb-73a28b292d0c · outbound

This paper cites Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.656160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.656160Z digest=sha256:8d645f5bdd37df43a5e2b2be5afd1414289df2b0e7ec11397a02ee1bc20b558a

Observation 66875d4e-66bb-4eda-9f02-e8ec46ad5fd4 · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.709715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.709715Z digest=sha256:7bb256ecf33049a861eac1fd8417926f392cf1c7f48a40c7988dead9005370b2

Observation 0b088345-0db1-4df2-872f-155d2b587552 · outbound

This paper cites Multi-modal data-efficient 3d scene understanding for autonomous driving.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Multi-modal data-efficient 3d scene understanding for autonomous driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.757264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.757264Z digest=sha256:31779b2fe1d4101169bc9e21a0d9d84c5c133fa22587d5d7b6c12a3da4ab5d06

Observation fae8e8b4-7b35-46bf-9ab1-af8a33b52cd5 · outbound

This paper cites Dme-driver: Integrating human decision logic and 3d scene perception in autonomous driving.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Dme-driver: Integrating human decision logic and 3d scene perception in autonomous driving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.826322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.826322Z digest=sha256:4d553714c98f9d6fa77d30324f602de143071b63f602bd6802685fe8a14fc2a2

Observation 0106f186-2589-4936-9d5f-189914c03086 · outbound

This paper cites Drivinggaus- sian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Drivinggaus- sian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.905153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.905153Z digest=sha256:71635efeddf597dd17eddba4cf02feec8cf4aa63437c1e730e90a6dd75815d49

Observation d46c641b-4510-45e9-98ae-738e371fd414 · outbound

This paper cites Editable scene simulation for autonomous driving via collaborative llm-agents.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Editable scene simulation for autonomous driving via collaborative llm-agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:12.979659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:12.979659Z digest=sha256:9ef1c3f09650df0cdfefc6d23584b419817838e7e448ca329784d4f5c154976a

Observation 93596bd4-263d-4455-8adb-1f4d7365138f · outbound

This paper cites How to enable llm with 3d capacity? a survey of spatial reasoning in llm.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts How to enable llm with 3d capacity? a survey of spatial reasoning in llm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.027357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.027357Z digest=sha256:923a10ceeb78ff77abd1abdb2702e5a2f1af821ec7ccd5f73441ba85d5b88544

Observation 53c7b625-754b-420c-a522-4ec004070c13 · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.059529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.059529Z digest=sha256:3a697bd0deb23864ec4a47468f84223bbb7f3ef52a2ee2930a58badd02608c75

Observation 8adc1b71-2718-4873-932d-b18a1222c50c · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.113531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.113531Z digest=sha256:29a02ae91e7d98b02c1c8709d27adf951ea036779a1aa4920a3c412c38324734

Observation 55ce2427-d730-489e-881a-b95412075880 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Grounded 3D-LLM with Referent Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.184192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.184192Z digest=sha256:2635f96279104e78d5c630c3d28fe89d0be578dee265d3502b63e373d13b9700

Observation 31d7a077-b571-46c9-ae5e-87c0b1ef5c3a · outbound

This paper cites Comp4D: LLM-Guided Compositional 4D Scene Generation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Comp4D: LLM-Guided Compositional 4D Scene Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.236790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.236790Z digest=sha256:8fd0a87e95a975437947c6d764934dd640ceff1b376a26efbc902fdead2a8b45

Observation 1f3c1414-9e44-478f-9482-700086ccecaf · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.286561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.286561Z digest=sha256:6eb3286ae469137ed5506802f2b8530a75c52c4d57bb345befee248bce9ff457

Observation 80c3b379-7431-4a6f-a63a-360908578eb8 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.319972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.319972Z digest=sha256:d277740e16ab39d53976a556c790dbc2053edcb59e8cc6fb7f4bd146b60d263c

Observation 2f8a92f6-16a5-411d-929d-e212a398ab36 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.372156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.372156Z digest=sha256:cb816704da9ee87c7c3e5213b3b2f295bf33b14aee4f35598be5f75c4963c85b

Observation 59c48915-716a-45af-a6a2-7de31346977a · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.435631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.435631Z digest=sha256:b3ef8046008a39c234cbcd419852d578d599919a472aa64a5e7f8ac3ee1c45ae

Observation 31acac97-89dd-432e-855f-aef3ea07a7bd · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.494993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.494993Z digest=sha256:038a82999e48e6bef220322f20d25da2326a55b8fb7c6041329839953a208bd3

Observation c01c3fa4-4a1b-4a66-967e-dce7171274df · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scanqa: 3d question answering for spatial scene understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.552321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.552321Z digest=sha256:c6502b8b9f03bbe9c06c5aa057f2bd5633967ceec7c5f977b8a97fed2a222c7a

Observation 5b4c3c6c-2eaf-420f-8276-b0542cabdb5d · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.583155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.583155Z digest=sha256:f613390bf841f55120b45711d1bd32c132222f4f1f0ff55da99c9323903a3df3

Observation 0c12676b-a1a9-41a5-b575-7c99e3bc2948 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts SQA3D: Situated Question Answering in 3D Scenes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.614106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.614106Z digest=sha256:44473f8c55253c92e1bda9ba34bbe26dae3f7cfc3b161946f690fc46a5cf6a01

Observation df72ca07-5be0-496b-8169-c3e3f170f8d9 · outbound

This paper cites Pointllm: Empower- ing large language models to understand point clouds.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Pointllm: Empower- ing large language models to understand point clouds

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.679474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.679474Z digest=sha256:e17e6fe3149f0729e73b6759b0417beeb79ec80c1e9fc495eac530d1372b3f4c

Observation 1fb51303-b186-4311-81fd-ed0387cc064f · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.736700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.736700Z digest=sha256:5624d1964be0872a7f684db2bba432f8935834bf5114ff20a780250b2fb4cde4

Observation fc4dd6aa-075c-4139-82bc-ad69408057e7 · outbound

This paper cites Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.798999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.798999Z digest=sha256:6088feb8fe2273ee43579669741d7a8af57af8ffaf55b36b6dbbf7a6d45a1fea

Observation 3b8ae0e8-de17-4292-8b25-e5a4f5566296 · outbound

This paper cites Objvariantensemble: Advancing point cloud llm evaluation in chal- lenging scenes with subtly distinguished objects.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Objvariantensemble: Advancing point cloud llm evaluation in chal- lenging scenes with subtly distinguished objects

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:23.502473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:13.876429Z digest=sha256:1f193baf95d0dc356c917d6c81d336348cfe28116df4528e8db3b672f4827137

Observation 56a74583-e0c9-4f22-b7a9-585c3006e47b · outbound

This paper cites Gpt4point: A unified framework for point-language understanding and generation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Gpt4point: A unified framework for point-language understanding and generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.932284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.932284Z digest=sha256:4b0ce0c539a0550849f4430e20c56227f0fe22b8e5ad49362e09235205950753

Observation a6c9804c-9211-426a-9349-80c5d5df308d · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Unifying 3d vision-language understanding via promptable queries

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:23.312411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.021149Z digest=sha256:89139cb55d2b6287d62910c348f02499150aac9d54756296192d9d9254599def

Observation b2b6703b-7b06-483e-a62b-1b25f607e2a7 · outbound

This paper cites Lidar-llm: Exploring the potential of large language models for 3d lidar understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Lidar-llm: Exploring the potential of large language models for 3d lidar understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:23.130721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.094545Z digest=sha256:64cf7ed97fc73f5f6568ef8cd73f7d0ac861253350879bb1fa327396cf38c927

Observation 2cbf34af-9941-4054-a2e0-7d34e679fe44 · outbound

This paper cites Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.135716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.135716Z digest=sha256:4ff188ec357343760295297b59a47dda211fdb6ce740d90115f1d154e8eeaf03

Observation d49dffac-20ce-4382-8ebd-b66ed401d2c1 · outbound

This paper cites Liba: Language instructed multi-granularity bridge assistant for 3d visual grounding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Liba: Language instructed multi-granularity bridge assistant for 3d visual grounding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.914612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.140769Z digest=sha256:0f2977b231dd0263eb0e90d14024fc34195cd10b53d305c10ed7dd4086992dd3

Observation 36cd438f-6601-4aec-b961-09c236eeb544 · outbound

This paper cites 4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:46:19.380964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.183457Z digest=sha256:87d379ddcbed750b37abaef6cfb1b4479b6d798b342e0c2eea9c2825b729a30f

Observation 71c8085a-3bc9-40f3-acc4-0a4af09b6440 · outbound

This paper cites Space3D-Bench: Spatial 3D Question Answering Benchmark.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Space3D-Bench: Spatial 3D Question Answering Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.247666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.247666Z digest=sha256:25b9ec40a836a163b057edc2460ee8ecffe766a78e3294c18d28e9dab5292f77

Observation dfa2a3c7-da90-4b0f-a8a3-343e810bad04 · outbound

This paper cites Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.341180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.341180Z digest=sha256:dab8b5119500911c337e38480bca4b00d71a26630d969259f4c2545a3189b4fc

Observation b26fd254-490e-46e6-b794-35981a26791a · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.423259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.423259Z digest=sha256:22bc7708ff46ceaef9c75b3b1da43b1f2b711faced7c4a316f2b3d0871cbc0dd

Observation 58d23528-1edd-4483-a4f0-a17c2c013d05 · outbound

This paper cites 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.473013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.473013Z digest=sha256:b18f38697f0a3302d04567ebfcafe90c12f3e7865c513dae59cb9a4fb949fc6a

Observation 5a9cbfaa-84a7-43f0-a20b-aa0e4d0e0786 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.529598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.529598Z digest=sha256:89a744a482f3710fa61ab7524d0dd1a6795efbb2cd193f6fb3e3d045497ad099

Observation c2fc252e-45b6-4916-941b-f3cca1a8531c · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision-language tasks.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Image as a foreign language: Beit pretraining for vision and vision-language tasks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.758988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.596999Z digest=sha256:cf4f7e112b3cd478b03b7a4d738039c0651afdc9392aad791ac10bedded49b1a

Observation 0abf9120-0f6b-4356-8a78-4edca9689486 · outbound

This paper cites Uni3dl: A unified model for 3d vision- language understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Uni3dl: A unified model for 3d vision- language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.578435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.681514Z digest=sha256:bfaa39349bbb5381363a48cf5f375ebf48f20cad2305ce38d24d2e88060cfb78

Observation c5ee0eaf-dda9-43a5-bfbd-d1e485fde89b · outbound

This paper cites Vision-language pre-training with object contrastive learning for 3d scene understanding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Vision-language pre-training with object contrastive learning for 3d scene understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.457441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.742528Z digest=sha256:4ddcfa97120e553b32d7f52f22c60ac16076cd7e5bda523c180ee57a52a95ba7

Observation cf7acec0-b0f4-4617-8774-0a8ffe20e460 · outbound

This paper cites When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.795239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.795239Z digest=sha256:8151a1ea610eb780e5451720e673cd51fa65c6282b9a2c18dcc4e55bd0e97a7f

Observation 974b4501-a380-4b02-837b-c29974481829 · outbound

This paper cites Mixture-of-experts with expert choice routing.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Mixture-of-experts with expert choice routing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.293795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:14.893770Z digest=sha256:8c176fa3a4e7be00287f18e1ab2f2d008b94ba8d8e1d056fe7112e8bb2c7e7f9

Observation f4c6c5b7-6629-4ee2-a0fa-36051bc411c8 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts A Survey on Mixture of Experts in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.957598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.957598Z digest=sha256:6ac9e196d04e9bac912c2569454b666582a18cf76648073904fd07390671b910

Observation 4e6411fb-704c-43ea-b942-30c9f3e33ffb · outbound

This paper cites Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.035924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.035924Z digest=sha256:d796e92a414bc0f7c3648df78fc980b4425c9a35a01edcf2b4d1487991a0a88b

Observation a6dc91f9-957d-438f-8e46-7b78ec8f0fbc · outbound

This paper cites ProMoE: Fast MoE-based LLM Serving using Proactive Caching.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts ProMoE: Fast MoE-based LLM Serving using Proactive Caching

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.078213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.078213Z digest=sha256:63c0cbce7d84a779237d5eccb6415be61aadcafec76be3e282464cc380ebaf72

Observation f67659c6-7273-469e-8866-95eb01f56ba6 · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.171955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.171955Z digest=sha256:49303cb4e3029ef6698ae4fabe4c78aed4c90a850fd9b56fab05bf8cc4f8dd4d

Observation a6b45c4f-df13-47d0-8e59-5727b1ed5834 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Vlmo: Unified vision-language pre-training with mixture-of-modality-experts

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.152608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:15.264291Z digest=sha256:0543787a7002978fb811eaba9704a241a53f00ccfc6d25747950be88af198a5a

Observation b25f7791-3c29-4b06-a152-e42c32f6c2c6 · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.339882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.339882Z digest=sha256:126f2c83aade62ef341352bd706634d845eff21130c65bcc668fe1ef4a08eb9c

Observation 44d5141c-8100-4e17-9fbf-3a6a0fe632e8 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.416411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.416411Z digest=sha256:b012d43a6b3669411e74c6c36dcac2ffb207ac7ffb1e92de26d55c3235d2f1dd

Observation a776cff8-8824-4d97-b82b-6dd8664a8bb9 · outbound

This paper cites Ada-k routing: Boosting the efficiency of moe-based llms.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Ada-k routing: Boosting the efficiency of moe-based llms

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:22.012673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:15.507691Z digest=sha256:dbe036801f04859dcb41982c8f7cc149194330f3beac153667a194e825a3060d

Observation 7f745a5f-f77a-49d4-a003-4884aa126729 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.586583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.586583Z digest=sha256:a1597fe8d777dcda2e281851ebe742f073d24610f2f7982f55450013e7562cd9

Observation dd29c3b7-32ed-4da1-be2b-23990d2161ea · outbound

This paper cites Llama-moe: Building mixture-of-experts from llama with continual pre-training.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Llama-moe: Building mixture-of-experts from llama with continual pre-training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.641775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.641775Z digest=sha256:201bdd57a43e892fa7473b77a7e19be0e2a64ce376824b0e89eb683a5c706241

Observation ab168ca7-d01a-44f8-9e8c-3d866fc399aa · outbound

This paper cites Uni-moe: Scaling unified multimodal llms with mixture of experts.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Uni-moe: Scaling unified multimodal llms with mixture of experts

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.874331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:15.679963Z digest=sha256:c13d7d80093c96e3bfd58ed27132908ad232e2d61ed9e872e933acf07a73c004

Observation ec264051-b15b-4672-ad89-4c4289841fd2 · outbound

This paper cites 3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.744253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.744253Z digest=sha256:ac1d839caa5ce0916711996a5e329b2d077746b7d30012685aee4e1768e15443

Observation b39e2d56-872c-4d1e-953c-087506ebb4cf · outbound

This paper cites Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.795066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:15.796253Z digest=sha256:8caaad0e557143ecc2be90509be0b07bc8620a3d8ed8b752779118a9a867cb10

Observation 9d0e7d08-ea5f-4681-8d38-ac38b5b22f3c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts DINOv2: Learning Robust Visual Features without Supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.889974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.889974Z digest=sha256:e99166a099394119d37a2180f39eac39284ac467bdce8143a74766482bae75ae

Observation 9900a261-95af-4cf3-96b6-77ae5aa9c000 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Learning transferable visual models from natural language supervision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.704411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:15.929697Z digest=sha256:ba269ff4b8ff192a49475c6607d27943b7503f64b5126e9c6a58443dbc9f1454

Observation a50b727f-c45f-49bd-a102-77120c739898 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.963506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.963506Z digest=sha256:229280ea00168d7a864f4f7494b7467bc69088a819116e7fe5e59cb3ad65e2e6

Observation 8dd7b16f-a899-4752-a93f-7586404e5a80 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.039764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.039764Z digest=sha256:08977d9851d6f6a41706d056f48c2795524aaa466ec5672caec9f90e52712c01

Observation d4fc414e-856d-4cc6-be66-fcf3f8a21dc4 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.132585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.132585Z digest=sha256:9454bf252651c18227d4b0b3eae18430b4076a02da359ab3e0138ac622e04186

Observation e62bd513-2842-4d28-9f94-4c3d807f284d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.200732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.200732Z digest=sha256:c4dcd93f3c7bd882ccbc29e4f139bcd5c5c92d2cec1e2badf0d3fb330807c1f7

Observation f138a48c-270f-4626-abab-eb9a6e134db9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Multi3drefer: Grounding text description to multiple 3d objects

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.544501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:16.245785Z digest=sha256:770cc58a3337b9db85be5a4883042eb314b67b495b49196c46d1cb931ced1071

Observation 41cd88d9-8585-4d71-b565-5a69a56e54f6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.300556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.300556Z digest=sha256:c84ac10e73855e45c18e82f5f2f64f57d8713df987c7fa17abd769920d96c833

Observation 8504db6e-dd20-4c52-a550-76c5b5c5afae · outbound

This paper cites Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.379978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.379978Z digest=sha256:ff6ef384b200fe9b6558ba30d48d0061a6d400b7dc2563e64b354e6539adc610

Observation 40031675-6144-4b68-a221-1d878f6e210a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Bleu: a method for automatic evaluation of machine translation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.425498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.425498Z digest=sha256:a3359d34262413f8b7a681d961cb95174458cef342eca8bbf3f436d9d9543c76

Observation ee138cc5-f44e-4883-ba44-1dbf94c066b0 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.457621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.457621Z digest=sha256:fc2a3ae5cf8e04d6e840310e851862b81d80d7aa5abbd36cbbf12c4876017fa9

Observation 38de929c-ac86-41a9-9fc0-e410279ffcab · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Rouge: A package for automatic evaluation of summaries

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.494360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.494360Z digest=sha256:ed4ca0fc49015d7d13b98032ef28fb0ce73b76560e4e419c74cbaaa8d112c0e1

Observation 7d724260-e5ea-4175-8c9d-ca14fe7e6c5a · outbound

This paper cites Cider: Consensus-based image description evaluation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Cider: Consensus-based image description evaluation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.515497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.515497Z digest=sha256:fb0bab477161bc93f2783ad309bdd750fc07518c9c70aa67b7598c8b3d8cd9cc

Observation fab14a4f-c41b-4382-9de3-42f4c6b02548 · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Context-aware alignment and mutual masking for 3d-language pre-training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.433344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:16.563766Z digest=sha256:d5709468d723881c696c52d28513f206c9349586e3b362550650ee647a2d3335

Observation a461d938-a693-48f4-b239-95aa8b05a8b0 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.625470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.625470Z digest=sha256:a5eff4a88db8cdea0a8e129c57b7c4849b6f2e58c3e71a1853375e50e1a189f1

Observation a11f7e75-7d53-4f32-8fc8-0fb0710c46c6 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.681525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.681525Z digest=sha256:c91ca2a6a39ae6526a4848c2803e684e376a4b343ed8b88c9359a20bb5b80e76

Observation 2d9e9fbe-b150-4716-a71a-1cf7ec52cda5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.756757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.756757Z digest=sha256:10fa14751a5e6b0720aa5125c863b3ffc4c01e4691163a5063c20cefcc1f7af9

Observation 123279b9-48f2-448e-98d8-9f2e0c7aed8f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.849356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.849356Z digest=sha256:d0c54d4e9aadc5b2e28199d8d2f59d587ce84534a3e418810473b0e08474307e

Observation 12c600fc-3b43-4561-a1ad-6e931994fde7 · outbound

This paper cites Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.316454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:16.948131Z digest=sha256:c898ba9d589b9498d41572e33c287d650c0a677a92247539f2f7a4ac609051a6

Observation fde67467-ac28-41b9-8fff-4ab2b418bd52 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3d-llm: Injecting the 3d world into large language models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.076369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.076369Z digest=sha256:0d0c26b2381ffae29e153014874b2999a75c8fe9f208a353659a8c900196f919

Observation f1d2de7e-81d4-459f-9cda-97874fd4f2d5 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.141359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.141359Z digest=sha256:506dfb6bd711b66931e89e16eb15ddbbf2893c482ba62a61a315f4567f1db6f1

Observation b896d076-9b48-4ca8-b16b-c194431e4349 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.190400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:17.215388Z digest=sha256:4515aa55a2a5df36fef428af28e561c5514514a3f60156d3ec61965ef6ad2c43

Observation 59aa549d-4755-4f82-9b60-f94b16bdc241 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts An Embodied Generalist Agent in 3D World

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.324846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.324846Z digest=sha256:6aab39c0f197ab2e408bf31f89192369ad7174579e6e70e4bcbc7af86e41282f

Observation 20c4dede-de2a-4823-9071-1ec931fa9b06 · outbound

This paper cites Principal components analysis (pca).

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Principal components analysis (pca)

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:21.080412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:17.434535Z digest=sha256:b4f327b230e8c6d5379545ecda2410ba4aabe68ace13aa57ce2a26469090c4fa

Observation 5b455741-a976-4bd0-921d-51a0d97dbf53 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.967618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:17.554741Z digest=sha256:8c20301897fa6f7db98fb5d6c445dcca3388d1c5759381fa0984dadaae1d1bee

Observation 4f9e6749-68c8-4694-aa86-00fe14be757c · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts End-to-end 3d dense captioning with vote2cap-detr

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.892441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:17.671882Z digest=sha256:0ba22a877210566b9313f1a14ceaa10442f14677451e5310a1ce4922a0ce9e1e

Observation 407a0d21-38d9-435a-b446-64c737fc95eb · outbound

This paper cites X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.765740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:17.746676Z digest=sha256:5b996046cdf3da9cdaee442fa8f683d7d1d782b2e9722349e6dbb44f5361f275

Observation 8cdb7196-3760-4303-a22a-0f95b1a3480d · outbound

This paper cites MVT: Multi-view Vision Transformer for 3D Object Recognition.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts MVT: Multi-view Vision Transformer for 3D Object Recognition

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.813022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.813022Z digest=sha256:3db70dee5b0c88fd504943247f4d1c2229357d95ba784d5e1a924b3fa7ed50da

Observation 7438619f-e002-4e53-97e7-4dff0e2a6e5b · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.891316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.891316Z digest=sha256:4686aecf47d0bedfa3ea64996fddc302aa2cd7a943fed74601a9ed70e0e666e8

Observation c3f666bd-9bc5-431a-83a9-3e99b6aa9ec4 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Language conditioned spatial relation reasoning for 3d object grounding

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.973560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.973560Z digest=sha256:18644133822367ae95e23b9a1954d2c15a5060e1a61402ce38bfcfa6b75f8012

Observation b1a8eeff-6005-43f8-a02d-35249b7e51d2 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.643575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.077406Z digest=sha256:6b9c5e519afd13c26b50488d5c805c17c079582c7c63bb388b209d5c52ee61ad

Observation d502179a-f70d-436b-beb5-3d0aed48b469 · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Text-guided graph neural networks for referring 3d instance segmentation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.435493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.161114Z digest=sha256:472a734d555dec63dc2bb105bcef76e25ab496eb04ae91aabe1c6b4916487b1f

Observation fb3fdcde-5000-4ef4-b0d5-4f987f3f9322 · outbound

This paper cites In- stancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts In- stancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.326600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.270692Z digest=sha256:c2975264bcd91573975af76aaa4668fa7356a39c2a966c4987381b5a94e2c247

Observation 19ee3c74-9ecb-4d4d-a1cb-2f1b8e355a85 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:18.389743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:18.389743Z digest=sha256:2dcff2cba7f14eac0e5b5261af2127848d7ec199224595cc39378b970d247def

Observation 9595cc4c-1e4c-4f43-b63d-3fcfb8f4460a · outbound

This paper cites D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.194643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.491669Z digest=sha256:0914c35bff8b4b9985755c87f81366cf9f3c5d95fe627a56fe17d07107d11d1a

Observation e307ec8a-2d47-4424-a4bc-f77ca43fa1c3 · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Bottom up top down detection transformers for language grounding in images and point clouds

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:20.036900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.602704Z digest=sha256:f16568eb74934d46057001264f05b4bcd6497ad8091d14c8b35d576230f1dc5b

Observation 7a183bd7-3cec-47e5-9641-643389b45c1f · outbound

This paper cites Learning Point-Language Hierarchical Alignment for 3D Visual Grounding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Learning Point-Language Hierarchical Alignment for 3D Visual Grounding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:18.684551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:18.684551Z digest=sha256:864563d2f97a9abdeca81ca8bc314ce69eacde38ff84d7353155e5260fdaa2d1

Observation 75982f4f-e0f7-44c8-aaef-8312fa98d3bc · outbound

This paper cites 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:18.790770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:18.790770Z digest=sha256:178a55d261b86c5ba02a983560fb590efbee18f0146f10c970960e2072ee5d7f

Observation 8fa273ce-e506-408c-8999-0442de193403 · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:19.857840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.827522Z digest=sha256:a9df5c94429d8ef1e299258f95e5751b778e7db1235d42249d24db57657ba393

Observation bdfbdf3c-1840-450a-abc8-6f0425535a05 · outbound

This paper cites the bed, which is rectangular in shape, is located adjacent to the door.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts the bed, which is rectangular in shape, is located adjacent to the door

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:19.736225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:46:18.874616Z digest=sha256:ed58a33fdf60ba9945689dc58317e590944152bc7529907d407b3931c9aebcbb

Pith citing papers

Observation 5b933114-d814-4c29-94b1-37285e3b4ebb · inbound

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts cites this paper.

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:15:22.052701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T22:12:51.364157Z digest=sha256:14f87850c6e764b4a1897bda1ec801850ec09326b778fe8938d85f0079af9f03

Observation 15f2ecf4-f69a-407a-b325-57b9f5e01fa5 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T04:28:28.887051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:28:28.887051Z digest=sha256:aa4e5e807bfad9e3570ace8a0e672295eef013b9a6b606014c40401cb9bcc0f0

Observation 7144489f-41c4-41bf-a46b-ea6c4bc40f60 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:41.136333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:41.136333Z digest=sha256:848ed24d674a349788347263077747f76cf4a9b64718c1ebeddb34dbc534c9cb

Observation cd59890b-4fea-4c10-b579-2d6ae6125f60 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:07.208379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:07.208379Z digest=sha256:9a1741dab66373db45f9ec7e6bd0a3dc0f63c0ad36839d16e9e746ecfac56805