Pith. sign in

Paper Citation Record · LEDGER

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.15054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15054 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:23:03.822300Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4afc75ac-6588-4b67-9378-2ffc5ef46892 · outbound

This paper cites GPT-4 Technical Report.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:00.973913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:00.973913Z digest=sha256:4f94fdfa3fc5dab2b96a8e07ce3c697d93f7573772e02d8982011be0f2712eeb

Observation f649197a-b618-408a-87f8-e57038564a1e · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.210993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.210993Z digest=sha256:4149e574c13ff7b9dd611f4f31c5eb265ca7d2b631486c976b3a6ce333a6eb5a

Observation e8b896e1-377a-4bf6-b5de-c6c188cc415e · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.362939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.362939Z digest=sha256:184fea4fd32ecc6183c555f83f934f59557cac6a5d2b438e4788ab70fed12b05

Observation 4675c47a-7b5a-4ab3-b24f-83a706477a0e · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.530176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.530176Z digest=sha256:3dbe584b55f84cdcf979347ed025f6a5738cb0244308b40ed3d6a33b40341f15

Observation 05707b40-b89d-40bf-8e03-39d0298d5072 · outbound

This paper cites GPT-4o System Card.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.696722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.696722Z digest=sha256:54ceb0f45f60c5d711c60ea43fd0331ae3ee2b8b391a6c8bf3c2fce38d1556af

Observation 87e64744-d2a1-4358-87b8-233278902a8f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.780480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.780480Z digest=sha256:1efffe667d6b824703e0c427c470c9c20688a84ac66ac53126f03ee7f198d00d

Observation 5e88b47d-566b-43b7-9ad8-e1851429c6fc · outbound

This paper cites Thinking with Geometry: Active Geometry Integration for Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.948198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.948198Z digest=sha256:42a329155b08b6ee9b1b63714586f5b7244b2d48ffcd05855a8c7f34735c83d6

Observation 2f189430-0882-4be4-8ef6-5b005ae1de66 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Depth Anything 3: Recovering the Visual Space from Any Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.051394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.051394Z digest=sha256:9c32e99e8e7940931c817a899898b48e32078f940b01601cdb6089f138e8274a

Observation 4a7fe89a-a414-4b7c-9969-fa9de197a021 · outbound

This paper cites Trace anything: Representing any video in 4d via trajectory fields.arXiv preprint arXiv:2510.13802, 2025a.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Trace anything: Representing any video in 4d via trajectory fields.arXiv preprint arXiv:2510.13802, 2025a

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.211902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.211902Z digest=sha256:54ff537b9daf2c770d445142497cfbfe4ba726bc7874c79244cc0b4043cd4e49

Observation bf8cbbc5-f5d1-4423-802f-21ab9d50271a · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.322345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.322345Z digest=sha256:d4ffd12c3dfdb52cbd333e6b8761d9e4f3242edc2e50a82d052372bcf84edd53

Observation d3b06a84-d677-4a0a-9a70-d483a1a880cd · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.432874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.432874Z digest=sha256:45dd2495b82fddacf701942594a41f5257c181cbacce9c349cd69dc8e6a830d4

Observation edf82e13-76f9-433c-902f-45a4153ecd22 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.519901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.519901Z digest=sha256:fc6c4d784be0d5079da7f4fbacc8d1047aff15914094a882de8841956b1b8e7b

Observation 2cac0151-ae00-4379-8ced-26b44dd6e06d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.675732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.675732Z digest=sha256:19b17af569c6f8c3c97cc9c1da769bad1086ea76db2d7c61c36e07fbf186fc69

Observation 21e08e6c-d6e8-46fa-afc1-3e2ea1de8f1b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.835659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.835659Z digest=sha256:b364041be0aeac659c66c0cbf45ddd1aba15a7666c8e6b1efff582805920d18b

Observation ce1e85c9-3eed-47d0-9f27-d9e51c7d7277 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.003520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.003520Z digest=sha256:799c00b10bb61ccdf638206a6797e418ece0d02fb2e8000d84bd2ad6cfccc106

Observation f91d52c3-b8c2-47b3-a065-f8004863312e · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.164292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.164292Z digest=sha256:9dc27bf5b8887e8820e6fb7f6b49277e384a936a5bb925ae5dd7bfe9d6ef4910

Observation a94d5c6e-44f2-4f55-b3f5-195fd52981b0 · outbound

This paper cites Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.278174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.278174Z digest=sha256:ce04cc5d8c3a6a9316cd9460af69801d0a938c7188722090a87d700651964aa9

Observation 7437ff82-5dab-4f3b-bb6f-57ef258a53b1 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.437035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.437035Z digest=sha256:97d7d0f6caaad2cd3d3fcdafd7fe944b8f78de27172fb141005aa890d4ab0bb0

Observation a305829a-724f-4163-949f-21405c8dc6f3 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.547032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.547032Z digest=sha256:d119925b591c3dc62a80bce72e3647443525b4711da2286badba083d5b5dd3d6

Observation 65a174ab-bebd-457d-9885-22696f41c3b0 · outbound

This paper cites Long Context Transfer from Language to Vision.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Long Context Transfer from Language to Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.658491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.658491Z digest=sha256:f8395c76f818744ddf98cafc52f59f6af6b0aee14bf116547cdee6292b97fd3c

Observation 9ffbde95-a8de-47a6-b3f6-8568ef4d81d3 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d objects

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.740289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.740289Z digest=sha256:1a394a3f7b35fffd64a5029a9dd46c8a3f24795b3ac24194dc44ecc2f04a178e

Observation c7f2432c-0cb8-4974-abe2-f3bc6d064307 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Unifying 3d vision-language understanding via promptable queries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.822300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.822300Z digest=sha256:04bb0e2aadda505fb3f63ce2f71911ef1dc47dff2f937e18055d13ad0138da92

Observation c891bda5-50ec-4552-b0e0-f6f66ad68479 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.864475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.864475Z digest=sha256:4c24072a683b0af7d0b977b48a263d273e96b61afa5217631c2f72a7d1768aa6

Observation c4ec22de-4f42-4007-8a85-dc08415a53dc · outbound

This paper cites Seeing through imagination: Learning scene geometry via implicit spatial world modeling.arXiv preprint arXiv:2512.01821,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Seeing through imagination: Learning scene geometry via implicit spatial world modeling.arXiv preprint arXiv:2512.01821,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.120140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.120140Z digest=sha256:ba2f14c4c284bf2ce3080bc233d4fc06c512f95ec2da0497f1daec0cf10b154c

Observation 99ea6e02-6a1d-4020-adf5-02c703efe42a · outbound

This paper cites Qwen3-VL Technical Report.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Qwen3-VL Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.048543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.048543Z digest=sha256:4908872b4e47071a67ad9d2102ee36d93d4b9ee4215604f4269e02b44471a5f5

Observation 71d197b1-e280-443c-b1c5-c20763002696 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.arXiv preprint arXiv:2506.01946,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding 3drs: Mllms need 3d-aware representation supervision for scene understanding.arXiv preprint arXiv:2506.01946,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.613098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.613098Z digest=sha256:bc493d59ca9e7ea86c281c4c9124796e2d9c7efc692378ede60d296298b197e1

Observation 4ba8b9c3-c4ab-43e9-9f7c-f875897ca14e · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.446109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.446109Z digest=sha256:85d18542a98fcafbfc6360a843778af05b6751ad6743a08658d4b299a264e1ca

Observation 3a027346-4e82-4a16-966d-da6358e1aa4e · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.277839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.277839Z digest=sha256:8ca384bb75de1bc10346072944274658cd01ed57c047167695625e514ad83c8a

Observation 8c4a1fa3-70d6-499e-b42b-d96f56b54eb2 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.979540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.979540Z digest=sha256:6de655890519c0b780606cc962160e25d88d52c400271ebf008579e2c96e3012

Pith citing papers

No inbound Pith citation observations are available.