Pith. sign in

Paper Citation Record · LEDGER

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2504.15037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15037 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:19.838336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:13:02.886381Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d53d132d-96fb-4e93-88c5-ec371e804f19 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:19.838336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:19.838336Z digest=sha256:c3ee51d61493c6f30281e3fe98725d911d62310bd10d12e3c2be8e8a67c89cea

Observation 82400e8f-f070-4736-8c8b-964758b22c3e · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.887709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:1ca8a50a309c40f45017889209c571c713cfd3ba53914d44689b90413163dd3d

Observation eeccf834-057c-4c69-be2d-ae4783029d6c · inbound

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning cites this paper.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.379020Z digest=sha256:0388c1fa80f06474f40144e0e00370f48568acafaebc643e7611fd04c8d527e5

Observation dfabdd0b-a398-4199-9685-ab6452281975 · inbound

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis cites this paper.

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:55.038720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:55.038720Z digest=sha256:f06507ad1c64f411269ef26ebfb84c9689f5dc43b6b26a244011caea768d398d

Observation 1d3e09d9-86c5-41c5-9316-9887ce4f4bc0 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 284

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:cd0abbde73e5e0e07e44d78bcf5b0bf0a24c83ff0875d6314125b3b47dd11d26