Pith. sign in

Paper Citation Record · LEDGER

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.15485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15485 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:32.315784Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T16:44:56.123044Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0f972e68-e574-42ee-a6ed-21d647e6ff52 · inbound

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models cites this paper.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.315784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.315784Z digest=sha256:accf3c535f2b12c41405cc78f678a798def9bee21bc27fdeced25220e68c29ff

Observation 89ed4db9-1ac1-465e-9632-203f668c3226 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.167072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.167072Z digest=sha256:208e83863820961d5472d365809b1f2131dbec5a6aebbdeb54615f62135fd976

Observation b1d2c762-7cdd-4a35-a45b-60977dcf6494 · inbound

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates cites this paper.

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:07.618065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:07.618065Z digest=sha256:45600adfdda447573193d3af95a0b273e6192310c872dc8ba592b75dbbd941c4

Observation 64a2ad40-79a8-4e5a-b180-c7fa679bd364 · inbound

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications cites this paper.

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T14:53:02.245247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:53:02.245247Z digest=sha256:ca07ea7c99daf2e3e7eb01be9112a158a1ef8b57a1e13761192fae5cb90fcfdd

Observation 818d1e8a-e5a3-4816-b936-dd8ead34c7a9 · inbound

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations cites this paper.

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T23:42:44.515158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:42:44.515158Z digest=sha256:318ca6389b63fbb25ab6be3aa74d906edea60691f2442e9f6a9854ba9b38f2ff

Observation 20a09eda-3925-4e68-8f68-0ab48bdece20 · inbound

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving cites this paper.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:15:22.749650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:10:32.522453Z digest=sha256:453107ec88c31aa5dfb164c461ed656e2492cd8a73d2b71433e8c80b23bf72be

Observation e513a6c6-598c-42cc-946a-022c513cc61d · inbound

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving cites this paper.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.124516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:7ef2d7a7dbce15803d759a4fafd1c2b365070f03c5d9ac07e945ac5238ef81ee