Pith. sign in

Paper Citation Record · LEDGER

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

As of 4 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2605.19382.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19382 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:52:46.575173Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact12
  • verified fuzzy47
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a708b0c-9974-4318-a485-72a96fcc4f9c · outbound

This paper cites Projudge: A multi-modal multi-discipline benchmark and instruction-tuning dataset for mllm-based process judges.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Projudge: A multi-modal multi-discipline benchmark and instruction-tuning dataset for mllm-based process judges

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.393151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:f754b87326f5d17102b4406d3752f041c5c40a3c9e1361e9fbd560f8f3aca416

Observation 3bbbd9e2-e93e-4e81-9665-4b0124dd63ae · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Kimi K2: Open Agentic Intelligence

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.344256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:55fe97b2ebffa6c26e375166e5e594f89602b09face806d15562c6825ab111f1

Observation 9b01d3c7-f0e9-44a3-ba96-c7b3af20e1e6 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.347056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:227343c1d1ef061e12771c0ceaca9aca236c11622e4e6fd10a43936814739dcd

Observation 59e15b20-0f79-4722-b640-d2e3b2cdd088 · outbound

This paper cites Claude’s extended thinking.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Claude’s extended thinking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.385671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:935852ba133f08a24d5086c1604bb4fe568baa7d262fd8665ec28cbb2f0061da

Observation bad6dd79-10d0-497c-ad4e-a233d9c5418d · outbound

This paper cites Claude opus 4 & claude sonnet 4 system card.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Claude opus 4 & claude sonnet 4 system card

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.391286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:2c1c422dfde2fb9aad729ad457b107c105f36f4bab2e6c31127f2b30a91946d3

Observation dc7e8086-4356-4f91-bde7-8d66c8487e2c · outbound

This paper cites Dash: Detection and assessment of systematic hallucinations of vlms.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Dash: Detection and assessment of systematic hallucinations of vlms

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.381983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:f27c0dc3a2470cd990879a87cbd836b2da993ea735e806f6fd7f42e1005755c0

Observation e23f9b1f-c494-4f4a-9586-4498d3a71d67 · outbound

This paper cites Tikzero: Zero-shot text-guided graphics program synthesis.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Tikzero: Zero-shot text-guided graphics program synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.383657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:60a769803123105ab1be548797362fe0be33b02e420f7f95c005423dd44dab6f

Observation 21f3be61-40a5-4470-b308-95134baa1020 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.378412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:40772db50987a37b2a52b576b4834724471c063e0ade41b30ab4f26cade27c58

Observation e6440870-8364-46ae-8199-b97ece5c3dcd · outbound

This paper cites Visualedu: A benchmark for assessing coding and visual comprehension through educational problem-solving video generation.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Visualedu: A benchmark for assessing coding and visual comprehension through educational problem-solving video generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.397015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:53023a145701e7912d689154ca00d6f039f29eb2853df2afd886275e4404ff27

Observation 36507277-95c5-42fb-8303-f092191772f1 · outbound

This paper cites arXiv preprint arXiv:2510.01174 (2025).

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning arXiv preprint arXiv:2510.01174 (2025)

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:53:04.341474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:d342eaab3d45dcaf7402c78d6a4cac05182113835ce9745dabd787ef6b400aab

Observation 21cadd12-d7ad-4208-b2f0-e2f848bf911e · outbound

This paper cites Wan-move: Motion-controllable video generation via latent trajectory guidance.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Wan-move: Motion-controllable video generation via latent trajectory guidance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.376439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:9db4a258902386876abfba298b8ced51e3704d17216c5e360dfb6c736b4bed27

Observation cd637ea1-60b4-41e2-9fa6-1b6c1cbbddc8 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.331841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:a83bdca2e074bbcbc4874bc619d2bedd14e06f663027da36f0dfe83ceeacb119

Observation 49eb2f2a-6888-47ed-815a-ebedb3fcab1d · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.338212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:3e3a8e636b95702926b7413cd47130791713ce19bf52b68b71b1104a7e13f4fa

Observation d66958c7-531c-45e4-b0a5-de628f5232a9 · outbound

This paper cites Codescore: Evaluating code generation by learning code execution.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Codescore: Evaluating code generation by learning code execution

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.372822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:8d55c7f017ccebc38cf465e728f74f983f5870dc66278b5caa9ce2b6c00edc1f

Observation e6650db5-0bf8-4511-9fcc-613539c46b97 · outbound

This paper cites A Survey on Code Generation with LLM-based Agents.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning A Survey on Code Generation with LLM-based Agents

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.321892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:a0cdf6feef585293ffe707819e83e219b1c0a26f06860a09ab7fff642c199610

Observation 60caa1ad-baa0-403d-9af0-90150ed57672 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.413600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:0146790204e1c2b9f4941e9d207d324f588a5e1d5f86918e1981e5c58c335da8

Observation a5dc679a-e866-426a-af43-43cd357fa5c4 · outbound

This paper cites Cad-coder: Text-to-cad generation with chain-of-thought and geometric reward.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Cad-coder: Text-to-cad generation with chain-of-thought and geometric reward

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.410111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:3fbdb3f3caf22bde4c86dfe8e0daa09d397c555617b5040cb37ecfff53366611

Observation df0e43c0-a025-4366-b325-49c1bbf6f3d6 · outbound

This paper cites Flaw or artifact? rethinking prompt sensitivity in evaluating llms.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Flaw or artifact? rethinking prompt sensitivity in evaluating llms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.371117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:5076a3d695a623d418495d237aa7dc8065dc6d1b81e4b2d2b15644bd52dc58de

Observation 5c6bdbbc-c277-4ac1-bff3-de6a2c415b79 · outbound

This paper cites Scipostgen: Bridging the gap between scientific papers and poster layouts.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Scipostgen: Bridging the gap between scientific papers and poster layouts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.374510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:f944dcea87ee4b864eee10b6fabb71cec2ead8a8f56448b6073a04bfe8d5c51b

Observation 21018879-d272-40f5-96ec-59373d4922c6 · outbound

This paper cites ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:53:04.328749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:f77bb8cb746d826c491bbfc1d56130b32d9e305c113d2b1302e335a4635c926e

Observation be2c55a6-1d44-42b5-8edb-ec8643d4db34 · outbound

This paper cites InAnnual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning InAnnual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.369323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:02a7b6b5c066cc40fae340925db8eb6b65cfcc44c85050bb6bab6505e5edc226

Observation 7af748af-0e92-47a3-8e17-24ce9ba8fc80 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.309218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:932a608aa8bbbf71f94a8df855dd1d82086f9d9e4c0de5f27e1d59bdad542e63

Observation 7117697f-41c2-44dd-a088-469919799f8e · outbound

This paper cites Theorem- explainagent: Towards video-based multimodal explanations for llm theorem understanding.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Theorem- explainagent: Towards video-based multimodal explanations for llm theorem understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.359821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:20cedb6afd3c5708046940dd3184f416a17cf3d479200139f1ebb71397f0df86

Observation 8e70c2db-b95b-405a-a12d-9fd0ad7f3dfa · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.361581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:8f478c75746256c0fd26899600eb734e6a2ae976748bb422c2aebd3e6b602013

Observation 0b67ef02-b2b5-4997-9b92-e8b790073409 · outbound

This paper cites A sur- vey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning A sur- vey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.363325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:5552927c000520be3f89a69f33ccc9d8e383f30c1b4a6c76221cdf1ba7173a57

Observation 9aa91fc9-6e96-4bde-b903-3ca44232f1cd · outbound

This paper cites Can multimodal large language models understand spatial relations? InAnnual Meeting of the Association for Computational Linguistics (ACL).

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Can multimodal large language models understand spatial relations? InAnnual Meeting of the Association for Computational Linguistics (ACL)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.356011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:87b770d3e57779a7e58327d4097976e6a0183074f13986e33a3f021fcceba5e4

Observation 8e54f074-d4a3-4308-8237-8717c9608525 · outbound

This paper cites On Robustness and Reliability of Benchmark-Based Evaluation of LLMs.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.325279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:810f6af443c3176e3e7b0fda452f2636476b4cb1b4218914f2b1855a4425b08f

Observation b94a7f1d-aeaf-425e-8af2-02f8391e5f9a · outbound

This paper cites Geogram- bench: Benchmarking the geometric program reasoning in modern llms.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Geogram- bench: Benchmarking the geometric program reasoning in modern llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.380283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:62c1dbc67940f04ddd36f5a6962d82b2725cc3211bae7151e1d854984848da45

Observation d1a9a95d-46df-4985-bc1c-97c23e3f14c8 · outbound

This paper cites Rethinking verification for llm code generation: From generation to testing.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Rethinking verification for llm code generation: From generation to testing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.389654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:5cca3f9434675cafefbec5c4177baaeb5cf9c2d7369a7cc564ed803cb60256b5

Observation 6e25f25c-4d14-476b-b802-ea1b49d822c9 · outbound

This paper cites ivispar – an interactive visual-spatial reasoning benchmark for vlms.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning ivispar – an interactive visual-spatial reasoning benchmark for vlms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.406210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:ae4eedd4efbc3a5ff594a65aa328ccdce828dd40c021ad1d22a377661c6e3b56

Observation 6ae5d81f-7de9-42b6-9de9-7c9511722349 · outbound

This paper cites CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.353109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:185ada50fe142359ca1514b7a343c0884d8ba1bf3e57e01e4453f247b0219bc3

Observation 32e647ca-9fd0-473b-b063-0b4517cf6489 · outbound

This paper cites Spare: Enhancing spatial reasoning in vision-language models with synthetic data.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Spare: Enhancing spatial reasoning in vision-language models with synthetic data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.430291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:6bf8f104bb4dc7cc64b299c95d58765bcf7b02d3cc5aed07a947dfd8618a4cad

Observation ad2a794c-8f60-4a35-88d6-f4b355e2ad6a · outbound

This paper cites arXiv preprint arXiv:2603.13251 , year =.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning arXiv preprint arXiv:2603.13251 , year =

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:53:04.350197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:0c2d118eb601fd7578c8a9388d4ff2ee4b1169ce119e57d98271b88ba0995b9c

Observation 6d2ece4d-cff4-449a-9dc7-d61de9681ef5 · outbound

This paper cites Gpt-4.5 system card.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Gpt-4.5 system card

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.415779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:c6f2797938e496c679d687d6bff6ee0d087e5af33ba163c1449d6bb1e63693dd

Observation 83077d83-d504-40be-bf59-07a9a7d18fc9 · outbound

This paper cites an unresolved cited work.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-20T05:53:05.424888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:a5c661bfd75f451fdd0bf16e332f17d49b5823edde3a5c228fe90499ef9ad118

Observation 4ba094f7-af21-4b9f-a288-c0e81494ebab · outbound

This paper cites Capture: Evaluating spatial reasoning in vision language models via occluded object counting.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Capture: Evaluating spatial reasoning in vision language models via occluded object counting

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.444205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:846e20d1c4e935aff7d04e1d2f005b0794e95a0587cb91cf43df211b28403116

Observation 6c7affdc-6cfc-4cef-9901-f87b4a05d8ff · outbound

This paper cites Forest: Frame of reference evaluation in spatial rea- soning tasks.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Forest: Frame of reference evaluation in spatial rea- soning tasks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.403976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:a2758c78a722d2621c5d6caad142dbc6e3ceace3e5e2a188635650d8ab7c234e

Observation 147db6c6-420e-4d16-ab4f-3fdd1fda9a6c · outbound

This paper cites Xiao, Katherine M.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Xiao, Katherine M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.402234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:71bb9e9c392271bb27a353b3ad871ed04632906f6179e923f76f905b7fcc1575

Observation dbba8596-7659-403b-991e-cd8f32d06207 · outbound

This paper cites Benchmarking spatiotemporal reasoning in llms and reasoning models: Capabilities and challenges.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Benchmarking spatiotemporal reasoning in llms and reasoning models: Capabilities and challenges

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.398766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:5e1d953ac1ba6d1d759fe7b1c885e2b328830e0034bb2a7ccb90afbfeea5cef9

Observation e6d4ea17-5c77-4410-85d7-ada5a29bdf33 · outbound

This paper cites Text2vis: A challenging and diverse benchmark for generating multimodal visualizations from text.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Text2vis: A challenging and diverse benchmark for generating multimodal visualizations from text

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.400500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:308796c28d7b259b42e9d1b770388dbe286a7d911d4a21c435bf14b83fed4d6c

Observation 4f575ed2-671c-4c77-a08d-5b6e9ae173f7 · outbound

This paper cites Brittlebench: Quantifying LLM robustness via prompt sensitivity.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Brittlebench: Quantifying LLM robustness via prompt sensitivity

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:53:04.315416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:7259b112be6063ad6f9d1ec3f204575cc83462d8a5a79893224ba5742ed58c80

Observation 1ea880ba-0999-41e2-be2f-18b53629c118 · outbound

This paper cites Design2code: Benchmarking multimodal code generation for automated front-end engineering.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Design2code: Benchmarking multimodal code generation for automated front-end engineering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.408297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:41f2200afb27c94389d624388206e36c610d60f7f40d761fed110c6db9480708

Observation c649528e-5567-473d-a1fb-313b252ec15c · outbound

This paper cites Training and Agentic Inference Strategies for LLM-based Manim Animation Generation.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Training and Agentic Inference Strategies for LLM-based Manim Animation Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.312281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:0355fadcadc6d8225fbbc528af748736ced8ae1a7c6a05a80628096acf3ca45e

Observation 3745eaf0-5980-4463-8fd4-23d26a985bc4 · outbound

This paper cites Manim - mathematical animation framework (v0.19.0).

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Manim - mathematical animation framework (v0.19.0)

Reference 44

Resolution
verified exact
doi, observed 2026-05-20T05:53:04.215760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:ea2e3e55ed97d15776949ec747737d597b8145e7a21b3bab0565ff820ecffe88

Observation ca2411c4-116d-4c4c-8033-b549db8f748d · outbound

This paper cites Ode: Open-set evaluation of hallucinations in multimodal large language models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Ode: Open-set evaluation of hallucinations in multimodal large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.394908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:da96afbdc57294343dd67cc1114bd79254b477aab1c4ac96f2a58138fc80a2e0

Observation 3d43a918-b822-4a0f-b95b-e795af0b9553 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Wan: Open and Advanced Large-Scale Video Generative Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.318855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:71569b8e12e349c51ce8527a55408c29ff67217be95d83b1fab66bc4beab0253

Observation 0c674adf-c100-4184-afd4-f015fcf96731 · outbound

This paper cites an unresolved cited work.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-20T05:53:05.357986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:278abee05da316cae4044ec81700e80d16390191b74fa3050c0c6924faf015b3

Observation 054c3384-13ed-4aa2-9a82-1dab1473e680 · outbound

This paper cites Ske- layout: Spatial knowledge enhanced layout generation with llms.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Ske- layout: Spatial knowledge enhanced layout generation with llms

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.367031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:cf8eba4c815987beb2229b086f2b15ee63212717820196ed984cc2cf0c86f751

Observation 913682d6-73f8-4832-84a6-2b876f08d630 · outbound

This paper cites Spatial457: A diagnostic benchmark for 6d spatial reasoning of large multimodal models.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Spatial457: A diagnostic benchmark for 6d spatial reasoning of large multimodal models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.365106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:535606ff69ffa5e4b47e214caea99a4f6fbc2e1f11d1a8653e25132d6834b840

Observation ef8e1f98-295d-4037-a3b6-c22a28af19b8 · outbound

This paper cites From words to structured visuals: A benchmark and framework for text-to-diagram generation and editing.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning From words to structured visuals: A benchmark and framework for text-to-diagram generation and editing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.387942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:72ddef2c3afaf3fc50c042ccacbae183d088eaee67cbb0a7b4b51a7d362e29a3

Observation 47296624-b4ad-463b-8dc4-eea1f442266e · outbound

This paper cites Plot2code: A comprehensive benchmark for evaluating multi-modal large language models in code generation from scientific plots.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Plot2code: A comprehensive benchmark for evaluating multi-modal large language models in code generation from scientific plots

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.411848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:617c187334067a0ff637c43e29e47e9e0f746a7bd0aa26e8288d576742d14d80

Observation 721ecf38-364c-4dcf-803c-ab9caa095502 · outbound

This paper cites PanoWan: Lifting diffusion video generation models to 360◦ with latitude/longitude-aware mechanisms.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning PanoWan: Lifting diffusion video generation models to 360◦ with latitude/longitude-aware mechanisms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.440467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:1b51ff67556354fce18dc241cee55fac575582db0830e2153e27a7b76a5dbb26

Observation e6dcf653-eccc-4a3e-9b05-bb46c8cbcd9c · outbound

This paper cites Core: Benchmarking llms code reasoning capabilities through static analysis tasks.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Core: Benchmarking llms code reasoning capabilities through static analysis tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.442234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:b07e05e63ecff912ba5e5a163e3219c18bc097b777f47773d2cbf05dc62d48db

Observation a4febb77-9b20-4063-b984-20d2606f10ea · outbound

This paper cites Empower- ing llms to understand and generate complex vector graphics.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Empower- ing llms to understand and generate complex vector graphics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.432061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:63b917546df9f93db3642cd2203cf861c97a7ac00150cba59743f5d4845a248f

Observation fb61ece8-debe-491e-83da-ae1c423b065a · outbound

This paper cites Defining and evaluating visual language models’ basic spatial abilities: A perspective from psychometrics.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Defining and evaluating visual language models’ basic spatial abilities: A perspective from psychometrics

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.434735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:78ca03d86008e5d05a32146c488aec8ce28099ffc76d2fa6e2da27c3b0ddf977

Observation 34f421d3-de84-4e8e-a301-812e1d8afbeb · outbound

This paper cites Qwen3 Technical Report.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Qwen3 Technical Report

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:04.335421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:3b4be3a6bd9faa07134fb07a8688901a27490731819a5ffd0360103e3626842a

Observation e600174e-2548-44cb-89e1-786278b865a7 · outbound

This paper cites Chart- mimic: Evaluating lmm’s cross-modal reasoning capability via chart-to-code generation.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Chart- mimic: Evaluating lmm’s cross-modal reasoning capability via chart-to-code generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.436900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:948a651f69f0d74bd3bf7d93c2b2b6da5881b405545e40d337bc6b11248073ad

Observation 5cca5557-48b0-4fcc-b284-8d7e616fa4bd · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.438534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:107d645c7e4a81e533c9593f6d9315f73486af4636c53a97d0c2ed4d1203c84b

Observation c3843e86-6945-411d-91e7-ff3f88983c63 · outbound

This paper cites Omnisvg: A unified scalable vector graphics generation model.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Omnisvg: A unified scalable vector graphics generation model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.446168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:71c10e3e16e01342d332f7331f35c86d610fc8ef41c5f6b28449529b6b929386

Observation dd593232-616e-4e04-8e15-73d4f9ac97b4 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.426689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:073b563d791e87eab5bfbefc127f67c2f7592bbcb7a6bad1a2a5e6a1d101e53b

Observation c0b96806-57f6-4af2-b9ea-aa75aa5778e8 · outbound

This paper cites Mitigating spatial hallucination in large language models for path planning via prompt engineering.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Mitigating spatial hallucination in large language models for path planning via prompt engineering

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.428497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:c91402a0557bd8d4928bbcc12c8d11f892312684eacebda26db6d174db0c0a46

Observation 0d515743-129f-4415-9540-794a405896de · outbound

This paper cites Sphere: Unveiling spatial blind spots in vision-language models through hierarchical evaluation.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Sphere: Unveiling spatial blind spots in vision-language models through hierarchical evaluation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.422848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:2ec7f055e003916bc3680c5a4dbd1483fa68f035e2d997e7b640206b5e545a90

Observation f00ea15c-b51b-4645-8b6d-7b837228a1f3 · outbound

This paper cites Chartcoder: Advancing multimodal large language model for chart-to-code generation.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Chartcoder: Advancing multimodal large language model for chart-to-code generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.417506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:45daea1db86b2fbad4e57c7e48ac312351da9db00fe2d183b385215aa8dc8227

Observation 044fddd4-91e7-46c6-b435-82d31e28044c · outbound

This paper cites Knowledge- enhanced large language models for automatic lesson plan generation.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Knowledge- enhanced large language models for automatic lesson plan generation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.419226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:b8c79d7e6d130d8683e8834c4ef9c3b8c649ca7bbe3616a271345bed4bda4870

Observation 63fdfdea-572c-4a8a-a0e3-710b12c07a86 · outbound

This paper cites Autofigure: Generating and refining publication-ready scientific illustrations.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Autofigure: Generating and refining publication-ready scientific illustrations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:53:05.420949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:8d9724e74a17ac6e4e5bf1f4b3619875633bbb2d7516b7e95cf7076d52a294fc

Pith citing papers

No inbound Pith citation observations are available.