Pith. sign in

Paper Citation Record · LEDGER

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 5 inbound Pith citation observations for arXiv:2507.20395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20395 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:40:02.915482Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:39:59.590925Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:43.323576Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · outbound

This paper cites MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:59.590925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:59.590925Z digest=sha256:53f714b9b60939c150c6bd9a13c3f1e6fd0c8d0074563e4bdf76df1ad1235054

Observation 99541be7-f270-4b91-aed6-9cbdbb8bbbcc · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:09.594740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:39:59.694094Z digest=sha256:6448838ef385d6a431bac84cb3c04ed6d27f5e0f8fd508c45c56a7dcf5c550da

Observation 0d21d79d-f7d3-4b02-82d3-c505c491a83e · outbound

This paper cites We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:08.444915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:00.204151Z digest=sha256:5b8c1787719b0adbf93869f5a53576a7c0d2034b076262a9f5d75768b22f5bad

Observation 41bb9291-c22c-4056-b142-7a36ac46df0c · outbound

This paper cites The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:08.254906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:00.343327Z digest=sha256:e91726ecab7ccc39f853d63d8a36ab9db3dc46e07c1ca3769189bcc65138aff4

Observation 434ccbdb-126d-429c-b277-edef3801563a · outbound

This paper cites illusion of thinking.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models illusion of thinking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:07.924848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:00.487347Z digest=sha256:3ce60cfb81e9c06303f5ab8d6791073877a205fdf999b1f34300141a7fef9ed9

Observation f50341c9-d0b7-4d17-ac02-04a6bfeb0dda · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:07.194752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:00.765301Z digest=sha256:68631dfc0130d075e5a0d607196120e0b0ba0216ce70dfaf005d7838e08c315e

Observation 36f47d57-c02c-48d9-9ad0-6130d5b6d901 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:06.673377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:00.958377Z digest=sha256:7881d5e8dbc65618d8d7b681648e1e627b68c63042d443682b33b5aecd06e2b7

Observation c22597bd-79f8-4566-8333-8d41d4303dd3 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.114830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.114830Z digest=sha256:920b327178ae4c60f08a4f27e96f83352c036d89c509eb06586719238c8129cb

Observation 17eaa702-14d6-46c0-b172-226e24bd879a · outbound

This paper cites Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.425301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.425301Z digest=sha256:45bc0b82e48ab967e02209d898f02084a7dafc27a99553ba6170d884f3464653

Observation 4d31b4ae-c2ec-4b68-988f-4c949a17cf66 · outbound

This paper cites Advances in Embodied Navigation Using Large Language Models: A Survey.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Advances in Embodied Navigation Using Large Language Models: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.594827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.594827Z digest=sha256:80012b3c183799cf323e2708aa25d21eefe92eeadd4f8f335a1d4702233962b5

Observation f5ecded3-8aff-4c95-9681-4f77893aac7c · outbound

This paper cites Exploring and Improving the Spatial Reasoning Abilities of Large Language Models.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Exploring and Improving the Spatial Reasoning Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.783017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.783017Z digest=sha256:8d299c8d29bb6320130bd65e8652153c6586c746bf64db6f28f70706556a72ae

Observation 1f793967-6f8e-417a-9031-75ebb38f585e · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:06.244842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:02.184839Z digest=sha256:9f1b8941685c7804f06b4b7a726918b2b5708d2381fd08999929b0aa97d1f9c0

Observation fcea007a-56e2-4d04-8b0a-8875684c0d63 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:05.837369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:02.384763Z digest=sha256:5a43c9a7c04d9a506b731fca19d59da10efad76b730e877cbbc17811e9522ade

Observation d0996122-b4b5-4084-ac0b-27e838eba294 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:05.374835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:02.604877Z digest=sha256:a9e40c80e379d72c9731a22a9d40838bec7fb60f2025b18d3a6e7de2024a5f56

Observation 23523d3c-1ce1-4d82-a7d2-8e4b7cfd036d · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:04.934759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:02.759852Z digest=sha256:3f8ae731cb60cfcebe66bab6bbbb40af4fc5d145252fd2df36455f59d5caf94f

Observation 7b9318bd-40fe-461a-8ce3-25d15b80fcc6 · outbound

This paper cites Choose a direction: north, south, east, or west.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Choose a direction: north, south, east, or west

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:04.577069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:02.915482Z digest=sha256:4aad92c078a6efee616635c8e422120a4a2706bc7415cc3d73f87f402ad8d729

Observation 71c1483e-9315-4cdc-9085-a1149bcdf460 · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.974813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.974813Z digest=sha256:19a9fb29294f11783aefd3a11d0072a1b92f758b2533b30bdceb326fc7548d19

Observation c45dc4ac-9566-44a9-a228-fe49abb6f589 · outbound

This paper cites TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:09.194750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:39:59.837529Z digest=sha256:d3bd79674950ee7da273a5ab848d65c684937442bd7c167d7a3ed218637b80fb

Observation a9adacda-781c-4bf2-975b-d2bcac3bfe2a · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:07.565032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:40:00.634891Z digest=sha256:5bbc0caff968906db8c64f46142e380d483bac24d6296e08c628c1356d4dbc16

Observation c2525298-d22f-4b43-ad67-4caf1691b83f · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:08.774753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:39:59.995856Z digest=sha256:e8db5e0a65de4cc2ea28d823c30be9aabbcb7be33c30e244298c3f21cd50e58a

Observation fa740c9c-f980-4591-a60b-95ef6745e02f · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.294846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.294846Z digest=sha256:27f2107882fac9b69cddd858d0691f1f008b3476757ef294874ade0fcecc71f9

Pith citing papers

Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · inbound

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models cites this paper.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:59.590925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:59.590925Z digest=sha256:53f714b9b60939c150c6bd9a13c3f1e6fd0c8d0074563e4bdf76df1ad1235054

Observation 2bfd5544-d885-4fae-a784-63ccd80d8a5a · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:18.991223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:a657958e8c4ff753526c05d74537f03f45864ed5b0cb684f186be92c311e4d7d

Observation 8a1c1c7b-fa67-4df1-a8dd-2da3ab196fa4 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:27.126726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:dff361991a93bc8782cc1c9c2ba72c19226e49507f75294f7d5ec6d412e63014

Observation b3db04d6-9c50-4606-b4a3-b4e5bc6bb440 · inbound

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning cites this paper.

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:23.712055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:02:14.499758Z digest=sha256:740d0545b244a843d35690284783a387c510eb3345fffd138666b4722694bac3

Observation c66416b3-7501-44d1-9685-c2c624d4d082 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:43.325606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:a311b42d3755f73c07da801095cedf19910d3b280d514071e69e1220f81071e9