Pith. sign in

Paper Citation Record · LEDGER

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2502.08859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08859 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:30:43.662827Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:56.639836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T08:15:33.926301Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82346c5b-96d1-42e8-81cb-8cd0681ce094 · outbound

This paper cites https://puzzledpint.org/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://puzzledpint.org/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.156381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.527786Z digest=sha256:2ddeef3227fcb9134eb2234b5c75994c30573f454bccca355a278ba1875ca86c

Observation 843cd5f7-0e66-41bd-a6cc-633ae07e1a93 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:44.145614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.531958Z digest=sha256:4f06b417c83000046943cc2882b501933af4194480e05bb982bc939fe5dcd764

Observation 5c814eac-f57c-4476-a4be-ec1b1dcaaf93 · outbound

This paper cites Puzzle Potluck.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Puzzle Potluck

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.133987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.535582Z digest=sha256:4553eac5f0e1689c942bdeda962732944356139b7522011b5c256574fbaa9afb

Observation 07555f79-d9b5-4272-9dae-060dc3c4e83d · outbound

This paper cites SOME PUZZLES by Mark Halpin.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges SOME PUZZLES by Mark Halpin

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.123189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.539340Z digest=sha256:a89b3af0c752626281879373cd8a9c895293fd120da7c5d3e189f292a79c8327

Observation 5126262f-5fa0-42a8-875d-6c8ad29cd47c · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:44.112500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.542911Z digest=sha256:c96f13029b348bc72b9604c6fe78d541fd0fdc4b527112b8cd299c9e42709b4b

Observation 8ae17add-ffd2-4fa9-b4b2-a6c2515e9957 · outbound

This paper cites https://puzzles.mit.edu/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://puzzles.mit.edu/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.100847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.546828Z digest=sha256:c5c8c83b08bd5586681a811c2dbf261ef8c7173d5ee8a83356ff9d7845736a72

Observation b2dfcc65-8744-4ff9-81e1-10de40e61f08 · outbound

This paper cites Grandmaster Puzzles.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Grandmaster Puzzles

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.090026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.550584Z digest=sha256:455398aab0f7517f460fac4735643bb02bc23210e65a517c20361c68575c9b9e

Observation 7cd582dc-3fd8-469c-a10f-7f7fbcb531af · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Measuring mathematical problem solving with the math dataset, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.554415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.554415Z digest=sha256:8b7f6e0f388231495fbaf8986830b306aab76722a6348e1acb7aff6ec247c5be

Observation 391adf69-e187-44ba-acb9-c27e1c67651b · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.557938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.557938Z digest=sha256:e841fc95d511692e0ea03384ce8afd63025dfb90546fd8b737bdbcecb82f2905

Observation 9d881114-a65c-439d-850e-23a6e4529837 · outbound

This paper cites Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.065829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.561371Z digest=sha256:44e17557e987032882f3bd2ffcdb5a2ee599e87368733a9f9914148714dec285

Observation 0cd4cd2d-5df6-4280-bfbb-f1383fef72a4 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.564791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.564791Z digest=sha256:3a6978b94c4446101a82997cb69bde78af5ef619be4ea7eb91d08b6203900c1e

Observation 7d64a4a5-0732-4d6c-bbf7-a631a43aa849 · outbound

This paper cites Humanity’s Last Exam, 2025.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Humanity’s Last Exam, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.046677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.568159Z digest=sha256:c6224d9b07c32028273ee2c084fa72d16b08b70926f549c4e57715eb28d6d405

Observation b177cc5e-d96f-4155-9f66-ae21f995b0b1 · outbound

This paper cites Measuring massive multitask language understanding, 2021.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Measuring massive multitask language understanding, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.571998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.571998Z digest=sha256:f71c9ccef1d5fd590ced1de3a72ca04d14bd753f8357effe4201ee58b6bffd62

Observation 8f5d7d3e-4517-4098-8eb6-4c80358f5981 · outbound

This paper cites Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.030082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.575607Z digest=sha256:b75b123a92016457bf1cca103d1ed561a2ffd72411bf67ed4d21734442a59f1e

Observation 65f7a8ab-ff2c-48cc-a61b-b9bc36381e13 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.578906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.578906Z digest=sha256:8563f38f14869a53de0266de186398f9eebc106809a28b421413fd474b51ad96

Observation 460bbc65-3f94-4551-9fc9-f93c875e9445 · outbound

This paper cites Vista: A rubric- based visual task assessment.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Vista: A rubric- based visual task assessment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.013743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.582228Z digest=sha256:51b85be193f7daf59d50c1e5ec8750c6e7ddb09e4b0b2c2468328f15ec5a1a7a

Observation cfc927e1-817f-462c-a866-d9bbabc01ebe · outbound

This paper cites On the measure of intelligence, 2019.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges On the measure of intelligence, 2019

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.585618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.585618Z digest=sha256:3bf0a6a58a5dbb5e9ac77b4804193fa19b43a0c78d0c3b6cf9a58943a38dfd89

Observation 8c0601a4-d84c-47d8-905a-eef165df06a6 · outbound

This paper cites Lanzendörfer, Yannick Niedermayr, and Roger Wattenhofer.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Lanzendörfer, Yannick Niedermayr, and Roger Wattenhofer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.997755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.588907Z digest=sha256:4c2a489a4727cdc82e5942246c2faa9613c9bbecacc8839c24f2c0dbc8be9155

Observation 5ac52958-4155-4a5f-98e0-eb7b092d71c4 · outbound

This paper cites PuzzlePlex: A Benchmark to Evaluate the Reasoning and Planning of Large Language Models on Puzzles, 2025.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges PuzzlePlex: A Benchmark to Evaluate the Reasoning and Planning of Large Language Models on Puzzles, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.987521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.592101Z digest=sha256:e00efcefadc914a9320a009e5f90814aa1ba7f0107825b325041f46c9623b9af

Observation d18500c4-f372-48a4-a7e0-2d5e38939151 · outbound

This paper cites FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.595386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.595386Z digest=sha256:7043c1e5b871c1c6490b0c8059b45e77f08b7dd9114926ae9545362baf8de5d0

Observation 46e6cb43-c7dd-4a51-8eb6-9c7d8dccd28e · outbound

This paper cites Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.599486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.599486Z digest=sha256:e821059381265729bf9711aefae513fe53900a73557a932a581911fde054a320

Observation d7a2bf39-b9b4-4d99-a75d-af9f8f300053 · outbound

This paper cites RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.603305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.603305Z digest=sha256:90bec5e913b4edb11a8ff5350b6ff7b3d7d4809c562e71d1ff75a64819d86680

Observation e6bf3606-0828-4f48-9380-73d8c2b2cc1d · outbound

This paper cites https://www.melbunimathsstats.org/puzzlehunt.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.melbunimathsstats.org/puzzlehunt

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.977237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.607095Z digest=sha256:28ffd29c8ba2a1e083ec68a34f660a68d3208256a8a825d5559e22d6f63b701d

Observation b1cee88c-c0d9-446e-b0f4-45668caf367e · outbound

This paper cites https://web.archive.org/web/20210725192741/https://www.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://web.archive.org/web/20210725192741/https://www

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.610435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.610435Z digest=sha256:16feac93b4ecdb2f8919b2ffe43a08a8f63535ee834778a6f83981043f0e4948

Observation 342ba2ad-cb5f-49c1-ae7e-e7f7828dfbdd · outbound

This paper cites https://harvardpuzzles.github.io/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://harvardpuzzles.github.io/

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.967490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.613709Z digest=sha256:4f6202dae4d25dbb3675cfc8223b02f04786654a4ffde1a5c5c7c4dfe895dcbf

Observation c7fdeb96-070a-4dc1-aaeb-f1ba1088d752 · outbound

This paper cites https://www.mezzacotta.net/puzzle/cisra/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.mezzacotta.net/puzzle/cisra/

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.957817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.616875Z digest=sha256:45671cdd647a3e2a43c9041419f84e34a860c7fa56e7ca85c74abbafd23251de

Observation 503937d6-4c14-453e-ac0c-00bc546d01df · outbound

This paper cites https://www.janestreet.com/puzzles/archive/index.html.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.janestreet.com/puzzles/archive/index.html

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.947801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.620473Z digest=sha256:209cf4b16ddc5e90d2a78a196b14df9f80825aa797f85cd7b8115c4c5ae722b6

Observation 3a1b91e4-717d-4d1e-af44-e612c8e06b68 · outbound

This paper cites https://gooooogol.theburninators.org/puzzles/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://gooooogol.theburninators.org/puzzles/

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.938099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.624023Z digest=sha256:0f34d7b4b393ae22e1fb396721fb223f3437b1bd5a19f38a5e784f55086965c3

Observation a1c18164-aca6-46c1-821a-3a837d60b995 · outbound

This paper cites https://playdash.org/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://playdash.org/

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.928572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.627592Z digest=sha256:3e40b85b6e95d5d530a97ccb5fa24827da8bb444f0b45fd75d1add7095bf1288

Observation 1776b629-139b-4d48-a687-d23cda9ee5a3 · outbound

This paper cites https://www.baphl.org/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.baphl.org/

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.918925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.631361Z digest=sha256:750f54e8e62ea0c02955e734aba3ac81585dcea4a4a74af2c4787aa22f223b42

Observation f509d110-5f11-4786-99f7-72c04cdb207d · outbound

This paper cites Forbidden actions.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Forbidden actions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.909101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.634785Z digest=sha256:8062ba80867dcd580b8ce12bd99ffa62a2d8ab7540217e7f2ec4d3b2c1358e3f

Observation 49d8ad0e-f128-40dc-9fbe-5178ccecdb90 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.899181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.638190Z digest=sha256:018d67cffac71c06f31bad0a2aacec8ef43741bc1cae8227616ffdd19be47daa

Observation de4a874d-eb72-42cd-87eb-a1f8537c0f43 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.889272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.641496Z digest=sha256:7a8196b7fc2f7e53bc6ff31894f0cc82986ea554c2ae2ac3172600c0373d465b

Observation 79fa53c8-27b7-4cb2-8e7c-6eb38e6962f6 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.878286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.645200Z digest=sha256:16d35a547889e0210b10a19678f25553a6614ef52c1d34b6596ac6453bdb5240

Observation 708a5c9a-11ba-41b1-9db5-99094dc2a39c · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.867120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.648403Z digest=sha256:b08c3e07314fe309f29c16f194e08796d056753677baa8b84b8ace57c763fbd8

Observation 25a673f7-2449-4694-8b86-389c4c22039d · outbound

This paper cites Problem Web Wage.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Problem Web Wage

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.856528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.651991Z digest=sha256:90833f784c7df827e5a29e63267d1ccb3111c9a5e25609ce35d389601cdc0c77

Observation ea411947-d5fb-4d7c-8130-9d9b09a33023 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.846067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.655828Z digest=sha256:3481fca6f6d22c8c1e46807e8e342e2cf8497efe4cf4463ebcd7c4deacde3349

Observation cd0e5dc9-42fa-4b02-8123-e1e335477d5a · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.835897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.659550Z digest=sha256:c8cb93d0880726a2d89561a9b898d5ec740bcdab41bc05ec2b4ec914e8a963c1

Observation d052f01f-dada-4f97-b0b3-42046345fa91 · outbound

This paper cites This structured approach to answer formats allows us to extract answers consistently and reduces ambiguity when comparing model outputs to ground-truth solutions.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges This structured approach to answer formats allows us to extract answers consistently and reduces ambiguity when comparing model outputs to ground-truth solutions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.824654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:30:43.662827Z digest=sha256:1d910de5f231cfa629fa08b868e187eb0ff1535f101b14da699e21b2f550a08e

Pith citing papers

Observation 94a3833b-a525-47d0-b967-5d248b0f689d · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.309833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:0c563630999a3b0189c11b66032e0a946e6f222374467a3e72c535de1094d918

Observation 57530e12-cda5-41fe-bf7f-6574601f213e · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.639836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.639836Z digest=sha256:4ab189391cd68da460f7750d3ec0e38c37c95531fc37f69cd91e8afdaef42cff

Observation e80c223a-4010-4c5b-ac75-717ba7a021e7 · inbound

Sudoku-Bench: Evaluating creative reasoning with Sudoku variants cites this paper.

Sudoku-Bench: Evaluating creative reasoning with Sudoku variants EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:30.334466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:30.334466Z digest=sha256:be1357345ff4ac276e41630d548413beac344750907d8fd7e5be52413c780ebc

Observation 3b4b129e-89df-44b5-ae33-826b8ebefe91 · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:47:15.164738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:19a4b51e65e6d0588b9e2d927246ef96872ad8c1abf160da3a9c01ef0542f127

Observation 048dd1b5-73d0-4c8c-aeca-b028b91d41fe · inbound

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation cites this paper.

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.930689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T08:13:05.328746Z digest=sha256:147fab239a55d7e0945912f518fa261d6470288ceecb8ccfbb326c52b53289df

Observation 8ee27dfd-b016-448f-9762-eac45709a9c7 · inbound

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning cites this paper.

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:39:47.813355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T05:35:32.806871Z digest=sha256:f4316f264b491ec6d69b793974e30e2b270be9c7bc0dcf6795eba5db797fcfc5