Pith. sign in

Paper Citation Record · LEDGER

lmgame-Bench: How Good are LLMs at Playing Games?

As of 11 August 2026, this Paper Citation Record lists 100 of 115 outbound references and 22 inbound Pith citation observations for arXiv:2505.15146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15146 v2

Coverage vector

measured 100 of 115 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:27:03.187728Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.015304Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 115 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved81
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a7e70288-ebc0-4e09-82a6-f0ea6f228c64 · outbound

This paper cites OpenAI Gym.

lmgame-Bench: How Good are LLMs at Playing Games? OpenAI Gym

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:55.977387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:55.977387Z digest=sha256:8d55398489b8523e71679a19434ae8c69073471b526e14c28fa1b3a56fd174e7

Observation 131c4169-85b2-400a-865b-adda2c05762f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

lmgame-Bench: How Good are LLMs at Playing Games? Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.072057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.072057Z digest=sha256:40908410aefe86cc72270f479c66e53ff2476312d8d570121fc39bf63762edf4

Observation ba224b0f-9985-4a21-bb0e-6fa7100a6c5c · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

lmgame-Bench: How Good are LLMs at Playing Games? RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.140833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.140833Z digest=sha256:323fa1b78cadab0bc28dabfe297146f780f50b8c7cb43fd4b65d5bcaa38a34dc

Observation 14a9e272-3755-4f82-8ebb-4c5d0b603de9 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

lmgame-Bench: How Good are LLMs at Playing Games? SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.255061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.255061Z digest=sha256:3de208a9bb9b73809babb622041ed4e27513f586c59282604992da82d8b6e2ee

Observation 1d9d2fe6-ad2e-4a8c-ae80-807da734ae0b · outbound

This paper cites In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., eds.: Advances in Neural Information Processing Systems.

lmgame-Bench: How Good are LLMs at Playing Games? In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., eds.: Advances in Neural Information Processing Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.363888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.363888Z digest=sha256:7290f23059c849addd6eb4cdd614925d0605b06aedd61920ae963c03cea37d4b

Observation 6d0d7fbc-84b1-43ad-a74d-6b37e571ae55 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.453904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.453904Z digest=sha256:1dd46f1bfa8808db32c92e60cfed3f4abc0a4749c51d6fddd335316f2eeea4f1

Observation babc8d0b-adff-4d21-9ae3-ee3d6a759645 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.552570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.552570Z digest=sha256:109eb338015f8f44a6c29b8f6a882526b1939c7b186c96d6fa0a7b099b4ce018

Observation 57530e12-cda5-41fe-bf7f-6574601f213e · outbound

This paper cites EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges.

lmgame-Bench: How Good are LLMs at Playing Games? EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.639836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.639836Z digest=sha256:06fc206efeb0d05a9b14eb2012be0404d402768a7b45ab7eaf88a11a6118c4cd

Observation a0ca5e66-1c30-491b-9df9-19c7522ef01d · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.720112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.720112Z digest=sha256:41601f1149d4b0470821a97f9f92f20576b8351cfa4b7400996208fafd48bffa

Observation fc84ddec-7530-45b2-aa78-08091fdbc4a6 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

lmgame-Bench: How Good are LLMs at Playing Games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.816724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.816724Z digest=sha256:76343054c4e5ae3833b79a7f9c7a847424e56f88324b1a0d68da82d86b80a24c

Observation 7e38bd61-6f22-4a25-8709-7c5227b6d909 · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

lmgame-Bench: How Good are LLMs at Playing Games? GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.894546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.894546Z digest=sha256:52425c952b553fa19b868ae698d94aa2aa8dd7cce4f00a546e08f4bffb78242b

Observation feb50f57-f0b3-4573-880d-350a08c99f3b · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

lmgame-Bench: How Good are LLMs at Playing Games? SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.020933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.020933Z digest=sha256:1d92aabea27cdb8fe51d22ea6e285a322c11f9a6f8687578138092c6c4ffdebf

Observation 6aa6771e-9efd-40a8-8ccd-f8cd596c8f53 · outbound

This paper cites IEEE Transactions on Games11(3) (2019) 195–202.

lmgame-Bench: How Good are LLMs at Playing Games? IEEE Transactions on Games11(3) (2019) 195–202

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.127873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.127873Z digest=sha256:dcf09d9aa6fb3a49bb8a80596058fb43411781f85667593522e11db96678cf68

Observation 97837f3f-1fdd-4b8d-9c47-82e019956ede · outbound

This paper cites AI Magazine22(2) (2001) 15–25.

lmgame-Bench: How Good are LLMs at Playing Games? AI Magazine22(2) (2001) 15–25

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.204676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.204676Z digest=sha256:5d00c0e381d3674a7e2706174bf8ad170dcf8a3a4dcaae22863cc136f23abef8

Observation 14041c05-4acb-4b6b-bb51-360566b36a2b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

lmgame-Bench: How Good are LLMs at Playing Games? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.341174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.341174Z digest=sha256:bba950304ee666d812281b18905b62f7da0a95de93e8506e8556f95c90c96d78

Observation 2607c706-b810-4f3e-a02e-76c5f1a9e56e · outbound

This paper cites Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games.

lmgame-Bench: How Good are LLMs at Playing Games? Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.413647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.413647Z digest=sha256:ead902b19e1792b5647e1da2ac4a20f36acfeeef3d4d8081bd1b700f85793c5d

Observation 42618b84-232d-4f90-b0d7-8e954d3735c2 · outbound

This paper cites Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot.

lmgame-Bench: How Good are LLMs at Playing Games? Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:27:03.886471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:26:57.507753Z digest=sha256:3529092025726611e84ef6606091c32bc4030760956110de7a6ca3333cc81f59

Observation 0e90353b-7af8-4f73-8b05-213dc2a49d13 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.604725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.604725Z digest=sha256:1912417122a40c80125fcd9f28efa87fe5b8fabec4666126357a996ebb982002

Observation 3c06fd2b-4bce-456f-ae46-6b275c790620 · outbound

This paper cites OpenAI o1 System Card.

lmgame-Bench: How Good are LLMs at Playing Games? OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.709482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.709482Z digest=sha256:d0de09aa5affd6839e26c44e549c6681b7404a2cc1003740b25d6b02c666cb0f

Observation e2613515-e9ef-4538-b0cb-848ce7429f6a · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.792901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.792901Z digest=sha256:09eda6388a36b92860f7b1c1db4ea06c1168122db1c896cb8ac590738779bbc3

Observation f282f996-6b9a-4e1f-8578-ad4a26aa3d85 · outbound

This paper cites In ICAPS.

lmgame-Bench: How Good are LLMs at Playing Games? In ICAPS

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.864760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.864760Z digest=sha256:5fa127b4a5b9151dcc41cb965dc5b83be6b94fff02f5262d6e6ef2d44a7b35f5

Observation 5ed9dfa1-d833-40fa-af0e-82ebf70154f0 · outbound

This paper cites Applied cognitive psychology31(4) (2017) 438–445.

lmgame-Bench: How Good are LLMs at Playing Games? Applied cognitive psychology31(4) (2017) 438–445

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.936132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.936132Z digest=sha256:2b52d343bdcb362881c088cdf047d9e30effa5892c774cc099782cb46a9d361e

Observation a3791a52-c555-4ec5-bb2f-a3ea2101db68 · outbound

This paper cites In International Computing and Combinatorics Conference (COCOON).

lmgame-Bench: How Good are LLMs at Playing Games? In International Computing and Combinatorics Conference (COCOON)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.046424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.046424Z digest=sha256:33e58bf522cf27e70f06706f5c3fee0b4f25fc0007aa419daf5ebc98daec8b27

Observation a661cdd9-5b21-433f-b165-7d7818203605 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.163042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.163042Z digest=sha256:9d54e5bf35a81d516a4d4dd53d083727f6220e82731838533d4f5449d96b8092

Observation 458995a7-b87c-48cf-892d-495c8af4cd2b · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.220901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.220901Z digest=sha256:ed150d70f9cd5d43c1901c9353a3daba723f5a1fc8dc76e9e596d1b29847c467

Observation 68d50081-a23e-4b11-ad68-4a9665b646da · outbound

This paper cites arXiv (2024).

lmgame-Bench: How Good are LLMs at Playing Games? arXiv (2024)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.280311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.280311Z digest=sha256:c38943cc6277e3ab7c04e8f3bf2ec690de6c50ddba1a5b8726f9b6661d5b91c6

Observation eadfffa4-9db8-4542-9f36-cc139a0416a3 · outbound

This paper cites arXiv (2025).

lmgame-Bench: How Good are LLMs at Playing Games? arXiv (2025)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.353668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.353668Z digest=sha256:4a9116897b55ab077978ea9c6ae1a54a71a2dd5844c27992207f64f18d17537c

Observation b9973ff4-17fc-4c65-a76e-ebac676753b3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

lmgame-Bench: How Good are LLMs at Playing Games? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.427826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.427826Z digest=sha256:bc19de4a21d6971e46e8ce357b8e711522c71a0113fd47eb220fed35efffc235

Observation 24d32840-27ff-4a3a-95c4-4e9206bb84f4 · outbound

This paper cites Computational Geometry 13(4) (1999) 215–228.

lmgame-Bench: How Good are LLMs at Playing Games? Computational Geometry 13(4) (1999) 215–228

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.487599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.487599Z digest=sha256:f5faed9ee41ff880e70bc4644913ad29781cd74fa1176514b4fbb9959f9c803d

Observation 315386b0-a926-4c90-b7d8-da2b71f4b8b2 · outbound

This paper cites A Simple Family of Analytical Trumpet Slices of the Schwarzschild Spacetime.

lmgame-Bench: How Good are LLMs at Playing Games? A Simple Family of Analytical Trumpet Slices of the Schwarzschild Spacetime

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:27:03.824241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:26:58.543471Z digest=sha256:04d2872f3cabc9a59106f617aae41b8d618bb3648f96336e5dda81b8aee8d372

Observation df103d95-f26f-4379-a55b-5e402af91911 · outbound

This paper cites Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models.

lmgame-Bench: How Good are LLMs at Playing Games? Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.602648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.602648Z digest=sha256:a93faf08619fe6354f4c57d4a5d4cf4352cfda14d9977415f6903922dca0b8f2

Observation b6f11591-6ffc-4ad2-a2a3-a1758280f791 · outbound

This paper cites The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks.

lmgame-Bench: How Good are LLMs at Playing Games? The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.673836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.673836Z digest=sha256:bcb0a0e7281f9b5ebc86e535b36d397bed7f36ad3abd6ed3bb061f20f4abb992

Observation 779eca80-188b-4c50-852e-178d48195594 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.734689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.734689Z digest=sha256:08d495f9229aa89824ebe010de421a40d5005fd7225642ced1ac64c9bf286071

Observation c90c83f6-52ab-4545-84d1-3304d47f63c3 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.819840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.819840Z digest=sha256:03b5aab34c5745fbcdb3ac15655f1eb276b0c4a3ba9f9ae9600b6ebc70a65c90

Observation 5a38610f-7acc-42ac-a56d-c9ab3c9d4908 · outbound

This paper cites Cradle: Empowering Foundation Agents Towards General Computer Control.

lmgame-Bench: How Good are LLMs at Playing Games? Cradle: Empowering Foundation Agents Towards General Computer Control

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.872288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.872288Z digest=sha256:f6eb7ea7f4df34af3a4fc909a0ae02678d8bd6e87b8be004207063fe8e24fbaa

Observation 091a4179-5910-422c-8688-a32d3cb84c84 · outbound

This paper cites In The Twelfth International Conference on Learning Representations.

lmgame-Bench: How Good are LLMs at Playing Games? In The Twelfth International Conference on Learning Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.941657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.941657Z digest=sha256:4f0853c7708340af1d4e7469d015c26ad71617b0c8285a6e382f81df1f928e51

Observation e14d8a20-afe8-440c-b5d0-b4b75a82f599 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.038784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.038784Z digest=sha256:bd7c9fdb9aeab9861f38a64cb55fee8581f2536f56af010890df18417debf516

Observation 1f4eaf26-1294-4981-812f-a1f1136a981e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

lmgame-Bench: How Good are LLMs at Playing Games? Measuring Massive Multitask Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.115491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.115491Z digest=sha256:9f583b01c92fc865e0514a1da30fabfadde0c9075efc47d8543564b61c869ff2

Observation 78264078-8bb5-473a-a8db-9465570d5a75 · outbound

This paper cites Humanity's Last Exam.

lmgame-Bench: How Good are LLMs at Playing Games? Humanity's Last Exam

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.189173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.189173Z digest=sha256:f6a0e5f2e2d08d185f9fbaac359b3b43e44e2a6ea7db9fa43c9cadad3a0fe293

Observation 5215fa44-fcc9-4428-ac26-30c80fcaa06c · outbound

This paper cites https://scale.com/leaderboard/ humanitys_last_examAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ humanitys_last_examAccessed: 2025-05-14

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.256804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.256804Z digest=sha256:669a1cf859d0ffb20d2d896c14ce3a19fc3fda1e75dc51246b6cdee46a5a5ca9

Observation fa21a2a7-b374-4ee5-8416-5e1e6fa90698 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.333175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.333175Z digest=sha256:cf8bfb941000c141ec95d3ff734d7a35b4ecb8b5df4373b7aa950d8ea5b54021

Observation 05197398-9457-479b-8f9b-c55b05938893 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

lmgame-Bench: How Good are LLMs at Playing Games? GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.405000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.405000Z digest=sha256:cfbb5df192a224ddbabc18f474531bb2fef76ada747634ce8f05cf84d2db815e

Observation 18401cb3-ae82-436f-a734-1c1b8cf03368 · outbound

This paper cites https://www.vals.ai/benchmarks/ gpqa-05-09-2025Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ gpqa-05-09-2025Accessed: 2025-05-14

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.464341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.464341Z digest=sha256:dbc29304e1897f8b89241189109f2903b2e946c3dc9d2e08324c1f7d566d6ef1

Observation c0b9f034-2e97-42c1-bb1a-d5525b541192 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

lmgame-Bench: How Good are LLMs at Playing Games? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.524310Z digest=sha256:3772f4b3a4a9986ddeb7243fe9ed4a37560f9b8d66db0f6466f74642c820e27a

Observation 8cafd74e-8f58-4fde-a158-02ea21ec09c9 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

lmgame-Bench: How Good are LLMs at Playing Games? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.566161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.566161Z digest=sha256:6427d3a656dc0f0cc3635c9c04321c2143246224e386485c6e1427fdf42d9596

Observation cce49f77-c6ae-4fb2-aab9-75b4ec1248d6 · outbound

This paper cites https://www.vals.ai/benchmarks/ math500-05-09-2025Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ math500-05-09-2025Accessed: 2025-05-14

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.676392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.676392Z digest=sha256:56c2e190c5f949594226d094febe467f2ad1f4e7a624b16cdb4238f0ef4f16e0

Observation 90b9f4e3-156d-4664-aee1-57fb91077fdf · outbound

This paper cites arXiv preprint arXiv:2410.03131 (2024).

lmgame-Bench: How Good are LLMs at Playing Games? arXiv preprint arXiv:2410.03131 (2024)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.788838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.788838Z digest=sha256:0a2426bc277e6c19d73b948a1ab98a8b4528510b3f7795efded1d56a4556cf80

Observation fd08dd1e-6fbf-48a1-98b3-78fd6a87369b · outbound

This paper cites https://www.vals.ai/benchmarks/ aime-2025-05-09Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ aime-2025-05-09Accessed: 2025-05-14

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.975697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.975697Z digest=sha256:6c15f89eccd8d05fa8ee6550dde4423af133bfee6bd30e568efcb5f6e77498e1

Observation b2b1ae1a-b149-46a6-8cf8-7810df23e33b · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

lmgame-Bench: How Good are LLMs at Playing Games? LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.152304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.152304Z digest=sha256:baf12e01009654371a0132b0b3104fdeacca4e2750c7be0cc60118d5f033c244

Observation 57ad3dd4-a046-4304-b642-c79511ba15da · outbound

This paper cites https://livebench.ai/#/?Coding=a& Mathematics=a&Data+Analysis=a&Language=a&IF=aAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://livebench.ai/#/?Coding=a& Mathematics=a&Data+Analysis=a&Language=a&IF=aAccessed: 2025-05-14

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.352797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.352797Z digest=sha256:03bb4fa09bc4d0719aeb24c5bc83ad92aaf5d4bd471b3e85a1d5a588878d7ef5

Observation 57249563-1ef7-4de2-a8ef-873d95490edd · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

lmgame-Bench: How Good are LLMs at Playing Games? BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.559765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.559765Z digest=sha256:f2c13d3a9ef8f6f5344608b9acc34ccb51e59ae32554f99a93fb3939469d8294

Observation 417548ee-d2f4-4b93-a41b-3768f88ef4a6 · outbound

This paper cites https://aider.chat/docs/leaderboards/ Ac- cessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://aider.chat/docs/leaderboards/ Ac- cessed: 2025-05-14

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.736996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.736996Z digest=sha256:e5d165dc233c29cff9f3cc1acffcc774a12aab9f7ea73f88843b2a308a5e3dfa

Observation f74dbb8d-05cd-47d5-805c-429e6fb2df25 · outbound

This paper cites https://bigcode-bench.github.io/ Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://bigcode-bench.github.io/ Accessed: 2025-05-14

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.875924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.875924Z digest=sha256:cbe7b3feb20239be26e0c1f705f84954ad02ca847b5d6716f7bab92a4748b4f2

Observation 5f981890-9fcb-428b-ba47-a854b244afe3 · outbound

This paper cites https://scale.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.977620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.977620Z digest=sha256:4e14364410dd77d023e4ee1d0f3c10c7443c2cd4852f609c3924e41b37f2b20d

Observation 1b3e83b5-e1af-4700-83e5-b546bd993e2a · outbound

This paper cites https://lmarena.ai/leaderboard Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://lmarena.ai/leaderboard Accessed: 2025-05-14

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.981839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.981839Z digest=sha256:6aa8da61d0756ed3ae419086be9bf851edd930903621f5fb2cbfd0f8d3a24ab9

Observation c8b4e03f-6aa8-40d0-9580-048940736906 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

lmgame-Bench: How Good are LLMs at Playing Games? MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.986636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.986636Z digest=sha256:403ec9d28c05de0b175c851f432e696eabbf12ad95872756a08be5c096b24953

Observation 46b6f099-93db-472a-bbb9-5991bf80b0b2 · outbound

This paper cites https://www.vals.ai/benchmarks/ mmmu-05-09-2025Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ mmmu-05-09-2025Accessed: 2025-05-14

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.031667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.031667Z digest=sha256:f98d279ab7088031c2e1c2aa9894f1770d022df45737640557fd91dfeffd74f2

Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.110876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.110876Z digest=sha256:fe9bcd50bb1a572712f0e4e90a1e9c752d4cc4bb66946a529387886835c6eef1

Observation 3e400887-44cc-4397-99cf-e6dcb36ebaa9 · outbound

This paper cites https://scale.com/leaderboard/ multichallengeAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ multichallengeAccessed: 2025-05-14

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.208322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.208322Z digest=sha256:8cab7150651aa32bbaaaf1c8ee8c7e487ff2542bc0097bf3f4cd4f2fec2d9155

Observation e543a205-cfd9-4a73-8081-570c10c774b0 · outbound

This paper cites https://scale.com/leaderboard/ enigma_evalAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ enigma_evalAccessed: 2025-05-14

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.299976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.299976Z digest=sha256:d2068b1620c0e0c6b4b0c00f2017bd288dae94d878d97c451aaf91fd1e152cb5

Observation d1b43e71-9dfc-4115-b0db-dde5f55ad797 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.362623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.362623Z digest=sha256:e2d3909856424a270c360968ba117b29429911b98a022d3eacdd9b40399a9cf6

Observation 9b944751-7bca-4f85-a541-36dd67dd2b0b · outbound

This paper cites https://github.com/mpSchrader/gym-sokoban (2018).

lmgame-Bench: How Good are LLMs at Playing Games? https://github.com/mpSchrader/gym-sokoban (2018)

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.442781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.442781Z digest=sha256:5434ccaf328a088ab438a7731bea67371731ba2ff6cb0d8bd3cf79474faaebcc

Observation 942416c6-c81b-4b08-a263-8b1b5a540611 · outbound

This paper cites https://github.com/jaybutera/ tetrisRL(2023) GitHub repository.

lmgame-Bench: How Good are LLMs at Playing Games? https://github.com/jaybutera/ tetrisRL(2023) GitHub repository

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.626893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:01.518032Z digest=sha256:72c99b43691859864bc2820bb684f0e00f26eb8bc25e04d4a6ed4767a3c5f364

Observation 086da939-2779-4dbc-9d8a-e528b44e84a9 · outbound

This paper cites Qwen2.5 Technical Report.

lmgame-Bench: How Good are LLMs at Playing Games? Qwen2.5 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.523920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.523920Z digest=sha256:69db724d9cb852c88cb5712bc7afee7ed356224d568e433b560a43589c0c0813

Observation cf891336-6b2c-40d8-abd9-19785f67c849 · outbound

This paper cites Communications of the ACM 38(3) (1995) 58–68.

lmgame-Bench: How Good are LLMs at Playing Games? Communications of the ACM 38(3) (1995) 58–68

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.609837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:01.528538Z digest=sha256:f628c045b4f341be45dbeb349321badc21054f99e98a5fe8c5e9175887cece3f

Observation b0d4f444-27e6-4395-a8e0-5ae73574d782 · outbound

This paper cites nature550(7676) (2017) 354–359.

lmgame-Bench: How Good are LLMs at Playing Games? nature550(7676) (2017) 354–359

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.592570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:01.544702Z digest=sha256:09f2ab04e238cac3b54db2b676a1a50d720effed8dabdb1c4e485fc37b707bc2

Observation 04da8a84-7fdb-4484-9b9a-ef622d288117 · outbound

This paper cites In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.

lmgame-Bench: How Good are LLMs at Playing Games? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.576974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:01.617277Z digest=sha256:7b58e9ff8dc28162e211f59b83354df87cbbb1201df008785508b5e501434d5b

Observation f9065042-7769-4785-a445-20b67486def2 · outbound

This paper cites Factorio Learning Environment.

lmgame-Bench: How Good are LLMs at Playing Games? Factorio Learning Environment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.689046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.689046Z digest=sha256:a46c7a5ced47df5ad4fd115e93720fdaa7b006fc650f4814f13da9b42c097f47

Observation 9a6fa030-734c-47ee-ad2d-e8c8af2764d1 · outbound

This paper cites In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.

lmgame-Bench: How Good are LLMs at Playing Games? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.560892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:01.766369Z digest=sha256:94687f98d27052bc43e728f66960dfb6daa785d8c96457cc886054caf9b307ec

Observation 16b11839-c133-4a36-bcd0-05ae76aa1262 · outbound

This paper cites TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning.

lmgame-Bench: How Good are LLMs at Playing Games? TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.826741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.826741Z digest=sha256:ed2ce2284a5104261a25de86e9c82a1923d84f37b662b69d20188bc4ca8d924d

Observation ef4dd5ca-d0f2-4830-a96b-26312327dea1 · outbound

This paper cites GameEval: Evaluating LLMs on Conversational Games.

lmgame-Bench: How Good are LLMs at Playing Games? GameEval: Evaluating LLMs on Conversational Games

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.892670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.892670Z digest=sha256:dcb9da43a41633c776ec4d6035ed0bac71e52a70dccc5e259eb24f586a77905f

Observation ff01265b-d150-4cf4-bf94-37cd01ea52a9 · outbound

This paper cites GameArena: Evaluating LLM Reasoning through Live Computer Games.

lmgame-Bench: How Good are LLMs at Playing Games? GameArena: Evaluating LLM Reasoning through Live Computer Games

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.953231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.953231Z digest=sha256:0fe85bed10b7081f2461e833644a90af973535b2faef2bc18bd63a8f8ada6efb

Observation 6a8c9fc0-bf4f-4d18-a4ec-336252fdc380 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

lmgame-Bench: How Good are LLMs at Playing Games? SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.019566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.019566Z digest=sha256:70da901182c5bc4d8051172a1a99a312ea690ce37a05111b00096228cc6214a6

Observation 93c0d409-ab27-49cb-8248-b6a7e5e39c44 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

lmgame-Bench: How Good are LLMs at Playing Games? WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.059918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.059918Z digest=sha256:35fb0619568456da3326587b0a42f6e9be3604a1b44f9ea7e9539bb932e67d94

Observation 71003540-eb94-4f0a-8628-80fdb703865a · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

lmgame-Bench: How Good are LLMs at Playing Games? WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.065525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.065525Z digest=sha256:699d0a66b5506f144439ecd8b2f6890132519787e404199edfe48740b388b251

Observation fbbce248-f0d4-44ce-9218-ff4a600c30a9 · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

lmgame-Bench: How Good are LLMs at Playing Games? AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.082501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.082501Z digest=sha256:eb6cc462bf67f05ffc7b9d3c412834be42644e8c089cbaa7b8d06502b899ad33

Observation 55c10a9a-0b26-4173-92d4-352510af82cc · outbound

This paper cites Advances in Neural Information Processing Systems37(2024) 52040–52094.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems37(2024) 52040–52094

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.545710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.131440Z digest=sha256:9dd974bf3cc6cbfdd3d6a1d3410cd485a3c012c4193f58961eeb52594ae7563f

Observation 23603768-7adc-4214-82f5-134ebcf65371 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:27:04.528964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.208401Z digest=sha256:b3ccd664a122cc53fa3d1579a7cbbd16c5ffa5b2002f87833cdd7b1ec295390c

Observation ab6ef914-7e02-48fc-ac4c-3a0923f26339 · outbound

This paper cites In The Twelfth International Conference on Learning Representations.

lmgame-Bench: How Good are LLMs at Playing Games? In The Twelfth International Conference on Learning Representations

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.513351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.293383Z digest=sha256:ec85ef0cb343d0d969a5990c1de4c697a02a98920d15f5efa42e3436732a721a

Observation ac079a74-203a-44ca-a45c-7e4e7b81e5a3 · outbound

This paper cites Advances in neural information processing systems37(2024) 110935–110971.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in neural information processing systems37(2024) 110935–110971

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.497010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.328494Z digest=sha256:8a9b5aa85faffe6b3d057051c633284d76aa6aaa57ccba64a07f3e69f3640867

Observation dd33cdcf-25ab-4ac1-8e0a-508f1a56a9f0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

lmgame-Bench: How Good are LLMs at Playing Games? Proximal Policy Optimization Algorithms

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.348919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.348919Z digest=sha256:f2efb88de221fec016c9a3e52b971505a5b636f2c2b48940b7f8be63318bc5ef

Observation f23d2550-8302-4b52-9bb2-70b39c762a73 · outbound

This paper cites Science10(3) (1995) 237–304.

lmgame-Bench: How Good are LLMs at Playing Games? Science10(3) (1995) 237–304

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.478783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.397270Z digest=sha256:60cfbb4fa235d718d1aea80d7009f3a63396f69b965e90f2f1e8ef0dfb214714

Observation d17c69fb-0447-4821-94ee-8b7e972aff5c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

lmgame-Bench: How Good are LLMs at Playing Games? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.445271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.445271Z digest=sha256:a00e33ad0647adc36be6e9209b471b93361351d32b75bdc4b715375de8db8e0c

Observation bc3d9d72-699d-4e55-a952-6d4994acc0ff · outbound

This paper cites Advances in Neural Information Processing Systems36(2023) 38975–38987.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems36(2023) 38975–38987

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.461822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.509094Z digest=sha256:7b14d7377e690642db48f075506f5118af51a450232a6d2b6a434c5bbcf92027

Observation 87ddd352-621e-4aab-a607-a901a790b5ca · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

lmgame-Bench: How Good are LLMs at Playing Games? Training Verifiers to Solve Math Word Problems

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.575808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.575808Z digest=sha256:7637cab593aebfbbc2462d8b065f4156908277c62970c8e1922bf7ad332ebf1b

Observation 102da256-cfb9-408f-8b4a-80014645a6b9 · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems36(2024)

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.444914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.602474Z digest=sha256:f39a93c428da959066ca54111276dbe24256ce93aae4588377e4da60ca2865bd

Observation 25bdbe78-156d-4464-be03-03384ac850ac · outbound

This paper cites Advances in Neural Information Processing Systems 35(2022) 20744–20757.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems 35(2022) 20744–20757

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.426651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.606960Z digest=sha256:ab984246067616407b0513821ce2dfefe6673c1bafa9c8782757eefb99dcc5c9

Observation 82404a63-dd5e-4dc3-bfea-47ec0b60f14c · outbound

This paper cites Biometrika30(1/2) (1938) 81–93.

lmgame-Bench: How Good are LLMs at Playing Games? Biometrika30(1/2) (1938) 81–93

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.409726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.679885Z digest=sha256:92b2dd78d2b696cd8df8c2d81783ea6a9d7f778bb9e1d155bacbf41b7f0959e5

Observation 49d0ee1f-ec40-4b19-b216-87c56e6c6c61 · outbound

This paper cites In Proceedings of the 19th international conference on World wide web, ACM (2010) 577–586.

lmgame-Bench: How Good are LLMs at Playing Games? In Proceedings of the 19th international conference on World wide web, ACM (2010) 577–586

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.392662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.740913Z digest=sha256:c4f5b9ef2c1d871af121e9824c336cfd5e2886606c361ef0da31755178ba7334

Observation 4f03dc3a-ea7e-47b1-8691-d7281719da5a · outbound

This paper cites Educational Researcher5(10) (1976) 3–8.

lmgame-Bench: How Good are LLMs at Playing Games? Educational Researcher5(10) (1976) 3–8

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.376853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.809560Z digest=sha256:f29909d643f4cbb2579baf00fdf56c964a6c84cbb0d06e6bd1f5f56834410b38

Observation fb942280-1434-4e92-8e42-927df2631386 · outbound

This paper cites Block 1 is on top of block 3, block 3 is on top of block 2, and block 2 is on the table.

lmgame-Bench: How Good are LLMs at Playing Games? Block 1 is on top of block 3, block 3 is on top of block 2, and block 2 is on the table

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:27:04.361377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:02.852880Z digest=sha256:83bf68a8bd03ddbfa67b8c435c4f5304f7a05a70b50df518b4f2e3ef757d1b55

Observation 616a293f-add6-4048-8061-6e2fee8b6378 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.911988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.911988Z digest=sha256:33e2524a24d0d8432a89b812d1121b07c8771334a5861e62866d1cd41af20b84

Observation 9dda3d17-5a70-466d-8d56-55981e5b0ef1 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.950794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.950794Z digest=sha256:dc0ab5f04bb99b30870fc86c64b7f0b7998f3c3db4a205800635e96db4ca58b2

Observation e1830c7f-0d00-443b-9b70-32409578e2b0 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.032310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.032310Z digest=sha256:a97e4091b35f60649388612643d58a8647dd1b07e6a50d2e8465ccff9650261f

Observation 56d4f2d6-0d9a-4a9e-8129-55c3b56a358f · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.119984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.119984Z digest=sha256:79bda64b8e28952b9e0d0159611427a7f4dcd9acdd9f0c4c779b811128c64477

Observation fff924be-933e-4c74-bb34-f0059920b182 · outbound

This paper cites up", "down.

lmgame-Bench: How Good are LLMs at Playing Games? up", "down

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.304373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T15:27:03.163759Z digest=sha256:0a140508fca2aaa405dae23654ebc50ec45388aa5252d5aff7e830cc4c6317f4

Observation ce3e4880-c2cc-45a0-ae20-063acc143f72 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.171436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.171436Z digest=sha256:35b7935342922bbc4c89a6d92d13408055d53b53994abb227784bc4408398452

Observation f08e1eca-d190-4a28-b8c3-50d5a956d03c · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.177177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.177177Z digest=sha256:68337c19f8c763f6a713b0c7ce2d68df932a2544611ba385c812cd06db22538f

Observation 1025f4c9-f81b-4b4d-b541-07f78456c050 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.182638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.182638Z digest=sha256:a485400d1322831cb4dcba6729b9e8d068e3ba7ef65ca77d4729b136af7dc739

Observation c58fb195-9556-4c8d-ae4b-178daf346b75 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.187728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.187728Z digest=sha256:9fb7148afcc9cdde0b5ae2b5f37ff2c47e6fa347fe5f1f92366ec1a1ece9c828

Pith citing papers

Observation 6eff8939-1f26-465a-9aab-d505c04f2523 · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers lmgame-Bench: How Good are LLMs at Playing Games?

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.015304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.015304Z digest=sha256:6cd394cbe2aeb3bda09efb63b957e357a738994f896fbb72050c7f580d67eb26

Observation 6b6def62-bd00-4966-905d-2beeba9e4b8f · inbound

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play cites this paper.

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play lmgame-Bench: How Good are LLMs at Playing Games?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:33:42.500139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:33:42.500139Z digest=sha256:3482d7075efac449a77e6076e4a672fbda7cfc4cafc53fde44737c6e64bb9a0a

Observation c87a4664-3ed6-421e-bf53-762d56b3599c · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.184183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:8870bb08da98810c11a2b79deac783f531c0a290d449188eff4ef25731ddecd4

Observation 27ec31b8-d11b-4952-889c-a66ceb5987ad · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models lmgame-Bench: How Good are LLMs at Playing Games?

Reference 203

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.526497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:cf2ab83f39a15791d4b9f498f5b35c3a0fd7140d2708fff144f33c1c3b4387e0

Observation f8567ab9-4edc-49b5-a74d-036f1e741261 · inbound

Gym-V: A Unified Vision Environment System for Agentic Vision Research cites this paper.

Gym-V: A Unified Vision Environment System for Agentic Vision Research lmgame-Bench: How Good are LLMs at Playing Games?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:05:26.044132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T10:05:09.049846Z digest=sha256:c622086ba9d200f611355d2f884c1008342aca0ed8020e74a68a8c68c0fba6dd

Observation 5f80c1c4-4c08-452c-a764-51911f301c29 · inbound

TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs cites this paper.

TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs lmgame-Bench: How Good are LLMs at Playing Games?

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:16:06.152542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T02:04:57.744754Z digest=sha256:2e8ce9989c2da6e3d5ee14cd44b2b19ae2e125ad5f4cd770cce0c3500ec8b90d

Observation e18a66e5-bba2-4612-b96a-11391a6046d7 · inbound

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks cites this paper.

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks lmgame-Bench: How Good are LLMs at Playing Games?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:14:46.428384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T00:14:07.017420Z digest=sha256:bfe8a3d0c5354baa9e58004ab4fa53839cd04405bb6c5bea261d9df188a90382

Observation 1fb4d952-c576-432d-96bc-4165ca85794e · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.364638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:f2d3175e7bac873289bb337adb16a652e5dbe6f54a87c670242bfbcfa676b0aa

Observation 70f766f4-694a-44b7-b27c-a41a16861d3c · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:07:51.419743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:05:36.511150Z digest=sha256:e9316ad5023f2e0b2b0793912664dc4af1e4453a9c85a2bdb6e2393223312e03

Observation 5d47647a-68a4-4f89-984d-0df532ab97c7 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:48.141281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T05:59:44.669877Z digest=sha256:569ab0e04dddb860016e9f397315103b181b10c1f17e71659ca6530a97a3041a

Observation 1d368f40-3266-41ea-8cdd-8b9fb6fa95df · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.559898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:45c11ed76ff4ff7da4ad1c29355c532f735000a1aa53dd55b5f385c7ba381bb0

Observation cfd7b7b2-11a5-41eb-8577-10df4fcd792b · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games lmgame-Bench: How Good are LLMs at Playing Games?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:17.261880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:5194faf5aad413cdd686f62ad03648a4924b8e438fa107ba2697e6c37045a5cb

Observation 2049d76d-8774-4384-a78f-ba5c6d631ed4 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games lmgame-Bench: How Good are LLMs at Playing Games?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.221897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:192190a0b7f4a5e827b4b4a91c5e3fa376feabcc5dc5dcc95eb65869460303c6

Observation cb5202f0-d7ec-4fae-b9ff-6d58f9651a63 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs lmgame-Bench: How Good are LLMs at Playing Games?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.196706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:4041afc7617dc3f51964e7cd079f4b6d03889bb61cc06172aa0b78a6f8544acb

Observation 23f86099-b0a4-427e-800a-d1b8d9bdeb5f · inbound

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? cites this paper.

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? lmgame-Bench: How Good are LLMs at Playing Games?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.593221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:21:49.763994Z digest=sha256:4f7dd31d05065dc36be457ac2c61ffc46cbeee721c82bfc4d8a1449c365de822

Observation cad8132a-3e76-4034-aed2-e4abc28a1d66 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models lmgame-Bench: How Good are LLMs at Playing Games?

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:11:28.899789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:f90feca4b127bf99d09daf28a20b1d6c1b2ae8a83f8d42ca1dcaba070229c4b6

Observation f4c3ebe9-e250-4a38-9fe2-063b44fa6a8c · inbound

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics cites this paper.

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics lmgame-Bench: How Good are LLMs at Playing Games?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:29.947294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T16:50:36.194650Z digest=sha256:efa8dc9a86c8d1cabf7be6254753068ab89428a1382b002294b987944f10be81

Observation 09f90fa0-57f8-4c7f-9529-c4a8dd5f4cc1 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application lmgame-Bench: How Good are LLMs at Playing Games?

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.495612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:2e879c4f6e22d4dc13e388fb0db1c7859836bb91f9af200a6f1d2e587de2c9f3

Observation 0be5c38e-8540-42a0-be0c-76e3c213b0e2 · inbound

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models cites this paper.

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models lmgame-Bench: How Good are LLMs at Playing Games?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.299890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T21:02:26.081166Z digest=sha256:6137bd8327e6542e595792890113b6cff6f00094918f05087945098fd0160f22

Observation 914f01fc-110b-4bcd-9f5a-dbb5aad4ede8 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games lmgame-Bench: How Good are LLMs at Playing Games?

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:19:13.718478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:866489f2bbe55cc352cca6141eacf71c07a34a46f4a9eba88cd8ea48bcf30bc3

Observation 0c2486de-cfd0-4b84-97a8-f3b4a71f6f00 · inbound

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents cites this paper.

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:46:11.447161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T21:59:25.449570Z digest=sha256:277261e1eaaf7c256e86942da4e47ad8deab047d79703813253838cd3bc92ad3

Observation c51221f7-15e7-4980-ac29-d6e0cb128700 · inbound

CAST: Game Solvers as Turn-Level Teachers for LLM Agents cites this paper.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:58:29.283205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:58:29.283205Z digest=sha256:796ea5191f1050b41e2acd7a2247a7421da7cd39ecedb76a43708abff05abf55