Pith. sign in

Paper Citation Record · LEDGER

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

As of 4 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 100 inbound Pith citation observations for arXiv:2206.04615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.04615 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 149 of 149 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 153 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:45:26.750295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T13:57:06.851393Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact35
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 499806ea-bbf2-4490-bbe7-dd3bcd80cd63 · outbound

This paper cites MathQA: Towards interpretable math word problem solving with operation-based formalisms.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models MathQA: Towards interpretable math word problem solving with operation-based formalisms

Reference 1

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.778350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:4665382c3ef1bc338d8224bac3fd098ffbade3d9ca0aa95863878130fae9e65b

Observation b3f33b82-cb8f-4a1c-8b27-ffb40ba3e906 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 2

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.505769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:fc7d70e9d6b046a973f918f2a39e24af789dcc834e1fe6c5801dae0680de6163

Observation 5b4d2b85-d7ba-4ca5-b9b0-fe4bed041872 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.518933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:c6c6bcf2dd1a08da285121fac198d80580c0bf5a1b7eca0b21c8b784b408e97a

Observation fccc22d1-56e2-4d45-a925-29e86d794ee7 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:26:25.529016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:5813129486acbf1f26cc68f6c94184489c770fbec8fd5fd37d5b556f8c53b127

Observation a1c6ef8b-9b33-45a4-9277-21817768be6d · outbound

This paper cites doi: 10.18653/v1/W18-6433.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/W18-6433

Reference 5

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.543638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:6696c3d4f29d149529f906d2a93cb01ace967506af5569546299fee7e8a0bb30

Observation 4146ebab-a66e-4011-adfe-a37e05701b55 · outbound

This paper cites Simplicity: a unifying principle in cognitive science? , volume =.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Simplicity: a unifying principle in cognitive science? , volume =

Reference 6

Resolution
metadata mismatch
doi, observed 2026-05-10T23:26:25.550296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:2088c60d7d9a4ab26673a0cea330302cbc5cab056b222de5a4f2a65a55f3e6c2

Observation 39fd2b3e-a7a1-45e4-835f-59cb7b62045b · outbound

This paper cites doi: 10.18653/v1/W19-3824.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/W19-3824

Reference 7

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.558222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:01f215cb52bb3fa7a81b86f7a7e78bfc2224b7d823bb42cc4a513941bf8445bb

Observation 8a79cc0b-eea7-4202-8267-ea89daf77dd8 · outbound

This paper cites doi: 10.1007/978-3-319-40566-7_4.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.1007/978-3-319-40566-7_4

Reference 8

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.562410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:2d078115fb3f9f773a75aa2bbfc359d9afb4ed2e3ea0f80e8cb545227a82b7ea

Observation 97543141-53c3-408b-a2a8-81373b042ea0 · outbound

This paper cites overinformative.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models overinformative

Reference 9

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.566922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:e600ad0f335f07442ccd8039dc029c6ef52fc09858d04b312c0acdb7949653f7

Observation d0676663-d3ef-47e7-8d71-97e29573918f · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 10

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.574294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:158e2d6b5fb0635f69a4ff38ad2379582b7600157a2d3b1b7e07128b763560f9

Observation bb376dd7-54e5-4a40-be04-f253ea6b370e · outbound

This paper cites Making sense of sensory input.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Making sense of sensory input

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.843131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:d3cb5bdf74728173fd0a7b55ba74e04f1dcb1e4ebd1cb716b1d4d98a0d18235b

Observation 2cf03759-18fd-43ff-8cae-3ad10a984262 · outbound

This paper cites doi: 10.18653/v1/N19-1395.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/N19-1395

Reference 12

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.618260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:3c530653103354b83c649134a40c3206ea85f47e66b57c30ee1ab7cc558390ab

Observation 7ca183db-c8a3-481f-af4f-623814213246 · outbound

This paper cites doi: 10.18653/v1/P18-1082.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/P18-1082

Reference 13

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.622610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:0c9a291e48cc9833ea0045c945c8ada0f4506d332b72442c60f0994c496321f3

Observation f6546cd8-dd12-4765-a18e-80c45a3645e1 · outbound

This paper cites Fodor and Zenon W.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Fodor and Zenon W

Reference 14

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.630445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:b610fc31289240dd79cc2c4c30aa885800debae6e2af8972cd65f934bcb70081

Observation 524977f3-62e1-4c1f-be93-0d118172103f · outbound

This paper cites doi: 10.5555/1625275.1625535.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.5555/1625275.1625535

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.636170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:647069ebf43d9f17698de303d74f165c7833e9efdb78ab6fe04a3fe377f5f64a

Observation b2f4698a-62f4-4395-ae66-d47c8c61a3f7 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 16

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.642596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:4dec11dd747ced917c4fc5ab35241f748dc778afd1d5ca8d1ea28a1f85b1bcc3

Observation 903fb3e9-d637-4c24-bcfe-a7d15167bd7c · outbound

This paper cites doi: 10.18653/v1/N19-1061.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/N19-1061

Reference 17

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.648277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:b19cfe69ae76c30fc44a4066c0eb57a338936e1930e6039f8811c42d92b4a106

Observation 5b1d8f6c-a77e-45e7-99c8-776eda943dbe · outbound

This paper cites URL https://doi.org/10.35111/0z6y-q265.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models URL https://doi.org/10.35111/0z6y-q265

Reference 18

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.658135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:89153e8def3e797747ab6179a139378dd04ef7302f1422ebcccc1eb07d4c2c18

Observation 761b008f-20c5-4bed-98d1-2e7d43ebf516 · outbound

This paper cites URL https://doi.org/10.1145/1925844.1926423.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models URL https://doi.org/10.1145/1925844.1926423

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.663352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:e6a6a1aee317d093a2f985201c25909f363bcd6db7e3809dcc0ec58e1c9e466a

Observation b87c8db7-c871-425b-aa99-3f54658d6a31 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 20

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.668517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:c74bd6d7fd067e24474be18c99e6297d34313df4a3acd485c88067c8524f8b26

Observation 7ce107d8-6f8b-4a64-bb32-ecb300aed9ad · outbound

This paper cites Henrich, S.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Henrich, S

Reference 21

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.671917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:4a27382936fdd1b1a4fc9f0a5cc918eaee1271c2b97950fb03e97e717f950e45

Observation e6a67e0f-d32d-4156-b59e-4aa5616ccb6f · outbound

This paper cites 29) China Household Management Research Center, Ministry of Public Security.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 29) China Household Management Research Center, Ministry of Public Security

Reference 22

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.679268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:98f252b967c47e9ef86ad2c7475db4fb0f6a5be932352e6a6e74133ec6e57db2

Observation f6f6bbaf-1890-4453-b1fb-fc4178b9beb9 · outbound

This paper cites doi: 10.18653/v1/2020.acl-main.164.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/2020.acl-main.164

Reference 23

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.684446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:c9297a47da1ea1d443673fd8c9207290668dfddd3d5938a38caafc9126f34dba

Observation a0c8357e-077c-456f-95f3-fa0a7f76b497 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 24

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.694224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:e1a00eefa71be39e9b8da3419cbc870127e13b41166129a3d714bd5091ac63e2

Observation b6f0dd71-dd83-4cce-9fc7-83a3b29def67 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 25

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.699993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:b09bcc1ee7c6ab40d5ff2042a105c2eee3d2b0b613ce356795e6a32b8f4c0053

Observation ca22ed36-c860-4d89-9fec-1bb709f5652c · outbound

This paper cites The N arrative QA reading comprehension challenge.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models The N arrative QA reading comprehension challenge

Reference 26

Resolution
metadata mismatch
doi, observed 2026-05-10T23:26:25.703976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:a7dce3d02f37f91952774ffd336d6dcf70e301d7a332ea1e5ca81f92b4a7a408

Observation 38a947b5-7079-4c93-88ac-6a87896fabc0 · outbound

This paper cites URL https://doi.org/10.1007/s10992-020-09581-6.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models URL https://doi.org/10.1007/s10992-020-09581-6

Reference 27

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.707647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:20452d30bb00ae659031dc0f1eb528e6f76392f6d443b196b31e2a8dd5de1dbf

Observation b6e2d620-b30b-4b4c-99d1-bb0b9819d509 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 28

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.717113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:8acbf7b19abba4cf5864f679cc638d0f8d41a3795b6e81d6fa27431221fb199a

Observation 67137944-82f1-4b14-ae6a-fbc76ee12077 · outbound

This paper cites doi: 10.18653/v1/W19-3005.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/W19-3005

Reference 29

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.723283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:594da6365cfd6c1ef12c86249b4f26f3de5040eadb63d7bd1112729263f1fea3

Observation 107b8f27-6805-4849-bd17-bf3c36680518 · outbound

This paper cites and Rudinger, Rachel.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models and Rudinger, Rachel

Reference 30

Resolution
metadata mismatch
doi, observed 2026-05-10T23:26:25.726946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:d47d02092e251e040d4919253a06b779fd5f980358782910129df78ae6ad9dcb

Observation df711b25-1885-4ae8-aba9-904d944a6714 · outbound

This paper cites Andere zeiten, andere lehren.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Andere zeiten, andere lehren

Reference 31

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.733673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:166f00bb8e69cd9848a412e795e865040284a357d6ebd21b66b714304fa0ab0d

Observation dfcc865a-7bcf-40cd-a00a-7ada0c416117 · outbound

This paper cites 31) David Milne and Ian H.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 31) David Milne and Ian H

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:26:25.865742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:2001f66cdd754e6d1a443cd454edc17b56f0f1426d5286b9816cf3934bc56092

Observation 49d99f63-600b-47de-9374-a6d684d8fef2 · outbound

This paper cites URLhttps://www.aaai.org/Papers/Workshops/2008/WS- 08-15/WS08-15-005.pdf.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models URLhttps://www.aaai.org/Papers/Workshops/2008/WS- 08-15/WS08-15-005.pdf

Reference 33

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.753871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:63d96a38f23f0e1bf87d4153e61a7be3ffa27cfe8f3b97d3220ae8c8c94cf144

Observation 9d40431e-0780-4aa4-993e-42d853c25105 · outbound

This paper cites The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.855730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:f896690ddf35766e59ca4aa14a29c1b263a78b2c1d7c9e4e8478f12fed5a3f07

Observation 20abb63c-278d-4e6e-8a6d-ed8e96ba45d8 · outbound

This paper cites Cohen, and Mirella Lapata.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Cohen, and Mirella Lapata

Reference 35

Resolution
metadata mismatch
doi, observed 2026-05-10T23:26:25.758060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:e774ab256ca84461c5ba070d9bcbc8f65cf9e9d82d4bae8bda95d0bf53bac767

Observation 980c77ab-506c-44b2-a1d0-7ee1e55d1df4 · outbound

This paper cites doi: 10.18653/v1/P19-1442.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models doi: 10.18653/v1/P19-1442

Reference 36

Resolution
metadata mismatch
doi, observed 2026-05-10T23:26:25.770353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:c5b75eeba590b19cb9b556242b5cf0d7bed93291672e4637df2a81fc61fcaa93

Observation 596f7243-5571-4f85-b020-89e74078c9ef · outbound

This paper cites URL https://doi.org/10.1080/02724980443000566.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models URL https://doi.org/10.1080/02724980443000566

Reference 37

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.474346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:dac9b3f51031dcc61cf20109d8c8cc21df6198219c44a4a5321fd6ae5cb9d340

Observation bf6af985-ea02-48ec-849b-ecfbb3a2c1f5 · outbound

This paper cites 32) Judea Pearl.Causality: Models, Reasoning, and Inference.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 32) Judea Pearl.Causality: Models, Reasoning, and Inference

Reference 38

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.785346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:f25bc54068f9abc0fd5d9add1d2c934c8fca1bebc4144a9bf9bd907e9e3d616b

Observation 41bf6c56-66ad-4d39-af61-31d4c40ffc41 · outbound

This paper cites 29) Tony A.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 29) Tony A

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:26:25.881584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:55c65cc32e7d4814f3016d237450292ce20b44b4b1bc935de20288f3796faeb4

Observation 51ff55af-49cb-4d4d-9d6e-1e4318f94bfa · outbound

This paper cites 29) Robert Plutchik.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 29) Robert Plutchik

Reference 40

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.792521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:a446743c76607fc396ad98ffbf5fe035fde8dc20543f819af450162d2d692000

Observation 245e1109-2da1-45d4-a2f6-03e8a52d39b9 · outbound

This paper cites URLhttps://aclanthology.org/2020.lrec-1.125.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models URLhttps://aclanthology.org/2020.lrec-1.125

Reference 41

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.805094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:3c3890046c0164e016de3535d7cfd64fe0a75f08d15f416257ef5c19fbd7d714

Observation 39592300-ccce-438b-a875-15853f495b46 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.814589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:6d17b8a1c2ab7d1a07f436ba37e0aa41ad794a4b4b536bb98178cd0ad6900e91

Observation e4e5305e-3980-4d06-84c8-40c1e1b746b4 · outbound

This paper cites 38) Zijian Wang and David Jurgens.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 38) Zijian Wang and David Jurgens

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:26:25.872973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:fc3ff47a0ffa8ca08700ceba1a11da502b999b5dfa12b11b2166b3d68493062d

Observation 3302fd33-ab6b-4d8c-993a-441d920dcd70 · outbound

This paper cites assessing BERT’s syntactic abilities.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models assessing BERT’s syntactic abilities

Reference 44

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.825599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:243c5f427eb74d21ffaa4dd49ac5c113d735fe2f184b687b986f6887d6d147c6

Observation a492a40f-8a0a-4550-b874-08fb972b1b5a · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.582688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:21d88e9fbfdbf3c1450c76882b204d74a6c502934eb9fc54c132a55b0201ac32

Observation bfec096c-fe87-4344-b1ab-ee0c9365a972 · outbound

This paper cites (cited on p.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on p

Reference 46

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.587000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:3410886a9cac08608c61e7fb08a5a05c922a967f403d9bd3750069db9987e810

Observation 2993704c-b483-4fe5-83e6-579323448f3b · outbound

This paper cites (cited on pp.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (cited on pp

Reference 47

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.597303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:dfa1ea527e02af4d51b258eef5258bff55805b42480d93fe61d97035f439fef2

Observation 09903080-9f60-479b-ac94-9d2368280315 · outbound

This paper cites 31) Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models 31) Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.604187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:d1615abc17ca84e270d36735699fe0495a3eaade4cf8d59e13a3ae63c6384f17

Observation fbb6d9ce-4ec6-4ee7-830c-64a5eebb8847 · outbound

This paper cites Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods

Reference 49

Resolution
verified exact
doi, observed 2026-05-10T23:26:25.608574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T23:26:25.323988Z digest=sha256:2314108cdc3ef975f335aeb083ce3d3a9a50f1402329a7eefc11183ffb082e64

Pith citing papers

Observation 8c9721d4-2582-4b35-b622-46b8cd1bcff6 · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:38:37.880668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:b6e4c458a4d03dc38e534a48d9e055eec088b21eaabbaf1d750b3ad078a1ef36

Observation 6fd99ee7-4c87-40a3-8c88-2e151787f520 · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:b3391c3fc8d6e3fe84cf021a367841d5ff7a3f6c492c4b74e32153ebd03c2bf3

Observation ca686e16-f7fd-4a9b-9aa8-209b6866ec99 · inbound

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them cites this paper.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.956384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:5c89637b5e1f94639c3460ed47f2b6d1b67f67107ccdae2856047b038194e41e

Observation fc9fc5b9-7fca-4383-8567-25f4e0934cd4 · inbound

Large Language Models Can Self-Improve cites this paper.

Large Language Models Can Self-Improve Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:00:48.297809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T17:00:48.167441Z digest=sha256:cfcfb82fb58fa2c96070f5340a11f472461203bf285e48a251f4966461b20ea6

Observation 3fa233c2-a16c-462e-85c3-a27abbdd2d14 · inbound

Large Language Models Are Human-Level Prompt Engineers cites this paper.

Large Language Models Are Human-Level Prompt Engineers Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:43:26.345236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:43:26.288866Z digest=sha256:76b4c25e1eae90ddf710a4d432a5baed9ea53c46ab5bf5eb93bf37528f5a50b1

Observation 7d0a8613-2704-41d5-ac87-4d3d39eb0b24 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 141

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.440917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:651f74a4b7915986ef97c93f7ab79eb890bd202e6641a056057baa2846ef26bc

Observation 81c0ff3f-125d-46a0-861a-946494458a01 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:53:22.029283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:a70ca5743f8d09059a96445b404a9290dcd7dba4838f109b5287e9b57c5d9561

Observation ba537a72-2496-4e2e-910e-77c977d362ee · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 237

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:53:22.261459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:0f65578f9792ff9d4f15ab501bb716b460cf8caf6b833ee333cc33c0343c9541

Observation 4648f200-ceaf-41e9-9635-e36dee31d635 · inbound

A Survey on In-context Learning cites this paper.

A Survey on In-context Learning Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T12:58:27.557343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T12:58:27.430374Z digest=sha256:16b316f582e25ae87d83508f8d667dfe8476a9147d320f9bf5171702c26731d4

Observation 81c0c17a-e935-4f62-8c74-08f14b8f96bd · inbound

Progress measures for grokking via mechanistic interpretability cites this paper.

Progress measures for grokking via mechanistic interpretability Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T21:52:56.172067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T21:52:56.040569Z digest=sha256:d1de5645606779727aa2fdf1825f5d6f0881bf0b171c49315d0a10fda1f37663

Observation e208446e-2900-4851-90ab-8b131f407bb0 · inbound

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning cites this paper.

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T09:14:16.494657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:13:30.054153Z digest=sha256:2032d0265cfff71c7cbb443eaafe3ddddb874d0aafd95aaf094b360fdb45c431

Observation 66c7c594-be97-4ee4-b1b0-a5f8e6b1f754 · inbound

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity cites this paper.

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T19:58:48.116991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T19:58:48.066562Z digest=sha256:50d7bbe3c139b8f638cf611827c11d44535c95bbde646acb2f2718a6c8155dc0

Observation 7023981c-6e77-473f-80a4-3a3747e58a41 · inbound

ART: Automatic multi-step reasoning and tool-use for large language models cites this paper.

ART: Automatic multi-step reasoning and tool-use for large language models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 152

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:03:06.252249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T19:03:05.597295Z digest=sha256:445f91eee9651127db525da862872fdb0693f8aaa40a1979c38293ed7033e456

Observation e0b65df4-9b71-448a-8545-f725732ae354 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 107

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:19:46.804424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:f3978d3d363d371a49b2462b4133081983003f10c67d86e0626a36041b763f5e

Observation 7da3c7be-ee12-4a6d-b4ec-5087c26ba17f · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:43196d3bb721d4d76e4630396a479022a5f29182c17a51bc87aa2e11b1ecc966

Observation ff7e6c23-fc25-4246-bf2a-fecdeafdedf2 · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 135

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:45:17.678544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:2a298c8235546b7dd99728f6aa821684efa2c90ff9037f7211a7e4dc93bab16d

Observation 3b189186-8328-4d9c-b3ac-e79d18a634c5 · inbound

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? cites this paper.

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:36:55.166008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T07:36:55.087443Z digest=sha256:e6aa9f96d1fbacbfa66e22551cc5e8fce99a01800f4416b975a04606c4528194

Observation 1651c060-1637-4d28-937d-eb8a4fc78425 · inbound

Towards Expert-Level Medical Question Answering with Large Language Models cites this paper.

Towards Expert-Level Medical Question Answering with Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T04:32:33.480460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-24T04:32:33.271634Z digest=sha256:de7d4336719d200fb4d9e53e017d441078593d85bb9e8c11951dbc778fac05cd

Observation 6539ad68-7843-4cdc-9428-3fdda8b441bb · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 255

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:59:27.339629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:4571559373c3d0e5fbfa1f8c0a355fea0eda5e87ac4bb54c93a3babe5b93d9ec

Observation c48315cd-a1d6-454d-9eb1-1eb430b1ae93 · inbound

Evaluating the Performance of Large Language Models on GAOKAO Benchmark cites this paper.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T12:28:32.430613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:d3bb77daae223cae5f5cd5fa449f736cb50405bc8f9fecd18b2d212ef095f133

Observation 9efdb8f6-6d40-4ddf-9f14-5b160c602898 · inbound

Improving Factuality and Reasoning in Language Models through Multiagent Debate cites this paper.

Improving Factuality and Reasoning in Language Models through Multiagent Debate Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:01:45.212602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:01:45.164412Z digest=sha256:b42535a16d3c636ca9478bb99c69829e02ce7116b42514e546690a96beb0a1e2

Observation ae218b2d-079e-4134-93ef-583a70e09dd1 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:35:21.380700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:6087b0f489859fa3eb4d9f92381905d42d8c7731587c28491d4301cd40350e69

Observation 7e1adb2e-beef-4bda-b5ac-906e746478fb · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:171c946b56321a5f4fe9eb9a00176e35e6b668272cd6de5705cb6b395bd9240c

Observation dac4623a-7cb0-49db-a419-f0ab19f0786f · inbound

Simple synthetic data reduces sycophancy in large language models cites this paper.

Simple synthetic data reduces sycophancy in large language models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:48:08.687462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T14:48:08.508109Z digest=sha256:50c053e2993d713941d6c29f8c9b6273402999c18bce43369df3e047801206e9

Observation e79ed1ad-f787-416f-aff3-fedc40d44e70 · inbound

Reinforced Self-Training (ReST) for Language Modeling cites this paper.

Reinforced Self-Training (ReST) for Language Modeling Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:59:55.956017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T07:59:55.849296Z digest=sha256:45ea642ede6bd48d0f035f2d7e7365da5aaaecb7ce932e194d6c3f18283e0437

Observation 3f47cb8c-b37d-47d8-93b2-3f79c3b0c7ac · inbound

Large Language Models as Optimizers cites this paper.

Large Language Models as Optimizers Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:04:31.357861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:04:31.212102Z digest=sha256:1ce2180fbd114f89395f884429b5ea907a8854462d08eb77627bf8605ed2b84a

Observation 05828090-91fe-4df2-b661-0556f42135f3 · inbound

C-Pack: Packed Resources For General Chinese Embeddings cites this paper.

C-Pack: Packed Resources For General Chinese Embeddings Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:24:32.114886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T13:24:32.084878Z digest=sha256:829fde80bcd628f7d737c1ed8aac1da6c8e5c399fa642e5cbd0360a81bf623cd

Observation 06822792-59bc-492c-9d91-e34fd508279e · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.661016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:399214b24dff0a88e614b5c76386fe37c8c4f3f6b0c383aa2e0bcf588ae408ca

Observation bd11d899-2e47-4818-8557-8e38eaf432fc · inbound

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs cites this paper.

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T11:11:21.637241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T11:11:21.460613Z digest=sha256:0541a0d4ee162998f040e7a24b3441622220f7b90c39b8220a66ae22365eb661

Observation fac13bf6-5b03-4344-ba03-85efd1e97681 · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:03:55.335301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:c0e737f50560d911111508914b87782d59613da0093d1144910e2f06f8b401f6

Observation cbfa9301-0159-4477-94f7-fca74eb99e7e · inbound

TinyLlama: An Open-Source Small Language Model cites this paper.

TinyLlama: An Open-Source Small Language Model Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:09:45.383248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:09:45.317768Z digest=sha256:983e9f55b3a14db44f0fa6c479bff532e3662d638ded2dafe6cb6f10525eb770

Observation cc313612-5b75-41e4-9f01-a621e387e9fb · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 114

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T10:36:18.222392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:368139f0c7d6a6dc9159fd6bd19a0d2667cbd3529ec3ae7dbdd126c0da0c0215

Observation 0cab335d-348b-41e8-b28a-fdcf70ca187d · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T12:17:53.570619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:9c13e82528052b72c3a5412cb45b60697458127b969d31ace4ec9f89dd525d80

Observation 4b20af0c-c671-447a-8d15-b69f1be61a32 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 138

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:04:44.508501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:6cfa2a3f24209268598bd925923cf13723e6af2f2628f1edeffe8eb1802ce633

Observation 26b56491-f2e6-4bf5-a2ac-0f9180e7a2b1 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:bcf05be294a67adf365070cf97a3b975a69dc9bc468d66f8eb1be228c277e5e6

Observation bc9ed97e-d90f-4a90-99fe-84a9c60fcef2 · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:44:28.879481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:5e75ace9c4dd8fb67a28bfe56496b8f44b1c6e10172544b5120d85c70f7de829

Observation 44b0fadf-d768-432a-9d07-5862ff58f81c · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:27d7db6da11427214b23af48a965a4444a8a45427b670d04ce49fc4775795a19

Observation f5a0b5e5-b0c3-4d9a-9b54-f14466339c26 · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 162

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:03:56.647244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:95ecaa8053b57e8a99bd19f10d5f08a6b1c68cb1eb85cd1627ddd91e91385527

Observation 66bdd477-3ef2-4772-a211-3ee151b39f68 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T18:44:49.728653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:e87b2eabb99e2a7fbf10f9055bc41c5eb2fcf968201373f967355d17f97c69eb

Observation 6dc02a4e-23d8-4905-8760-e3009390288a · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:05.108235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:ab7e1be710a2d6115bd18e20fbaa7486c8f8f6e8bc646a30f0d7c016c48e2344

Observation a1a0b8c6-96f0-4693-b834-2bcc59cefcec · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:08:09.678351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:5b4bd4fdee729d4a4a137c8d3c6c9287071c73d68f2704752d5a5483a50e6af4

Observation 298e6435-eb7e-4b40-96b6-51b5bf4bc50e · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:48:26.484783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:168c0f12eeb0a05ce92cc04becbb12c71bd3191c5c0380ba9383c2bc64fcc660

Observation f506b820-157b-43fa-9113-400f359ecf1b · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:36.697127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:43c02c52de006a06dc31d148199184f231f2efb306d1d23894d4e955d6914ab2

Observation 3de9ad2b-c6f4-493e-8eee-02f35626dc37 · inbound

In Context Learning and Reasoning for Symbolic Regression with Large Language Models cites this paper.

In Context Learning and Reasoning for Symbolic Regression with Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-23T18:53:21.163360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T18:50:40.378720Z digest=sha256:32f684bb9899d616b4795b9d9a6d3e5e11ff9139b7adcca1f68a5e07ce2b38f4

Observation 8b070f90-0ee9-4746-900e-91bf360f1dd8 · inbound

Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models cites this paper.

Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T18:05:44.713031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T18:03:19.212619Z digest=sha256:d215cddfa331bf91f0b8d992b5e60f3de616260d4b7e0fc390a8e74343705a11

Observation 42f5d02c-d974-48c2-8473-d4b4e8d1beab · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:35:44.194739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:fe9c7da313fa6693be465971cfa9980bccc250a3ed71786875f5855ac4425878

Observation 56f5e2ec-d951-42eb-9e15-3fbe4fc8bd70 · inbound

A ghost mechanism: An analytical model of abrupt learning in recurrent networks cites this paper.

A ghost mechanism: An analytical model of abrupt learning in recurrent networks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:32:38.822685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T06:31:21.006691Z digest=sha256:78f6df1e1e0c3dfd7f8c389df175e1c05eed5cb446871ccffe40b66d45d1c5d5

Observation 697f7b4d-c536-4372-a270-c84f7110c643 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 120

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:52:37.669278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:e71439625945ed0f520541fcfab9a60245e39ce9c4ff4eb173b2f45329820236

Observation 412b4df4-37e4-4546-8fe9-9d919cadaec1 · inbound

Do generative video models understand physical principles? cites this paper.

Do generative video models understand physical principles? Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:47:05.894142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T12:47:05.825659Z digest=sha256:86d58dc3d86aa956bd21f126fd6b5082b39093f8b604358cc67b68133d9d3d69

Observation af5e1950-1d1b-4cf1-9108-b245c7f21bcd · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 138

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:20:59.288892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:4c8483fbd2279df48e2054b3c90218e8c10e23262a20a1d791f2356b1ca2a1d8

Observation 7ceabc31-aa91-46dd-97e1-43cef4c1a185 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:350a602d4b9c2b9ab16a386294890890a1ad61e0e56253b8a62d36fb02e79e6b

Observation 06984b8b-98b9-438c-aa10-b52545e32b64 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:31.211761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:21134d1e495a8114bff343684f0ebff056c2c6490764540172570083b7d69e3b

Observation a023a95d-9ed1-41f0-8143-ffd0049b7396 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 122

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:44.826710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:ad7a1f32a9c48b88a2a0e4dc8bd1d969139b5e0f842b3029f89058ef18c97d27

Observation 2ad8cbf1-a96d-4bee-a255-1e061f6aeede · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.139480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:ddbe63c844a828695103ec084398b665b67463902a4ba3f2077001b203311c73

Observation 0ddd7ef1-3f11-48bd-bd19-29aed95615ce · inbound

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR cites this paper.

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:32:04.650901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T20:31:34.074705Z digest=sha256:6ca5436d387ef46944ea2188ee45b630e7220ecc5a33be1b98dfceba32cfb65d

Observation 75e71827-cdf9-4db2-967e-915f3839ce01 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 130

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:38.497701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:e7f45c1556c4e9d3813e32c916ebf6b725a134d34428cde8534c8d827f7f312a

Observation 2361ad68-2484-43e5-966c-7c8206d5f9a9 · inbound

BacPrep: Lessons from Deploying an LLM-Based Bacalaureat Assessment Platform cites this paper.

BacPrep: Lessons from Deploying an LLM-Based Bacalaureat Assessment Platform Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T11:22:16.448303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T11:20:42.807705Z digest=sha256:117aa2d519729523bb99bc0fc778974c91e04a3d4b37a2fd6110f6cfa9c62dd2

Observation 204b904e-2e3c-45b2-8acd-7be9f63e6e8a · inbound

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention cites this paper.

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:27:14.679600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T09:24:04.289799Z digest=sha256:614370e79bba1d9487ac5c8e3891956d6f15f60299e7e717976e9f2a636fd035

Observation db9e0a50-6b49-499d-8873-ce9c82f4dc3b · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 162

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:42:04.860990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:5abe4ae1c26f74cf905fc57174b66b55d5429ff9e1c9b0dd20f3f73b9a1daef2

Observation efceb3b3-d0ca-4cf2-af61-b6978a548905 · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:12:55.782077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:1f8426b7a3906d3ada3a092f403e8e657404e31386a402284b1ede464c8d548b

Observation e7d87052-6883-4476-9245-600fb22d3202 · inbound

Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies cites this paper.

Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:32:50.792744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:31:49.857895Z digest=sha256:94249bd3fd2d49a654ac6bd720e6c1ecbbc4c1d4df54181c13c08683aac06f68

Observation f6bd6156-db78-4c68-96ac-890b8f526955 · inbound

Designing Psychometric Bias Measures for ChatBots: An Application to Racial Bias Measurement cites this paper.

Designing Psychometric Bias Measures for ChatBots: An Application to Racial Bias Measurement Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T22:12:51.645229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T22:12:15.876447Z digest=sha256:b3b747902dc1ed8f827725e8f9466e67dcbbdd78e26bc1bc5612f6dd3444cdc1

Observation e7b4900c-c7d6-4aac-a18f-5b204c5e0ccd · inbound

OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models cites this paper.

OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:11:14.049897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:09:08.307199Z digest=sha256:30ad7090356ffe2c05052457e08adc688e55f97a9045e06a63968ad459cfffc2

Observation cde58f09-7efa-495d-8171-387b557312c6 · inbound

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting cites this paper.

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:45:26.750295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:45:26.750295Z digest=sha256:2537097bffc15271334be2e7a9c30dfbe55b355b9f6d162a90e84311fec4946b

Observation 80be55b5-0748-45e0-a2ec-247eaa29da5c · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:06:13.954720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:40674dc39e7fe0c55440a8847c9f37087797797829a9998bf8091444d3148974

Observation 319e9b58-67a0-441e-9d78-f2a7366e8fd7 · inbound

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization cites this paper.

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:58.226873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:58.226873Z digest=sha256:ca9a6f9b55961ff54034ed39a9035b19449bfb46e8d03b4b4f9cb9b0add1f7fa

Observation a3e4bc57-bbd7-42e9-a1b9-7b1ca3204f87 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:14.034600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3849c2dea74132568e61566229fd86a969e3f22b08521d65f4924788866e61d9

Observation 854d13bd-fdbf-443a-a3c1-f9df365e5b8a · inbound

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism cites this paper.

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:05:48.081120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T03:05:06.069642Z digest=sha256:1d3e80f47d1e5245b4f93e25044cc729104d8c3d32fe129a253e66f2e32a0c1a

Observation 3b88ff94-1e23-46d1-a5ca-4b9005d66098 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:20:34.413999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:de9ae4058ba0302cd28a3dfdf5a6943e5483834b79004d9a396e6989fda6b69a

Observation 0c49015a-43ae-444c-bd7e-88131d4a7d83 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:52.162655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:52.162655Z digest=sha256:b6102f73040c3e285c0ad12470312d00f2135ce24bc6a0c9d181816ff9f24a53

Observation 783aef56-55e9-4758-a1c9-d04e478c10a4 · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T11:01:17.135129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:cd0eee091de9b8eea89380c76645682bc94098727c697dbdb998c3c5e33003aa

Observation d7849b73-2840-4e4a-a39d-98d78d63ae77 · inbound

Contrastive vision-language learning with paraphrasing and negation cites this paper.

Contrastive vision-language learning with paraphrasing and negation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:38.915404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:38.915404Z digest=sha256:5e5e066e8af48aad77c2b86f4348d75f1033df9e624cfd5d32c169a5987daeca

Observation 9910e596-a17e-4a09-8ee8-bdde9bcf70d5 · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:18:20.401592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:020df35882b62a00d7cf6ddf8eae0068cd26b8d34f371dfab18a2f0b76139e37

Observation 8cbe9bf7-32e8-43ca-aa7a-22c6c3e1a805 · inbound

Beyond Context: Large Language Models' Failure to Grasp Users' Intent cites this paper.

Beyond Context: Large Language Models' Failure to Grasp Users' Intent Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:11:13.634496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T20:09:25.827452Z digest=sha256:74bd2498b0443ad89bd479a1a8aeb064a2937bc9583728c85ebc283daefcfbb9

Observation d1f6dc17-0cf5-437c-a7b8-d7322cc96c0c · inbound

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems cites this paper.

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:53.638767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:47:28.248540Z digest=sha256:014cf58cf31ad918e789c17b37ff2540b27315758416ecea943378af287825ed

Observation 52b1f8f9-cd37-4b3d-9806-07d0c9eaf17f · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.348309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.348309Z digest=sha256:2c15f3360ef354cd5eebd35bf293288a4571bac247731a5fd7b8ae7dcc50365d

Observation 1f21422c-1069-42a7-88db-dca543e86cba · inbound

Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning cites this paper.

Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-03T04:23:24.116611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:23:24.116611Z digest=sha256:de5dce03e246662b894a2283ab003fea77b4a92c839edf5a57543c66a159c7ea

Observation 272130d4-2c34-4bf2-8dde-3596f4889729 · inbound

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty cites this paper.

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:56:43.635300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:56:43.635300Z digest=sha256:d1e95434fff6d5fd9ee907968dfb84c68f64524b4ee04ae90bd77e76ee0a415a

Observation c140785b-4a83-416f-b3cd-a0029e39f001 · inbound

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers cites this paper.

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:42.809515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:42.809515Z digest=sha256:dff280af34e2c607500c77cad429312714644458bfa28c6b03f1ea5458ea5ff2

Observation 4c0aaf51-26a8-4fba-95cd-e6e0708aa674 · inbound

CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching cites this paper.

CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.091238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.091238Z digest=sha256:250808fac6de8e69372c585b8d629a9c57b4748281fb28ec970c87966aa16f6a

Observation 04320c7b-90d2-45b4-a803-b26a1c80f824 · inbound

Graph Property Inference in Small Language Models: Effects of Representation and Reasoning Strategy cites this paper.

Graph Property Inference in Small Language Models: Effects of Representation and Reasoning Strategy Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:10:18.559695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:07:07.522017Z digest=sha256:2bc35652a4d3dc74be77a845ca37573dc33998e696b40b540d57a5d75c88f3fe

Observation 301c55de-fbe7-4ca3-b28f-d157955d06e9 · inbound

When Models Know More Than They Say: Probing Analogical Reasoning in LLMs cites this paper.

When Models Know More Than They Say: Probing Analogical Reasoning in LLMs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T16:52:59.652395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:a796751ececb8bb56e17e5069ad425105a40619585e21698ee6fd41e96f5c5c3

Observation 0c175797-94bc-4917-bc83-75eac0c795d4 · inbound

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks cites this paper.

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:49:32.983778Z digest=sha256:6b1a9ea2c8133caf6954f9b6c590d69833e379aaf0257b74101c9c2e3d503f0f

Observation 54d18fb5-ad1e-498f-b34e-6e2e2af628a4 · inbound

Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs cites this paper.

Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:11:35.346010Z digest=sha256:daa47ddbb46e7fb00c7c602088984acbb4421da5a41ab822f6cc866cc8d0c07b

Observation addbe7bf-e89d-4757-9468-bbcfaa47bc1f · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:41:02.922816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:23:26.128263Z digest=sha256:afb170c1359e49e60185d7055e537dd9e89807b3b45fc039ecb0c8e03f71b136

Observation cd5c6406-627a-4879-a478-656a49e76187 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T21:35:10.173253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:35:10.173253Z digest=sha256:b6f7759d085829d1edff68d3f261846f343721acf02a48db73ee85bf5bb3c5dc

Observation ec686702-ebf9-4b0b-97e6-cfd86ecedae2 · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T10:21:01.053272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:3edda9c17f9da8652271d89537d770dcace077bdce78cc46bf8430dbc8333801

Observation bd2831a6-a122-4d3f-afd0-a44c195896cb · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:fceade5847c915821d943bb66dd6cd543d412c472e3084e1dfedeb4c2522fc49

Observation faf54622-7ee2-43d9-8b27-785d8f0b568e · inbound

Consistency Analysis of Sentiment Predictions using Syntactic & Semantic Context Assessment Summarization (SSAS) cites this paper.

Consistency Analysis of Sentiment Predictions using Syntactic & Semantic Context Assessment Summarization (SSAS) Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T10:55:20.435471Z digest=sha256:74b4c5304ad9f7adfb5b3585d3fe156de3e9212109381ab7b3a68bda62422a45

Observation 09f64f09-602a-467d-b4fc-b9500cd8316f · inbound

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild cites this paper.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 57

Resolution
malformed identifier
local_arxiv, observed 2026-05-16T11:27:47.797554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:bd7bb357525a2b1041900aca2977df687d185e177a605d642ca8dc3c2afa2f49

Observation 7fcd0e40-c8dd-4b2a-90ed-bdcd4c874f6d · inbound

Measuring Representation Robustness in Large Language Models for Geometry cites this paper.

Measuring Representation Robustness in Large Language Models for Geometry Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:38:10.517729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T19:35:32.531660Z digest=sha256:8732d114c7eabf6b47bf7330bc0eda1fff79f44ec1c3858d805149c70372723b

Observation 08069325-db17-4714-995d-f9947b29d698 · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 155

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:ce4d29b8eb4e52f3bd73163e85bb195740d5cb462d69abeaa6943203029eb80a

Observation 0a6f6160-40b6-4173-99cf-22707a57ec45 · inbound

Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier cites this paper.

Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:02:07.838745Z digest=sha256:72587910397f978c082962e343255f4b4d237b6535e04e2da2b8972d9fb3ff7d

Observation d6df45f4-c7f2-41e9-8119-468aab917164 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:56:29.680852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:4e098035a6a282717b3f39dc08d835ba6163273bc86e45c874ab60c21e508a1e

Observation 9c0beaa9-fa61-49c3-847b-da814739ba65 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:35a39365f3d3b87fbe09574a1ff3b367ad08274f985af40d2ee6fa7dc08c1600

Observation 72c5f6d7-45bc-4714-985c-a48ff089c4f6 · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:4f48cdd5e2f283b0924bdffbc998b2118470bab2d4dadcb96ced7b6f7f3b58a0

Observation 9f57a9cf-c60c-4b9b-bd09-3b9fd8b18829 · inbound

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation cites this paper.

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T20:13:06.376037Z digest=sha256:2940c0ac33f96b0899226244b2cd0646f3cc063560dc3ec8dfdd864cac6a4c11

Observation 049bb1c8-3a57-416d-af92-7a0377c98006 · inbound

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework cites this paper.

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:04.431882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T14:02:08.970042Z digest=sha256:436f82fe4cf0b8d7cd385085688cb36db6e2c058182c317ae1617beddb7c0d06

Observation 5fd3f028-5201-45c2-8662-d19a15e5ce82 · inbound

Complexity Horizons of Compressed Models in Analog Circuit Analysis cites this paper.

Complexity Horizons of Compressed Models in Analog Circuit Analysis Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:26:10.460133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T16:37:05.972866Z digest=sha256:d5c962a14e6b8c562bf25812988c0301f852a3ae174110a54989fa82106ba916

Observation 213dbb2c-6a05-4eff-bc1c-29ccfa692f48 · inbound

A Meta Reinforcement Learning Approach to Goals-Based Wealth Management cites this paper.

A Meta Reinforcement Learning Approach to Goals-Based Wealth Management Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 299

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T18:42:50.962120Z digest=sha256:78d4c31292582bbafd5ff5dd16416b27b914a12ccc6df60ffd6b6461ec0a0ec0