Pith. sign in

Paper Citation Record · LEDGER

CryptoX : Compositional Reasoning Evaluation of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2502.07813.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07813 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:34:26.820347Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:26.867470Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:54:44.144191Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact4
  • verified fuzzy4
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ccb7eefd-a373-4959-8053-d187c253a598 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.701322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.701322Z digest=sha256:be77285c387f042427f73f18c581bfa5fab26872e2d0afbdbad56529ac875a52

Observation 2681304e-4f9d-45bd-a739-a3b6aa88da22 · outbound

This paper cites an unresolved cited work.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:34:27.318246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.817457Z digest=sha256:65e7338dcb051267343880ac3b95a0346d0439c11f4bc7991c84dfbddd91f294

Observation 4d60f8ba-7cfb-405e-99b6-77643daf5c7a · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.709405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.709405Z digest=sha256:095f3978ddeacd42b3bc4980edf1ae6498cbfa80261a71d456e5cec92594a483

Observation 05336bf8-4c47-4db1-8a32-1b9d12459245 · outbound

This paper cites LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.713246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.713246Z digest=sha256:d312cd52b88d1b41d34014f5e8a7d1264b7d448b6c28252a1c4a09c014ae3634

Observation 25176236-cd10-4602-85f3-52cccce66007 · outbound

This paper cites FOLIO: Natural Language Reasoning with First-Order Logic.

CryptoX : Compositional Reasoning Evaluation of Large Language Models FOLIO: Natural Language Reasoning with First-Order Logic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.717301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.717301Z digest=sha256:7eb2d72f39533223ae7f2a938b22f8ceec053431e638a183bdf3a7d910ac79d1

Observation 805903e1-86c1-4771-abe9-5859b9512c9c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.720730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.720730Z digest=sha256:076f07aa276a08602f191d22fd7d355e1c86ab4a1f3ce78dc5da5eb51219225b

Observation 84af3a71-7359-4ca4-b4e7-e7c83c596b4c · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.735307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.735307Z digest=sha256:524d5fa0309e1807896fbc075bf54c56d56079c775ab4b9993fcd895f63981ef

Observation ab3a0cd4-2022-4d65-9e2f-96e067fe4eb5 · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.738811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.738811Z digest=sha256:8e9cc637f33efc2b23acc8acc35fd98cc30dcd0b354865fe1e61e7af74635a44

Observation a1fdbbcd-5da7-4f93-bd9f-ec4bf45b17e8 · outbound

This paper cites Understanding and Patching Compositional Reasoning in LLMs.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Understanding and Patching Compositional Reasoning in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.742368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.742368Z digest=sha256:7cec2c796524ef7141a789861034367e8730126445a9ce5f97268bd71b3461a6

Observation 4e8ba3d3-1d6f-4555-a55d-665076f171aa · outbound

This paper cites KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks.

CryptoX : Compositional Reasoning Evaluation of Large Language Models KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.750021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.750021Z digest=sha256:efddf35b61e0341c888a3f6393a2c430f992f06111c1c1243fd113105551b8af

Observation 652a2938-4f87-4910-8a2f-cf95a524bc88 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

CryptoX : Compositional Reasoning Evaluation of Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.754527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.754527Z digest=sha256:06f38acc05aa16ce64a2e6d160bc2580332beb564c27d8ed86161a29256a480d

Observation 86c95cac-c9d8-4708-8375-4f91d80da6d3 · outbound

This paper cites CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge.

CryptoX : Compositional Reasoning Evaluation of Large Language Models CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.761666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.761666Z digest=sha256:8ef4350a62f47f3669c2a51e5dcc6fbf2470041c359c40d24fd12d5e4217ce2a

Observation 936b3ffc-bc94-46b8-958c-f6d1a55b483f · outbound

This paper cites GPT-4 Technical Report.

CryptoX : Compositional Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.765639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.765639Z digest=sha256:ebb8bee3eade8a4c022cee39722e2089d732cecff4cbc2c9378f4a850ae09cb5

Observation 9e0daa46-0ed4-4942-b1a6-267367916437 · outbound

This paper cites GPT-4 Technical Report.

CryptoX : Compositional Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.769019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.769019Z digest=sha256:f93b380a33a6918b0e7dfbada05e3bf5fcc40d883279e4972e53d8fd36ff5627

Observation 68e73fb4-2b23-4b32-9e9d-74a7a73040ba · outbound

This paper cites In Domain_Words, Words denotes the number of words encoded in the given question.

CryptoX : Compositional Reasoning Evaluation of Large Language Models In Domain_Words, Words denotes the number of words encoded in the given question

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.349044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.809026Z digest=sha256:971be74ec75ad4f1d7585f66ce7d6539cc277c597af37d584ae2f1d2f0137542

Observation 518baa32-b162-4b38-b82b-87fef3d11aff · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-08T18:34:26.775997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.775997Z digest=sha256:1426add0d24f1c96eed375fb18111d0bc4cdb0a6d4e53f38ffdca670688cf18b

Observation 15e6df3c-a392-48ea-ab19-f338dc71fe53 · outbound

This paper cites Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:34:26.969627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.779685Z digest=sha256:34f1a8e0a15d7c0ef0fcf3d1bb5d9165ed1d59ac7688b4b18ae622e626bd0b95

Observation d93e3111-19dc-4ca9-b6e5-662e7f598c04 · outbound

This paper cites COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences.

CryptoX : Compositional Reasoning Evaluation of Large Language Models COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:34:26.954943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.783168Z digest=sha256:5679168c3f20ef9f5b1467bae6160b505a3d9acb555790afb66c59fca1d06105

Observation f2c5ac7a-a6e2-4931-a46e-a312c619e32d · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.786749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.786749Z digest=sha256:97373b196863167e2427919a2c10c06631b4593c76f1782b88d4a00e6732caac

Observation 60fb0b32-6fdd-4301-9796-28c7e711193a · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.790561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.790561Z digest=sha256:0dee543ec4e9189d0d11cd5641052d6421b321741c0c0d102dfbd9b044c1f293

Observation a3e2b4d0-6647-42c4-8a7e-ec9194fab777 · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.794369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.794369Z digest=sha256:e22cbb2ef797df694c0cc3d669880064b60c4320ee1c3b8267af827aa1285f97

Observation 62f1b971-13f2-4aa7-9d9a-f6353a2c4c5d · outbound

This paper cites Qwen2.5 Technical Report.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Qwen2.5 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.798112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.798112Z digest=sha256:4afe6ca272dac54856205f8cd340fba4fefb97e79ff06cc9d82b05f9b64f59a9

Observation 8baab08c-8576-4f14-86ab-0fc728ceae9d · outbound

This paper cites Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.801743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.801743Z digest=sha256:5ae0fccefb82b047d57911c7905451fed0b571c6bb92131634e655646980228b

Observation 10966f3b-5996-410e-af9f-27accce9a829 · outbound

This paper cites Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:34:26.855372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.805435Z digest=sha256:d6a605516407e2eac3475c705c2ea088868378cfeeeb82b189a791fab69c37cc

Observation 04568401-4326-42b3-a0e7-5d7d9dfb16fa · outbound

This paper cites In Domain_Words, Words denotes the number of words encoded in the given question.

CryptoX : Compositional Reasoning Evaluation of Large Language Models In Domain_Words, Words denotes the number of words encoded in the given question

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.338667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.811860Z digest=sha256:71fb9925d14f0d261adf6fe6da5b98f0f630a101ea544ccf7304f43906482e88

Observation 6a4687c2-516c-4ca2-bdd8-a90eccde5856 · outbound

This paper cites gold standard.

CryptoX : Compositional Reasoning Evaluation of Large Language Models gold standard

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.328482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.814690Z digest=sha256:2e2e6d7252b76c92c27e8d3b086be515c11426de8be5868aa27f583c7f6c2d1b

Observation 9e895137-0b86-46ec-8ca1-f56957b41db5 · outbound

This paper cites WHICH FOOTBALL TEAM WON THE WORLD CUP IN 2018 AND THROUGH WHICH LENS OF CHILDHOOD EYES DID THE SOLDIERS?.

CryptoX : Compositional Reasoning Evaluation of Large Language Models WHICH FOOTBALL TEAM WON THE WORLD CUP IN 2018 AND THROUGH WHICH LENS OF CHILDHOOD EYES DID THE SOLDIERS?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.308028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.820347Z digest=sha256:8d888b7bd171335e5dc1f031c9a9faa28b38ff3751e16eca31ba4d74cf6e6595

Observation 1c184897-55e0-4151-b955-7fb4c9ef0f5a · outbound

This paper cites an unresolved cited work.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2002

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:34:27.361364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.772608Z digest=sha256:3af6d6865d35ec5fe68abfbaa1eb2a640098d4ebd6ea17c3013a9e248f82505f

Observation f167447c-541c-45e9-8052-3bb18fb2021d · outbound

This paper cites MdEval: Massively Multilingual Code Debugging.

CryptoX : Compositional Reasoning Evaluation of Large Language Models MdEval: Massively Multilingual Code Debugging

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.746312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.746312Z digest=sha256:87953304186958bf6692cf6f5c33c386054dc6a892420409868f29c700c0a216

Observation 1257f69c-e07f-4bb1-aeac-b723add7ebc4 · outbound

This paper cites Nostalgebraist.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Nostalgebraist

Reference 2017

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T18:34:27.189377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.758125Z digest=sha256:be44e2e2df2da9963bb106964dc0ff0d96b0a1043258e2a6b422e944e8f90e34

Observation aea7e8c3-cfe4-41e4-ab65-61c26e5b097a · outbound

This paper cites an unresolved cited work.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:34:27.371913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:34:26.730433Z digest=sha256:6e8381bf67d9a69d76af7b08384dc56de8886c94294b791f898a751ac2fd5372

Observation 819165ce-9185-4f63-90b9-ad869203cfcd · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.724372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.724372Z digest=sha256:ca6bdb2ef2c177e94a3889dfa818f93dadee47095e2f19a40ab154484c3a5673

Observation b990aa96-800a-4fad-a30c-3b72dc7b3148 · outbound

This paper cites The Llama 3 Herd of Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models The Llama 3 Herd of Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.705289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.705289Z digest=sha256:3b175c4a25cd7a267d0a4308d64d1ce0c6db2edc05a774db546f0749403c37ea

Observation 2f6b8b88-8861-4f05-9c91-f210d82fe01a · outbound

This paper cites Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.696989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.696989Z digest=sha256:7af3e33f6632e0269285b63500238b90100436c3cde4a4897210425d23639565

Observation 4b70aba3-ed82-4839-b29d-c718607ce05a · outbound

This paper cites Program Synthesis with Large Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Program Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.692438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.692438Z digest=sha256:2228e41f468a32edf4fde4095fd5680e88d040c5e68200f14e0f765ee378b328

Pith citing papers

Observation 21535db2-bb2d-4c3f-a2ec-3be81e92849b · inbound

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation cites this paper.

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation CryptoX : Compositional Reasoning Evaluation of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:26.867470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:26.867470Z digest=sha256:9711bb025368251b1227b74545c63ad0d8ff218b0ce4dc36ef04dac99b32d166

Observation e24ea313-392a-407b-8bb5-786d85133471 · inbound

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data cites this paper.

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data CryptoX : Compositional Reasoning Evaluation of Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.146308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:45:38.305767Z digest=sha256:b3c0fa4fb963d3caaf7abdeeef30b03494ea584fae9ba04affd16034a1516f63