Pith. sign in

Paper Citation Record · LEDGER

CryptoX : Compositional Reasoning Evaluation of Large Language Models

As of 22 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2502.07813.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07813 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:34:26.820347Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:26.867470Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:54:44.144191Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact4
  • verified fuzzy4
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ccb7eefd-a373-4959-8053-d187c253a598 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.701322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.701322Z digest=sha256:e2f23ab049aa1dcd2e38c5532e72d40c5f286d5176b917e3da83c59bdb12f69f

Observation 2681304e-4f9d-45bd-a739-a3b6aa88da22 · outbound

This paper cites an unresolved cited work.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:34:27.318246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.817457Z digest=sha256:c6db7c310df077e40344f636f7ff4abce1d786e53f3cfe285a7063a8a48c61d2

Observation 4d60f8ba-7cfb-405e-99b6-77643daf5c7a · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.709405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.709405Z digest=sha256:28731b9e1ee30d9a118f271d86215c17f3e90a26609a3300299de5bd3c3e6289

Observation 05336bf8-4c47-4db1-8a32-1b9d12459245 · outbound

This paper cites LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.713246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.713246Z digest=sha256:293b6aedcc296af6216dabaed559e48fe87056211cc02d99b3e3863db0541482

Observation 25176236-cd10-4602-85f3-52cccce66007 · outbound

This paper cites FOLIO: Natural Language Reasoning with First-Order Logic.

CryptoX : Compositional Reasoning Evaluation of Large Language Models FOLIO: Natural Language Reasoning with First-Order Logic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.717301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.717301Z digest=sha256:9d8b785bcd2028aad3e7a5ae35062b899b23ad729cf41be0842e552cc5b2505b

Observation 805903e1-86c1-4771-abe9-5859b9512c9c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.720730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.720730Z digest=sha256:78e5b834aa9318a09765160829d3a67be2d3d87511381cccf40f0aed49a929b6

Observation 84af3a71-7359-4ca4-b4e7-e7c83c596b4c · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.735307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.735307Z digest=sha256:f9e56c87ede0e554e5369e13969653da823c24a498b36fcde98bf6d2df05e3c5

Observation ab3a0cd4-2022-4d65-9e2f-96e067fe4eb5 · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.738811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.738811Z digest=sha256:bbe0354f7e460e68ae3ab3d390ffd52e804736d0f80ed6cf4431e47358c7e6d3

Observation a1fdbbcd-5da7-4f93-bd9f-ec4bf45b17e8 · outbound

This paper cites Understanding and Patching Compositional Reasoning in LLMs.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Understanding and Patching Compositional Reasoning in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.742368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.742368Z digest=sha256:f4964d56ab5418daeb82845bdafea126682af8e555b7d67ecd4a6aeb8ff89285

Observation 4e8ba3d3-1d6f-4555-a55d-665076f171aa · outbound

This paper cites KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks.

CryptoX : Compositional Reasoning Evaluation of Large Language Models KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.750021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.750021Z digest=sha256:5bd21c93a3587f2827f6f5459f5dd2f6edda2a44863bfd4f51d75c85a6f3abf3

Observation 652a2938-4f87-4910-8a2f-cf95a524bc88 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

CryptoX : Compositional Reasoning Evaluation of Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.754527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.754527Z digest=sha256:0b1ef216c3f341c811549d5cdcf11e8f25e9d667e36592663e8490dee622f2eb

Observation 86c95cac-c9d8-4708-8375-4f91d80da6d3 · outbound

This paper cites CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge.

CryptoX : Compositional Reasoning Evaluation of Large Language Models CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.761666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.761666Z digest=sha256:dfb0211deb6ad4b249dcc1704e824dc0df39191624a3c79fa870f84a41483ba1

Observation 936b3ffc-bc94-46b8-958c-f6d1a55b483f · outbound

This paper cites GPT-4 Technical Report.

CryptoX : Compositional Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.765639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.765639Z digest=sha256:dc831524685aa5bfd8af9a3d8a60dae2f44328fe9175c9349e79e189de692eba

Observation 9e0daa46-0ed4-4942-b1a6-267367916437 · outbound

This paper cites GPT-4 Technical Report.

CryptoX : Compositional Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.769019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.769019Z digest=sha256:aede8b95940c06ed3ee21977bdf7f720370eae6d4ccc0c17823fc38816b153df

Observation 68e73fb4-2b23-4b32-9e9d-74a7a73040ba · outbound

This paper cites In Domain_Words, Words denotes the number of words encoded in the given question.

CryptoX : Compositional Reasoning Evaluation of Large Language Models In Domain_Words, Words denotes the number of words encoded in the given question

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.349044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.809026Z digest=sha256:18d1434562326ce4923eb842601b7d9230be0be117176b21d406d6cb635b2156

Observation 518baa32-b162-4b38-b82b-87fef3d11aff · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-08T18:34:26.775997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.775997Z digest=sha256:600a68a8db8fc7eee3c1bbdddfeb32997b20182e0230615ab83d135d0cd3b088

Observation 15e6df3c-a392-48ea-ab19-f338dc71fe53 · outbound

This paper cites Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:34:26.969627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.779685Z digest=sha256:3a86914bd691a9224980dca2c06137a80ccc29520d8bd621530dac5789c3a157

Observation d93e3111-19dc-4ca9-b6e5-662e7f598c04 · outbound

This paper cites COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences.

CryptoX : Compositional Reasoning Evaluation of Large Language Models COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:34:26.954943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.783168Z digest=sha256:a4cb34c3f3900256c8ab304c2fd6deee5e90a6639b20035b480f2d333fcf41da

Observation f2c5ac7a-a6e2-4931-a46e-a312c619e32d · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.786749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.786749Z digest=sha256:fea51d7680c060180ee71dab39487c9611ca5f528f8567a270d42e7b1f85580e

Observation 60fb0b32-6fdd-4301-9796-28c7e711193a · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.790561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.790561Z digest=sha256:788c577783ef9f23bf1c0e250346b3a77a7a2ab0bc6782ecf99793bfa61e156a

Observation a3e2b4d0-6647-42c4-8a7e-ec9194fab777 · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.794369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.794369Z digest=sha256:f627d8bc004c02209f0985128ae7abeb203def908a455e8d7874f83582bb35fc

Observation 62f1b971-13f2-4aa7-9d9a-f6353a2c4c5d · outbound

This paper cites Qwen2.5 Technical Report.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Qwen2.5 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.798112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.798112Z digest=sha256:0f869b7a3b38b7246af1c113a20ab13f9c60c0ccd4b9e252f722f68140ac3453

Observation 8baab08c-8576-4f14-86ab-0fc728ceae9d · outbound

This paper cites Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.801743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.801743Z digest=sha256:6122ea11f565934aadd2bb2ad1d2a7571fcc34fa643c9e038e03ad158bc44e7e

Observation 10966f3b-5996-410e-af9f-27accce9a829 · outbound

This paper cites Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:34:26.855372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.805435Z digest=sha256:899f741c8e36a614c33884c013fcadde7c8b36d47cfd68586133223961abf27e

Observation 04568401-4326-42b3-a0e7-5d7d9dfb16fa · outbound

This paper cites In Domain_Words, Words denotes the number of words encoded in the given question.

CryptoX : Compositional Reasoning Evaluation of Large Language Models In Domain_Words, Words denotes the number of words encoded in the given question

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.338667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.811860Z digest=sha256:4936cde328f55c888e5f1b6f4a427860a00c52d9e0b7d459009f2a9d9adfd9b1

Observation 6a4687c2-516c-4ca2-bdd8-a90eccde5856 · outbound

This paper cites gold standard.

CryptoX : Compositional Reasoning Evaluation of Large Language Models gold standard

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.328482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.814690Z digest=sha256:6b27d99d2cb898dd4f4718aa18edf09d75be8d886d229ec10559e43b1fd2da16

Observation 9e895137-0b86-46ec-8ca1-f56957b41db5 · outbound

This paper cites WHICH FOOTBALL TEAM WON THE WORLD CUP IN 2018 AND THROUGH WHICH LENS OF CHILDHOOD EYES DID THE SOLDIERS?.

CryptoX : Compositional Reasoning Evaluation of Large Language Models WHICH FOOTBALL TEAM WON THE WORLD CUP IN 2018 AND THROUGH WHICH LENS OF CHILDHOOD EYES DID THE SOLDIERS?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:34:27.308028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.820347Z digest=sha256:2bb780d68b59a70db22667c058830f25133a105ae4e21c34c41f09ab82d3c807

Observation 1c184897-55e0-4151-b955-7fb4c9ef0f5a · outbound

This paper cites an unresolved cited work.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2002

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:34:27.361364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.772608Z digest=sha256:93ffc61c63c6f7e9927d45e965e66df1a0d64445094b39d1f23ae92f98bb5bb2

Observation f167447c-541c-45e9-8052-3bb18fb2021d · outbound

This paper cites MdEval: Massively Multilingual Code Debugging.

CryptoX : Compositional Reasoning Evaluation of Large Language Models MdEval: Massively Multilingual Code Debugging

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.746312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.746312Z digest=sha256:232c977d21a1c962070a91751af4602a39ae5a0277c7354803742d619e13b922

Observation 1257f69c-e07f-4bb1-aeac-b723add7ebc4 · outbound

This paper cites Nostalgebraist.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Nostalgebraist

Reference 2017

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T18:34:27.189377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.758125Z digest=sha256:e1ca53633251b5e884f11b1c93cdfccfc33e9fd0a8fbac21ff28372460a9fdff

Observation aea7e8c3-cfe4-41e4-ab65-61c26e5b097a · outbound

This paper cites an unresolved cited work.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:34:27.371913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T18:34:26.730433Z digest=sha256:83cf1b485f1bb9428689916d54fde7a183f25dc1486fd6da5868707aaaed51ef

Observation 819165ce-9185-4f63-90b9-ad869203cfcd · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.724372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.724372Z digest=sha256:df23665e1e7ae86083fc08ff3b362260f6fcb0b9a75dcb7c11193b21f6f3d230

Observation b990aa96-800a-4fad-a30c-3b72dc7b3148 · outbound

This paper cites The Llama 3 Herd of Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models The Llama 3 Herd of Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.705289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.705289Z digest=sha256:f69834c92d539073a58d7be1e36071626a0fa28a59594578e5167ed106fd3a41

Observation 2f6b8b88-8861-4f05-9c91-f210d82fe01a · outbound

This paper cites Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.696989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.696989Z digest=sha256:e35d628b6f728f9dd1f81c3625016f192d1df56bbb7839653f1561b5f06f1814

Observation 4b70aba3-ed82-4839-b29d-c718607ce05a · outbound

This paper cites Program Synthesis with Large Language Models.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Program Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.692438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.692438Z digest=sha256:14190f02c28b476ba67ad9ffde598ab6b86d11cbaba943d3348a959661b280dd

Pith citing papers

Observation 21535db2-bb2d-4c3f-a2ec-3be81e92849b · inbound

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation cites this paper.

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation CryptoX : Compositional Reasoning Evaluation of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:26.867470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:26.867470Z digest=sha256:8e6f244d53edcf7c4d1355ab1888c55ff8fa56e670b02364bc08e13ad53f7d4f

Observation e24ea313-392a-407b-8bb5-786d85133471 · inbound

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data cites this paper.

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data CryptoX : Compositional Reasoning Evaluation of Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.146308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T13:45:38.305767Z digest=sha256:ee9609e3e14320da284745be47a1105c8799a91bd9ee17dc289cbf0f247a1e7a