Pith. sign in

Paper Citation Record · LEDGER

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models

As of 12 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.17970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17970 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:12:43.381851Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e321e14-b6e4-4359-9ae2-ae5a99b319c6 · outbound

This paper cites Which modality should I use - text, motif, or image? : Understanding graphs with large language models.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Which modality should I use - text, motif, or image? : Understanding graphs with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:12:43.761628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:12:43.278589Z digest=sha256:701dc2ebf16f1925320f20296aaa3f5b24c2a5a2b465da7bbb7760ef3cc4d5ed

Observation cce26841-3931-42dd-a016-4695e79722ef · outbound

This paper cites GPT4Graph: Can Large Language Models Understand Graph Structured Data ? An Empirical Evaluation and Benchmarking.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models GPT4Graph: Can Large Language Models Understand Graph Structured Data ? An Empirical Evaluation and Benchmarking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.284336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.284336Z digest=sha256:3c4b59ede9967eb121056a970a113f8db4f6620b495319637cd218715b034e7f

Observation ddbb3e12-ee5a-4089-9ef8-8fe37ba65873 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.296105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.296105Z digest=sha256:9019eafea2fa2d49a2c78f637914a0a4d8e29cf9f5da7f64ed327c78d7681938

Observation e7abd8f6-3351-44fa-8e1f-cfe5ef9951e2 · outbound

This paper cites Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.308046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.308046Z digest=sha256:27210023c2116c4967e3721a3103779314131f72bd05b946c0e398b06c2765ef

Observation efe915b9-89b2-45f7-a9c2-a6c82c355041 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.313289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.313289Z digest=sha256:018eeb89410fbe0eb4c3ed3b293cfb0a408fc890e4979538731709676f958427

Observation 410715b7-b360-46f4-8897-74b558352d82 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.323516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.323516Z digest=sha256:7fd3564926a7a12e76d402f9d0c1adc3faf4df0a990e8be6d3a8a3ea9804a4c0

Observation 7ec14147-348a-42af-a786-60faa50d9ce3 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.329468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.329468Z digest=sha256:b7cd08b882e73b7ee387ee2520978ff492fef92cf7f7ebdaffe2c68c8a7da6e8

Observation 4e205a35-27de-4950-8d71-61de680c88c3 · outbound

This paper cites CELLO: Causal Evaluation of Large Vision-Language Models.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models CELLO: Causal Evaluation of Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.334604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.334604Z digest=sha256:a4b55fb903b27372cd5785271d5abfe20563c09a72de25ed2d340dd8e02e69b0

Observation db5fc4f9-a60e-4e3f-8d4b-8368119f7a73 · outbound

This paper cites doi: 10.1162/tacl a 00446.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models doi: 10.1162/tacl a 00446

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-11T05:12:43.339710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.339710Z digest=sha256:b4bdbddb43a6d86aaac29e35ceadf8fd3f78ea4bc27e0afd7ecf8582fbafad10

Observation bbdea996-a866-4be3-bbbd-a7927e5fec55 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.355171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.355171Z digest=sha256:88a022043bfe1d6a865c7aebae9134594a88d650b7ab4e10a2a5c1df3b25403e

Observation 882aeca8-3b24-46d1-bc10-8f1058cb3e44 · outbound

This paper cites LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.360460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.360460Z digest=sha256:e1a47509c7adaac7eb106659ae444085b0628e7be769ac348de4ff65fdd4d7d3

Observation 03a001b8-ab05-49cf-8762-f2a6f9e3f65c · outbound

This paper cites Causal Evaluation of Language Models.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Causal Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.365536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.365536Z digest=sha256:5a4a45246be364d5ffde38c1920db4eb8af070d1f41d8a6192e0b95b567c84d4

Observation b9d36e33-eac4-4069-b3f4-806d1e71fc7b · outbound

This paper cites URL https: //www.kaggle.com/m/3301.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models URL https: //www.kaggle.com/m/3301

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.370932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.370932Z digest=sha256:54e92b7c76adb3d59f40c0ef0335eccd2c259c9df86d656060777c2c465e2947

Observation 6cc21b9c-a814-47ff-8612-ebe2951d6a21 · outbound

This paper cites Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.375998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.375998Z digest=sha256:b6bbd4c8815dceea4bd96ed28af8dc71a0c6a1b467ac9f0b33a65553a2ca1fcb

Observation 3a74b270-bf06-45eb-afd1-f0f17112696a · outbound

This paper cites GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.381851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.381851Z digest=sha256:0beb2ffe11baa60e9297de055c946615c7afc91bcd1577a3aefc0960e5c6be4d

Observation e8517cd0-0800-479a-abca-ef29594ac560 · outbound

This paper cites doi: 10.3115/v1/P15-1142.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models doi: 10.3115/v1/P15-1142

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.349932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.349932Z digest=sha256:2197d444df12984d43a43115248f2feb28db7de9c0d5b83e653e2eeec3310489

Observation b50c1fa0-d44f-413b-9fca-8e79188792fe · outbound

This paper cites Mistral 7B.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Mistral 7B

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.301848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.301848Z digest=sha256:630eb0d97a0cddd68ab85874c233e5b1c7c3c04e6761b9387b26bd765d6e57a2

Observation 44acbef8-94bc-49bf-896c-66c3d8321ee3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.272640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.272640Z digest=sha256:ad88fdeebe2ef8cdc5cd9586d89917cb65ecbb341aab9ee8974ec58cd93c7dde

Observation 1388eede-7ba2-46c9-989c-07deafe7eef6 · outbound

This paper cites doi: 10.18653/v1/2020.emnlp-main.89.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models doi: 10.18653/v1/2020.emnlp-main.89

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.344468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.344468Z digest=sha256:8d79797aef058d39802bb7efa12c54dd146492f36b7d3a3483b2669230f40e55

Observation 839995f1-f280-4122-8ade-a246f0db9a3c · outbound

This paper cites Qwen Technical Report.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Qwen Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.260483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.260483Z digest=sha256:feaafa25f332a964fa0e115629d88d22dc860534aafc0291143d256f269df07e

Observation f6152efc-e87d-4250-ac75-cc54b726fc34 · outbound

This paper cites Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.318435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.318435Z digest=sha256:7d4b0c58c74194034673d98d99327eb1e1ad8dfda398ffefac3052ee86d11b8c

Observation ac627b00-9b4d-48eb-baa4-a5df380691f4 · outbound

This paper cites The Essential Role of Causality in Foundation World Models for Embodied AI.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models The Essential Role of Causality in Foundation World Models for Embodied AI

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.289901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.289901Z digest=sha256:424d52d11a3ec0177ca1ac9c4623d5180634f7527d19db1ac9ba3ff7cef479b7

Observation 54801bf9-0f58-4a8c-87cf-6bfa13526d57 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.266768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.266768Z digest=sha256:886cc230b69e30f4aadff8a34e9beaebc74c13cbfb3352f0b25fe779b441277f

Pith citing papers

No inbound Pith citation observations are available.