Pith. sign in

Paper Citation Record · LEDGER

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.17970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17970 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:12:43.381851Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e321e14-b6e4-4359-9ae2-ae5a99b319c6 · outbound

This paper cites Which modality should I use - text, motif, or image? : Understanding graphs with large language models.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Which modality should I use - text, motif, or image? : Understanding graphs with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:12:43.761628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:12:43.278589Z digest=sha256:17cab027029cdf00288ac1b9b3f32d0b7c716ec7971ef33a86b78e7dac05e57a

Observation cce26841-3931-42dd-a016-4695e79722ef · outbound

This paper cites GPT4Graph: Can Large Language Models Understand Graph Structured Data ? An Empirical Evaluation and Benchmarking.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models GPT4Graph: Can Large Language Models Understand Graph Structured Data ? An Empirical Evaluation and Benchmarking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.284336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.284336Z digest=sha256:3c4b59ede9967eb121056a970a113f8db4f6620b495319637cd218715b034e7f

Observation ddbb3e12-ee5a-4089-9ef8-8fe37ba65873 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.296105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.296105Z digest=sha256:9019eafea2fa2d49a2c78f637914a0a4d8e29cf9f5da7f64ed327c78d7681938

Observation e7abd8f6-3351-44fa-8e1f-cfe5ef9951e2 · outbound

This paper cites Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.308046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.308046Z digest=sha256:a5dd8c09e37be9c284bac4de47fafddafbfe6dcebcd8c139dc7d87cfae046ee6

Observation efe915b9-89b2-45f7-a9c2-a6c82c355041 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.313289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.313289Z digest=sha256:018eeb89410fbe0eb4c3ed3b293cfb0a408fc890e4979538731709676f958427

Observation 410715b7-b360-46f4-8897-74b558352d82 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.323516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.323516Z digest=sha256:7fd3564926a7a12e76d402f9d0c1adc3faf4df0a990e8be6d3a8a3ea9804a4c0

Observation 7ec14147-348a-42af-a786-60faa50d9ce3 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.329468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.329468Z digest=sha256:b7cd08b882e73b7ee387ee2520978ff492fef92cf7f7ebdaffe2c68c8a7da6e8

Observation 4e205a35-27de-4950-8d71-61de680c88c3 · outbound

This paper cites CELLO: Causal Evaluation of Large Vision-Language Models.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models CELLO: Causal Evaluation of Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.334604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.334604Z digest=sha256:1eb12e39a669a00e8a3d5a7ad7939d03db86fa2f2705d04488cf666981624a41

Observation db5fc4f9-a60e-4e3f-8d4b-8368119f7a73 · outbound

This paper cites doi: 10.1162/tacl a 00446.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models doi: 10.1162/tacl a 00446

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-11T05:12:43.339710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.339710Z digest=sha256:b4bdbddb43a6d86aaac29e35ceadf8fd3f78ea4bc27e0afd7ecf8582fbafad10

Observation bbdea996-a866-4be3-bbbd-a7927e5fec55 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.355171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.355171Z digest=sha256:88a022043bfe1d6a865c7aebae9134594a88d650b7ab4e10a2a5c1df3b25403e

Observation 882aeca8-3b24-46d1-bc10-8f1058cb3e44 · outbound

This paper cites LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.360460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.360460Z digest=sha256:4aff707f179a0db33c91ec771df55f1e8196d95e0c728cd0a6ba358842f36762

Observation 03a001b8-ab05-49cf-8762-f2a6f9e3f65c · outbound

This paper cites Causal Evaluation of Language Models.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Causal Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.365536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.365536Z digest=sha256:24757525fcf55bcb2ea46f2c578de50535203d173fc0f8cbb287dea6625b3337

Observation b9d36e33-eac4-4069-b3f4-806d1e71fc7b · outbound

This paper cites URL https: //www.kaggle.com/m/3301.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models URL https: //www.kaggle.com/m/3301

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.370932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.370932Z digest=sha256:54e92b7c76adb3d59f40c0ef0335eccd2c259c9df86d656060777c2c465e2947

Observation 6cc21b9c-a814-47ff-8612-ebe2951d6a21 · outbound

This paper cites Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.375998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.375998Z digest=sha256:9d0149e39690c9d536d061f2b6a606969cf09ede7dc4e03bed2d86315cf45c9b

Observation 3a74b270-bf06-45eb-afd1-f0f17112696a · outbound

This paper cites GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.381851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.381851Z digest=sha256:27b96b238d4b5e59ed7bb3bded400ffd6cfc091c3b1664acc740cb6a7fb3dc01

Observation e8517cd0-0800-479a-abca-ef29594ac560 · outbound

This paper cites doi: 10.3115/v1/P15-1142.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models doi: 10.3115/v1/P15-1142

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.349932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.349932Z digest=sha256:2197d444df12984d43a43115248f2feb28db7de9c0d5b83e653e2eeec3310489

Observation b50c1fa0-d44f-413b-9fca-8e79188792fe · outbound

This paper cites Mistral 7B.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Mistral 7B

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.301848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.301848Z digest=sha256:630eb0d97a0cddd68ab85874c233e5b1c7c3c04e6761b9387b26bd765d6e57a2

Observation 44acbef8-94bc-49bf-896c-66c3d8321ee3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.272640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.272640Z digest=sha256:ad88fdeebe2ef8cdc5cd9586d89917cb65ecbb341aab9ee8974ec58cd93c7dde

Observation 1388eede-7ba2-46c9-989c-07deafe7eef6 · outbound

This paper cites doi: 10.18653/v1/2020.emnlp-main.89.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models doi: 10.18653/v1/2020.emnlp-main.89

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.344468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.344468Z digest=sha256:8d79797aef058d39802bb7efa12c54dd146492f36b7d3a3483b2669230f40e55

Observation 839995f1-f280-4122-8ade-a246f0db9a3c · outbound

This paper cites Qwen Technical Report.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Qwen Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.260483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.260483Z digest=sha256:feaafa25f332a964fa0e115629d88d22dc860534aafc0291143d256f269df07e

Observation f6152efc-e87d-4250-ac75-cc54b726fc34 · outbound

This paper cites Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.318435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.318435Z digest=sha256:f7cff4149d18e99352f377f72d8873e13516aa2409e8cfca6ca767778600970d

Observation ac627b00-9b4d-48eb-baa4-a5df380691f4 · outbound

This paper cites The Essential Role of Causality in Foundation World Models for Embodied AI.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models The Essential Role of Causality in Foundation World Models for Embodied AI

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.289901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.289901Z digest=sha256:355bc3e01c2a5991f7f42881d3709a1454aeb1d61c7c662a499cbad585bae0dd

Observation 54801bf9-0f58-4a8c-87cf-6bfa13526d57 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:43.266768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:43.266768Z digest=sha256:886cc230b69e30f4aadff8a34e9beaebc74c13cbfb3352f0b25fe779b441277f

Pith citing papers

No inbound Pith citation observations are available.