Pith. sign in

Paper Citation Record · LEDGER

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.17180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17180 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:31.684757Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c955ac35-2da0-4220-bc05-89b0955cb2b4 · outbound

This paper cites Assessment in science education: A study of teaching effectiveness.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Assessment in science education: A study of teaching effectiveness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.856619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:27.638495Z digest=sha256:f39f94582dedc1e9ef3db533309b6ae3ca6cd45899747ffcf691468c0427bfe3

Observation 6c8dde00-4a1f-49cc-b6f3-e5def3c6f121 · outbound

This paper cites Cbse assessment framework for science, maths and social science classes 9 and 10, 2020.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Cbse assessment framework for science, maths and social science classes 9 and 10, 2020

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.643752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:27.847326Z digest=sha256:63d39f09d56e87b1d39303aad6ef41630d443e569d32ef206a0e23f1c2cff38b

Observation f70a2faf-7ce5-4f9d-9d97-80effefa7046 · outbound

This paper cites Bloom, Max D.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Bloom, Max D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.467416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:27.999453Z digest=sha256:20d3a58c7e9fc66c8b5cf9f358871a3383efc76dcae6393551884b68bca716bb

Observation cc3aac27-7414-47c1-af97-732c5a7629cf · outbound

This paper cites Anderson, David R.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Anderson, David R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.304831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:28.216326Z digest=sha256:f7d1df7945b5df5840c34b6efc66e1a4cbba5b9b32d5caa88020e533f3a28546

Observation 4e0c1e0d-557a-4186-b9a5-aec58b7c7245 · outbound

This paper cites Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.115224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:28.287009Z digest=sha256:2b94fd4c4f326d2fb86f6d82cab335784919df51c0ae202d834549e094c42bea

Observation 5fa84c2e-fb64-4a47-ba99-f6ef48232ab1 · outbound

This paper cites Thinking, Fast and Slow.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Thinking, Fast and Slow

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:28.419985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:28.419985Z digest=sha256:2a653f69ba7f6eaf45a173dd9bfe29239052bbea4a7a2c108a7eec41f6560c27

Observation 299d020f-5a79-4b5d-bbbf-7ddbf98535b0 · outbound

This paper cites Dual-process theories of higher cognition: Advancing the debate.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Dual-process theories of higher cognition: Advancing the debate

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.916792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:28.635181Z digest=sha256:64a961c7090292320adc9bca29e23246d5941e22b89e26b5334d9ef4c2d93ec5

Observation ebf56e91-d8d5-4222-b94a-f513544a5419 · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.740210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:28.734129Z digest=sha256:85e4ca480a4257995b12b3b5b58ecf7aca790eda9dbfc70be9466e4e2b40aeeb

Observation ddecb367-5cf1-4aaa-993a-984ad2228a55 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:35.621043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:28.877039Z digest=sha256:be1fdac255c0018d5a13757886662bc65eaf31c3fbc5ed31787bcc87ff20af02

Observation c2bccd31-ea7c-4933-9deb-3f9c479135be · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models SQuAD: 100,000+ questions for machine comprehension of text

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.456528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:29.049684Z digest=sha256:b0222823e426211cc9c03fa0fe1bf7f05e99e2f92a4333812d0d8e0e11a97044

Observation b4ec3bf0-a908-4e57-a880-2bf82157f0a2 · outbound

This paper cites Hutchinson, and Richard G.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Hutchinson, and Richard G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.200748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:29.223595Z digest=sha256:7df7a3dae505aa960445d48635fbf9a05384288d2cee5e0bf513f0cf006a60db

Observation 77217a69-f736-4e16-8949-2b0be8431aa2 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.960660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:29.379705Z digest=sha256:a5e8d8be2b3ab91c1e87e719c57d2c83d7f8479a675761163482d20d059969c3

Observation 6c0ef4bf-850f-47b7-b6fc-291edba5b592 · outbound

This paper cites WinoGrande: An adversarial Winograd schema challenge at scale.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models WinoGrande: An adversarial Winograd schema challenge at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.641689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:29.555304Z digest=sha256:376b865373019e432b5e2b8ed2403e6eaf8471e4b08015de9ed408ef1bce300f

Observation c78751fa-86a9-4360-a5a7-7734acbf13ae · outbound

This paper cites FEVER: A large-scale dataset for fact extraction and verification.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models FEVER: A large-scale dataset for fact extraction and verification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.476987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:29.641483Z digest=sha256:05a8fe633e2cd5c3151ab0946ba258005cbbc829882289a4556bb38ebd9aab25

Observation 591bbfb5-3252-403c-af8e-d41d352e88e8 · outbound

This paper cites Fact or Fiction: Verifying Scientific Claims.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Fact or Fiction: Verifying Scientific Claims

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:29.816345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:29.816345Z digest=sha256:2f97056a1fb28cad48b43d82bb3f024091044112ef97b6183f7f4f1364c2b25e

Observation 35b332ff-9fc1-4faa-aa6b-2531285d7808 · outbound

This paper cites The Llama 3 Herd of Models.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:29.937166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:29.937166Z digest=sha256:daa2125a6785c2058c7da43b532ecbe792631eda0da75ad141732ca1115184ad

Observation 32d3aeae-e6f9-4c04-9e40-fdcf7673f323 · outbound

This paper cites Qwen2.5 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Qwen2.5 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.032234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.032234Z digest=sha256:32c032ff8b8925fbcd46d847497d95b87f7d74c514cf9f93d32923c457efed60

Observation f3837aa7-8513-464a-bf76-ca83285f287d · outbound

This paper cites Qwen3 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.125545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.125545Z digest=sha256:ba3d05e0e3f91ba79d819132a7ec376b8428a5bb7786b63aba952080847eaec5

Observation 96a14359-8ee3-4192-b5b0-ed03107b1a3a · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.266416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.266416Z digest=sha256:e4b74fc15616d48b955da29278ed39f98e9cfb0d13c2bd034673850dbda284fe

Observation ff585210-453b-490f-8e1d-f95af7e25606 · outbound

This paper cites Phi-4 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Phi-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.385458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.385458Z digest=sha256:6c0a288a3592c2940ba2c809aad6c0d6aadb9847cfe483241e21fd14fd484e7f

Observation 6b1b6237-cd99-4231-83e4-e464c67bf3dd · outbound

This paper cites When one model casts doubt on another: A levels-of-analysis approach to causal discounting.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models When one model casts doubt on another: A levels-of-analysis approach to causal discounting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.297459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:30.554878Z digest=sha256:ba86462aea88b2be8a40fd635c8a6454b43a69db64d26a8bc6bcac82a534810a

Observation 055c84a6-a36c-421e-ae31-884271148fc2 · outbound

This paper cites The most common error is incorrectly classifying option (b) as option (a) (176-214 instances across models).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models The most common error is incorrectly classifying option (b) as option (a) (176-214 instances across models)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.094387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:30.735614Z digest=sha256:c88b2d86b0b39feb2df017b46a66bf7f1f224e25be445c0ffbfaaafd03b788c3

Observation 871c92f1-1083-405e-a5ba-60a540d572f7 · outbound

This paper cites Even the largest model, Qwen3-32B, misclassifies option (b) as option (a) in 195 cases (27.1% of all true (b) cases).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Even the largest model, Qwen3-32B, misclassifies option (b) as option (a) in 195 cases (27.1% of all true (b) cases)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:33.847413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:30.887249Z digest=sha256:16f69d22d88f5c15468ff6ce80f04ddd71c0fe1751806f830f12f446fdf550c9

Observation fe6cdd58-1579-48e2-95e8-1f11e5c126c6 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:33.658462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.013302Z digest=sha256:19b9cf06cfe836c768be95b9c33a4d65e5dbedb7af36d92c82fbf3a8d57868cf

Observation 4f8b2fdb-3e17-4b0c-99bd-2b73f6e7ca35 · outbound

This paper cites Model family differences suggest that some design and training strategies may better support causal reasoning, particularly in complex domains.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Model family differences suggest that some design and training strategies may better support causal reasoning, particularly in complex domains

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:33.339491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.082574Z digest=sha256:13541d569a8897c3fff7f1622edb0e03edd71f925e016d6a4339f4f72630b566

Observation b4062b7d-825e-4ea3-841f-3f0893eead36 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:33.017115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.206215Z digest=sha256:99007c4e0105a18624cc089743f6129e9178792e6a27f22443f5072183516619

Observation 47c042e4-5365-4b94-8326-4dc3c76dc8b1 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:32.804755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.284075Z digest=sha256:d49fc28973757df35bd895b7a7d22e984e7750086526c41e96e171070afe5787

Observation d5df085c-b832-4137-8289-6784299fc9e8 · outbound

This paper cites These t-tests determine whether the differences in similarity are statistically significant or could have occurred by chance.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models These t-tests determine whether the differences in similarity are statistically significant or could have occurred by chance

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:32.563860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.431743Z digest=sha256:242ff9a7d9b22b016bc8c058ee84eb0d5d7c3a803895fc75d5e25eb4f8435a1a

Observation 58740287-4290-4ce3-8d25-b045a34a1198 · outbound

This paper cites Yes" versus all cases where it predicted.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Yes" versus all cases where it predicted

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:32.318389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.532676Z digest=sha256:4614de8ba28d3a33d8f05f57a3482515738369fb38a537ffad13982c5dfffec2

Observation 82369732-23bf-4c32-b2c9-944a32f473f5 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:32.123936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.612422Z digest=sha256:7a369c71d0c1deb78f93a4f6d3937c50cd991413b13c9356992be3e046fd4f63

Observation b7dd928c-1127-47ae-bb7d-6a017bf98679 · outbound

This paper cites This pattern represents a fundamental confusion of correlation (semantic similarity) with causation (explanatory relationship).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models This pattern represents a fundamental confusion of correlation (semantic similarity) with causation (explanatory relationship)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:31.976303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:31.684757Z digest=sha256:e2c14e39bae608cf0d8e89189c70e4c9addcf739192c1a5938937633b8507d57

Pith citing papers

No inbound Pith citation observations are available.