Pith. sign in

Paper Citation Record · LEDGER

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.17180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17180 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:31.684757Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c955ac35-2da0-4220-bc05-89b0955cb2b4 · outbound

This paper cites Assessment in science education: A study of teaching effectiveness.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Assessment in science education: A study of teaching effectiveness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.856619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:27.638495Z digest=sha256:ce160bffc52762bb970d7b9631ea5b0f0419ef0a2db2bef7cdeb00926c35ead8

Observation 6c8dde00-4a1f-49cc-b6f3-e5def3c6f121 · outbound

This paper cites Cbse assessment framework for science, maths and social science classes 9 and 10, 2020.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Cbse assessment framework for science, maths and social science classes 9 and 10, 2020

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.643752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:27.847326Z digest=sha256:bf99e04ffe13e87a2be7ee76e1873f51d08c778348fcd549a07ea9d97a570811

Observation f70a2faf-7ce5-4f9d-9d97-80effefa7046 · outbound

This paper cites Bloom, Max D.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Bloom, Max D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.467416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:27.999453Z digest=sha256:810744168e6bdf3da1a54589ea21f8d0dd681e2f8302a84995db139e5814548c

Observation cc3aac27-7414-47c1-af97-732c5a7629cf · outbound

This paper cites Anderson, David R.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Anderson, David R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.304831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:28.216326Z digest=sha256:4985033f04aa9882db2441f10a5b8d4e7aa77d45479f5289aa946d75fdb7550a

Observation 4e0c1e0d-557a-4186-b9a5-aec58b7c7245 · outbound

This paper cites Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.115224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:28.287009Z digest=sha256:8499173803de92c6109760afd51e7011a497ecc445ff293b466bd6444fea812e

Observation 5fa84c2e-fb64-4a47-ba99-f6ef48232ab1 · outbound

This paper cites Thinking, Fast and Slow.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Thinking, Fast and Slow

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:28.419985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:28.419985Z digest=sha256:6b12546a6b1c3b0156ffff57170deccdf08029f2c4c95a83a0f71757d52f588a

Observation 299d020f-5a79-4b5d-bbbf-7ddbf98535b0 · outbound

This paper cites Dual-process theories of higher cognition: Advancing the debate.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Dual-process theories of higher cognition: Advancing the debate

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.916792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:28.635181Z digest=sha256:593f05352fd15c6f78d1451b089e58fcf6d95bd905a3d4fd589985d1930f9fdd

Observation ebf56e91-d8d5-4222-b94a-f513544a5419 · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.740210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:28.734129Z digest=sha256:ed3a99a3c9a8241cdafb9081f3b4f3f1bf19d75ad2c698994ccec2fb5955bc16

Observation ddecb367-5cf1-4aaa-993a-984ad2228a55 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:35.621043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:28.877039Z digest=sha256:5c5bb37224cbd1fda67b1929a62cd3c1123ea2e118779777f9516d884eb63d2d

Observation c2bccd31-ea7c-4933-9deb-3f9c479135be · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models SQuAD: 100,000+ questions for machine comprehension of text

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.456528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:29.049684Z digest=sha256:e97ee9d7a0b53cc01a2b41e17e8757ffab1ba38b3ce7d1d8bdff38b205037a46

Observation b4ec3bf0-a908-4e57-a880-2bf82157f0a2 · outbound

This paper cites Hutchinson, and Richard G.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Hutchinson, and Richard G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.200748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:29.223595Z digest=sha256:f4dac31696ae0a0801be7fbb21829374812b86516c78ae04c857d9d18c55cb50

Observation 77217a69-f736-4e16-8949-2b0be8431aa2 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.960660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:29.379705Z digest=sha256:84c5dcbc6012c63957994b8231094ea6d68650d06b0175b2e7bfdc1180a98975

Observation 6c0ef4bf-850f-47b7-b6fc-291edba5b592 · outbound

This paper cites WinoGrande: An adversarial Winograd schema challenge at scale.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models WinoGrande: An adversarial Winograd schema challenge at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.641689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:29.555304Z digest=sha256:2ba77950e0c981b73fe383991eaf7f823e563537952f349a00d7871c71b61051

Observation c78751fa-86a9-4360-a5a7-7734acbf13ae · outbound

This paper cites FEVER: A large-scale dataset for fact extraction and verification.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models FEVER: A large-scale dataset for fact extraction and verification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.476987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:29.641483Z digest=sha256:6fcd3c931f25c0166b146022430f4944b80a2240caa579f2c7c7b5bb0f73ea25

Observation 591bbfb5-3252-403c-af8e-d41d352e88e8 · outbound

This paper cites Fact or Fiction: Verifying Scientific Claims.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Fact or Fiction: Verifying Scientific Claims

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:29.816345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:29.816345Z digest=sha256:81801b150c8653309c7f79015065237e8440d8ce9758fbf66531409caa1f0298

Observation 35b332ff-9fc1-4faa-aa6b-2531285d7808 · outbound

This paper cites The Llama 3 Herd of Models.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:29.937166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:29.937166Z digest=sha256:2759ea753033c617fda49d6ae8eb2222b92a327821c33692a6bbaef41df04610

Observation 32d3aeae-e6f9-4c04-9e40-fdcf7673f323 · outbound

This paper cites Qwen2.5 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Qwen2.5 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.032234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.032234Z digest=sha256:c32d0738e9e821cf9ee211b96efe8b584bd076a66500fa89772b651c217aef49

Observation f3837aa7-8513-464a-bf76-ca83285f287d · outbound

This paper cites Qwen3 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.125545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.125545Z digest=sha256:df92f670a0025589af0c8263405af6208bf69bc895bb85535fa789712a311184

Observation 96a14359-8ee3-4192-b5b0-ed03107b1a3a · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.266416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.266416Z digest=sha256:248419ba213bdb795426b2a2038bd3794a86d9ac8f2e7d940b85c86ce75d93e9

Observation ff585210-453b-490f-8e1d-f95af7e25606 · outbound

This paper cites Phi-4 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Phi-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.385458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.385458Z digest=sha256:a98a349fe5d40be7ab069d955f6e05c2b040a649849680fc4438599f7da1ee0c

Observation 6b1b6237-cd99-4231-83e4-e464c67bf3dd · outbound

This paper cites When one model casts doubt on another: A levels-of-analysis approach to causal discounting.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models When one model casts doubt on another: A levels-of-analysis approach to causal discounting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.297459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:30.554878Z digest=sha256:063c39bed4215c42acebdcf3866288ca95b2dcdbd4e665a4fa154c478f1ae145

Observation 055c84a6-a36c-421e-ae31-884271148fc2 · outbound

This paper cites The most common error is incorrectly classifying option (b) as option (a) (176-214 instances across models).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models The most common error is incorrectly classifying option (b) as option (a) (176-214 instances across models)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.094387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:30.735614Z digest=sha256:0277028ffeb7a1be3dc8541a864dafdfdec7d992242f3280cfef5d74ab99e65b

Observation 871c92f1-1083-405e-a5ba-60a540d572f7 · outbound

This paper cites Even the largest model, Qwen3-32B, misclassifies option (b) as option (a) in 195 cases (27.1% of all true (b) cases).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Even the largest model, Qwen3-32B, misclassifies option (b) as option (a) in 195 cases (27.1% of all true (b) cases)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:33.847413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:30.887249Z digest=sha256:1f22f4b5281a93f9e5c0d7fc041c6cc45006663479ed2d749d714414cce5d382

Observation fe6cdd58-1579-48e2-95e8-1f11e5c126c6 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:33.658462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.013302Z digest=sha256:3f31391eed668c5cbba7886ac3d6a78d43f02cde44c326e8abd0df3bbcc55ac1

Observation 4f8b2fdb-3e17-4b0c-99bd-2b73f6e7ca35 · outbound

This paper cites Model family differences suggest that some design and training strategies may better support causal reasoning, particularly in complex domains.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Model family differences suggest that some design and training strategies may better support causal reasoning, particularly in complex domains

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:33.339491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.082574Z digest=sha256:91174b217b4d225da506f4846405266358f2dfc2fd2efc037b5714bea53549cb

Observation b4062b7d-825e-4ea3-841f-3f0893eead36 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:33.017115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.206215Z digest=sha256:1f812665e5274fad9e3f2cdc383332f0780e05d2d7d1e9670fbb34c1d3471d0c

Observation 47c042e4-5365-4b94-8326-4dc3c76dc8b1 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:32.804755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.284075Z digest=sha256:0d33980408c0030c4cd202f447fecd173ddf7d05de7f88d60b793d9dc6ec2940

Observation d5df085c-b832-4137-8289-6784299fc9e8 · outbound

This paper cites These t-tests determine whether the differences in similarity are statistically significant or could have occurred by chance.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models These t-tests determine whether the differences in similarity are statistically significant or could have occurred by chance

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:32.563860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.431743Z digest=sha256:7c0c8280f4c25815cb44ae9f5d05712d4b76edca82d3b9ea207538c5f8e1eeb5

Observation 58740287-4290-4ce3-8d25-b045a34a1198 · outbound

This paper cites Yes" versus all cases where it predicted.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Yes" versus all cases where it predicted

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:32.318389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.532676Z digest=sha256:e2fa63bd83197788bc4142411b0eb3aef09589082993402d6b22e2dc9460497e

Observation 82369732-23bf-4c32-b2c9-944a32f473f5 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:32.123936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.612422Z digest=sha256:1c24efd279de9705d603d21714dfa2a2ee21dae4ce876adc38c0d6200840f786

Observation b7dd928c-1127-47ae-bb7d-6a017bf98679 · outbound

This paper cites This pattern represents a fundamental confusion of correlation (semantic similarity) with causation (explanatory relationship).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models This pattern represents a fundamental confusion of correlation (semantic similarity) with causation (explanatory relationship)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:31.976303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:31.684757Z digest=sha256:5e9b2a2759c683a4ac20d5dec6caa9d3d43e3444e68cdd890944165a217bff17

Pith citing papers

No inbound Pith citation observations are available.