Pith. sign in

Paper Citation Record · LEDGER

THiNK: Can Large Language Models Think-aloud?

As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2505.20184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20184 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:37.575936Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:53.998837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved15
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a7c8880-5fd9-469d-bcb4-c7d5e50c4e4d · outbound

This paper cites By analyzing which concepts are invoked, we assess the knowledge dimension activated during problem-solving and whether the LLM navigates these domains coherently.

THiNK: Can Large Language Models Think-aloud? By analyzing which concepts are invoked, we assess the knowledge dimension activated during problem-solving and whether the LLM navigates these domains coherently

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.254905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.403974Z digest=sha256:49061dfd53a86ad74f254ba984d6926ab6aea99a46298bddddfa6e8999e900eb

Observation 0fa765ff-d851-481b-a821-0373e1944790 · outbound

This paper cites Thilo Hagendorff, Sarah Fabi, and Michal Kosinski.

THiNK: Can Large Language Models Think-aloud? Thilo Hagendorff, Sarah Fabi, and Michal Kosinski

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:44.055481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:34.621929Z digest=sha256:da07680a6a1521f1660dd98a3726f9e18942b0b6eb3e6418f69dc91f8064967c

Observation a4aa73a5-1796-43c5-aa1f-a76b5d6ed06c · outbound

This paper cites Representations are critical for logical coherence and traceability in reasoning.

THiNK: Can Large Language Models Think-aloud? Representations are critical for logical coherence and traceability in reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.892838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.570182Z digest=sha256:80be08f10174d043dccbf9712e870bba87d4ae696310cf6ddf127d18ae695feb

Observation 2f68695c-0270-47e6-81cb-67b468eebf67 · outbound

This paper cites A model’s ability to adapt its reasoning across such variants reflects generalization ability—an essential attribute of HOT.

THiNK: Can Large Language Models Think-aloud? A model’s ability to adapt its reasoning across such variants reflects generalization ability—an essential attribute of HOT

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.649080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.638056Z digest=sha256:8fc340833784166b5609c69b81198622c45c20fc2ceb6f62e36dacfd48688f1f

Observation d9d1a54d-e351-4c68-a515-5b41cb3329af · outbound

This paper cites Five Keys.

THiNK: Can Large Language Models Think-aloud? Five Keys

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.396321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.760718Z digest=sha256:c828037e3014543e4d507b0d7601f1c395c55a9597c906d7e9d6ab5054f202ee

Observation 30a63fb8-e0f4-41a6-a468-e77e64a0ac3f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

THiNK: Can Large Language Models Think-aloud? Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.989895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.989895Z digest=sha256:4059faf9c3cd70ada7d044f5b7f1b2b105823a1e3cdd385429bcb315339c2665

Observation 6f285043-ceae-4a7e-b7ac-a08237cb27f7 · outbound

This paper cites Assessing and Understanding Creativity in Large Language Models.

THiNK: Can Large Language Models Think-aloud? Assessing and Understanding Creativity in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:35.188891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:35.188891Z digest=sha256:17f6e5e0b10042e6a993d23565d1e2d58b9dc775ab932f22a24b7c8c1b2b745c

Observation eda686bc-321f-44f8-b9d2-085ef43b9912 · outbound

This paper cites Five Keys.

THiNK: Can Large Language Models Think-aloud? Five Keys

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.413353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.274686Z digest=sha256:06328b171f7c9802b86c393592e010ef17c55089e3abe3bb0cc6ac526189c87b

Observation fc6c6ce0-d87c-436f-b349-e6e86497d3cf · outbound

This paper cites These skills serve as proxies for prior knowledge and inform whether the LLM draws upon relevant background competence.

THiNK: Can Large Language Models Think-aloud? These skills serve as proxies for prior knowledge and inform whether the LLM draws upon relevant background competence

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.501537Z digest=sha256:c52603b2c1db921996ba9dfd69a3eff9fa5abbf0c22b375f477808f3c78bfdb5

Observation bc43f647-0c80-4f11-92ad-366e129e0133 · outbound

This paper cites You should try to understand and retrieve the specific mathematical information in it such as facts, patterns, objects, or contextual information, and decipher these meanings.

THiNK: Can Large Language Models Think-aloud? You should try to understand and retrieve the specific mathematical information in it such as facts, patterns, objects, or contextual information, and decipher these meanings

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.080713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.856691Z digest=sha256:2f75e040556270f877c8030e14df1aef69d1f28a7a68e266c20938367404d543

Observation f80c2551-b0ae-441d-87f4-f27be7b9cd08 · outbound

This paper cites It includes understanding and organizing information, analyzing relationships, drawing conclusions, and distinguishing nuances.

THiNK: Can Large Language Models Think-aloud? It includes understanding and organizing information, analyzing relationships, drawing conclusions, and distinguishing nuances

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.709139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.926296Z digest=sha256:dddeb3da93321d9307932df3d81a1874176845e7d22e1da0da0014de8c11ddc5

Observation 632569ac-4052-4187-86a6-dc67e329e041 · outbound

This paper cites These new expressions should have the same form as the given expressions in the previous generated math problem.

THiNK: Can Large Language Models Think-aloud? These new expressions should have the same form as the given expressions in the previous generated math problem

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.478200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:35.991143Z digest=sha256:924ca7a186c4697c80a93fc00879eec690a2fdedd22ec484bd7c862f8085f545

Observation d90fac4d-e73b-47fc-bc1f-49e61c612496 · outbound

This paper cites The generated stories must be a mathematical word problem with the corresponding expressions.

THiNK: Can Large Language Models Think-aloud? The generated stories must be a mathematical word problem with the corresponding expressions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.250540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.112437Z digest=sha256:c6f54f0ccbb4932fbb1b145967c03e2ce27c41bdd2ec000d1a55fc69ed80ac57

Observation 4ae966b1-69d3-4983-b64b-2753e21a22c2 · outbound

This paper cites question.

THiNK: Can Large Language Models Think-aloud? question

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:40.942421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.212697Z digest=sha256:7411cf25d3891f75a75ff5a35722ced4be664e61d5d78bcc2cad2e71f81d06a9

Observation d7221b4a-69f1-4292-9c7e-b07d16682968 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.731850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.299113Z digest=sha256:fb84f9b71b46f18984d63f17c8c4dc8b42b685227ffb9a9a5422e7888ca237db

Observation 097e0f7a-e8c0-4aa7-8a53-1112802d3640 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.516278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.371403Z digest=sha256:9ff4b7e76156b29cbf22941e2ef1d5ecaaa7ab13bfdfca33f20d6a51dbee702d

Observation e7cca71a-fa57-4018-81a1-996ea4f8a584 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.346332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.467020Z digest=sha256:c1ccd2c74af653060b8212a9c7efae7250334b9533637fa16f68f2092080ec91

Observation 71cfc3bd-c73a-4638-a83b-201870d4932c · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.119981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.557122Z digest=sha256:7710ea0a9b20dd0cfa65509d30f19e933df2f32d753bb0bc38e0d3ade38cc03d

Observation 0bd16435-849c-4b77-b423-35024e0a2a0d · outbound

This paper cites ID": null,.

THiNK: Can Large Language Models Think-aloud? ID": null,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:39.859244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.650966Z digest=sha256:2323d5c34cb8da1d78c219652be9a79dabcd02e5d620e5e7455af663b7df5434

Observation 905ea987-80e9-4efc-ac45-90783e9ccf18 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.683488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.725952Z digest=sha256:cffc4d6b796ab52c67ba050cb94d4366614df35dc702a61c8dd67b214d5e94d4

Observation 373900a1-c4ed-414f-acab-352e4fb4840a · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.422480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.802182Z digest=sha256:33d4236c484585160f47db1958eba80db149c0b9e9e2efaea04d6e101a92c3d3

Observation 28bbc1f8-ccd1-4e0f-a597-8346676859cb · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.228336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.871190Z digest=sha256:c178782ca435d684f7bf78a3a095272f9f230d7c68bc712247464c9dad50c716

Observation 79f727b5-a1b4-4900-b207-266a76810f5d · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.013489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:36.981345Z digest=sha256:fb6d8b5c027bfc866b7b914664d82608bbc858150d107765cce6d4702f4ff316

Observation fba5fb64-1ef4-447e-8f46-3851e00a2fc2 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:38.843275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:37.076085Z digest=sha256:9160324b04a1e9e176b76e77813721059b31b4d213dc00b36f334842f62d8d4a

Observation daaf974f-20f5-4e54-9b92-d6d3073a526b · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:38.604568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:37.196578Z digest=sha256:d404e430ac51dc36a6262aa7d9f3abac968224c11f10399e306bae934cc55f08

Observation 88110793-a66b-414c-aaa9-dd9f02d9fef1 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 31

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.445619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:37.286441Z digest=sha256:908c667b57b621cd07e8c524fcd2d69ff891da7fefa412fd98d357579e4d8008

Observation acfc7e07-2c28-4479-9c14-bb5374b74b76 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.191305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:37.367133Z digest=sha256:c61b56dd2c4d2d8394636945265d58c7838d09b048779306f1128fde45a2ed54

Observation 237518b0-cf69-410d-8cfa-93d6a34f5c0c · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 33

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.000547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:37.458862Z digest=sha256:6745ba4e31b032e00b0abf94aed45a78f31b25c735226cae7333b74fa1707074

Observation 7735d87f-a2ae-410b-84eb-b1424ad23de3 · outbound

This paper cites performance_score.

THiNK: Can Large Language Models Think-aloud? performance_score

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:37.798360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:37.575936Z digest=sha256:55a63e94e4dae9b21d8d7f795072539b2c6f13530a0c6cf576a33ae828ef9059

Observation 6e3bc5a6-54ea-4d38-a129-9d44eba41060 · outbound

This paper cites English language teaching, 3(4):237– 248.

THiNK: Can Large Language Models Think-aloud? English language teaching, 3(4):237– 248

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.595243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:34.885086Z digest=sha256:d6aaa7d45b52a8a94eca17397afcf55fd66e79bd018f978089bc4db84a9ae212

Observation 4f4a0605-e26a-479d-bef5-6bda3f3c3127 · outbound

This paper cites MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models.

THiNK: Can Large Language Models Think-aloud? MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:35.106825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:35.106825Z digest=sha256:5a3149c0bd3a18866920629ad4453c7dcea72126626f3edde8b691a85c164a7b

Observation 93bfd8cb-0986-4884-9228-76953960d739 · outbound

This paper cites HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.

THiNK: Can Large Language Models Think-aloud? HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.719074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.719074Z digest=sha256:41558339b954b416c8bacbd29f558c824b05d59eed27e3dc1cedc477a5f8939c

Observation b6c0dea8-bcd4-4dcd-8b03-81fbd12f337b · outbound

This paper cites TOFEDU: The Future of Education Journal, 3(5):1488–1499.

THiNK: Can Large Language Models Think-aloud? TOFEDU: The Future of Education Journal, 3(5):1488–1499

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.817646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:04:34.807017Z digest=sha256:25e948020d0a1747af13ec3b09bcc926d9f92a55713e51c3f15c73c2b5aa0f9c

Observation e335ab1c-82d5-46d0-9b99-a5dfdcb741e1 · outbound

This paper cites Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks.

THiNK: Can Large Language Models Think-aloud? Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.565123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.565123Z digest=sha256:95da776605c9e54d209879d152481afa00a8556c655b7a54e79f62eeb2082a7e

Pith citing papers

Observation 13673311-eda7-400a-b9d2-8c8c39141232 · inbound

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy cites this paper.

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy THiNK: Can Large Language Models Think-aloud?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:53.998837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:20:53.998837Z digest=sha256:148e4e81f5232963e9f99ca0e64ce5d2df97c01cda360856608a912c89f4b318