Pith. sign in

Paper Citation Record · LEDGER

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks

As of 23 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2507.23146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23146 v4

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:06:08.015456Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact6
  • verified fuzzy23
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d54002f8-0522-4c64-82c3-6fc6fa672aac · outbound

This paper cites Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models

Reference 1

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T11:06:09.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.460274Z digest=sha256:6c762586d4e0a9608e0e3d4ef10fec50c61fbb49debc4eeb0eecbf30ec83fef1

Observation 5e62f23b-f03b-4b46-be2f-432198371a61 · outbound

This paper cites Towards automated phenotype definition extraction using large language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Towards automated phenotype definition extraction using large language models

Reference 2

Resolution
verified exact
doi, observed 2026-08-06T11:06:09.037740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.504790Z digest=sha256:9967a53147184fb390a2fb4f9fa4c58be397b71bb54084f42a8ffc61e5d81256

Observation ecce1537-4bc9-4df4-a0ef-6e49733c6165 · outbound

This paper cites Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network

Reference 3

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:06:09.737003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.551733Z digest=sha256:dc691f840d6d57c12ca7d2273dd7a6673635c5bb2c8d4a320a4e00d74c58572d

Observation 60002c1c-ee0a-45fe-b3ce-cf618ee97445 · outbound

This paper cites A general framework for developing computable clinical phenotype algorithms.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A general framework for developing computable clinical phenotype algorithms

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.810249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.601759Z digest=sha256:99a77fb7643c3ea1e982e4397ebc80d8507514175eb6a6ad220d23c7bd9ed8be

Observation 3e8d64af-19ab-4855-86b7-148a74d8cde3 · outbound

This paper cites SHREC: A framework for advancing next-generation computational phenotyping with large language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks SHREC: A framework for advancing next-generation computational phenotyping with large language models

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.636195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.617860Z digest=sha256:4449d3b9e4aa30ab0780a0156d4008ffbb41ca873d9ccee2a882d3ba29042408

Observation 0a146695-c30c-4ffe-99bb-97119e4c1640 · outbound

This paper cites Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:06:09.489695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.656771Z digest=sha256:cf951ee6e5224e562ecb5bb40d747c0efbf3ef5607dcc67ce3c5aae5e48135fa

Observation b616fc55-cd06-4d26-aa42-b5ff3b0a8acf · outbound

This paper cites Deductive Verification of Chain-of- Thought Reasoning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Deductive Verification of Chain-of- Thought Reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.822496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.701412Z digest=sha256:1dcd4b5873052493e8ef109de1792de3882759d0f3be48ac8ea6a0c6172738c1

Observation 7d356364-cca8-44a5-952a-e662e2c37094 · outbound

This paper cites Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:06.741394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:06.741394Z digest=sha256:0404d92a50e0324e511e3e689cd555445e6cdccf5d625cfac686b835405828ed

Observation abda25da-8ce7-42d6-8127-16a5f79c8862 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large language models can be easily distracted by irrelevant context

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.813349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.819302Z digest=sha256:0529e0dee4727780cf7fdd3b9d59f752683cb61336ea294ea66f01b84bfc4d11

Observation ecb8e2b1-4c21-4515-bfb6-8dfb8bea19a6 · outbound

This paper cites Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.804372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.857730Z digest=sha256:72774bbf1ee067ac7989942ab9df183efd14090afcc41c0eae37e7fa492153d1

Observation b8782a17-710c-4bf8-9506-823e92a9e840 · outbound

This paper cites Language Models Are Greedy Reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Language Models Are Greedy Reasoners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.794725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.899616Z digest=sha256:342e5eae8e030721ce65b7f8c6b58475f3d0f32a5d30bd8491050818cce6a8fb

Observation b8b0937b-036d-4321-8968-deee035c40ce · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.785754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.916852Z digest=sha256:ca8ba050449db4f8769bd933e1f34c6cf84a71c105cafebf8130c1fd864344da

Observation 5b9ef066-84a6-4d58-9623-7783d4479da1 · outbound

This paper cites Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.777191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:06.960032Z digest=sha256:27be9c2ddc7431b1c5c507c50a1d41a85caf3d43afdb1d938d9727bf085cf106

Observation 3184707f-4722-46b9-90f4-3ff422920433 · outbound

This paper cites Chain-of- Thought Reasoning in the Wild is not Always Faithful.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain-of- Thought Reasoning in the Wild is not Always Faithful

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.767840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.021933Z digest=sha256:144746702c9bf8a9a01e68f00267e793b603f70ffcc0788ba240dc5c95a40f36

Observation ef73b372-82cc-494e-bfc6-1fdb4ff1aad1 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Reasoning Models Don't Always Say What They Think

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.060768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.060768Z digest=sha256:5d8fbc8b522fa399000cd54ca272cefdfb087ac12a62daa03fbfec1f6169c7d9

Observation 7f52cf14-0354-4663-9143-5919abe82b67 · outbound

This paper cites [cited 2025 Jun 10].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks [cited 2025 Jun 10]

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.758678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.094504Z digest=sha256:79c4fca874b257af7e6b4f8df9700cab62dce13b26b70be29620c7c9152a1651

Observation 024f890d-d398-4dd5-8c37-feb8dc44ad1d · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.155650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.155650Z digest=sha256:82c379a48a25cd305ce928b0b4b67adeeac357ba1be4156881ebcb3a937c7b06

Observation f893180c-571f-442e-9f2e-92dc02c39bad · outbound

This paper cites PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.749502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.193010Z digest=sha256:35bff1e3109ffa9e2545834cec90f4a3bf881e4842bc7521fcc6a2f3aef3f1ea

Observation 891c4f96-23b5-4df3-b540-349cd65ad081 · outbound

This paper cites Rule-Based Cohort Definitions for Acute Respiratory Failure: Electronic Phenotyping Algorithm.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Rule-Based Cohort Definitions for Acute Respiratory Failure: Electronic Phenotyping Algorithm

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-06T11:06:07.220218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.220218Z digest=sha256:ba1b07f753d883626ff4e088b91832e9d598d99050e50fb785c3351f33d0462b

Observation 123c2014-cb64-4304-b7a1-370b2f64a11e · outbound

This paper cites The eICU Collaborative Research Database, a freely available multi-center database for critical care research.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks The eICU Collaborative Research Database, a freely available multi-center database for critical care research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.263225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.263225Z digest=sha256:a330c83da76e9e530556cde4e9f170bc75edde8d75839a8b5b4d8c513a73bf8a

Observation b02f6982-1d95-44ce-ac39-6fdde679ab57 · outbound

This paper cites Ollama; 2024 [cited 2024 Dec 18].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Ollama; 2024 [cited 2024 Dec 18]

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.629792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.294351Z digest=sha256:32e945a3ab175bf6a05c748a755287662fc09f06f2819b30a712a4e9d13b4182

Observation 7778794b-f615-4645-9215-924a56d0266a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.310546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.310546Z digest=sha256:0b564e3afc00bac19091fcfec08e2c5c6ca058bff9600e61cec46dcc1393943d

Observation 4d48fb61-40d5-4023-a5ab-870f5cdf9f9e · outbound

This paper cites Interrater reliability: the kappa statistic.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Interrater reliability: the kappa statistic

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.476809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.354644Z digest=sha256:399f6bc2d85600f5672d7880755852d86855940017e0fbbd897aa0f268e508c9

Observation 4bddb60d-a483-4e33-b631-11d2be39d3cb · outbound

This paper cites Emergent Abilities of Large Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Emergent Abilities of Large Language Models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.353567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.391617Z digest=sha256:833991b816bb02bc9ed81c8bbdc52d6dd4f262506587374681f668647e7351e4

Observation da6d56fd-43bf-4ae9-b22c-afc316fc28a2 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks ReAct: Synergizing Reasoning and Acting in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.424864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.424864Z digest=sha256:0b8dbfac8351ca5ec48e147a04643859fc73b400116bd45455239eaa212ad95d

Observation b5cbca71-b2bf-4625-ac33-4dee552f2ab7 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.153034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.465902Z digest=sha256:5886432b0eadd3d986d092c23140279b0823fcc7a25fba4c5d4176edb8f96c33

Observation 69bb715c-ee4a-479f-acbe-59de0fd428f0 · outbound

This paper cites Large language models are zero-shot reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large language models are zero-shot reasoners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.061431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.509894Z digest=sha256:98c2046b4dd67aa09a37c5331180d85d33f9509f227997072e40e08625f46196

Observation 1056e776-a356-423a-997c-8a9764e62143 · outbound

This paper cites Large Language Models Still Can’t Plan (A Benchmark for LLMs on Planning and Reasoning about Change).

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large Language Models Still Can’t Plan (A Benchmark for LLMs on Planning and Reasoning about Change)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.960175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.536971Z digest=sha256:984cfa340b3b2fa687e6590d399c2f302a04ae3a47cd22349d41d0061d241e9c

Observation c2544d17-6930-4b8d-88d5-f83cfb1cdac9 · outbound

This paper cites GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.559969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.559969Z digest=sha256:9aa2c463ad94e1f2d55bae4a02ee8ef5799b19943de66b78b2a599565d90ef75

Observation 888dab80-b151-4868-8594-17139b86394a · outbound

This paper cites A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems [Internet]

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.423912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.597872Z digest=sha256:be4ac7b91ca4fdf3e3dce3a3452d6930322dd6794548d6f7a8c2c2c6ccba60ed

Observation fa68a6f3-8dcf-4edf-8fa5-97cb8720bf9d · outbound

This paper cites LLM-based agentic systems in medicine and healthcare.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks LLM-based agentic systems in medicine and healthcare

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.642134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.642134Z digest=sha256:2337685a9951780fff064dfb85e2c83b3a04afe69a145ff22300febef0882f9a

Observation a3afde07-30b4-4c87-9655-91f2c8b96af5 · outbound

This paper cites Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models [Internet]

Reference 32

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.277622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.686497Z digest=sha256:b3192d8f6c4888476e695a87849583661c1bc7d079e4fb7d24912540e6f4a873

Observation 3edaae83-e511-42fa-9c1c-11e2599f7b9f · outbound

This paper cites Understanding Reasoning in Thinking Language Models via Steering Vectors.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Understanding Reasoning in Thinking Language Models via Steering Vectors

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.802590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.751682Z digest=sha256:5c470a097deef3d4984d0ceed1dc593ae220bbad1ec0ba201579701c05ebae94

Observation 4807b643-07c5-47b0-9b0d-a06a4faff031 · outbound

This paper cites Transformer Circuits [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Transformer Circuits [Internet]

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.695705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.768976Z digest=sha256:717cf344d755e966ff16aa8107f08a9dd3ab7fb574c47b081e5fdbdfe0ca1320

Observation 9914b539-2384-4531-91b6-7ef5e1c56318 · outbound

This paper cites Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.565561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.797726Z digest=sha256:feab2e658b7ba281e2dab63836d41c281b5c7371379222e28ada3524020b8347

Observation fcfb08dd-3b77-4e4b-a4ce-7492b59e8e6f · outbound

This paper cites Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.443518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.837016Z digest=sha256:7c963ba4aa1e9d0dc6d30b955b4be6b5a0925d05d44b23008b479ceda796dd4b

Observation b1ab383c-93a7-4d8e-85fb-177e7e7fc905 · outbound

This paper cites A Methodology for Generating and Optimizing Chain-of-Thought Based on Knowledge Graphs.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A Methodology for Generating and Optimizing Chain-of-Thought Based on Knowledge Graphs

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.130096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.858668Z digest=sha256:391a6cfa53db709aa7d29851483eada3fdcb73ddb1857e4f749e015ca77ca003

Observation ed20780e-9f5f-4627-ab4f-07f56ea36138 · outbound

This paper cites Training language models to follow instructions with human feedback.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Training language models to follow instructions with human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.322042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.903626Z digest=sha256:2a86c5db296a1599fb04d194d7dc28075e4f7abe5d6e640c3fc6378e5d550e0e

Observation 5dbf6a72-a40e-4f07-bcc7-ac3a419b23c0 · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks STaR: Bootstrapping Reasoning With Reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.189166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.943266Z digest=sha256:8f63432e5e1287cd6c4da843bb2742b5d334304cd7aa624aa7543d64bbe6af2a

Observation 20faaa8f-1841-409f-b20a-4059bee80af7 · outbound

This paper cites Leap-of-thought: teaching pre-trained models to systematically reason over implicit knowledge.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Leap-of-thought: teaching pre-trained models to systematically reason over implicit knowledge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:09.967053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:07.980640Z digest=sha256:816634d1f3f129e9e5cc279ef982425cb28741acee02a135bd6e29d4d72e8331

Observation 136d5028-6b67-4198-b405-940d93b9e715 · outbound

This paper cites On contrastive learning for likelihood-free inference.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks On contrastive learning for likelihood-free inference

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:09.853077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:06:08.015456Z digest=sha256:69789e36084c0eff0623dd24fc2062b7194d2c3ca6fc26954171e3f682d3531a

Pith citing papers

No inbound Pith citation observations are available.