Pith. sign in

Paper Citation Record · LEDGER

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task

As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2604.12426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12426 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:34:02.796850Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact16
  • verified fuzzy8
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce70fbc0-6252-4ac6-be41-790364502c07 · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:54:37.831311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:17be1ab23e38b23dd45875436fe4628ffe0f1b87e04d43dd1f589ec5ca13cf62

Observation be5aafc9-b8d7-4015-a724-00a650b1d822 · outbound

This paper cites LoRA Learns Less and Forgets Less.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task LoRA Learns Less and Forgets Less

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:07.435143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:4453adc2f9b4f9e9bf5ef7333b620b60cb4c80566e58c4c7bfcdae2bf270218d

Observation 827c7083-46af-43f0-95fc-48bc0aa77d24 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.052943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:dd7a48ee461e8a6ebbe2d55896017fa4238512d255cf43fd8e1ddb0f1f7cc43e

Observation 826809d5-cfbe-4d39-8b51-dd857d0c8c14 · outbound

This paper cites InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:07.448865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:b863e9aa3d1895f0a8a7d66d7324aa64cbf0f1208856f6763095c01997e402d4

Observation 0f598c69-f104-4c4f-aab3-1b5aac7e2269 · outbound

This paper cites Universal Transformers.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Universal Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:04:31.534239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:338ea206bb44857c994ba74336b3a61c10357660d1a6c7f3807e13c7e0c96f08

Observation 441218a9-2ed3-4888-a8ae-b08199911a77 · outbound

This paper cites Looped Transformers for Length Generalization.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Looped Transformers for Length Generalization

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.428196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:5c20726c9ec3e8eed4a6cca72ce73edd688762f461c06f4a4bde02d02f3de30b

Observation e2ad1504-d358-41d6-ad21-1cae28825b56 · outbound

This paper cites NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.456977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:cc73e83efa955d46982b377c0426d4d4616f917307d1ef89a34d034f871f0f76

Observation c3e3afa1-a371-4125-8327-5cc9424ff23e · outbound

This paper cites Predictability and surprise in large generative models.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Predictability and surprise in large generative models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.056948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:9c751140f43975b9237fd2cbc8d2dbde76b7265a2fbf3e3691008d7981adad9a

Observation 5f2dfd27-ea50-438f-9d5d-94ec11faa3d5 · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.045950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:c7c90cf654b15a643c6f1d57bfdeded5f157a636e448986c310caf521b07d24c

Observation d33c3ed4-d16b-4b5c-9096-fdb75cea654a · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task The Unreasonable Ineffectiveness of the Deeper Layers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.468222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:3702b167fc263a6382f5671493980588bc7d925323c49a1d50361941261e7bf9

Observation 1f6eae49-f688-4d33-9b9e-7f8389eb3ac0 · outbound

This paper cites How do llms use their depth?arXiv preprint arXiv:2510.18871.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task How do llms use their depth?arXiv preprint arXiv:2510.18871

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.511536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:11968ee951852cb901b4b113785d3e50fc9904a08bb76f6d784a8b8fbc08dd41

Observation 96dfb280-c5f2-4c4a-9f11-91cdeff79d60 · outbound

This paper cites Overthinking the Truth: Understanding how Language Models Process False Demonstrations.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Overthinking the Truth: Understanding how Language Models Process False Demonstrations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.478892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:a813e23fa953c411e13e87844a8bf8185a9efd67ec058b5aa407f8d7cfccf91d

Observation 76acdfda-3577-4f59-b503-954295a0467d · outbound

This paper cites What affects the effective depth of large language models?.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task What affects the effective depth of large language models?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.586348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:190369d36b9614242052aed95ba718346577e3839cd17b8183c5dbd9346a8bcc

Observation 3d28ba6f-b72b-4a56-9343-362ffd474b46 · outbound

This paper cites The Remarkable Robustness of LLMs: Stages of Inference?.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task The Remarkable Robustness of LLMs: Stages of Inference?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.389361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:4a35edd39d9806c3c24b15f5874aadd57ea5e41c675ac7817208652b6352c3e5

Observation c1a9323e-886c-45f2-87a7-87a31a45dca7 · outbound

This paper cites Racing thoughts: Explaining contextualization errors in large language models.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Racing thoughts: Explaining contextualization errors in large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.049581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:d67f9dc2334d962c566d6fa0ff1eb4e29f990656cee46825eb09a042d72ad45f

Observation 7d15240a-58ef-4f0b-9bfc-44abbfcdd8b1 · outbound

This paper cites Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.332182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:a575b96c247e4f0b72112537e9d14b3eab9a988a49db83278665763cc040666b

Observation b412cea5-9305-4354-9159-b7ed227b7048 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task The Expressive Power of Transformers with Chain of Thought

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:07.576268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:4d614820d8b160bd8f75df0d3b0b2783bb67706e951e65679b11cbd0e0c7c6dc

Observation 2591b39e-bd0a-46f2-b869-1e2c176a639f · outbound

This paper cites A little depth goes a long way: The expressive power of log-depth transformers.CoRR, abs/2503.03961.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task A little depth goes a long way: The expressive power of log-depth transformers.CoRR, abs/2503.03961

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.407695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:ea120fab2a8ab7d4a594a3b4cb9eac61cb49207041b981ace48552ef75336686

Observation 13da4258-01d0-41d1-b289-ecbca14f3bb4 · outbound

This paper cites Language models implement simple word2vec- style vector arithmetic.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Language models implement simple word2vec- style vector arithmetic

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.039148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:6a29f59b794d8fd6b52d29abf96950e0ecff27df488926b7ec9010cbcd8fabee

Observation 7a613b11-0755-4a0c-b9ff-f1f06654101d · outbound

This paper cites arXiv preprint arXiv:2510.06477 , year=.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task arXiv preprint arXiv:2510.06477 , year=

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:07.553758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:6e37f25fcd4e1716ab9f95bff911e0e6418d6b6de6b6799ee3a09ff2b941217a

Observation b985359c-6e21-49ca-b091-11eeda09e607 · outbound

This paper cites Understanding transformer reasoning capabilities via graph algorithms.Advances in Neural Information Processing Systems, 37:78320–78370.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Understanding transformer reasoning capabilities via graph algorithms.Advances in Neural Information Processing Systems, 37:78320–78370

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.042779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:4d5a33ed48363c5efcc215bc4d120d73adff948a848b06b7a54e8e487b14f033

Observation c46d9110-389b-429b-b60f-141fa562ffd6 · outbound

This paper cites Reasoning with Latent Thoughts: On the Power of Looped Transformers.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Reasoning with Latent Thoughts: On the Power of Looped Transformers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.361356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:f3b58ef5bc99a933d14748822f2fe092b706afaede259a9820a1bcae08b66e2e

Observation 265b363a-d799-4ec7-a651-a82870d5be20 · outbound

This paper cites CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.545195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:292c501c45f52874012bf6b472d490fd8d041aacee40ec242414d0a3d3825e31

Observation f15ec79b-23bf-4366-aa9c-963d2699961d · outbound

This paper cites Emergent Abilities of Large Language Models.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Emergent Abilities of Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:16:07.521598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:6487a6274544b512d6e60343f2bb7507ce19e1f502df8f07c8e8f8288e0f186f

Observation 3337a277-a7f9-4a89-b829-3ab35601ceeb · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Transformers: State-of-the-art natural language processing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.067600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:a0c5457aa541b03590bc88df8f17ce35e52d0dc9b5fa69d0f8d28048c1871a11

Observation faa8ded1-7880-4a76-9615-01489e70ff03 · outbound

This paper cites How Do Transformers Learn Variable Binding in Symbolic Programs?.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task How Do Transformers Learn Variable Binding in Symbolic Programs?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.537276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:de3fb8c0ac2a418154b537376f4d94a6d38825201acd7a0410d21519efc2cdf1

Observation 60d57098-151a-42fb-9c05-7ac244a1f871 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:56:11.458226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:88a75d09707b6e27ba97224a595d6cf847b12ebb491c36961407e437610dd84e

Observation ef8ae413-68fb-45d2-9daf-e5e2e9585bc6 · outbound

This paper cites an unresolved cited work.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:02:06.060407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:6cd57812f392e53d0c9ff07a4d0860b1fa38812ab9a2f078a55f9478d598c44a

Observation 8695217e-518f-49da-8c7d-4847ea6f42a2 · outbound

This paper cites Table 1: All pretrained models used in this study, with HuggingFace identifiers, parameter counts, and number of transformer layers.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task Table 1: All pretrained models used in this study, with HuggingFace identifiers, parameter counts, and number of transformer layers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:02:06.064022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:6c36cd3629de1cb58c15afad0ba4a2517985066a6ec2facd30d4909d271f4d73

Observation d61d8620-5b7b-4231-914b-b3b0c9439e6f · outbound

This paper cites The first row is identical to Fig.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task The first row is identical to Fig

Reference 30

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T20:02:06.035798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:8571f1ec0c14f6a4f0569bb6d688824f08a461eeaec36a61e4781f0e9ab4e071

Pith citing papers

No inbound Pith citation observations are available.