Pith. sign in

Paper Citation Record · LEDGER

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2602.10431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.10431 v4

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:10:01.889411Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 189626e3-626f-4219-b01f-3f530f891696 · outbound

This paper cites Table 13.Execution Ratio and threshold results for LLaMA2-7B Layer CSQA Bits Execution PIQA BoolQ SIQA ARCe ARCc Winogr.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Table 13.Execution Ratio and threshold results for LLaMA2-7B Layer CSQA Bits Execution PIQA BoolQ SIQA ARCe ARCc Winogr

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:10:01.889411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.889411Z digest=sha256:4ac65f654414d96f20d7368a07a2a0c0a6469c3a73dd86278be3f65dc990851d

Observation ebedfd77-61c2-4f38-b2fa-4839a843bb09 · outbound

This paper cites The Llama 3 Herd of Models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.051255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.051255Z digest=sha256:2fd9c9be767c14f42af1097663c4654c339e91e69f6852ffb3484d77dac74f4f

Observation 96f307ce-9c26-403d-a3e6-bdf56fed6a96 · outbound

This paper cites Diffskip: Differential layer skipping in large language models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Diffskip: Differential layer skipping in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.420590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.420590Z digest=sha256:14f312a7d27c084294e36f987dfa8f538ea785fe4fb1bc3a4600cdc6626396f1

Observation 570d3115-7e5e-4515-9b33-c714a60fcead · outbound

This paper cites Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.552555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.552555Z digest=sha256:6e7d9023c6b79e7a2f7927af8284bafe3d6dd12c4008528beeca98f8a40893fa

Observation 69e51996-ac64-4924-a9f6-469c55b5e0ee · outbound

This paper cites Compared to QTALE, structured pruning provides better memory efficiency because redundant layers are removed entirely.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Compared to QTALE, structured pruning provides better memory efficiency because redundant layers are removed entirely

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:10:01.718466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.718466Z digest=sha256:173a06a72095d6f0a51c1b7f8d5580ce622e6db90b31b7e4fe1a02155963203a

Observation 7b24ac29-383e-475c-a7c2-2bcbb635c621 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.851512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.851512Z digest=sha256:94dc49d22dbbac8b16eaf1268181d9f72530c3037702b05f8c9ca12ddbf54e20

Observation 137591b0-1f4c-4330-a3b5-3e24df962f63 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.002656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.002656Z digest=sha256:90b85270b32594410d8735de99c0d0176a4ab167bc22a7c60fe75983324cb0fe

Observation b02cc9e9-a320-45b0-8a9a-c75ac17676d8 · outbound

This paper cites Qwen3 Technical Report.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Qwen3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.208569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.208569Z digest=sha256:98fbf1a38db820f5da2494bee00c3f860c50de1aec901d34e7b7a036238fc1b7

Observation e54c6126-01c7-4758-bae2-16ad5c902b93 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs OPT: Open Pre-trained Transformer Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.386927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.386927Z digest=sha256:21f00afe51d402f955d38593020eff401986214e03efd483000578e00652082d

Observation f59be0af-9ced-45b2-9bab-88ada9fa19c1 · outbound

This paper cites The dashed line indicates the trend line.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs The dashed line indicates the trend line

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:10:01.561370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.561370Z digest=sha256:26c2e02731d6e025911a7c4bd85de25aa9c41f6ca63d0664bc6f810066274dc6

Observation 34f8a773-1b07-4cbf-acbe-a5c06228b9a5 · outbound

This paper cites As shown in Table 11, even under these demanding conditions, QTALE consistently provides stronger quantization robustness compared to D-LLM.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs As shown in Table 11, even under these demanding conditions, QTALE consistently provides stronger quantization robustness compared to D-LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.831301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.831301Z digest=sha256:8eff75150e9703f499aa21ed0ee1249d8d97b54b54c3ab4d0f497cbff16c3ff4

Observation d48f9557-afd7-4392-917a-4745206b7e9e · outbound

This paper cites FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.287006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.287006Z digest=sha256:95e6308911203987f30b92c2fb28d4b1daba933ee6cfd82f5a6da078f20af11e

Observation f9d686ef-61f8-41bd-9c6d-53bf6b1067af · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.784748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.784748Z digest=sha256:1f1fb08bc30298bef98f5a7b54381da11c57b8cd0f7846d9f72a1216a4f4d1f6

Observation 4aa0d380-7b66-4292-81ac-28e41fb7212f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.650638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.650638Z digest=sha256:7396cd354838c6df58df9833250c8b570ac6ac0b55d02b72cede054bcd70dbf1

Observation 9f9e2749-9eb1-4098-a98c-af2be8b89df7 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.542418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.542418Z digest=sha256:3f2de853ec64fe9cc5ae382527dffc3e7b6e942c9e19b702a61decbc83bed6a9

Observation 9ed4adc5-c9cd-4fb7-bbe0-6557e7402923 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs SocialIQA: Commonsense Reasoning about Social Interactions

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.713286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.713286Z digest=sha256:8ac1e813e86f84be10867755a61ec71309313b7d0f08fefa076ffe529a5bb8ff

Observation 8cc00bc8-0130-4489-a8b3-f6cc33c6276a · outbound

This paper cites Appendix A.1.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Appendix A.1

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.492840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.492840Z digest=sha256:8c3d0bc855138282b7c5d02ad243e8e1b5492aa5f174aafbcfd63e6f90de1cdb

Observation 6987f4af-12b5-4d0c-b06b-04618912e65c · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.925845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.925845Z digest=sha256:e20041168c4e0785beb5392e66ff3f61d3b09500810544bf2a941d9634d2ae9a

Observation 3e695c57-6c7f-4887-8a3e-0044a15c9bc9 · outbound

This paper cites AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.166572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.166572Z digest=sha256:002529959bb195a60af069c49f86a65ee248bdc6cb8e87482001127edc609c08

Pith citing papers

No inbound Pith citation observations are available.