Pith. sign in

Paper Citation Record · LEDGER

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs

As of 11 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2602.10431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.10431 v4

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:10:01.889411Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 189626e3-626f-4219-b01f-3f530f891696 · outbound

This paper cites Table 13.Execution Ratio and threshold results for LLaMA2-7B Layer CSQA Bits Execution PIQA BoolQ SIQA ARCe ARCc Winogr.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Table 13.Execution Ratio and threshold results for LLaMA2-7B Layer CSQA Bits Execution PIQA BoolQ SIQA ARCe ARCc Winogr

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:10:01.889411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.889411Z digest=sha256:1ef614a24e23fbf5b780cc731fe3bd343cbf4eba2eb406c4e3c3f4386487ef45

Observation ebedfd77-61c2-4f38-b2fa-4839a843bb09 · outbound

This paper cites The Llama 3 Herd of Models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.051255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.051255Z digest=sha256:78f5367cba636c96abaaf66bd45067807f42bb25ac8eb063aeb4de2a22eb98fe

Observation 96f307ce-9c26-403d-a3e6-bdf56fed6a96 · outbound

This paper cites Diffskip: Differential layer skipping in large language models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Diffskip: Differential layer skipping in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.420590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.420590Z digest=sha256:84609db54710b9fa0d99900d003316d4b52af1cf9e686954aa1bbaaaa2978cbd

Observation 570d3115-7e5e-4515-9b33-c714a60fcead · outbound

This paper cites Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.552555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.552555Z digest=sha256:cbb2addcf4742b933ace1994ba130297e62c695a231ea51a42462f939f25b8f6

Observation 69e51996-ac64-4924-a9f6-469c55b5e0ee · outbound

This paper cites Compared to QTALE, structured pruning provides better memory efficiency because redundant layers are removed entirely.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Compared to QTALE, structured pruning provides better memory efficiency because redundant layers are removed entirely

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:10:01.718466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.718466Z digest=sha256:9614a127986fa3ee51c8ece3719a0b3b391af68a634f0d0e54628a68bad83481

Observation 7b24ac29-383e-475c-a7c2-2bcbb635c621 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.851512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.851512Z digest=sha256:f3869d24baaf3b528c4911fbefdad06c8096fa3e02c5ea3aa3382a51663bf7b8

Observation 137591b0-1f4c-4330-a3b5-3e24df962f63 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.002656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.002656Z digest=sha256:f9053a99736941a10b2178c23fa365237d500f1dc5495640f41bd80739e41667

Observation b02cc9e9-a320-45b0-8a9a-c75ac17676d8 · outbound

This paper cites Qwen3 Technical Report.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Qwen3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.208569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.208569Z digest=sha256:41d24853408d6ee011770088844ded016729e67d4626a3d5d2e774113e0766d7

Observation e54c6126-01c7-4758-bae2-16ad5c902b93 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs OPT: Open Pre-trained Transformer Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.386927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.386927Z digest=sha256:a8e2b18b6ccee93b99200a36c2ce1c6e54c96cbd5a5cea59fda1408b6c16ae7d

Observation f59be0af-9ced-45b2-9bab-88ada9fa19c1 · outbound

This paper cites The dashed line indicates the trend line.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs The dashed line indicates the trend line

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:10:01.561370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.561370Z digest=sha256:2a605dd7506603dd0927d334430c14adabb0b93333fe31488931a3ee36cf6c1b

Observation 34f8a773-1b07-4cbf-acbe-a5c06228b9a5 · outbound

This paper cites As shown in Table 11, even under these demanding conditions, QTALE consistently provides stronger quantization robustness compared to D-LLM.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs As shown in Table 11, even under these demanding conditions, QTALE consistently provides stronger quantization robustness compared to D-LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.831301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.831301Z digest=sha256:e27348cce95e1c7cb7e8db103e2b5b65bef368ab29c3ddf52a714222cc581cf6

Observation d48f9557-afd7-4392-917a-4745206b7e9e · outbound

This paper cites FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.287006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.287006Z digest=sha256:061e5c170771d2c98d4c0a355d5c414f13d556693980aaa72a1736f348d47b71

Observation f9d686ef-61f8-41bd-9c6d-53bf6b1067af · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.784748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.784748Z digest=sha256:655e5e6ab4abd64c52640b0085eaddc4920d2fe32f6fe4acd1bf46126ee4875c

Observation 4aa0d380-7b66-4292-81ac-28e41fb7212f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.650638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.650638Z digest=sha256:fd5ed06df379812f7695611ee45653869c814d47e31cf182ad8a375b5fcd5890

Observation 9f9e2749-9eb1-4098-a98c-af2be8b89df7 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.542418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.542418Z digest=sha256:b78d9bd4a71d1898318ecdef606240480474021b00ae38a22e41b7527f8dc0ef

Observation 9ed4adc5-c9cd-4fb7-bbe0-6557e7402923 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs SocialIQA: Commonsense Reasoning about Social Interactions

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.713286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.713286Z digest=sha256:b24f0474a6d802a69116c3631000814fa993cbd2b7f43eec38499116a077229e

Observation 8cc00bc8-0130-4489-a8b3-f6cc33c6276a · outbound

This paper cites Appendix A.1.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs Appendix A.1

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:01.492840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:01.492840Z digest=sha256:7f7fe69e35aee9e699b69161a78843fc98d8494b8a0c3d395206827089bf2b35

Observation 6987f4af-12b5-4d0c-b06b-04618912e65c · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:09:59.925845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:09:59.925845Z digest=sha256:033d48f89d3e8085cb1411d7795cc5ee04ddb5d8cf8985fa8fe491622a064d3b

Observation 3e695c57-6c7f-4887-8a3e-0044a15c9bc9 · outbound

This paper cites AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference.

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T01:10:00.166572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:10:00.166572Z digest=sha256:cf445ee40c93642efb7684aa0a284bb9e486c1b9d732cb3f43d54cba6554113d

Pith citing papers

No inbound Pith citation observations are available.