Pith. sign in

Paper Citation Record · LEDGER

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models

As of 11 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2501.12956.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12956 v3

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:39:33.130911Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c517ffd-14d7-4958-8c9b-a7514b41cea5 · outbound

This paper cites GPT-4 Technical Report.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:32.975219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:32.975219Z digest=sha256:dd0e7986dd94bdb699020a1c006b77684fb06fa5c2d9d97496243a730854f866

Observation aa2789a8-956f-424c-b61b-ee1253462a75 · outbound

This paper cites The Llama 3 Herd of Models.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.020391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.020391Z digest=sha256:7e69e34dc0a89b54ec8e42a60ce0ae39c34520a545eeb89e2b9395dcc4ee64b5

Observation ae6a6cfe-e605-47a4-8ad3-6bef7fe0b62b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.032799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.032799Z digest=sha256:b4a3886a5c45dbb7b599c69dba7386c9cacd45248a2fba9537ff45bce6631c1a

Observation e1a5d41e-b01d-4b3e-85eb-d8f6afae6bbc · outbound

This paper cites Fast matrix multiplications for lookup table-quantized llms.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Fast matrix multiplications for lookup table-quantized llms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:39:33.885673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:39:33.044885Z digest=sha256:93b86662981c8b7c465df66298542f1d74c95e5f768e75a73037332ac2c9e5a5

Observation 8bc6ef43-27ae-4835-8f6b-7ac349944999 · outbound

This paper cites Scaling Laws for Neural Language Models.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Scaling Laws for Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.061039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.061039Z digest=sha256:08d781f71b7049bd7e9a47551e804835fa495944d4faf3313c9e8dd2b1cf49db

Observation 2ea6260f-dbee-4083-9a4d-6e25561ee95c · outbound

This paper cites X., Nie, J.-Y ., and Wen, J.-R.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models X., Nie, J.-Y ., and Wen, J.-R

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.070637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.070637Z digest=sha256:6fb41a51f54f0fda3a14893dce544915e49ceb5645521f342ab4de8cbbb15dcc

Observation 1b6b94ea-19b0-4aca-8210-3fec4b411edd · outbound

This paper cites Llm-qat: Data-free quantization aware training for large language models.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Llm-qat: Data-free quantization aware training for large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:39:33.853814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:39:33.077198Z digest=sha256:55d5e831c890042e0f56272fe6188bf820f639adc3ba91635968e2c6f4c88317

Observation 5826ddfa-8a06-47a9-9793-ec4bda5a408b · outbound

This paper cites A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:39:33.829834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:39:33.084416Z digest=sha256:e1aec3393d69c28a5964f84ff79dd5f6f06197b8b5700d596422951c77194085

Observation a4025208-1a45-4d00-bec7-8d4464fd1dad · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.097499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.097499Z digest=sha256:aac48c8880c1264089623cd2a04e98cb9f2f48996ed598d1861470f6dfdfa658

Observation 814adf65-71dd-4d0a-b1c4-b592457988cd · outbound

This paper cites Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.103957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.103957Z digest=sha256:bb0f264c8767e9f9e8d3b080b8454c5b1fdea6175c58a6c3e32cc4ce640234f9

Observation fe1ec67c-a8aa-4c2c-99f5-bdbbcfd8c571 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.109217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.109217Z digest=sha256:3dbbbf4f38350283ea93d6ea9ae0a805708ab27350ef52ad59cccf653c2cfaca

Observation e18376b8-e0dd-4e98-8c91-0559c3b0abc2 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.121158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.121158Z digest=sha256:1f26c0e1d2c0f48aa9c2a857cceb687c578feea954d5fbbf8a5efad019536af8

Observation 42c6e2f1-8283-46da-8389-668bb169e234 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.125674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.125674Z digest=sha256:ff8a726fcd01c13e515b8ad5efa4855d010e46500a3b22ec4712fa11576278fe

Observation 6121e131-c2a5-4321-9291-395e8aa041ae · outbound

This paper cites an unresolved cited work.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Unresolved cited work

Reference 23

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T16:39:33.292416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:39:33.130911Z digest=sha256:fc3edcd5941f4d8418c22425059c621b0f91a26f9030b1755310854594fd1b07

Observation 4d280e49-f035-4377-9229-8007e5dd0d6a · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 1925

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.026316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.026316Z digest=sha256:ed1136ab763ec1245be40ce2b0141dd1b1fcbfe97f0cf1dab764800864c78d74

Observation e5befff7-2b77-46eb-9120-8a117f26f522 · outbound

This paper cites LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.091094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.091094Z digest=sha256:7fa5857e1090e9abbf0431e20cd19e22e3b48508f2149aa366c1550dfaf8df2c

Observation 8e3cd454-0649-4207-a733-76ab915a6aa4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.006003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.006003Z digest=sha256:42f437553f94e3bbf1e91eb387cde6437c00571c49f27e0d01272589d3809194

Observation aa5f5249-3049-4a86-9084-2bacfb5e58a6 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:32.993144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:32.993144Z digest=sha256:7a05c6b0c503546aa51b2316516c9e61ac7f0b78c9f37b787149d74f5b13cb55

Observation 7d3debeb-5733-4a17-a01c-95f6d60c1710 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.011960Z digest=sha256:ed7bd36559d432a58375f049be37a879a7ed908e61cf7e1485778462b85df641

Observation 6a79fdb1-6923-4457-901b-1c455ce25e1e · outbound

This paper cites NUPES : Non-Uniform Post-Training Quantization via Power Exponent Search.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models NUPES : Non-Uniform Post-Training Quantization via Power Exponent Search

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.116771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.116771Z digest=sha256:0cb0cb7c94cc4c9cb42ed933217c3597d421c32c62d7ba5dfae8587fd7d40c4f

Observation f582e29a-cd3e-4074-bc11-45edd6b1c7bb · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:39:33.944337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:39:32.986306Z digest=sha256:4dc81fb0f4a52b47ea3cd2cfcf12edff648022f6aac4881a99a36a1c902d6dff

Observation e8e4dac9-0063-4553-a53f-0c832cb86a8c · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:32.980810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:32.980810Z digest=sha256:1044bda939b3d88de33290bfe264226f0869200657b80f5eb6991c157de5ed5a

Observation d8f54638-188a-4ceb-a98f-3d4c61185778 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:33.051539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:33.051539Z digest=sha256:72f0e6380cd98363ab88078c42654cc4933c8770fc653c7056f6306fa1e172ee

Pith citing papers

No inbound Pith citation observations are available.