Pith. sign in

Paper Citation Record · LEDGER

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs

As of 23 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2504.19746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19746 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:52:21.268115Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6751774a-8f98-4987-9cbb-8bf2f4250342 · outbound

This paper cites Attention is all you need,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.139977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.139977Z digest=sha256:226da86b0fad28281bfffb2e43187a83697f7ffe4ce49ec8d2d09a9d809d8e77

Observation 638ce25d-5a80-4655-8c64-39e2374a410a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.144776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.144776Z digest=sha256:885e5c03cb0c1bb9b2d3d06deb64283276a8615494b91a817ffec0488aa50cd5

Observation 05e56dfc-d151-491d-af8e-5d6043306928 · outbound

This paper cites ubrain: A unary brain computer interface,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs ubrain: A unary brain computer interface,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.775631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.149618Z digest=sha256:401e342d75702491c9ec4b6d036590aacd8e8d28948f80984355516a5606be7a

Observation c83e3c9b-cc66-4d61-868d-0b9ebd906c2d · outbound

This paper cites Up or down? adaptive rounding for post-training quantization,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Up or down? adaptive rounding for post-training quantization,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.760776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.154108Z digest=sha256:5d73d33ec8964e74f8002d7f7005515574cc499a16c42acc720af2a588034434

Observation eb4f35e2-de4f-447a-8f54-c965a1f6daab · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.158660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.158660Z digest=sha256:2db8dd39e96ef469baadc744fc3db6743e8416a6b4aa31a31e509cd39d633dde

Observation 4456aaf4-3e8a-4940-a502-aa661fe3f3ab · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large- scale transformers,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Zeroquant: Efficient and affordable post-training quantization for large- scale transformers,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.163454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.163454Z digest=sha256:ba30020416533e5324b61e5dbee55a3e5ad3779d6877eba9ecd5d2af2cddd6a8

Observation 9c7da2ac-8051-4cb3-85b5-2855d55cd0bd · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.168361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.168361Z digest=sha256:7f7f3452930891f4a39dc678bc4ab682cad23a03ad8693e1e42577a2b29e69ad

Observation f608a497-daff-4357-a017-dcb07b099ff8 · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.173072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.173072Z digest=sha256:abc2c72e85fd306b4553783956766db09fbdc486a18234a214bfcef6f3725c32

Observation 06a58058-8e0c-40d1-9a29-f14523a891d8 · outbound

This paper cites Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.727028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.177560Z digest=sha256:dcf7986b5203244cf8f79ae560e0f79530399848923ed1b35c7d8d7879d0eaaa

Observation 62505196-e02f-4ea2-808b-c0f77d7744ec · outbound

This paper cites Llm-mq: Mixed-precision quantization for efficient llm deployment,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Llm-mq: Mixed-precision quantization for efficient llm deployment,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.711484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.181818Z digest=sha256:c83656b20b0295c754939db9c8a9ca608df85df1ef7d970d98ed9b1513d701b7

Observation fc1bc929-847a-484a-a88b-f37e9b9c8f78 · outbound

This paper cites PB-LLM: Partially Binarized Large Language Models.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs PB-LLM: Partially Binarized Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.186348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.186348Z digest=sha256:0f555da4fd76772c570fa86538aad5bff6d53bd788d06342efbba24810f1e6bc

Observation 3527ea1e-b722-4025-9e9b-c938f18f1eeb · outbound

This paper cites Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.191287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.191287Z digest=sha256:f8354c24ad68dda8aa15957168313151b9f62b1856bbf06d7386803e85d09c3a

Observation 9f2d7edb-3921-4959-865f-af6472d1851e · outbound

This paper cites Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.195688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.195688Z digest=sha256:bcf160342b60abe4d1f05146e577233f376ccf0679752f6b37c5bff6bfc6b9e2

Observation 4811a75c-1c52-4c44-9085-3a46c9967802 · outbound

This paper cites Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.200180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.200180Z digest=sha256:a022a9c7e25c443ff35299a4831fc750c0f02d2adbf41a682c2f6127702fd36b

Observation e28c5483-fe6d-4372-bfa6-545fc1be7c74 · outbound

This paper cites Abstractive long text summarization using large language models,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Abstractive long text summarization using large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.677369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.204411Z digest=sha256:67669acba58faacc4bff38a3edd9be1dc48b767504fe8a069dbffcd834fb0491

Observation 3ef05d73-7f2b-4db3-ad4c-49de2d765f03 · outbound

This paper cites Transformer models used for text-based question answering systems,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Transformer models used for text-based question answering systems,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.662057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.208631Z digest=sha256:6b4ad479b4be10b9450f8d7c43888acfb661d5de0a3e9eb65e57e8de6400dc51

Observation 438f8120-5cef-4a00-9b57-a056c1abd8b6 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Efficient memory management for large language model serving with pagedattention,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.212800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.212800Z digest=sha256:3f651b462c529807802b986294580d27775b8456bc0419004793af1e5411c916

Observation b2296042-e679-4ffd-8581-3eab77aea4a8 · outbound

This paper cites Pointer Sentinel Mixture Models.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Pointer Sentinel Mixture Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.217234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.217234Z digest=sha256:7de25fa77efaee93e414b94acbb0060a32b816045ff36ce059a138f46ed142a5

Observation 429d3f48-cfa7-43a3-a3ee-10f6c5f1f093 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.222919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.222919Z digest=sha256:faebca4893ac0ba39ec4868b7f95293b93943e6d629608363862a6a2dfaabad9

Observation c9b1d251-a023-4c40-8716-8f7cb0d88ab0 · outbound

This paper cites Kurup and T.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Kurup and T

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.627195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.227188Z digest=sha256:6a54eb1930c0c5fb72566ee4d1878f801c758fcedbf95c4ef3fc19049e4d2d44

Observation 7626c832-3930-4635-975f-c69ffffbbc7a · outbound

This paper cites Ascend-freepdk45: An open source standard cell library for asyn- chronous design,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Ascend-freepdk45: An open source standard cell library for asyn- chronous design,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.611495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.231562Z digest=sha256:ec3751447b795040bee92a29034ebade6bc1783e80f6ec9c7b6efdeffef500c2

Observation 4122b31b-5767-496a-92b4-3d0843d4a72d · outbound

This paper cites Efficient processing of deep neural networks: A tutorial and survey,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Efficient processing of deep neural networks: A tutorial and survey,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.235917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.235917Z digest=sha256:c4dc75b2d6fb809406ebce2a7d8c74d26e1041baa8616e380b6fa80ac512348e

Observation 74ce057f-56fe-4998-8e35-af6173c88172 · outbound

This paper cites A Survey of Large Language Models.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs A Survey of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.240324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.240324Z digest=sha256:94ad259e656302e964251b7dbbca28df5a29fa444e730d2c9250e00a2f7bff81

Observation 8c4c853e-fc2a-4b1b-b7df-37648d77476c · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.244939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.244939Z digest=sha256:adfb1f48fbdf4c32b1a1be2677dc2c7d5215047cf5c19f701e212d5963736d23

Observation b972c15e-35e2-4a06-877e-6224ca043421 · outbound

This paper cites APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.249448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.249448Z digest=sha256:6b83c8b96ca7cf31661cbd64eaa6f535d75963f63af309fe8d11e3d992839a6d

Observation 7674783b-f7ee-4006-b928-3368671632ba · outbound

This paper cites Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.253975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.253975Z digest=sha256:9d6d6db66042b97ee296343acd8406529018d4451876ab62c2d08013575a3c83

Observation 4a266efd-cef3-40aa-9bef-13ecc57e9ac1 · outbound

This paper cites Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.258649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.258649Z digest=sha256:a4fa8629b2adf210283731d000daed6769f9b98b4f0b702f9a14e2167bd49cba

Observation 25103a6c-4a20-48cc-9945-ffff86ceb0a7 · outbound

This paper cites Mobile and edge evaluation of large language models,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Mobile and edge evaluation of large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:21.575887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T05:52:21.263315Z digest=sha256:e02cdebe9d5acb8110bdbb7e89ce993da966a3e2466d7a10c36c23409d4f6e71

Observation 8fbfc193-ed8d-4c28-bc7c-1808fd5cedab · outbound

This paper cites Hardware-aware parallel prompt decoding for memory-efficient acceleration of llm inference,.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs Hardware-aware parallel prompt decoding for memory-efficient acceleration of llm inference,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.268115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.268115Z digest=sha256:f90a7343bce5a315637beeaa99a1169ff82383f6a4287bb122349d0faafddc90

Pith citing papers

No inbound Pith citation observations are available.