Pith. sign in

Paper Citation Record · LEDGER

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.14638.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14638 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.850300Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 236c17b8-9918-47c9-aa7d-3e09aa1d7f80 · outbound

This paper cites Learned Step Size Quantization.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Learned Step Size Quantization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.491210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.491210Z digest=sha256:9132b217fa77473a132b3ce8c0f7e1397f405c67f8fc7d8c7a150b51b171eb4a

Observation 94b1badb-5fc8-4880-ac05-496904854c23 · outbound

This paper cites Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.082662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.498781Z digest=sha256:0e8ceb01e70b8184d45e8a484394681e72348c554611890b849fda8f456fc009

Observation 2c1403f1-ef5d-45e0-9004-6e41e8b4630e · outbound

This paper cites Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.056103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.506094Z digest=sha256:f9f49986785dc93fc07d2b22219154883b4ac81ade2be8b1bd3a359a59df1741

Observation 1c5fd896-d33d-4a13-8799-4efc3d93369b · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.513839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.513839Z digest=sha256:99b40c77bc168d19e0127c5adc82944771e28f0f860c5e9027dac8d2ab760795

Observation 89a56030-cdbd-4dc1-94f1-e2a0167bf1b3 · outbound

This paper cites Data-free quantization through weight equalization and bias correction,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Data-free quantization through weight equalization and bias correction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.030194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.520642Z digest=sha256:62eed3125ba9105d6e0a0d85913dd09d83d2c2a9c0540b3acbe0dfb28133e161

Observation e0c27305-2a9d-404e-a070-7be9039062a8 · outbound

This paper cites Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.527071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.527071Z digest=sha256:a4f182716b322be0a7d43ef43fe33a57060392151dca33293a40bec2f45c23ff

Observation 13814e29-fb9e-4965-9a5a-ca9090c6622c · outbound

This paper cites Accurate post training quantization with small calibra- tion sets,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Accurate post training quantization with small calibra- tion sets,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.009418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.534639Z digest=sha256:05407c299e8f3ac4f14156740012ec0fee98fcdf168b4a59ec0ed9b70a06d454

Observation c581bb25-a830-4af4-a641-a331f5e4e420 · outbound

This paper cites Post-training quantization for vision transformer,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Post-training quantization for vision transformer,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.987699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.540267Z digest=sha256:0600db1f0076401168fb450469dc2e74a9cf9ceb052b35576539cd0fba8930b7

Observation 0475f67b-30fc-4443-97e9-829742ba1e62 · outbound

This paper cites Loss aware post-training quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Loss aware post-training quantization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.970339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.545529Z digest=sha256:1bc37c6e2981d75985c85c6e6f0082e99dc2fa72c1d1222f8fd764a61be6327d

Observation 080915ef-394e-4e00-845d-b555b13a803e · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.953361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.551616Z digest=sha256:ce863d3eda18bae387cae63ed3f5a34a1613f291b8350180ffaddf2c04358f67

Observation 5e6e918e-5112-4acb-8998-657589850c16 · outbound

This paper cites Up or down? adaptive rounding for post- training quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Up or down? adaptive rounding for post- training quantization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.935823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.556800Z digest=sha256:0072b830c1ac888a9c14c2704777d3c35bb3ddaebdf8fb33ee489b3b4b116ce9

Observation 26777c0e-02f8-4ef8-8aca-08f2bc01fa26 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.562492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.562492Z digest=sha256:68708695cec14f10fa344471e006d327012bcb472ffdd3f4562b1d1b48915ae5

Observation eb585fe7-2705-479a-84eb-424dbe2362c1 · outbound

This paper cites Awq: Activation- aware weight quantization for on-device llm compression and acceleration,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Awq: Activation- aware weight quantization for on-device llm compression and acceleration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.918359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.569140Z digest=sha256:b03a18209f6b7a7cf5df330693a0ba36f7d70f6ccb7165b85d2a715d2bde85f2

Observation d959deeb-80d8-4ca0-9b45-3fab8ec4f7a9 · outbound

This paper cites Q-vlm: Post-training quantization for large vision-language models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Q-vlm: Post-training quantization for large vision-language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.895542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.575244Z digest=sha256:0ac9ee7d83b79fbb5c585cfee09d9f7dce31d681236f567d2d06887ac9ad4796

Observation d0470db3-9213-4550-b031-f30ff1dead43 · outbound

This paper cites Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.877051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.581813Z digest=sha256:69d4f1d18b9ed677f5120d33e2ea1ca3682d7de0c3599cfe2a7677ae9e67ca14

Observation 71c0371c-cbd3-4427-9d06-8d0ae17c2a4b · outbound

This paper cites Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.857168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.588245Z digest=sha256:7434588b7330261b4d4518220c7d9c74a466bdfcc33ba7cd2fb7ba6f3b381411

Observation 246f09ef-fe89-4379-8f6c-7cdabb71c9fb · outbound

This paper cites Ptq4sam: Post- training quantization for segment anything,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4sam: Post- training quantization for segment anything,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.834265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.594252Z digest=sha256:081b9d641ad4cbc6f90e298e7f7a60b7251d1897908ffe768c05cb5f94355287

Observation 16077d71-deda-48f9-8cb1-53adac73f52d · outbound

This paper cites FP8 Formats for Deep Learning.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 Formats for Deep Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.600094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.600094Z digest=sha256:e85cfe28a6a6c4d6fa3c2f94034f8712513ff711a3d9a883d0322040b2cf2342

Observation 1bd83e60-759d-4ada-924b-038855fe6362 · outbound

This paper cites Faster Inference of LLMs using FP8 on the Intel Gaudi.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Faster Inference of LLMs using FP8 on the Intel Gaudi

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:06.326317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.606709Z digest=sha256:62800a61280d3e7545662e31eec8e93ec7bb4608e2f0efebbcc75bf24b75a8fe

Observation 838360a2-baa4-4155-99d9-7c9a44345f88 · outbound

This paper cites Fp8 quantization: The power of the expo- nent,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Fp8 quantization: The power of the expo- nent,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.814266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.613578Z digest=sha256:874d0f9fc243bc3e60e059699a36034cde508b16def67f702f70eca31d77e1d1

Observation c798dbe1-7d59-4c39-bb16-ae3069122cd4 · outbound

This paper cites An Inquiry into Datacenter TCO for LLM Inference with FP8.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference An Inquiry into Datacenter TCO for LLM Inference with FP8

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.619868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.619868Z digest=sha256:85389adca2c2c437d302fd7902c003c6ce2e21fe754d1e2b23ec5e1f84db3ad0

Observation 2eb5b9e2-2d15-4d7e-a109-12e46ba393d6 · outbound

This paper cites FP8 versus INT8 for efficient deep learning inference.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 versus INT8 for efficient deep learning inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.626348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.626348Z digest=sha256:275cbd8800c5bafaac18bb6fef79bde05d01ca5f18134c666fb27d0210d0b04f

Observation 737d520a-c7aa-464c-97e5-ec41adcffdce · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.631849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.631849Z digest=sha256:71a62e00cf3dc88fe0aa320304084c670295dc30ecf29bc664804b889991d871

Observation 2041b6a0-c748-49ed-bdc0-c59c99c7781a · outbound

This paper cites The Super Weight in Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Super Weight in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.638811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.638811Z digest=sha256:40097047e4ef3fb64e7a18190eaa45d72a73db79e93aa47c9bb2b2a0c5edd169

Observation 256a2434-d6c1-47b5-af54-542cc5f7bff8 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quan- tization for large language models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Smoothquant: Accurate and efficient post-training quan- tization for large language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.792801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.645160Z digest=sha256:62c6cc17317937d782eeee2f00a36ab067642b4f607d5240ee4d2db87719e2ea

Observation 7ff6de07-f271-49c8-89c7-15423dd1ddcc · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.650647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.650647Z digest=sha256:77e031e210fa3de350e4425f6ffe006c50e038b964d3d56219d22146d43d4a3e

Observation 776bbb39-0683-4ba8-a428-617b2fe31cc3 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.656359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.656359Z digest=sha256:4ec7657a63b22247dc9a464b8338cf43ed335bc9789bd59188df517a8284f1c9

Observation 3930d6bd-754d-4e32-b8df-b762c1043d1b · outbound

This paper cites Towards accurate post-training quantization for vi- sion transformer,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Towards accurate post-training quantization for vi- sion transformer,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.774110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.661934Z digest=sha256:278ac51a5a8b1f48e92b26392bea9b3998fd3b1165e36cf32bd9ea63b6ca42a6

Observation 95cf1fa8-c207-4e47-bc0a-7d556ba4f0a5 · outbound

This paper cites Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.752996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.666862Z digest=sha256:a6e4ad3d438830c2199903b866698996a4b78ca801e21a1ccc638a33d2b5a6f3

Observation 6c4e8709-e7a7-4e5d-9709-47d47966f795 · outbound

This paper cites Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.672891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.672891Z digest=sha256:b83eabeaa5a00cd4f44bd14d1afff846cfab12f7198a14225b6a38ccd947775e

Observation 0f4a9bee-0eb8-4885-ac3b-a90dbc2e4367 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.678322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.678322Z digest=sha256:19cae30cb384d4c6d59ad85cddd90d161762aa91f4863800059a8273ea2432bf

Observation 288a4b30-0a7e-44a2-9498-3fe66a8649ad · outbound

This paper cites Half-quadratic quantization of large machine learning models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Half-quadratic quantization of large machine learning models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.730280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.684564Z digest=sha256:975518d063b30bf247409cd7c8b1fc95e1890fb7663dcc8346c0371723c436ec

Observation 0acc76a7-1570-4324-9422-4a8600415e8c · outbound

This paper cites Pd-quant: Post-training quantization based on prediction difference metric,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.706490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.690691Z digest=sha256:bee18c8bcdb5ca93ad0fbee857e9a011f7a5b497e257b93b6252affb93f27809

Observation 93cd2ddc-5142-4fe7-b545-2e32eeb6e68b · outbound

This paper cites Pd-quant: Post-training quantization based on prediction difference metric,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.686680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.696624Z digest=sha256:f0ba6d769432b964c6eda4781e7c1115a3193c6c5f213497ffd4537394eb9e26

Observation 1bc2262d-b2c5-4f0c-ba8e-28df625beb99 · outbound

This paper cites Optimal brain sur- geon and general network pruning,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain sur- geon and general network pruning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.667965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.702137Z digest=sha256:b0d9dc27f76fed40b8ebd30e3d6638aa3acdd233e8d3a48844b4b947cf3b4d46

Observation fa35d4cf-af9e-4a6c-b26c-467877ef1d27 · outbound

This paper cites Optimal brain compression: A framework for accurate post-training quantization and prun- ing,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain compression: A framework for accurate post-training quantization and prun- ing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.647598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.707688Z digest=sha256:f7362ffcf39b46af490305c7c9f167f0baebf0aa6281754203be18a7411011f5

Observation 24f09a99-4a9d-4cae-bad6-bd6e87e8ce7c · outbound

This paper cites Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.623799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.715369Z digest=sha256:2aced2e62995cc830715005dbb636dac5f37e558a7d8df68739d18c13d2dcb68

Observation f168e218-e86b-4e7b-932e-e33431ef6247 · outbound

This paper cites QQQ: Quality Quattuor-Bit Quantization for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.720980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.720980Z digest=sha256:648f6b364d4b4651ecaa53ec932df379a2b6b4f6759eda932f9b8bb7a3a17125

Observation fd3c5138-7ee3-4af2-8310-3fa74fe23239 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qlora: Efficient finetuning of quantized llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.601626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.726742Z digest=sha256:6df16297753596263b55f2ea9e7b89737ed442f69f47150d2873b9d1d8560d89

Observation 45deaeb0-6ad2-445a-a43a-de4b52f0b1e6 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SpinQuant: LLM quantization with learned rotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.731877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.731877Z digest=sha256:3955319044dbd0516ab0b3c8280db8c2db81698d2dfbb6ad828e42123d9a8e31

Observation 6f2c8c5c-774a-4a46-9200-b00f149dacea · outbound

This paper cites Massive Activations in Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Massive Activations in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.737732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.737732Z digest=sha256:85ebabc780a7e8dd091f232234ba44bfcfb1d31fae8eb17eb9917c1ec76b2f42

Observation b8afbdc7-d703-48ec-8fb4-9b0bbc7b8aec · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.581618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.743472Z digest=sha256:99720ef89c887183fbc6d2e876e15acb8fe6231e2172288902fa4dc34f62ae17

Observation 4574bcbf-a4c9-4416-8eea-b5ec034b1866 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmbench: Is your multi-modal model an all-around player?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.561697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.748996Z digest=sha256:dcab2f66df380cff43c807c927d3ed61704555746e75f15f10812a79af0dcb82

Observation 3745ee05-ec80-4ae7-a844-0a136a981a5c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.754358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.754358Z digest=sha256:e9ae5cc05f17306b7a986fb6419b86258b1de644cec96fcdf0b8f8eb4643c360

Observation 00b31de8-d9bc-4270-b39f-07fc5ad2812a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.760223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.760223Z digest=sha256:1147c108430b63272a2de204109e6f927f4b0292dc3bc26acf304c7906eae510

Observation 42cebc4f-9923-4baa-be4b-2fc40dc95cfb · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.539573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.765099Z digest=sha256:a0e6d7296cf9c5841976d9b7dbbd711b4a822a7b655710b5cd2d6768960030d1

Observation 4bfbd2ce-a4ad-4242-abdb-26744edb88a7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.771908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.771908Z digest=sha256:77e807f563cd182d19f71207ce356713d77f7be73f903c45932e8e667821251d

Observation 7e06e537-d460-4808-8db1-67c7902ccab0 · outbound

This paper cites The Llama 3 Herd of Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.778165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.778165Z digest=sha256:787f674fa8700d08d4503479fa3113dc0bd1e192b0dbd9f209b796ce175f9d88

Observation 6344a6fe-09f4-452f-8914-209142fd5ef1 · outbound

This paper cites Pointer Sentinel Mixture Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pointer Sentinel Mixture Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.783715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.783715Z digest=sha256:6dec4475ca04bf072fbc296ad2046785552bd2ba0b6b57035baf4cec85fa22cd

Observation 7c1037a0-f90d-49be-a677-3cf71978fae8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Measuring Massive Multitask Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.790299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.790299Z digest=sha256:8542acf608be4c34a1da3be1474a145aa59c6b7749c84600847106034d90a6eb

Observation 2d12ea00-6193-4ffc-83d4-73c7f518e560 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.797383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.797383Z digest=sha256:627bec920347cb3942ec8cbc07f10b9f80e45b0135e70666ae3ce6e1e0755a30

Observation bff160fa-db2b-4996-9a0b-bc3ed1d1035b · outbound

This paper cites Language models are unsupervised mul- titask learners,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Language models are unsupervised mul- titask learners,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.517796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.803578Z digest=sha256:f86260012aed82d7fb65a8a15507171ad5e285f78fe24a680a49545cd209e493

Observation 200b87c5-d284-4a8d-9322-28264c0cb5e2 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.809663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.809663Z digest=sha256:276c10425ab22dd3c7411725ebf7a3f34a187cff5bf9d263fecc5fb2af84be17

Observation d74b1fd4-21b7-48a3-be62-12be5560fa0b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.815524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.815524Z digest=sha256:3b8df9c69c2eabddf1c92e99a5c8b09fea0b4fe52810804a897f4705bfd06bc4

Observation 773700e1-cc7f-484f-9068-9c7681b95810 · outbound

This paper cites Reason- ing about physical commonsense in natural language,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reason- ing about physical commonsense in natural language,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.494418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.821393Z digest=sha256:fdc6c5e3749f5124a5c32a7a5997c4c3fc49d720d7fbafd036d7c3f85f97bed7

Observation 307dd09e-be6c-4d85-8b7f-2c63168013cf · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.826444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.826444Z digest=sha256:92c49b8f17a61c23ce71ef23fb385148b76410c4bb18bf0079c78b7faa52d2c7

Observation d2267353-2366-4298-b1d5-bf7637c81f46 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.831954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.831954Z digest=sha256:a6931263d87b056362a6a2bc35034951a199e95c7959abdf65594aefda0fed7a

Observation 14025751-8455-4e36-a77e-90ac1c071fa5 · outbound

This paper cites A framework for few-shot language model evaluation,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference A framework for few-shot language model evaluation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.476599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.838070Z digest=sha256:3ef90676af51d35eaf1d7eca9a9bf3c758ffb8159a64dccf98ae484e71ec0cc0

Observation 3973de9a-8f95-417c-9b88-ddac976c38c2 · outbound

This paper cites Semantic parsing on Freebase from question-answer pairs,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Semantic parsing on Freebase from question-answer pairs,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.455038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:05.843955Z digest=sha256:2d4e22c10252b041caba2a4c7ea07a65ad2e8c9e15c90a7f83ed8ca1b74c364b

Observation 82a53ce8-7103-4aba-baf1-5aa3ba9503b1 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.850300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.850300Z digest=sha256:8a96fda6d2df23eaf29d8abdf62e8cca5855afa5098388944fff133012888a43

Pith citing papers

No inbound Pith citation observations are available.