Pith. sign in

Paper Citation Record · LEDGER

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference

As of 19 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 0 inbound Pith citation observations for arXiv:2607.17733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17733 v1

Coverage vector

measured 100 of 136 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:12:37.492741Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 136 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0389ef3-da40-46f6-b694-d9dd588926c7 · outbound

This paper cites Training DNNs with Hybrid Block Floating Point.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Training DNNs with Hybrid Block Floating Point

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.151541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.151541Z digest=sha256:fd04f7d7af79ffc5882fefc965bfadd2c5c2c18c8fa6338fac55da0a20c29227

Observation e210193d-bab0-49e1-b61c-e1d60621ce0e · outbound

This paper cites Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.226905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.226905Z digest=sha256:85327550855ab8399fb73252f3587c5ffa5c7eedac21a90b70de697eb860a252

Observation 6ad65d40-01ba-469f-a601-1e9fa4083ded · outbound

This paper cites 2021 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2021 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.293408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.293408Z digest=sha256:62ec5eb048c7855314b2bb6c050594b626c0ba37c34025efc0081ba5cbd19b21

Observation dee5b13f-7326-4bb9-a742-9a0e250b63b4 · outbound

This paper cites All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.357886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.357886Z digest=sha256:e3bd80423ac834c7630024e404f6ca8e747c14b1817fd3aeb11c59f2c694b28b

Observation c1bb3eca-1d65-4534-ab2a-c7672dd1587e · outbound

This paper cites Understanding and Overcoming the Challenges of Efficient Transformer Quantization.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.428735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.428735Z digest=sha256:5c0a708df6a233461867c3d9e8c54014e7bfea7826dae6f228cba7b9f8642843

Observation 0ab90303-b202-492b-bdd9-e5f065d5e817 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.497511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.497511Z digest=sha256:f2fb3b37ad771f23be652392b6cd5256670d5f63fa7a7185b0a4f73cfc3bdd15

Observation be13ca87-9323-40a3-bad4-990ba4de8185 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Microscaling Data Formats for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.562005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.562005Z digest=sha256:b46b6032fb73f2e0e0b04fa853f78772b81ce2fd67ca5efbcac72062a01d56a3

Observation 09757515-49e5-465d-a2b6-4d90a40f9eea · outbound

This paper cites Proceedings of the 36th ACM International Conference on Supercomputing , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 36th ACM International Conference on Supercomputing , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.623831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.623831Z digest=sha256:695d9a20a2daa126beadbef432750765245b47a4a633c0d214680635ee8106e9

Observation c164e073-01a3-4777-ab6f-479d2ada1105 · outbound

This paper cites Proceedings of the 52nd Annual International Symposium on Computer Architecture , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 52nd Annual International Symposium on Computer Architecture , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.685221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.685221Z digest=sha256:16acad0158ab80f1c7ade22635ed362197dc8b307bd1cc2cbb35fc65804aff49

Observation ca3d3686-578e-44f1-a518-051294a96911 · outbound

This paper cites MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.746603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.746603Z digest=sha256:eaf7bd8e7f2f820bf72d57350aa0f538c1c8bc6631143d3490f551d0b25533c7

Observation ea9fbdea-66b5-4ed5-893a-3c13ad9b828f · outbound

This paper cites Outliers Dimensions that Disrupt Transformers Are Driven by Frequency.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Outliers Dimensions that Disrupt Transformers Are Driven by Frequency

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.817771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.817771Z digest=sha256:36dab18ae741a2de5454a77240f34ea63e8025476902fcd46a01e0569915f169

Observation 81a639f2-4965-41dc-a50c-f5979d8c4a27 · outbound

This paper cites LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.877147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.877147Z digest=sha256:a75f1a7da58adecd8b44892e9bd21b43deb95df8ea05576b18f417a45fcacf19

Observation 25b40a43-1701-4822-b146-c3fd2fdc511c · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.941291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.941291Z digest=sha256:da82c7e441eaf5edf03af390166c68bcb7eacdaa922efccfc247831ac4f7ae7e

Observation baea74a0-fbe7-4d83-a4f2-5bd54d12bb5e · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.018798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.018798Z digest=sha256:132be9844097b0acc480a4595293a2a11756e14967e409992fb199243ce20c00

Observation 60856a8d-2376-4ff0-ae21-e3b971402ee1 · outbound

This paper cites 2018 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2018 , eprint=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.080747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.080747Z digest=sha256:f7c7ada41dce4d80767ea8f210b1c852509b614a2648018baecf5158e156e68d

Observation b0f61271-4cc9-4183-9a57-8f0aa716cb7b · outbound

This paper cites 2019 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2019 , eprint=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.165900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.165900Z digest=sha256:c2dee81ba671e087416067f7266fa069103012a3ba6e9dd95bd0fd6a7fa60b7d

Observation d150366b-146d-4864-8500-6d7060c83dfc · outbound

This paper cites Microscaling Data Formats for Deep Learning.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Microscaling Data Formats for Deep Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.283144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.283144Z digest=sha256:10862d5032822b308001ea9db24eb26330acd6fd6edac0bfed4d26631dabfa37

Observation a7eaa3de-7ed0-40dd-9210-432173e44e2a · outbound

This paper cites SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot , booktitle =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot , booktitle =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.431686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.431686Z digest=sha256:513b6c89b3516c734963ebefbebc48f047729c9dae471f83196dd25ae3a6674e

Observation ce162646-48aa-4be1-9ce6-a3778d250254 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SpinQuant: LLM quantization with learned rotations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.552379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.552379Z digest=sha256:117c1a87d6ae8cdc2064447c034caece9c056875a5714972a6ca2383854c85ef

Observation 25996b48-6dd8-422b-96bc-e1f2fada2255 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference A Simple and Effective Pruning Approach for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.673389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.673389Z digest=sha256:45e09ea0fb37e64a48f5655e75e05aa34382603829305f74a03bc71766b63ee6

Observation 9e82efbf-98d6-4a66-9323-68cde424840e · outbound

This paper cites Accelerating Sparse Deep Neural Networks.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accelerating Sparse Deep Neural Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.796302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.796302Z digest=sha256:5969a979cc6ae7dbecbea9fa94aa9faf7f809bdee22aee4e696898b4c45724b0

Observation 966a36ac-594e-4460-85d5-7fc95c47b4af · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:28.913039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:28.913039Z digest=sha256:9886f485d310951369820b3fb1dca1a64040b33a6d48c9fb4f2ff99a6a615449

Observation 2ce3fa1a-6ef3-410b-825c-6da4cac418d2 · outbound

This paper cites Proceedings of Machine Learning and Systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of Machine Learning and Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.029903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.029903Z digest=sha256:22f65cfe7abe231b77fe9b5fd45a27bc7417edab553cc301aad0b622ff9053bc

Observation 6dc16e26-a030-47d0-bc34-98fe85a0e521 · outbound

This paper cites Proceedings of the 50th Annual International Symposium on Computer Architecture , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 50th Annual International Symposium on Computer Architecture , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.203862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.203862Z digest=sha256:09a25c0b389f25ffd308dad2737832f4002f8afd5f3257ecd357c4f0e7ebef02

Observation ffd64e2b-1494-475b-af17-45feca4df541 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models , booktitle =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models , booktitle =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.352243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.352243Z digest=sha256:fa53bfa5f46f5cc3c08becfa449e104b6e50f4ecb685a78c803a36a7b101f6c4

Observation e26c0881-9a2d-4871-9292-808ac00247b1 · outbound

This paper cites 5th International Conference on Learning Representations,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 5th International Conference on Learning Representations,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.460631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.460631Z digest=sha256:f3d3fee80e790ef0043890bc8c2a1c6a31e791d9c358b6bf78e7f296f98af934

Observation f333cd00-d380-48cd-bdc3-4a4424b391ee · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.579510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.579510Z digest=sha256:ed4b307623522d9f9773ab9fdba101ceb7bd551e35781deebaed9c0505b9023b

Observation e6804e4d-f97a-4505-bab9-91bb0299cfdf · outbound

This paper cites Wikimedia Downloads.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Wikimedia Downloads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.693882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.693882Z digest=sha256:2dbe7cb08019019135eee1a1f5c4e1211c6152e09fd85961e28713485c8dd88a

Observation 8d3df61a-3d72-4515-a04f-de8cec42d0d2 · outbound

This paper cites The IEEE International Conference on Computer Vision (ICCV) , month =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The IEEE International Conference on Computer Vision (ICCV) , month =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.834874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.834874Z digest=sha256:e58929cecd36cd4d10538e5d3bf8ee9a181d02e15e62d68b877966e739da5877

Observation c68e3a6b-6601-45e1-b8d1-b1c1da5fab5a · outbound

This paper cites 2009 IEEE conference on computer vision and pattern recognition , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2009 IEEE conference on computer vision and pattern recognition , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:29.960689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:29.960689Z digest=sha256:2753cfeecff33c5675f6c14ade6783bb77b26c5389af910ddb2b14f21494b7a0

Observation f6c9bd04-9fae-4297-887f-17543f335fa9 · outbound

This paper cites Gomez and Lukasz Kaiser and Illia Polosukhin , editor =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Gomez and Lukasz Kaiser and Illia Polosukhin , editor =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.030535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.030535Z digest=sha256:4cc2760ab187ef8782c5a342d78118350d36cbeee5ec094519475be3869bfb49

Observation def36568-02ca-476d-86f0-fb1192916241 · outbound

This paper cites Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.114261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.114261Z digest=sha256:a8c7b81acddfd88a3d6ac7333c8008f72b98bde6c5de442db97c33e48be8575f

Observation 644754d7-4896-4a11-8dce-b379667eeb6c · outbound

This paper cites 9th International Conference on Learning Representations,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 9th International Conference on Learning Representations,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.203691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.203691Z digest=sha256:d46283ba35a5ffca56d285f2f9e974d9e3f3b50b850d12b7a0a00ad6f08626fb

Observation 474a638d-3154-4452-a01e-77602f0cde8a · outbound

This paper cites 2019 , file =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2019 , file =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.285158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.285158Z digest=sha256:ef660dc7fd2d8682fecd3449dca76834e357db7d71edf196d066637242012430

Observation a877c2d5-a1f7-4888-9baf-33308fbe9605 · outbound

This paper cites 2024 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2024 , eprint=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.372935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.372935Z digest=sha256:a026d955be63522a34560fec14a94347d81d291281b06ba89bf533f703b4a40f

Observation 1e7a4a44-a773-460e-928b-ad3ba0d012ae · outbound

This paper cites Mixed Precision Training.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Mixed Precision Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.460917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.460917Z digest=sha256:0f8901b8caed733f5a2db640f345641d7352f68e32f865982c928da281b5e697

Observation 78f2d4e3-27db-43b5-96ff-c32b9f855a35 · outbound

This paper cites Chung and Zhaoxia (Summer) Deng and Sam Naghshineh and Jongsoo Park and Maxim Naumov , editor =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Chung and Zhaoxia (Summer) Deng and Sam Naghshineh and Jongsoo Park and Maxim Naumov , editor =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.542653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.542653Z digest=sha256:6431716692ee4c8d94abb3beceed2a968d7a0958f55ac9eb62dd04a83162fb3b

Observation 10c5b0b5-783b-4aaa-a46e-b81edeb8e1e0 · outbound

This paper cites Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point , url =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point , url =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.653340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.653340Z digest=sha256:4c54c2bf086db7502a6c138b7a1703251fda1321fb9390b906d622229e02ac31

Observation e1f1a71c-be6a-43d8-89dc-7a4fdc419470 · outbound

This paper cites 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.710712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.710712Z digest=sha256:a33ea044ba4372f11fc64931fc524b29a76126ced5cb22b1b1cb17fd45fc922e

Observation 29762dc7-fc3d-43a3-be61-016b011f42e6 · outbound

This paper cites FP8 Formats for Deep Learning.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference FP8 Formats for Deep Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.799204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.799204Z digest=sha256:a3aec3f411aa7b3bdcfd9bed817f646b434f0f3129da03dd9ab943aec26cba82

Observation 42b3fb62-35de-4c9d-ab99-2c785ed18b7a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.911126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.911126Z digest=sha256:95addd23b3ddc60e45d862106cabeca0b989783ac323306c248304433a3d73f4

Observation 05fb1707-c374-4d69-884e-ca9c5228a565 · outbound

This paper cites 2024 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2024 , eprint=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:30.991729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:30.991729Z digest=sha256:bf02cebffcfba76027e42bc5925e3ee99b9ffa7fa1c716b873ade383d3164b82

Observation 305d4f87-174e-49f1-b38d-6a611044d583 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.046211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.046211Z digest=sha256:e155a90cf3ba9f22d70ec1597346f344e937d526df9ae3467c8ee3650825929b

Observation 9e1c64f7-e564-4713-ab81-3c9b78de1a97 · outbound

This paper cites ACM Transactions on Embedded Computing Systems (TECS) , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference ACM Transactions on Embedded Computing Systems (TECS) , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.128667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.128667Z digest=sha256:83e63de5047cdab16ea771bc18fd5a9acd7f9439b7008e0b4d4a4174eb3d4364

Observation 8fe9537c-75da-4cbd-8d2b-47820e39d969 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.201090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.201090Z digest=sha256:b074ceb6dea0d651bd4ef84b64df6e200a72fd1624e0db4bcc2d50a11d1488cb

Observation 64484466-22d4-4f93-b236-c19aa6ff4dd0 · outbound

This paper cites ArXiv , year=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference ArXiv , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.262683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.262683Z digest=sha256:9b9c5dca1bb881c902b7f90a94b76036d04ce2fd20eef48fc4ff8fca3618ea48

Observation 1815c434-9e51-456b-a152-bd5f2a7291e7 · outbound

This paper cites International Conference on machine learning , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on machine learning , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.291334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.291334Z digest=sha256:0a98e0112b43d0bb8a2d4ece3bc03754a70c9f1de9d2d37efe5382aaf624f92e

Observation 797337d0-c682-466f-bdb8-ee247590fa9a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.414720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.414720Z digest=sha256:a45eeb3752a82daf9ec8713892bbb66d0ee435658bd858ce2245beec812041df

Observation 6424e814-38e0-4037-939a-c5b3d31e7335 · outbound

This paper cites Joint Pruning & Quantization for Extremely Sparse Neural Networks.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Joint Pruning & Quantization for Extremely Sparse Neural Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.547507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.547507Z digest=sha256:d99897228cfdcb2b13d1d182d872fddea222d6559d015640f8c94b4320f48706

Observation 32dce6fb-050d-4af8-ad01-1edb704f03f4 · outbound

This paper cites Towards Optimal Compression: Joint Pruning and Quantization.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Towards Optimal Compression: Joint Pruning and Quantization

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-01T17:14:13.463510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T17:12:31.711966Z digest=sha256:458a79bbd52a0981b4d4446c3e7ef0e25c8aeba8033bb48101594a89f2fa8e80

Observation 8446f3c8-35a7-4497-9763-2116970a036d · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:31.870818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:31.870818Z digest=sha256:4b3a1f8d1c82fbb4e93ad63749b4cd92b7b03cc85e407c1eec70469992369629

Observation 7941fbd8-78b7-48a4-8d00-31924986a027 · outbound

This paper cites Deep Compression of Pre-trained Transformer Models , booktitle =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Deep Compression of Pre-trained Transformer Models , booktitle =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:32.020992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:32.020992Z digest=sha256:a6949dbd8a6260ff2ff7aaf34f9532f8ccd9099a9ef97e2e7885b30bb11b9c09

Observation aba23c4a-96e7-4f66-b483-04d77f2d208b · outbound

This paper cites Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:32.251422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:32.251422Z digest=sha256:4172fdacbc841bdcddc7b2065ce76e0a49533eeb8a99d61c9c59ce1ccc37e021

Observation 65dcef89-67da-49b8-99cc-08dedd0ad554 · outbound

This paper cites Frontiers in Artificial Intelligence , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Frontiers in Artificial Intelligence , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:32.440054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:32.440054Z digest=sha256:44ea41cd7b2cf764cb963c13b6e0c776aba02178f4e900a575fcb10aec0fb07d

Observation b8c5b75c-c143-4f4f-981e-bf7e805fc667 · outbound

This paper cites International Conference on Machine Learning , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on Machine Learning , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:32.631618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:32.631618Z digest=sha256:c909ccaae6599b5e18798b9d2d8664e03d608c57bce9c2d5091cfdf8ccded980

Observation 95579736-d00c-4d24-92b0-f59d40975db7 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:32.793626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:32.793626Z digest=sha256:26061879ebd13f9c983bd1070ea838d7a6dc6b1d3e47275ca5261d0172f505c2

Observation c751855a-07e2-4107-bd3d-766c41280517 · outbound

This paper cites Advances in neural information processing systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:32.906537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:32.906537Z digest=sha256:94b9bdbb85badbb44a4056dde275ed850f732408ed59388573e551555c6486e8

Observation 0be79ddd-3208-48c3-a2a9-c320f0bd02c8 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.043803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.043803Z digest=sha256:8481fa6227b7902c9ffdcbfe2a7e4b81cff509698b328d9a6cb13cffdb8ddecf

Observation 5fed8344-2da6-4942-9f8c-b8a8c30f01b6 · outbound

This paper cites Advances in neural information processing systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.180569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.180569Z digest=sha256:d77fc987cfc64025853956b8f75bc8d7758bfd4d7370af4c767c98256b6438b1

Observation 47b0d489-c94f-45a7-955e-f63e5af53786 · outbound

This paper cites Advances in neural information processing systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.314659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.314659Z digest=sha256:8b7127ef7175318ca018d58bf163aae5b8830e2a7231d9a968d2c10d8e2c51d1

Observation 13597980-b2fc-4589-86e2-9cbe858d1848 · outbound

This paper cites Advances in neural information processing systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.442558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.442558Z digest=sha256:79a2923b70b57bad7d3a510d95c35466ebc2955715829f22387ea31fae58745c

Observation 241dab0e-83f3-4a53-900d-5974b2afe19e · outbound

This paper cites IEEE international conference on neural networks , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference IEEE international conference on neural networks , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.514064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.514064Z digest=sha256:24958bf1c0c8e3f1c5b2124dff84abfe68bb36d79aa6bf7547d07ba0401005ba

Observation ada5906c-7e26-4b42-85b2-b6f5fd44572a · outbound

This paper cites Advances in neural information processing systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.611359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.611359Z digest=sha256:442f34631cdb5dee1ffdde2627fe49aab63d31faa52c4f909c20575676f147fd

Observation b4da7d07-508c-4ae8-9598-85399e46d19d · outbound

This paper cites 2024 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.714207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.714207Z digest=sha256:49c7b108794a6b214e7d0b35c6eb913562fd0b1f036da07d769a10ec9f7ebc90

Observation 3d797cb1-e33a-4b9f-96dc-1652136e8d05 · outbound

This paper cites The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:33.850789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:33.850789Z digest=sha256:efbe879362bdb32a1e58d8d42dc1f83216bcd50d36a8fe8e8a65bdf55858af9e

Observation 8746ab1d-cedf-43d5-99ce-89b99cd35aa2 · outbound

This paper cites The Thirty-Third.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The Thirty-Third

Reference 66

Resolution
verified exact
doi, observed 2026-08-01T17:14:13.147953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T17:12:33.984135Z digest=sha256:aee3295477ea09ae7696947db4beb4b28618d8f4b8af14d610c6f3e4cdc841fc

Observation 373a98bd-fdc6-4aaa-b8fa-48ec2557e0fd · outbound

This paper cites Accelerator-Aware Pruning for Convolutional Neural Networks , journal =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accelerator-Aware Pruning for Convolutional Neural Networks , journal =

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.100741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.100741Z digest=sha256:cf045dc3801da7031d9e3d047c61422f2fddbf913a4297101bb7c59bc8f1d4b1

Observation 0a810e8b-7652-4b8a-b392-75ac3fbafffb · outbound

This paper cites Channel Permutations for.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Channel Permutations for

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.266577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.266577Z digest=sha256:b5f914513069ce6c080baeb16672b61a7f06a59aa2d77ed3cc5f89895e2eaa3f

Observation 343f5256-7909-482b-9295-dd4d1bfa9e39 · outbound

This paper cites 9th International Conference on Learning Representations,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 9th International Conference on Learning Representations,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.408624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.408624Z digest=sha256:0a59e6ee07128c4394278976681860d6420c4c91c39e166832e640a945f62bac

Observation 745da703-9490-4c08-b19e-4464f14a3ac0 · outbound

This paper cites 2021 , URL =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2021 , URL =

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.483305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.483305Z digest=sha256:3d42f6a4cb69e36e69b2be1c7e111b0daf6cce06610bec1387c035bb15cfc17f

Observation 4300319b-01b1-4228-b901-88b0544c573f · outbound

This paper cites 2022 , URL =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2022 , URL =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.554232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.554232Z digest=sha256:50cf13dbcab59822df262ec37d0cd275470601cbf3e6a3f892b9b451a5b84a28

Observation 70e9241b-bc2c-49aa-9631-a3ed4689380b · outbound

This paper cites 2020 , URL =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2020 , URL =

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.637548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.637548Z digest=sha256:b591535e47163c525b7618dc46fe946f1ade0e7078ddf4f33397fff2ce4e444e

Observation bd403b5e-f19f-4d4a-b375-b4a06b601831 · outbound

This paper cites an unresolved cited work.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.735064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.735064Z digest=sha256:a1086cd53abea64a7a5e35860170cba7dd6d8f6e910aadc0c8a672b463cb013b

Observation c0e05e7b-dd3a-45d1-a6a1-b47c56e23fbe · outbound

This paper cites Accelerated Sparse Neural Training:.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accelerated Sparse Neural Training:

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.815551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.815551Z digest=sha256:b9f49c3cb60c5790d2a425ffb0fc67587f4991907a9bbd01be4531a227780113

Observation a0aa4e79-4833-4aff-b50c-ded34aa84d66 · outbound

This paper cites International Conference on Machine Learning,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on Machine Learning,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.917531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.917531Z digest=sha256:a486229135509e14f53e6a6cb81660a9693a7c6231b85c439c95b3820e821230

Observation 867f4a22-91d2-41b7-859e-50d858f61792 · outbound

This paper cites Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:34.994910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:34.994910Z digest=sha256:3c977627232f41a6f641208344119872663f0109586fd3d444bb5a9338a8f7a1

Observation c9560cc0-c5e4-4a2c-a9ac-65191242fae9 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning , pages =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 37th International Conference on Machine Learning , pages =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.100065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.100065Z digest=sha256:d01c21408ef30aadbe4b670be01e4b87236a90ccdabd5b01f69430d12701f73d

Observation 4b85aa1d-4cae-4590-9534-acc867ec750c · outbound

This paper cites 2023 , eprint=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2023 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.176718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.176718Z digest=sha256:9ec5707af3c94f163d09ebbfbd4b9cdb3d31424ecd593a6bfff25b62c230398f

Observation 61014f32-1bc1-48f0-8eb1-7ab69133c344 · outbound

This paper cites Proceedings of Machine Learning and Systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of Machine Learning and Systems , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.291683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.291683Z digest=sha256:b350648f38f9f923fb64323b1f2ba82ab3c42aa64a57d5533211bf9187bc88ae

Observation 07bcc682-d597-4f52-b2f9-1927492a6882 · outbound

This paper cites Advances in neural information processing systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.375508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.375508Z digest=sha256:acf9b38b9515763186576e26eb9e036d5156726c31142986f3eab098ec61f0e2

Observation 69adbd2e-2e04-4f3a-a401-c2f24fb86b1b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in Neural Information Processing Systems , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.476974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.476974Z digest=sha256:584a06d661e70caa9fabcc2e910d4101b409f50831a3b28a453862c111bec4e2

Observation 4e17c489-6900-4efc-816d-db83e14640d2 · outbound

This paper cites International Conference on Machine Learning , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on Machine Learning , pages=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.582690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.582690Z digest=sha256:844ef654ae922ca6c8813f33e3b6462bf7371f91a2724ff951d44a547abf3871

Observation b58cbc9d-c761-451e-aa26-f54b78697daf · outbound

This paper cites Scaling Laws for Sparsely-Connected Foundation Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Scaling Laws for Sparsely-Connected Foundation Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.695223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.695223Z digest=sha256:458f965265049cf9ff1bc24fdac087687cce882bb7786058ef23cc70f53884fc

Observation 98c21efa-e11d-4d5b-aa82-16d3e1bf8ecc · outbound

This paper cites Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.769810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.769810Z digest=sha256:6848216caae914976b78ea80a93bcbde5390ce17cb08526bd67d0d4565faf9c8

Observation ebd3c216-b495-4a92-bb44-6cd0934186f0 · outbound

This paper cites Dynamic Sparse Training with Structured Sparsity.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Dynamic Sparse Training with Structured Sparsity

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.883139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.883139Z digest=sha256:d272c63274b05403352dc4efa6d7c7e61c0be31c889de4d3601972e2d37bf57c

Observation f7af72ee-4284-4e6a-81e0-01d3a8ee0427 · outbound

This paper cites Efficient Processing of Deep Neural Networks , series =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Efficient Processing of Deep Neural Networks , series =

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:35.977544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:35.977544Z digest=sha256:2c64381d407e357ced2ef17bced5ceda48b7f001ca907d5999d8d7cac38a4a05

Observation ab0625bd-c4f3-4105-b3b6-67e5e2ef27ef · outbound

This paper cites 7th International Conference on Learning Representations,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 7th International Conference on Learning Representations,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.026630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.026630Z digest=sha256:14d349be25eb57c86bca51b256c9126e38d11f1101228350681e3aaeb275eb5f

Observation cd48744e-1df3-435f-984b-5707532ed539 · outbound

This paper cites Intelligent Computing: Proceedings of the 2021 Computing Conference, Volume 3 , pages=.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Intelligent Computing: Proceedings of the 2021 Computing Conference, Volume 3 , pages=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.065756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.065756Z digest=sha256:7f1d5751f33b50962fe86d89940d8c32d6be128fe183e32e2e7a8df30fc57d53

Observation 27dbc204-fc95-4273-8515-b26b13b2a05c · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning,.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 37th International Conference on Machine Learning,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.119748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.119748Z digest=sha256:952ec835840587d2612fe2c0f589176f885b657991dcdbf639c78f7d8babe661

Observation 02e80b84-3ca6-4e61-a0fd-1b9befb9d992 · outbound

This paper cites USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-08-01T17:14:12.531140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T17:12:36.198632Z digest=sha256:e9e4770c8d696e301173afafa37cbaa6f8d98633ea20d43341f74f755652b641

Observation 0798e960-4843-403a-afec-510fafde30f7 · outbound

This paper cites Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-01T17:14:12.239179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T17:12:36.315756Z digest=sha256:a810319af01e64d8ac127924107e0a868867e6c8e3587f0f62966094447b05f0

Observation 8fb65693-9341-4d79-b5e9-98c056c4a409 · outbound

This paper cites Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models , booktitle =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models , booktitle =

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.476138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.476138Z digest=sha256:302222e7bb713af30a3c422233e044f5b0da525002cd89f5a15c5d6d00baf280

Observation 6d56828e-22be-413d-8a7f-eeddc0f13cef · outbound

This paper cites Prune and Tune: Improving Efficient Pruning Techniques for Massive Language Models , booktitle =.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Prune and Tune: Improving Efficient Pruning Techniques for Massive Language Models , booktitle =

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.636279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.636279Z digest=sha256:f1f70205736d55c9c3470a2837cac5b8a2c4f05ca312d4181f3e6023b537ff6e

Observation d2938dbe-005b-49dc-80d4-389d659c0a9c · outbound

This paper cites an unresolved cited work.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.735369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.735369Z digest=sha256:09578c4e4260d61c696494ba4c9065d42997c8189fc9fa7ab60dd229d17a57c4

Observation e7cdbae7-6958-43f1-9bb5-4bc88e2f5aeb · outbound

This paper cites Mistral 7B.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Mistral 7B

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.742446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.742446Z digest=sha256:0d0d35b0d7d6fa3da96f1680c191e86b0cdb5b328608079649f54872a75b868f

Observation 1ee4cc51-b346-419c-9899-a56708e9e1ce · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:36.861472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:36.861472Z digest=sha256:8fa758231d17fa15f9e005d34bbbbc51159190e1319fdd0590ab8418de1f284f

Observation 46bde947-e50b-46bc-adf8-7158ecdea99a · outbound

This paper cites The Falcon Series of Open Language Models.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The Falcon Series of Open Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:37.076833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:37.076833Z digest=sha256:48c7c13f7872ace5fe108d674f11ae41d1999472768d2938bc64dcb58d54c966

Observation 5e3f08f2-4f8f-4a72-9a3e-13b8900b7457 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:37.200752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:37.200752Z digest=sha256:545654c8107a06766821650222c95209410dee6a3b54ca4ec345022e69aeed16

Observation d329632a-b6e8-47b3-8adb-85009046346e · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:37.341730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:37.341730Z digest=sha256:bb723aaad538f74a65d79f871591e493e93cec3c78c74717af871400181d57dc

Observation c28d77a2-3c32-42fd-8617-a38ec0732f0a · outbound

This paper cites GPT-4 Technical Report.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference GPT-4 Technical Report

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:37.492741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:37.492741Z digest=sha256:490fb72b7574fdf61597f305a82f22f8e937f34eba230acabfcdc02038aa888d

Pith citing papers

No inbound Pith citation observations are available.