Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:12:37.492741Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 0 inbound Pith citation observations for arXiv:2607.17733.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:12:37.492741Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 136 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0389ef3-da40-46f6-b694-d9dd588926c7 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Training DNNs with Hybrid Block Floating Point
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e210193d-bab0-49e1-b61c-e1d60621ce0e · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad65d40-01ba-469f-a601-1e9fa4083ded · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2021 , eprint=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee5b13f-7326-4bb9-a742-9a0e250b63b4 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1bb3eca-1d65-4534-ab2a-c7672dd1587e · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ab90303-b202-492b-bdd9-e5f065d5e817 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in Neural Information Processing Systems , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be13ca87-9323-40a3-bad4-990ba4de8185 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Microscaling Data Formats for Deep Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09757515-49e5-465d-a2b6-4d90a40f9eea · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 36th ACM International Conference on Supercomputing , pages=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c164e073-01a3-4777-ab6f-479d2ada1105 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 52nd Annual International Symposium on Computer Architecture , pages=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca3d3686-578e-44f1-a518-051294a96911 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , pages=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9fbdea-66b5-4ed5-893a-3c13ad9b828f · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Outliers Dimensions that Disrupt Transformers Are Driven by Frequency
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a639f2-4965-41dc-a50c-f5979d8c4a27 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b40a43-1701-4822-b146-c3fd2fdc511c · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baea74a0-fbe7-4d83-a4f2-5bd54d12bb5e · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60856a8d-2376-4ff0-ae21-e3b971402ee1 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2018 , eprint=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f61271-4cc9-4183-9a57-8f0aa716cb7b · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2019 , eprint=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d150366b-146d-4864-8500-6d7060c83dfc · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Microscaling Data Formats for Deep Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7eaa3de-7ed0-40dd-9210-432173e44e2a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot , booktitle =
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce162646-48aa-4be1-9ce6-a3778d250254 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SpinQuant: LLM quantization with learned rotations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25996b48-6dd8-422b-96bc-e1f2fada2255 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference A Simple and Effective Pruning Approach for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e82efbf-98d6-4a66-9323-68cde424840e · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accelerating Sparse Deep Neural Networks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966a36ac-594e-4460-85d5-7fc95c47b4af · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce3fa1a-6ef3-410b-825c-6da4cac418d2 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of Machine Learning and Systems , volume=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc16e26-a030-47d0-bc34-98fe85a0e521 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 50th Annual International Symposium on Computer Architecture , pages=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd64e2b-1494-475b-af17-45feca4df541 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models , booktitle =
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26c0881-9a2d-4871-9292-808ac00247b1 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 5th International Conference on Learning Representations,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f333cd00-d380-48cd-bdc3-4a4424b391ee · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6804e4d-f97a-4505-bab9-91bb0299cfdf · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Wikimedia Downloads
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3df61a-3d72-4515-a04f-de8cec42d0d2 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The IEEE International Conference on Computer Vision (ICCV) , month =
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68e3a6b-6601-45e1-b8d1-b1c1da5fab5a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2009 IEEE conference on computer vision and pattern recognition , pages=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c9bd04-9fae-4297-887f-17543f335fa9 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Gomez and Lukasz Kaiser and Illia Polosukhin , editor =
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def36568-02ca-476d-86f0-fb1192916241 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 644754d7-4896-4a11-8dce-b379667eeb6c · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 9th International Conference on Learning Representations,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474a638d-3154-4452-a01e-77602f0cde8a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2019 , file =
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a877c2d5-a1f7-4888-9baf-33308fbe9605 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2024 , eprint=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7a4a44-a773-460e-928b-ad3ba0d012ae · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Mixed Precision Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f2d4e3-27db-43b5-96ff-c32b9f855a35 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Chung and Zhaoxia (Summer) Deng and Sam Naghshineh and Jongsoo Park and Maxim Naumov , editor =
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c5b0b5-783b-4aaa-a46e-b81edeb8e1e0 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point , url =
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f1a71c-be6a-43d8-89dc-7a4fdc419470 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , pages=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29762dc7-fc3d-43a3-be61-016b011f42e6 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference FP8 Formats for Deep Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b3fb62-35de-4c9d-ab99-2c785ed18b7a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fb1707-c374-4d69-884e-ca9c5228a565 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2024 , eprint=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305d4f87-174e-49f1-b38d-6a611044d583 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the AAAI conference on artificial intelligence , volume=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1c64f7-e564-4713-ab81-3c9b78de1a97 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference ACM Transactions on Embedded Computing Systems (TECS) , volume=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe9537c-75da-4cbd-8d2b-47820e39d969 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in Neural Information Processing Systems , volume=
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64484466-22d4-4f93-b236-c19aa6ff4dd0 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference ArXiv , year=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1815c434-9e51-456b-a152-bd5f2a7291e7 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on machine learning , pages=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797337d0-c682-466f-bdb8-ee247590fa9a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6424e814-38e0-4037-939a-c5b3d31e7335 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Joint Pruning & Quantization for Extremely Sparse Neural Networks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32dce6fb-050d-4af8-ad01-1edb704f03f4 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Towards Optimal Compression: Joint Pruning and Quantization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8446f3c8-35a7-4497-9763-2116970a036d · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7941fbd8-78b7-48a4-8d00-31924986a027 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Deep Compression of Pre-trained Transformer Models , booktitle =
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba23c4a-96e7-4f66-b483-04d77f2d208b · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65dcef89-67da-49b8-99cc-08dedd0ad554 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Frontiers in Artificial Intelligence , volume=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c5b75c-c143-4f4f-981e-bf7e805fc667 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on Machine Learning , pages=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95579736-d00c-4d24-92b0-f59d40975db7 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c751855a-07e2-4107-bd3d-766c41280517 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be79ddd-3208-48c3-a2a9-c320f0bd02c8 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fed8344-2da6-4942-9f8c-b8a8c30f01b6 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b0d489-c94f-45a7-955e-f63e5af53786 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13597980-b2fc-4589-86e2-9cbe858d1848 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241dab0e-83f3-4a53-900d-5974b2afe19e · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference IEEE international conference on neural networks , pages=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada5906c-7e26-4b42-85b2-b6f5fd44572a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4da7d07-508c-4ae8-9598-85399e46d19d · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2024 , eprint=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d797cb1-e33a-4b9f-96dc-1652136e8d05 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8746ab1d-cedf-43d5-99ce-89b99cd35aa2 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The Thirty-Third
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 373a98bd-fdc6-4aaa-b8fa-48ec2557e0fd · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accelerator-Aware Pruning for Convolutional Neural Networks , journal =
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a810e8b-7652-4b8a-b392-75ac3fbafffb · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Channel Permutations for
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343f5256-7909-482b-9295-dd4d1bfa9e39 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 9th International Conference on Learning Representations,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745da703-9490-4c08-b19e-4464f14a3ac0 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2021 , URL =
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4300319b-01b1-4228-b901-88b0544c573f · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2022 , URL =
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e9241b-bc2c-49aa-9631-a3ed4689380b · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2020 , URL =
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd403b5e-f19f-4d4a-b375-b4a06b601831 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e05e7b-dd3a-45d1-a6a1-b47c56e23fbe · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accelerated Sparse Neural Training:
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0aa4e79-4833-4aff-b50c-ded34aa84d66 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on Machine Learning,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867f4a22-91d2-41b7-859e-50d858f61792 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9560cc0-c5e4-4a2c-a9ac-65191242fae9 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 37th International Conference on Machine Learning , pages =
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b85aa1d-4cae-4590-9534-acc867ec750c · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 2023 , eprint=
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61014f32-1bc1-48f0-8eb1-7ab69133c344 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of Machine Learning and Systems , volume=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07bcc682-d597-4f52-b2f9-1927492a6882 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in neural information processing systems , volume=
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69adbd2e-2e04-4f3a-a401-c2f24fb86b1b · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Advances in Neural Information Processing Systems , volume=
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e17c489-6900-4efc-816d-db83e14640d2 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference International Conference on Machine Learning , pages=
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58cbc9d-c761-451e-aa26-f54b78697daf · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Scaling Laws for Sparsely-Connected Foundation Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c21efa-e11d-4d5b-aa82-16d3e1bf8ecc · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd3c216-b495-4a92-bb44-6cd0934186f0 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Dynamic Sparse Training with Structured Sparsity
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7af72ee-4284-4e6a-81e0-01d3a8ee0427 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Efficient Processing of Deep Neural Networks , series =
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0625bd-c4f3-4105-b3b6-67e5e2ef27ef · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference 7th International Conference on Learning Representations,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd48744e-1df3-435f-984b-5707532ed539 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Intelligent Computing: Proceedings of the 2021 Computing Conference, Volume 3 , pages=
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dbc204-fc95-4273-8515-b26b13b2a05c · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Proceedings of the 37th International Conference on Machine Learning,
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02e80b84-3ca6-4e61-a0fd-1b9befb9d992 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0798e960-4843-403a-afec-510fafde30f7 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8fb65693-9341-4d79-b5e9-98c056c4a409 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models , booktitle =
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d56828e-22be-413d-8a7f-eeddc0f13cef · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Prune and Tune: Improving Efficient Pruning Techniques for Massive Language Models , booktitle =
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2938dbe-005b-49dc-80d4-389d659c0a9c · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Unresolved cited work
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7cdbae7-6958-43f1-9bb5-4bc88e2f5aeb · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Mistral 7B
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee4cc51-b346-419c-9899-a56708e9e1ce · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46bde947-e50b-46bc-adf8-7158ecdea99a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference The Falcon Series of Open Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3f08f2-4f8f-4a72-9a3e-13b8900b7457 · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Gemma: Open Models Based on Gemini Research and Technology
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d329632a-b6e8-47b3-8adb-85009046346e · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28d77a2-3c32-42fd-8617-a38ec0732f0a · outbound
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference GPT-4 Technical Report
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.