Pith. sign in

Paper Citation Record · LEDGER

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

As of 17 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 22 inbound Pith citation observations for arXiv:2501.13987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13987 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:13:20.004756Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:54:26.511410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T14:47:14.541739Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9e28a0a-43f5-4102-b60c-d2bf42b495e4 · outbound

This paper cites GPT-4 Technical Report.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.891667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.891667Z digest=sha256:80641e5a27001ceb7455de21fd4fe1368156e1b6a0c1dc40fa965c5d6bbb91bd

Observation d96730dd-3b9b-4b4e-8f45-535fc4726c1a · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.895149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.895149Z digest=sha256:5b9c20012adcd1ae86b318343560eac5970a19644b9c44b7f73370008e3fd2c8

Observation 5a7f3044-ddba-40da-a6d2-5f8d7a476db1 · outbound

This paper cites Riemannian Adaptive Optimization Methods.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Riemannian Adaptive Optimization Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.898135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.898135Z digest=sha256:bb9939c726c9487a25c896917300d9a5bb0955584772e7f992b65c73db92de32

Observation 1a1eb707-f59b-4400-890e-dc51fa180a0a · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.901788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.901788Z digest=sha256:ea7279461e648815780ad9566cbd97aea4b676c3fb0c07087889ad99f8fc4c2e

Observation 87526d87-580b-44a3-ac45-b82f1efe4d3b · outbound

This paper cites A Systematic Classification of Knowledge, Reasoning, and Context within the ARC Dataset.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting A Systematic Classification of Knowledge, Reasoning, and Context within the ARC Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.905063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.905063Z digest=sha256:cb381a117d1b9a5ad1b7c1c91e517ac3d734304cffa4095ac087a41272296e2e

Observation 50759f5d-a546-42c4-888f-72e3fee4cbcc · outbound

This paper cites QuIP: 2-Bit Quantization of Large Language Models With Guarantees.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting QuIP: 2-Bit Quantization of Large Language Models With Guarantees

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.908404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.908404Z digest=sha256:19c296c03c8890c61a9781f4967d546112e6f587109e69c9bb420d8bc568f30b

Observation b983fb3d-502e-4f32-a4e8-cbd9a9de49d8 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:13:20.451293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T16:13:19.911984Z digest=sha256:347f84d1f4d821686e9cc26cd892c2a0d410fc11f10f97acfaedc36ad5d1aae7

Observation 8aa98808-749c-480d-a83e-082e4fe38242 · outbound

This paper cites Distributional quantization of large language models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Distributional quantization of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:13:20.440804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T16:13:19.914847Z digest=sha256:8838a6dce047efe42b1fd45061e0164b52931d6fee3c637857b2df102cd03f52

Observation 2e74fef4-d385-43d9-8b0c-1f008c0f3e55 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.917627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.917627Z digest=sha256:5ad7e11bbc0beffad31cac34b7a17b77c9475757ebe75bbf09d03536263a845a

Observation 4a5ba079-b1c3-4606-816b-65b18b35b9a4 · outbound

This paper cites an unresolved cited work.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.920908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.920908Z digest=sha256:55c0465c5349704593d7c6ed4b819bd5e9239149d9d841a11d702b04b847d3d0

Observation a0be1eaa-e935-424c-b151-c8778f59b39d · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.923928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.923928Z digest=sha256:6ff707b38fe3c17359b8b38ec948f8ede5ab640ecc96838a22fc134e9397e792

Observation d158af50-efbe-404e-a5b5-dd08d3199988 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting A framework for few-shot language model evaluation, 07 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.927047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.927047Z digest=sha256:7a2b3afface19ad2296e1f87f18c23187003bc379ff9f9246463e992d90fb8b4

Observation aaaa6f37-4691-448f-b13b-563a8c370dec · outbound

This paper cites I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.929890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.929890Z digest=sha256:7005cea261e2703f4ce93d6d27d149855aebde68e1edfc0032db77ff13fb4d3d

Observation 173008cf-539b-4b48-bf75-2b117e72f090 · outbound

This paper cites The topology of Stiefel manifolds, volume 24.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting The topology of Stiefel manifolds, volume 24

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:13:20.424841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T16:13:19.933027Z digest=sha256:668d7defff336c5f9688f007eaa877de510125d867645176da9a0f361f00a6de

Observation ecc961d3-0e04-44d8-9d9b-79cb1e4398bc · outbound

This paper cites Adam: A Method for Stochastic Optimization.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.936085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.936085Z digest=sha256:5b625420caeb071e5003916a645a4242f748594234584464d864d8c67a24609d

Observation 57493104-f9ce-4cd2-a527-6bba40b14ace · outbound

This paper cites Geoopt: Riemannian optimization in pytorch, 2020.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Geoopt: Riemannian optimization in pytorch, 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.939014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.939014Z digest=sha256:9baa25aec41fd93015eaa02f563377856096d2f78a8184907e0b5eac79671312

Observation 8099866e-fca7-4610-b48f-6d02622d2c7e · outbound

This paper cites OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.941751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.941751Z digest=sha256:d8ce6d188ccdd7c7390837b93acd3af31f357f0fa6cd121fab41df09bf4b6edd

Observation 0f5b7e03-ee08-4fc2-afb7-9614d5a8acaa · outbound

This paper cites Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.944930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.944930Z digest=sha256:d61b12dcf2e9f1a032026ed4b67d8465cd5009d3b18088633d5dc29047be64b8

Observation 930d2d93-5223-4a70-aad2-3647360bb88c · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.948946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.948946Z digest=sha256:89ae9fc081cca840b7c69f13d1ab6440db3bc60b90fde4dac30a4b6751771df6

Observation 9a126cef-7a74-45cb-aa0b-3c97758022b9 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting SpinQuant: LLM quantization with learned rotations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.951975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.951975Z digest=sha256:fc74f52a78f2b5278cb9ab9afa8c829953f7cdab1c05474ffa0e388385ec737c

Observation 166a56b1-ef19-463e-839d-1cb62672ca48 · outbound

This paper cites AffineQuant: Affine Transformation Quantization for Large Language Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting AffineQuant: Affine Transformation Quantization for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.955208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.955208Z digest=sha256:587203ca00077cdb7f1d2acd5f0785e3909cc761ae9b0ba0a00ce670d211163f

Observation 6a645625-2a43-4db3-818a-4bdb75671c5b · outbound

This paper cites Pointer Sentinel Mixture Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Pointer Sentinel Mixture Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.958622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.958622Z digest=sha256:027ffd42849311b34382a6e2ee71dbfd8709a14608de70853aa3866a1b4f0524

Observation 01b1de57-5097-44b0-b57e-c2821d653482 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.961767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.961767Z digest=sha256:d9913f62e5ea404743571807f8b2166eb73aac87c444fc3ce30f173186cd9c88

Observation 7c0038c3-37cd-4fd9-bc2c-522ec5ac7ffc · outbound

This paper cites Language models are unsupervised multitask learners.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Language models are unsupervised multitask learners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.964432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.964432Z digest=sha256:25737c02c69eb23e81449092717db6daa3cfa79b697adb13e247212cd763d3d2

Observation 1fd6c393-d5d7-46df-b17d-270372abe28d · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Winogrande: An adversarial winograd schema challenge at scale

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:13:20.404531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T16:13:19.966892Z digest=sha256:8c46c677edfc50fc06bb78490d344fe986642db0216210bfa15bda92a19974f0

Observation 67d41117-7f90-49cc-8211-4eeed2a23e34 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting SocialIQA: Commonsense Reasoning about Social Interactions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.969429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.969429Z digest=sha256:29099e3a86fcf36bebcd350a7ada4bf20cfda08a6711c4c8ea00f886fad42cfe

Observation df928ba2-f7dc-455a-b9b7-4c19c81a4607 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.972488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.972488Z digest=sha256:352bf64a4f45a909a86e11bcef0eb628601c9dc2b398191ad2824c146670f51f

Observation 90b7234d-c566-4752-880b-594c5e760a1d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting LLaMA: Open and Efficient Foundation Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.975474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.975474Z digest=sha256:e20db001480f9e392008336675f1aea9f475e769796a6f8b458ff3dba7e1487b

Observation 71a20098-ff60-41a4-bd5f-56be22d2b7a0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.978710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.978710Z digest=sha256:5282374915c5bac5b4401fd54b18dc6cfa093ab3a0afe2f8560a9976250b2557

Observation 6537f63c-1669-4b9b-a292-4bb7a9402611 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.981883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.981883Z digest=sha256:a529eb6b548b4cfe2052936489efd7e120da9a43f05ed9774d4726336eccc289

Observation a8e874f3-822b-48d7-af15-c9534422461c · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.985136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.985136Z digest=sha256:6b2a642b3ec155515820f6118de14833022c52e20bb2fab902c0f7c615cebce2

Observation d1c187c5-2f9a-4c8e-b164-77642734c8de · outbound

This paper cites ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.988135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.988135Z digest=sha256:6a7c42ed072e0793c2baae3920aac52ee30c1adce157505e4af0c1cb65b17ecd

Observation 1e4030b3-c143-4615-9710-868a7e384154 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.991146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.991146Z digest=sha256:d450a79f22f3f3b5906fa9596ec330d7e450b278edb738156dc615e364d1c27f

Observation ff7a0792-c45d-43dd-beb7-3ba3753d3629 · outbound

This paper cites write newline.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting write newline

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.994255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.994255Z digest=sha256:c47e9f7322c9fc94d9b99b82cddedd04a860270eb3be0d106bcd2dec9676fc7b

Observation 68bf4dff-eb60-4194-876c-414bde825a80 · outbound

This paper cites @esa (Ref.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting @esa (Ref

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.998256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.998256Z digest=sha256:a173ca99223fdecd35280df92fe5e9e9c0989392c13d1b62aa697dd8d65e5013

Observation 853873af-147e-4eea-bcc5-073c0c79443f · outbound

This paper cites an unresolved cited work.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:20.001515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:20.001515Z digest=sha256:c64a609822067a270005bee9c2d96b42d6afa212f74d5a90ad912c0c5ac214b5

Observation 72d7de02-d8d0-49d8-94a9-0dbc2585bd05 · outbound

This paper cites an unresolved cited work.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting Unresolved cited work

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-10T16:13:20.004756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:20.004756Z digest=sha256:8032d7e00fbe311f39c0e6ea9d9b6bbefa6567321ee6ecb91f69fcf9cae5a761

Pith citing papers

Observation 5fe2956b-78de-4d3e-ba8c-ce8102e30ea0 · inbound

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation cites this paper.

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:23.849109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:20:23.849109Z digest=sha256:1c9f0f9eca607e630fd96884271438741023ec134ee245cd84a0eac445287ca7

Observation a559fc8b-f7ee-40ab-b741-7b5fc3d2db5f · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.082348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.082348Z digest=sha256:b1288539001ef997118e0d9ae81ca4492ff4ba60bf4f7667c6ca9d3420eaaf44

Observation 76dddf82-afb5-4854-bccc-a09f17123e90 · inbound

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs cites this paper.

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:06.797957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:06.797957Z digest=sha256:b5d9c4d9c4b81eb41f6f9655b121e7dd11eed90ee1b2e942faa1ed75c3c8801c

Observation 67e9cea5-60b0-4f83-a1c7-19a6a6315861 · inbound

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache cites this paper.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.965959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.965959Z digest=sha256:01c0c9fd7f0b649e486abcaec27574d554dcd01324981e0fd0cf88adbaa8415d

Observation d79a8ae5-d216-42f4-a1ae-10e1599b4b39 · inbound

TAH-QUANT: Effective Activation Quantization in Pipeline Parallelism over Slow Network cites this paper.

TAH-QUANT: Effective Activation Quantization in Pipeline Parallelism over Slow Network OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:27:15.970801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T11:24:36.046052Z digest=sha256:d7e39ca618657c73e9ac25fae827a45d6a1ac128882f0c9c77f31b355a388b0d

Observation 4f427d71-a399-4caf-ab73-9d9024678da8 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:40.948535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:40.948535Z digest=sha256:2448802beffe0b5a714c277d7221b6e13030c3a8ab44b7ad8e2cb08366a35213

Observation 9c216eaf-d8bd-49af-87fe-7753206914e9 · inbound

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook cites this paper.

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:07:20.549231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T14:03:35.214840Z digest=sha256:7f0539851ce1dfa5eca9b6113559be4800303f21bded45bc18d71c7ddbebbf02

Observation 6d56403d-f536-4b30-b255-95fc7fe68239 · inbound

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models cites this paper.

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:24.512952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:24.512952Z digest=sha256:ef0ed4875c6bcd388b6e89280e448adf262cd7b5139e8f84001d5485f798c859

Observation 332b1750-a916-408a-9724-90631d4a2c20 · inbound

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models cites this paper.

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.353194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T20:20:10.435886Z digest=sha256:22ca5317b5463cb6640c0e595c6a81c588b1f8e48d9b5934dac5939326dae65d

Observation 7a238a67-7853-49d2-9596-37552a95dd6b · inbound

CoQuant: Joint Weight-Activation Subspace Projection for Mixed-Precision LLMs cites this paper.

CoQuant: Joint Weight-Activation Subspace Projection for Mixed-Precision LLMs OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:51:24.365884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T13:39:31.422613Z digest=sha256:af95819dfbeed62c91dd73be5f002e482c3929c6ca321e024a79955d943f788e

Observation 9d909944-23ac-45db-bc4c-45acc74a9ede · inbound

LoopQ: Quantization for Recursive Transformers cites this paper.

LoopQ: Quantization for Recursive Transformers OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:50.823348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:41:55.787556Z digest=sha256:6f9663cc421dbb1d47b0eb6fdaa518adc17cf33d6df9b930d0c0174c4ef7b928

Observation 1f2f61f0-e9d5-4cf8-9ad8-f9a6ec79edf2 · inbound

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets cites this paper.

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:16.794228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T12:25:39.417436Z digest=sha256:a1c36da81d052a3e521466202c37fbfc46d24e4c092982d57e0b972a0cc5cd92

Observation f79f4af6-d087-43fc-9b2f-f646131dfb32 · inbound

Theory-optimal Quantization Based on Flatness cites this paper.

Theory-optimal Quantization Based on Flatness OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:09.902771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:38:06.888665Z digest=sha256:f801ffb753546587177937a0d79ec59eeefc76c365e9814c92308c2400b13217

Observation b600cc3a-2a90-4409-8fc9-e66fa77b68e6 · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.619341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:faff76032b5892ad23ad94588dc3f186051136a2e64d8531e8a6e3fb9ff4e2d7

Observation af2a470a-7516-4cd6-93cf-417adf65cc40 · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.333381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:4ba812fd1f8ec9374870d84d2682d8eb54fd6d091aa02a393ae0c6535f3543a2

Observation 6a58618d-e155-46f8-80b3-7768f06d4160 · inbound

HoloQ-VLA: Uniform W4A4 Quantization of Vision-Language-Action Models cites this paper.

HoloQ-VLA: Uniform W4A4 Quantization of Vision-Language-Action Models OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.946776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T12:44:50.624778Z digest=sha256:79b2d1d7f584b3323d48a1505c3916d9b302d8fb299e80ef62949fbcec76c90b

Observation 83fea8c0-1853-41f1-94ad-521d2a547e1a · inbound

MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations cites this paper.

MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.216780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T20:17:16.036226Z digest=sha256:c40780b9b3d88d1a9f3fba930b31089eaa12720ed877521ed92625392738a664

Observation 7542aa95-3ac1-4b1c-8cea-3e8ba3d8b8e6 · inbound

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers cites this paper.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.376068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T14:56:10.553212Z digest=sha256:ae3324f6e1dd36fdcf93b00367b94bfecffe16d52bd3c35dcd1b9e8941ae6136

Observation eb0e06c6-54a2-4af6-843d-8aee223ec01b · inbound

KronQ: LLM Quantization via Kronecker-Factored Hessian cites this paper.

KronQ: LLM Quantization via Kronecker-Factored Hessian OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:47:14.543000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T14:38:16.781357Z digest=sha256:d44101803a3809f93c029e0a2d0e5b5f9e6566adb1777667885f6528f63e8038

Observation ba66d4ad-6293-4e9e-ad0b-b288e8cebc40 · inbound

Break Through the Compression Bottleneck: From Theory to Practice cites this paper.

Break Through the Compression Bottleneck: From Theory to Practice OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T14:29:36.125062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:29:36.125062Z digest=sha256:3cfe646d899026f3abed20232e8ee29ea8e8775b11ecef7b7f23a6dd0dfb15c4

Observation a0da8510-6be4-4114-a83f-c86f6c1a7264 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.815346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.815346Z digest=sha256:c3d908dc3f29ec16d920b8c05ec2398f3963e8ec12d4f3d6d5905234309ff7e6

Observation ab679e37-2974-4fad-ac33-4266e21ecace · inbound

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation cites this paper.

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T12:54:26.511410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:54:26.511410Z digest=sha256:4fffcfc5978f34d656382ba7e75d1fb58f5798852ff1a02828553f260446b766