Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:23.218413Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2505.08981.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:23.218413Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dbd54209-21f3-4fed-9405-838acd6c5b68 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Rae, Oriol Vinyals, and Laurent Sifre
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6beb585-68ae-4fd4-a4b5-558043069dde · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition M4bram: Mixed-precision matrix-matrix multiplication in fpga block rams
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b1d16bfc-9ecb-420f-b000-ed94bec445ba · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Msd: Mixing signed digit representations for hardware-efficient dnn acceleration on fpga with heterogeneous resources
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dee71ee8-47f8-45c7-afa4-e18926d0dec5 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Democratizing neural ma- chine translation with OPUS-MT
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad9b9f58-fb11-4ec6-9f3c-1c534041a094 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Omniquant: Omnidirectionally calibrated quantization for large language models, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3fff66a-3173-4584-9548-a39e5e29db75 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Qllm: Accurate and efficient low-bitwidth quantization for large language models, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6f3fc791-5796-4577-a26e-910f529db4d2 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Efficient arbitrary precision acceleration for large language models on gpu tensor cores, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4898226-3c6f-4410-be4d-42511c865935 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Q8bert: Quantized 8bit bert
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aac596a7-35fc-41dd-b2d3-efa5553c8959 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Q-bert: Hessian based ultra low precision quantization of bert
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 09eb866a-4af7-47ea-b2e4-40e00e3fa77b · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 96c12863-bda7-490f-a7ff-1c3cdec005a2 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Hawq: Hessian aware quantization of neural networks with mixed-precision
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ac93e7-a401-4204-8c39-840f13031454 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 51426ebb-e62d-4c45-9137-6b920764db9b · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Optimizing bit-serial matrix multiplica- tion for reconfigurable computing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 60f872e6-3a1f-403a-a971-02abb297cd71 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Hihispmv: Sparse matrix vector multiplication with hierarchical row reductions on fpgas with high bandwidth memory
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad0b3b13-9f5f-4fd5-97c6-07321e878da9 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 90de80ed-2264-47ba-883d-ce89c42b9af4 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 275824aa-a5a4-4876-a287-29f5aae853ee · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Streamsvd: Low-rank ap- proximation and streaming accelerator co-design
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5ba8875-fc91-4098-973f-55f9f27a509d · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architec- ture, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2473c3b1-39c4-4243-b2b9-d609f08395be · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Film-qnn: Efficient fpga acceleration of deep neural networks with intra-layer, mixed-precision quantization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ebbecc7-6dbf-4877-ae6d-e4c0ae4d8643 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Understanding the potential of fpga-based spatial acceleration for large language model inference
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1fe1fd53-c6bd-44b1-a568-66745d185d29 · outbound
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.