Pith. sign in

Paper Citation Record · LEDGER

MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2408.11743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.11743 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:14:01.423333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:45.492485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3d38cf8b-ea40-411f-9334-1878b369d20b · inbound

Is (Selective) Round-To-Nearest Quantization All You Need? cites this paper.

Is (Selective) Round-To-Nearest Quantization All You Need? MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:01.423333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:01.423333Z digest=sha256:fd0f7eee3bd6140c636ba829f240bbd0844e2acf88536c2f6b999019d8178870

Observation 22abf7b4-91e0-48e1-b5b4-8b0ae3a4152c · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.650576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.650576Z digest=sha256:86c66b442d5502816c5d6bddeaa8772eee2a02632ebf93723a2af82401f6d14e

Observation fe734b39-930f-4176-a525-f89adae7241c · inbound

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models cites this paper.

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:32.827124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:58:32.827124Z digest=sha256:14269a37121cc058f59dc8b6e05f995d76b7bd57387e10af948d403e131e627a

Observation dba3f5c5-a497-44c2-aed6-0a5112f34dd6 · inbound

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning cites this paper.

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:30:22.035971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:27:49.943714Z digest=sha256:2877cefa00b486eef5a81078b64cd370679893ab887ca108405f70decad9e0c2

Observation cc2268d3-fe29-4278-bbfc-c56238e32119 · inbound

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate cites this paper.

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.871774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:15:25.783283Z digest=sha256:915ee901dc721cba1035f9f5ed96226122c4034982d32d2028d2ecd8022e032e

Observation 748daa29-4c86-42e9-b226-daa366409178 · inbound

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks cites this paper.

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.037549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:21:30.101748Z digest=sha256:7f886beae3b7339d48eca14b89a81508e19bdcaaff04791d3a59637fa843ca0f

Observation 0383d94a-7265-4073-a26f-8b039f84430d · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.389954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:69d923a68483ee1ec947de49a3b9c21b45cde8dc05b1dac17e4c746ee0b1b80b

Observation 0b29538f-c8c9-48b2-a311-c077cbef8d77 · inbound

Statistically-Lossless Quantization of Large Language Models cites this paper.

Statistically-Lossless Quantization of Large Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.085341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T16:22:13.399819Z digest=sha256:4640e2161472d74fef0156b9f8de6d5d54a8006f9a6a0f3d4c8fc4b0212f0252

Observation 00ce7d43-f1d6-40fc-80e7-5d319d9656f2 · inbound

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference cites this paper.

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:58:34.202730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:57:18.944617Z digest=sha256:81fefcc447cc57984dbe7f00897f7307ec7666eb7840e6b02fbe1b96ff00a2ea

Observation 99b5642a-a6df-4c1d-8d29-951f569a962d · inbound

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models cites this paper.

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.737221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:41:57.350416Z digest=sha256:43df446db383cf6b66e01523ce8f0e37b8f2242d97115f75f5a87f1360987de4

Observation dff7ba71-ad0f-482c-8229-0f5facc5dc09 · inbound

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis cites this paper.

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.436235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:59:11.821440Z digest=sha256:adac89177c40d94ad6fa04f4fbed32f44b3715f21431c163af3b0b5655cef6db

Observation 0387748f-437f-4a66-821c-29a58aafe98a · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:57:44.946716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T11:26:56.230283Z digest=sha256:66567289bb9cf66d051d12fcf49dcfcebbebbed6f4f21bd749fe558f5a9cd248

Observation 3a4a5887-8b73-41cd-8e0f-03d3901bdf78 · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:02.894834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T23:37:25.391853Z digest=sha256:4a209c2927c0163d1bf2d4343946eeef97ad6e6ef0caab7c1c006f56716d095e

Observation 41709bc4-e638-4950-a218-3121d2e44dbc · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.493866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:4f3bc3198e215f20254295929e6d143df3b0368b49198d10d73498344e436606

Observation b43e328a-fb3f-472c-96ff-8c1c3a949f1f · inbound

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study cites this paper.

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:15:48.293856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:15:48.293856Z digest=sha256:beae64292b0f3ee4b4872a555926f41abe4195e5ab8134e5aac145e8545ed933