Pith. sign in

Paper Citation Record · LEDGER

MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2408.11743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.11743 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:01:20.081983Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:45.492485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bae30d21-2384-42e1-abe4-7ed254df34aa · inbound

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem cites this paper.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.150804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.150804Z digest=sha256:bc956f835a403fbb84bd3195b95d0cf13ac330a634b2a2189407035d86243054

Observation 13c6b8ca-56a4-4744-9da8-a3e1a6b31f5f · inbound

iServe: An Intent-based Serving System for LLMs cites this paper.

iServe: An Intent-based Serving System for LLMs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:14.215382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:14.215382Z digest=sha256:78ac608a9d516b430bd5227da72831d570e9b7ecebd085882ef769bbe005d3b6

Observation 317a89ed-31ea-490f-b39a-6968f2e8019b · inbound

A Proximal Operator for Inducing 2:4-Sparsity cites this paper.

A Proximal Operator for Inducing 2:4-Sparsity MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T01:15:20.402556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T01:15:20.402556Z digest=sha256:d735d5ddeabb346e67946359e7764225553f6d073228234d215a0452fbbd3ff0

Observation 3d60be69-3fb9-49a9-b18a-05b46e5b8c58 · inbound

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design cites this paper.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.081983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.081983Z digest=sha256:056a41e2d6ad8467416661355464bc357c44f8dd589648ac95d4b27ff67e37e9

Observation 3d38cf8b-ea40-411f-9334-1878b369d20b · inbound

Is (Selective) Round-To-Nearest Quantization All You Need? cites this paper.

Is (Selective) Round-To-Nearest Quantization All You Need? MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:01.423333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:01.423333Z digest=sha256:c17d7ba0cf2c166b88c1519c9304b09f463ecd48b2a7cf7a0e71a1ed8df3f47e

Observation 22abf7b4-91e0-48e1-b5b4-8b0ae3a4152c · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.650576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.650576Z digest=sha256:000e94cbd7761d90df1ee0167bcbf837fb2516fd49bd93293a88514838d20665

Observation fe734b39-930f-4176-a525-f89adae7241c · inbound

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models cites this paper.

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:32.827124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:58:32.827124Z digest=sha256:97d42e3749ce5daa0789c619d73df445e729f7c1f54aaa8c663165511ade6742

Observation dba3f5c5-a497-44c2-aed6-0a5112f34dd6 · inbound

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning cites this paper.

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:30:22.035971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T22:27:49.943714Z digest=sha256:83d1d88336c2d1da35e40569627a03f2d37c9d8978cea6bff0bcb8180dbcde3d

Observation cc2268d3-fe29-4278-bbfc-c56238e32119 · inbound

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate cites this paper.

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.871774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T14:15:25.783283Z digest=sha256:24f968f5b5185dbda3081499e645578b94aa9bb37e020653ac7bf657ad096044

Observation 748daa29-4c86-42e9-b226-daa366409178 · inbound

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks cites this paper.

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.037549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:21:30.101748Z digest=sha256:ac1d7f5c1cf80aa057576366266f4bda63af48f357f80c2047cacb8073a6cb82

Observation 0383d94a-7265-4073-a26f-8b039f84430d · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.389954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:4a7db4cf3b8f8f110242b96edc92816a6303c746b5cbf3c958d0b3ff4d0549d7

Observation 0b29538f-c8c9-48b2-a311-c077cbef8d77 · inbound

Statistically-Lossless Quantization of Large Language Models cites this paper.

Statistically-Lossless Quantization of Large Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.085341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T16:22:13.399819Z digest=sha256:bdbd7672182c4395fbc4e8e46250da72d8238c26379d4a2f5587c8eb30ae1710

Observation 00ce7d43-f1d6-40fc-80e7-5d319d9656f2 · inbound

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference cites this paper.

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:58:34.202730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:57:18.944617Z digest=sha256:469f6e1459ed389a5720991336b61203122c5e616be3ccc633ef0e2eb62d7d8d

Observation 99b5642a-a6df-4c1d-8d29-951f569a962d · inbound

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models cites this paper.

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.737221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T19:41:57.350416Z digest=sha256:0f9e8e537acfad130abd3f7d0a6dfbb9012c74becb9270809c118b11a353d9ff

Observation dff7ba71-ad0f-482c-8229-0f5facc5dc09 · inbound

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis cites this paper.

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.436235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:59:11.821440Z digest=sha256:c64c02d9f20c6f0d5351409727920e488dbecdb014164c6ce671a69ada2f2776

Observation 0387748f-437f-4a66-821c-29a58aafe98a · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:57:44.946716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T11:26:56.230283Z digest=sha256:25a862bcc6ed0bcc88ba52947bc8e00df86f9c6f6bcb85020882f58ba04df6ff

Observation 3a4a5887-8b73-41cd-8e0f-03d3901bdf78 · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:02.894834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-03T23:37:25.391853Z digest=sha256:5fca9f6eda0deb322d6569b97f1d7130ce1a76c8d3abdc6025e234741c8ea659

Observation 41709bc4-e638-4950-a218-3121d2e44dbc · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.493866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:bb41deb518976d482e0ddd6d607d12504fdb2ae3ca5cae93719e8effb1ca7de7

Observation b43e328a-fb3f-472c-96ff-8c1c3a949f1f · inbound

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study cites this paper.

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:15:48.293856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:15:48.293856Z digest=sha256:208c7e27853d23f66330e78340a1b5fd83abd55ce5f119826bc5fe535051722c

Observation 018de6fc-3bc7-4061-a741-0e320dd5ff01 · inbound

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights cites this paper.

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:51.932311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:51.932311Z digest=sha256:a524dedb4f51cb2523693a8b1a9996dd92d7edb19a6368ff239321b1d67d26fa

Observation 6d6caa7a-ce97-49a4-bd0f-58e754eecabf · inbound

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving cites this paper.

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:09.295942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:26:09.295942Z digest=sha256:7fc20b6a8e9fcbd9f6ceb1e9b7ab38ce89b871660e5b595acc7add9491c2090b