Pith. sign in

Paper Citation Record · LEDGER

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference

As of 17 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.03867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03867 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:30.919036Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9c12499-2e2b-41b5-a52a-a149e625ea0b · outbound

This paper cites AMD Instinct™ MI350 Series GPUs,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference AMD Instinct™ MI350 Series GPUs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.254496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:28.061738Z digest=sha256:57e24d9b1d677c799c715a379df5b394f9434823939e766440016a2c639a80f9

Observation 38f6080c-a363-4330-adc0-61b81412f0f8 · outbound

This paper cites System Card: Claude Opus 4 & Claude Sonnet 4,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference System Card: Claude Opus 4 & Claude Sonnet 4,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.245420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:28.148840Z digest=sha256:2d3be154c56eb68b3edf855586111cc5f64c7c936ef4b0b226c95588a8f1cab6

Observation 83a2abac-bb0c-4fd5-a249-0cb396aaaee0 · outbound

This paper cites QuaRot: outlier-free 4-bit inference in rotated llms,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference QuaRot: outlier-free 4-bit inference in rotated llms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.236474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:28.229286Z digest=sha256:8136b29125173f23277650da1ff2c2cdc034f7a2a2b8ad5f42c5fabefff32a10

Observation 949f57d1-a545-4f46-8424-af02b4034e2e · outbound

This paper cites Cacti 7: New tools for interconnect exploration in innovative off-chip memories,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Cacti 7: New tools for interconnect exploration in innovative off-chip memories,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.263418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.263418Z digest=sha256:9efd06bd40f983c4ffb80ad5254ee24f64a076eda593886b70d1c125c891ae53

Observation d98bff83-0c8f-406b-ba1f-0fbe50a22354 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Piqa: Reasoning about physical commonsense in natural language,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.227307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:28.329216Z digest=sha256:7d3b00f3a85818c0a55fab635460ce5866d92b27cd5bbc1f11e2e7622817445f

Observation 8fc5a00d-0a28-4996-9cd6-cc39ec0b7450 · outbound

This paper cites Int v.s. fp: A comprehensive study of fine-grained low-bit quantization formats,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Int v.s. fp: A comprehensive study of fine-grained low-bit quantization formats,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.401800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.401800Z digest=sha256:8f5a775017d844a5ee5ef5bb4f5ddde608491aa0cfbbbe7ac07b55e607aa2674

Observation 5c2c8913-fb8f-4491-af78-b5ec33739032 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.471068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.471068Z digest=sha256:b51e84078a18dae293e6e196637a3045bd83bb0f90768aa0ca1d374fd8f8f607

Observation 48277872-e5ae-49e4-8e7e-4aab138abb8c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.580769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.580769Z digest=sha256:5a5b6d9e41488bd3b05361cec54f2db0b0578faf2aedaf9f60a11571ee4c27eb

Observation f9a7d5da-ac93-47c1-997c-294233991dce · outbound

This paper cites Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.676839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.676839Z digest=sha256:82d69c4c8ddc327ad2415226c7feb836fb45b881e0617f5c38c972627d0b9f7c

Observation a83d7654-d208-40e5-858c-cd7b06bc6112 · outbound

This paper cites Efficient precision-scalable hardware for microscaling (mx) processing in robotics learning,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Efficient precision-scalable hardware for microscaling (mx) processing in robotics learning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.699322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.699322Z digest=sha256:ce6a602f99ff34b76fbeb6a38eeedca67d3a4aa549eed014fe689ae118721e50

Observation d0573c81-314f-4b28-a94e-022414ad12b4 · outbound

This paper cites With shared microexponents, a little shifting goes a long way,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference With shared microexponents, a little shifting goes a long way,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.792184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.792184Z digest=sha256:ec3585c2a4ee97dcbd3ab472ed8ca27bd1e85272a524a47f262b997235809f21

Observation 479b6f3c-623a-4d54-bc5a-2de962b1b321 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Deepseek-v4: Towards highly efficient million-token context intelligence,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.907764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.907764Z digest=sha256:0b6888d2b67805ac068576919b995d9aa4aa8fa8e90df86cabfda554a6a15495

Observation b762538d-8fec-4633-999c-c990af96b4c5 · outbound

This paper cites LLM.int8(): 8-bit matrix multiplication for transformers at scale,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference LLM.int8(): 8-bit matrix multiplication for transformers at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.212025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:29.005235Z digest=sha256:05be4bb213bacaa74aae1209b59fc998af5b6034f26acf992ff68435a79df096

Observation 0d52327e-a994-4712-a432-e8aace3f64cc · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.065972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.065972Z digest=sha256:bf1fb1aa93d65c19cc141668e57fae0faa666fb8ab978c34a70e8f7c2af4ea75

Observation b83e9f0c-927d-4c3f-8660-7c00c0204d1c · outbound

This paper cites Extreme compression of large language models via additive quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Extreme compression of large language models via additive quantization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.202728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:29.143132Z digest=sha256:189e937b5880f10e08be1409996f2b99b29c2daa49f51fd359e1f53d778284b9

Observation 35899f3b-85b2-4c86-97ea-076263ccede1 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.303924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.303924Z digest=sha256:adbcde15924c2033a70d9ced2553a479fc471d2957f427e9cb61364c91205a1e

Observation cf638634-a821-4b27-9813-0eabcd400783 · outbound

This paper cites The language model evaluation harness,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference The language model evaluation harness,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.417903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.417903Z digest=sha256:fe127917ed338b174006d95abe2112c3149f6e664aed2b463bfd8f8e1728d9ce

Observation 2a8c3788-d6d7-416c-a0aa-37bb6bd0f321 · outbound

This paper cites Gemma 4 Model Overview,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Gemma 4 Model Overview,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.193835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:29.550359Z digest=sha256:cac1066d3f3dc812425092956a92007159a26e5831cd7b44451e44cc0599e57a

Observation e7b05792-47c7-4ea6-b42d-3ec3fdfafa03 · outbound

This paper cites ANT: Exploiting adaptive numerical data type for low-bit deep neural network quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference ANT: Exploiting adaptive numerical data type for low-bit deep neural network quantization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.184673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:29.637719Z digest=sha256:8d8a96231ef36dcea421c80236e9b932650ca72b0ae3cb96b491b5fe9a31be62

Observation 8029f34d-20a4-4469-bd82-df0ecdd958b5 · outbound

This paper cites BBAL: A bidirectional block floating point-based quantisation accelerator for large language models,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference BBAL: A bidirectional block floating point-based quantisation accelerator for large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.770682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.770682Z digest=sha256:54f8635311ee1b3babe758c09b4470fd4271e257f8a30d6959a243785c4e846d

Observation 125c4c63-2891-4ec1-a961-9cf9aef6f499 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.896018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.896018Z digest=sha256:241bdafdafd7c57f666157da47e7231d53a7b7626a5bd1cdc728eb17878f9bc9

Observation d5f515fe-86d0-4f6a-870e-042bcc1a7de5 · outbound

This paper cites 1.1 computing’s energy problem (and what we can do about it),.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 1.1 computing’s energy problem (and what we can do about it),

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.016464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.016464Z digest=sha256:f565dd364ded45074e9eddb5e0ca249a79c32831e1fa1a927ad0f8c8bbbbf7b7

Observation 17875c65-6793-40a3-bcbd-cb21bf5e4199 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.113712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.113712Z digest=sha256:5c0fbc684976fc0f0ddc8c62ae2a9183f0abad6560e7d16e25ef4f5bd576455c

Observation 4770c4e1-84c9-40a2-820f-f86281e27bae · outbound

This paper cites M-ANT: Efficient low-bit group quantization for llms via mathematically adaptive numerical type,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference M-ANT: Efficient low-bit group quantization for llms via mathematically adaptive numerical type,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.169110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.264778Z digest=sha256:b2395ad0e7d3c491c023b76e9c16837d4b9237a48c493218ba7b1ccfaa96e0c0

Observation 8a9df075-b593-4db8-b489-565517638e7b · outbound

This paper cites M2XFP: A metadata-augmented microscaling data format for efficient low-bit quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference M2XFP: A metadata-augmented microscaling data format for efficient low-bit quantization,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.375670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.375670Z digest=sha256:490cfd0bb5c789f4a65b664cd4898dcbe980bac665e93498e02ccd94b9d91fcf

Observation 3c1f572b-6924-49b8-b55a-265e253bf2fd · outbound

This paper cites Blockdialect: block-wise fine-grained mixed format quantization for energy-efficient llm inference,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Blockdialect: block-wise fine-grained mixed format quantization for energy-efficient llm inference,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.158601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.535766Z digest=sha256:e4e1f883a33717362ee89194c9511aa87f3c015ae2f59765e044cc6b10215546

Observation 1942b81a-3f95-47fc-ac60-9edbe07e590e · outbound

This paper cites A diagram is worth a dozen images,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference A diagram is worth a dozen images,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.147849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.700003Z digest=sha256:1f27d7259f7e2235feee783effd8d84e92e3097aca8ea6b74304a11f6b419dc0

Observation 5fb49d37-4664-4019-8491-513e06c21bb3 · outbound

This paper cites 14.2 a 16nm 216kb, 188.4tops/w and 133.5tflops/w microscaling multi- mode gain-cell cim macro edge-ai devices,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 14.2 a 16nm 216kb, 188.4tops/w and 133.5tflops/w microscaling multi- mode gain-cell cim macro edge-ai devices,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.138096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.799634Z digest=sha256:2658b5047eddc49d7b7aabe231f046c30bc522b9c992bf956523990c0c940210

Observation 9293ecb2-7281-40f4-9ea4-f69183e2d117 · outbound

This paper cites SqueezeLLM: dense-and-sparse quantiza- tion,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SqueezeLLM: dense-and-sparse quantiza- tion,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.127987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.802603Z digest=sha256:7256ec73fcb1f8f0793ab99732300c50aeb9fe0c2610e2a0b0307e67f908259e

Observation 1f446c27-ee3b-44c2-b0aa-87ce14dce278 · outbound

This paper cites Tender: Accelerating large language models via tensor decomposition and runtime requantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Tender: Accelerating large language models via tensor decomposition and runtime requantization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.118081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.805468Z digest=sha256:e163740d28703cdd11cf7447053a9a5af7d0ffa25f22644f1349d6fcc5a42632

Observation 68b802df-f590-4a37-8d47-e15787d70166 · outbound

This paper cites MX+: Pushing the limits of microscaling formats for efficient large language model serving,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference MX+: Pushing the limits of microscaling formats for efficient large language model serving,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.808079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.808079Z digest=sha256:7f7272a2a5cee68635042a559edd1e2956e0744d46a2afca95d2f39a23025dd4

Observation 9b42c4bc-c813-4322-9aeb-87e3d7147472 · outbound

This paper cites AWQ: Activation-aware weight quantization for on-device llm compression and acceleration,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference AWQ: Activation-aware weight quantization for on-device llm compression and acceleration,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.810712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.810712Z digest=sha256:7f6e97f7f89f5194db08897ab6f9da3b9c9ea454c49463d3e6b737ed3809df25

Observation 7d175522-191d-485b-864f-3e336d846344 · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.813720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.813720Z digest=sha256:9814f1ad560be6701f3db80ff2a9265596e10e17988bc6c521202a461fb4d421

Observation fe303a0a-7967-48ef-bba0-6916b10b4dd8 · outbound

This paper cites Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:54:31.243463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.816676Z digest=sha256:3ebb162babd67022f26df4e196393fa75376e8c8be630c525bc941697f3bd3e4

Observation da2dd973-057b-4547-84f8-d408236063a4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.820251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.820251Z digest=sha256:46dbfe584158f8930360b2f08fddb3ad1473f1e6e1baa9e73ba9f636d4ade041

Observation dbf06264-3899-4c5b-ba41-699d6a484938 · outbound

This paper cites Pointer sentinel mixture models,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Pointer sentinel mixture models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.095906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.823013Z digest=sha256:aee7a0c03381a8c112411d049adf64a4db2ba899d3b787b6625ad9ad020bfcb0

Observation ce95d21e-1406-4669-98a4-0523f4c90fca · outbound

This paper cites The Llama 3 Herd of Models.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference The Llama 3 Herd of Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.825891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.825891Z digest=sha256:4244ca2606e53e9b4b541808d8235ace0d2d674cfbab92449a37be8823826eb2

Observation 28354834-da05-47a8-a1f2-f1e269bb85d5 · outbound

This paper cites Four MTIA Chips in Two Years: Scaling AI Experi- ences for Billions,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Four MTIA Chips in Two Years: Scaling AI Experi- ences for Billions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.086557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.828823Z digest=sha256:7e042cbd282e00df0046fe86686bd873c99255df848fd05da3cd6de67502c864

Observation 4524645d-eafa-4379-bd0b-c0b96d1facf0 · outbound

This paper cites Recipes for Pre-training LLMs with MXFP8.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Recipes for Pre-training LLMs with MXFP8

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.832174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.832174Z digest=sha256:544f9c522770fb2919746956c2fee1837973a337001aa32724ddc1d740c82473

Observation 89dd0de1-3794-4a65-81dd-6b76de128c2f · outbound

This paper cites NVIDIA H100 Tensor Core GPU Datasheet,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference NVIDIA H100 Tensor Core GPU Datasheet,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.077146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.835513Z digest=sha256:c0e527d31a2d2840f43dbf5c677567f16c74cb0501fdb8c0f224761d55259234

Observation ede43b36-648a-4bc5-8422-66a65e77d04c · outbound

This paper cites NVIDIA Blackwell Architecture Technical Brief,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference NVIDIA Blackwell Architecture Technical Brief,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.058539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.841679Z digest=sha256:b7900f5b76505b92edeff6a1120acb057d5228c937001b1dba37969bef24735d

Observation c378b2c6-7f03-425e-8096-0634641038ae · outbound

This paper cites OCP Microscaling Formats (MX) Specification,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OCP Microscaling Formats (MX) Specification,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.048394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.844766Z digest=sha256:e48eabec202c7a940b00f3b827321894ac565f50ddd8f2a3d3c41a3fffa5e3ff

Observation 96f60da0-4c40-4a88-a26d-c01fe9b2c159 · outbound

This paper cites OpenAI GPT-5 System Card,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OpenAI GPT-5 System Card,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.037525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.847687Z digest=sha256:858dcad8ef6a8ff5eb0446c36b16159d5a3add8046be5e9e250353302a9d8a42

Observation 9235f464-15d1-4964-82b4-eab6716076e7 · outbound

This paper cites Paszke, S.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Paszke, S

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.850566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.850566Z digest=sha256:770b0fb9cc440e6a6b5f1e90b661c248489114cfb27084c1218911dab051ff6c

Observation 026e5800-ad1d-4268-b9f8-e02e0ec0b35a · outbound

This paper cites Scale-sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Scale-sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.019795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.853377Z digest=sha256:d81efd9e8a5c566eb81b814a9ed0c68e7904c203d9096f5e56b9f4ec7e18496a

Observation 6faf28f5-d8e9-49bf-8d03-f2d1878ab052 · outbound

This paper cites Microscopiq: Accelerating foundational models through outlier-aware microscaling quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Microscopiq: Accelerating foundational models through outlier-aware microscaling quantization,

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T10:54:31.211877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.856127Z digest=sha256:9955ed627081b85b9fdea70e24491f188e266c875ca2595ece37020a2e2603b7

Observation 988dd783-6710-4168-9800-8fa23fba956c · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Gemma 2: Improving Open Language Models at a Practical Size

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.858903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.858903Z digest=sha256:2b53da08e18b9c2dda1efbadf44791817375e91494e1578fa80cfcba9efd8436

Observation 0d5b7539-4204-4fde-a6af-5cf1f3177c2b · outbound

This paper cites Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.008859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.862614Z digest=sha256:e2bd1dc91e2ee05b2d9fb14d181fd44a47e3e3c399aec0ca958a4c629fb20e32

Observation f8149f9c-a0c1-4b1e-84be-2c3e9c5b26d4 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Microscaling Data Formats for Deep Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.865642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.865642Z digest=sha256:87f4ad2a9e714a2913b0dfe5619b045e5bd04bd291e6b9d278636ad205a9e5f2

Observation d1b68ba9-8e7e-4d30-9af9-23f8051ab664 · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Winogrande: an adversarial winograd schema challenge at scale,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.869033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.869033Z digest=sha256:e5b9a5150176130ea1e98b65487dd667cd0d750d805e5e3560a29a5faa6a3d27

Observation 3c3f1d7f-18c7-44ab-b77e-62787b4c2eaa · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.872438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.872438Z digest=sha256:ebe4730fea96190cbdb3948b9117fe1e2a3e324afff18e3939e46cd806a405d9

Observation c817380b-ca78-47b7-a649-d41cb3bb13ef · outbound

This paper cites Towards vqa models that can read,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Towards vqa models that can read,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.998117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.876298Z digest=sha256:850ca7954bceac0fb5c3b47ef4b085f90563519bea2c93b15183e301f5b04780

Observation ea0eed34-34b1-4b9b-b1c9-15b74de6e25a · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Commonsenseqa: A question answering challenge targeting commonsense knowledge,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.988088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.879465Z digest=sha256:d42d0b86118a7310384896f5c170b9aa884ad6b20f7012227164b89cbcebf4bd

Observation 92f0e344-78fd-40aa-ac6a-384dabf9c983 · outbound

This paper cites A microscaling multi-mode gain-cell computing-in- memory macro for advanced ai edge device,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference A microscaling multi-mode gain-cell computing-in- memory macro for advanced ai edge device,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.977638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.885660Z digest=sha256:b718de8513b029b28b7bbe1fb673bb3d2cdd38179cb25adb1d1a7388348bb9c2

Observation 0f9c46a0-8a74-49b2-ad29-c854ada80fd8 · outbound

This paper cites Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.966233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.888361Z digest=sha256:28a1752dc46113cec86af1b799d3a09b626bad3d9c7b13a0ec1f07f532e19696

Observation c8c0dbc9-5fda-4415-a49b-b34e1914798e · outbound

This paper cites 30.1 a 28nm 127.54tflops/w mxfp6 and 117.42tflops/w mxfp8 compute-in-memory macro with adaptive- preserved-bit-width and serial-dual-bit-sliding schemes,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 30.1 a 28nm 127.54tflops/w mxfp6 and 117.42tflops/w mxfp8 compute-in-memory macro with adaptive- preserved-bit-width and serial-dual-bit-sliding schemes,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.956423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.890967Z digest=sha256:c41a61d7569c34aa3038ce6ef7e45720bd1bed125a8561544ff481d434ac9023

Observation e41a0196-9109-443b-97d9-14fa97783159 · outbound

This paper cites Accelergy: An architecture- level energy estimation methodology for accelerator designs,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Accelergy: An architecture- level energy estimation methodology for accelerator designs,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.894125Z digest=sha256:9ae9fe0a5a134e9058ca6be166205e3332e0455808eb8c619af54b5936d24eb7

Observation 63e2bea2-e0a6-42cc-b28d-a2181327298e · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.939793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.896961Z digest=sha256:493fb8625ec975d133e454a4f9c46cb01df4a5a4122f96511bb705e6e7c0b8f5

Observation 14e1b5ac-698f-466e-81f7-f6a81c81dbbc · outbound

This paper cites Inside Maia 100: Revolutionizing AI Workloads with Microsoft’s Custom AI Accelerator,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Inside Maia 100: Revolutionizing AI Workloads with Microsoft’s Custom AI Accelerator,

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-05T10:54:31.096960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.899617Z digest=sha256:bf9067de27f6f2ebbcfa89203a9bebcefd0ebcf4c15c6974f50590757b4a6038

Observation ce701b2b-ff15-4e2e-ac69-3e6a1a2eefa3 · outbound

This paper cites An empirical study of microscaling formats for low-precision llm training,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference An empirical study of microscaling formats for low-precision llm training,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.928622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.902228Z digest=sha256:565856e7f0b693bfcfe9c8bccd88cd6bfa66803790f216d1398ece0954b1d819

Observation a94af574-4fab-41b3-be30-29247e664c8b · outbound

This paper cites Qwen2.5 Technical Report.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.904865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.904865Z digest=sha256:1c800afb13f1d2dba1e1b6a7e2fd60d2159858b34f99f9ab059ac025a7cf25f6

Observation e4480135-ffae-484b-a485-d5b94840eeef · outbound

This paper cites ZeroQuant: efficient and affordable post-training quantization for large- scale transformers,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference ZeroQuant: efficient and affordable post-training quantization for large- scale transformers,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.919045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.907783Z digest=sha256:56609803c52e3113d8d3e7b731ca34dcb07049b41fd40e184e9fd7be89f80daf

Observation c371b86c-9450-4eaa-aa06-56caf514d875 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.909352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.910696Z digest=sha256:7c7b6699be808f3fe340c5ad826bfa105ba814858c860e55edf77e52e354bc3a

Observation a5732eb2-2ab1-4f99-a65c-9ab9668675d8 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence?.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Hellaswag: Can a machine really finish your sentence?

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.899313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.913180Z digest=sha256:440cfb70e394251437affd2ffb25812457dc258dd735246ff5f8edb11c862e01

Observation 2fedea90-72f7-44d1-b563-6394105f8d4b · outbound

This paper cites Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.916143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.916143Z digest=sha256:a9a97958680b5591a0326ebcf57647c66a079d6ea32ddee6976c8cd8b976ee9a

Observation e0425dbc-440e-4ee4-84bc-020d68b61a3d · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.889954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.919036Z digest=sha256:511dd605f4e5a5886211a9f09df6f3f3929880598f1df56586596cef8dc5dc32

Observation ebe5e4cd-69d1-4586-a98f-f71b09d188b4 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.882544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.882544Z digest=sha256:f934780cd27a3038126e8422ec9df668402e0bc5df7916281fd875833c4bb5f0

Observation 2d2a2b2c-93a8-4238-84da-dbd2ecbd8d5d · outbound

This paper cites Available: https://resources.nvidia.com/en-us-gpu- resources/h100-datasheet-24306.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Available: https://resources.nvidia.com/en-us-gpu- resources/h100-datasheet-24306

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.068016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T10:54:30.838691Z digest=sha256:ad8c4945c8a3112f71762a6e2ea374bccce14eb83af9ca5d72a38cc1184f3ee1

Pith citing papers

No inbound Pith citation observations are available.