Pith. sign in

Paper Citation Record · LEDGER

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

As of 13 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15982.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15982 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:46:01.321581Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy65
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb908304-59b4-4082-8b4e-8ca13af78c73 · outbound

This paper cites Resq: Residual quantization for video perception,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Resq: Residual quantization for video perception,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.916127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.916127Z digest=sha256:6f7740947a776f1d07e5e1bf5e9dc0f32c52435c730b7c10ccec83ca59caa5e5

Observation 07b6e85e-10eb-4742-ad69-eb44c17f2a83 · outbound

This paper cites Bit-pragmatic deep neural network computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit-pragmatic deep neural network computing,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.924633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.924633Z digest=sha256:b94129782f38d3590c157d4c49d4947b935d6c9cc45967704b7ff0a312080074

Observation 109156cc-ec88-4c3e-b740-788a6f30fc11 · outbound

This paper cites Explaining neural scaling laws,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Explaining neural scaling laws,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.928808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.928808Z digest=sha256:8fdcabdfa0268f1fd081e8383ad99e71d783fd9430009ebf1c71fc31cd08a553

Observation f22e992c-6df8-4344-b35e-d24c47398f35 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Longbench: A bilingual, multitask benchmark for long context understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.933063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.933063Z digest=sha256:defbffecb5791c4e47cba8108ad233f114eecd04004d8811007a23375da0c5f1

Observation c2fbf310-6db7-48f4-b945-3546b592e458 · outbound

This paper cites Demystifying chatgpt: An in-depth survey of openai’s robust large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Demystifying chatgpt: An in-depth survey of openai’s robust large language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.937312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.937312Z digest=sha256:80a496ba3ad9d1c6291861aea67ce31dc5608dffe978de7fd7866b77c37dd511

Observation 512cd457-8d26-4023-b16d-7a04113f5a87 · outbound

This paper cites Genus synthesis solution,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Genus synthesis solution,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.941946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.941946Z digest=sha256:56b78b9a09ae097a025fc0094dce2e0f7fc89c8cc669c80ead59252d7d4b96a5

Observation 5c53dec2-f8da-4aba-a7c7-4e8f421615c9 · outbound

This paper cites General purpose deep learning accelerator based on bit interleaving,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format General purpose deep learning accelerator based on bit interleaving,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.946229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.946229Z digest=sha256:9fe14285911da6662de818e5d455a9b68008a2553aec29e596f3c270ae718fed

Observation f864ed24-45c9-401f-ae54-5c8143377e09 · outbound

This paper cites Quip: 2-bit quanti- zation of large language models with guarantees,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quip: 2-bit quanti- zation of large language models with guarantees,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.950195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.950195Z digest=sha256:e4a0c51dd00838bbd56b5d85268248daf14c57052c574b32fbc7c18270479b16

Observation db8e462a-25cb-46fb-b2b5-1b5a23573c3e · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.954375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.954375Z digest=sha256:a7b1271f9848d6f53977a092ee584e0ebd7713e4f31ffe5e226b850a189db050

Observation c08c524f-0c14-4d61-b5d6-f0409de0d6e1 · outbound

This paper cites Nacl: A general and effective kv cache eviction framework for llm at inference time,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Nacl: A general and effective kv cache eviction framework for llm at inference time,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.959299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.959299Z digest=sha256:2ddd33053cd8631eb35294abcf4d270e441b352e7cdf66523a0874de2774b97c

Observation 3e140eb0-9c01-4fba-b3db-04b06150bd15 · outbound

This paper cites Palm: Scaling language modeling with pathways,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Palm: Scaling language modeling with pathways,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.688552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.963367Z digest=sha256:93076b4b7fc8f53a6ce9a706e77b9b5e57b7b5066d0f727cd29119e062f8ac3e

Observation 0e391695-d4ee-48c0-a897-a30f0f232543 · outbound

This paper cites Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.673582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.967701Z digest=sha256:59221bc96cd49bb7ec1188fea7c1f5dc005d701a056f73d7b812cca90cb76fab

Observation cbfde174-f710-43b3-bc1e-c0588e604377 · outbound

This paper cites Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.660384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.971544Z digest=sha256:dd72792b5b9efedbf1916d9c5217c6791bc8332f1c88ae3680456aaaebba5fc1

Observation 987c2874-d2d8-4184-a3d2-f795bbdf96e0 · outbound

This paper cites With shared microexponents, a little shifting goes a long way,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format With shared microexponents, a little shifting goes a long way,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.647418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.976155Z digest=sha256:1e36aa0a8114e270f64547e41d56d325de632adba96eba2dae8116a6e33a12b5

Observation d0088e03-3677-4f8a-8050-e4d1ccb8efc8 · outbound

This paper cites A timing-driven approach to synthesize fast barrel shifters,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A timing-driven approach to synthesize fast barrel shifters,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.635399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.980180Z digest=sha256:caed685aa3d61728dc4c26030980b8398c01d633063d133e563f5d26a4f13884

Observation ec788826-bdd6-4922-b1b5-c725344a1534 · outbound

This paper cites Llm.int8(): 8- bit matrix multiplication for transformers at scale,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llm.int8(): 8- bit matrix multiplication for transformers at scale,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.622334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.984391Z digest=sha256:4f00af0c72ae1540e557f04a2b5785057a966e1ba022df1b377179c5cdbe0238

Observation 8480b9a4-4201-4813-a0d8-68b85334e80b · outbound

This paper cites The case for 4-bit precision: k- bit inference scaling laws,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The case for 4-bit precision: k- bit inference scaling laws,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.606453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.988162Z digest=sha256:bdef7ab2d6023d9edc01cfb02d4d5a42dbfec2bdf3691ee371b8e37783254aa1

Observation 00ecbf6f-8afa-4325-9fb0-89afa02558c9 · outbound

This paper cites Hawq: Hessian aware quantization of neural networks with mixed-precision,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Hawq: Hessian aware quantization of neural networks with mixed-precision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.588365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.992180Z digest=sha256:f069af9aa5c82f5f9887056a0fa1302a43112a49e15b06f2a464a0dac4dcb95e

Observation e6510551-5932-40cc-b624-1173024cb859 · outbound

This paper cites Training dnns with hybrid block floating point,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Training dnns with hybrid block floating point,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.571178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:00.996132Z digest=sha256:5b619633201589c505be205f913be7e0b11002822ea9ff37679d8a17dde9917c

Observation dce06533-ae2f-40b8-af4b-eab6c87e6b06 · outbound

This paper cites Skvq: Sliding-window key and value cache quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Skvq: Sliding-window key and value cache quantization for large language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.557598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.000362Z digest=sha256:3739aee2ede19057ce607bc12794979c787dbbbd62a8c3b74b6de75ba16f508c

Observation 178aaf15-301b-40f6-8b13-bc2fbe940d64 · outbound

This paper cites Extreme compression of large language models via additive quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Extreme compression of large language models via additive quantization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.545100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.004503Z digest=sha256:5e14a657b0a7385854b8965af13fcea30be52e809d2884cfda739156018deea2

Observation f389ac12-1c55-47f5-80ad-99753d11967f · outbound

This paper cites Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.532602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.008439Z digest=sha256:10fc552b13725376abbcb445bc7888323530158dec9f8be5d26cded2c1e48c64

Observation 01002d37-24f7-4610-8d4a-c2d9c229c290 · outbound

This paper cites Static block floating-point quantization for convolutional neural networks on fpga,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Static block floating-point quantization for convolutional neural networks on fpga,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.520801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.012376Z digest=sha256:399e8dd0012b025c0b83072413cdc14f8560ff578e5dc21b7980c7ee6ef2c124

Observation 553aa229-a25a-4769-a67b-c312da294cc8 · outbound

This paper cites Optq: Accurate quantization for generative pre-trained transformers,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Optq: Accurate quantization for generative pre-trained transformers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.506326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.016683Z digest=sha256:c9b6ce69636314f8d8e9ee8dfb10a4e1ab97e936f683cfe99da1eb0ca8061f8d

Observation be801d61-b0fc-4bb3-a27d-e66b430b3b5c · outbound

This paper cites LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.020686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.020686Z digest=sha256:1435f00dd723194b086c2aa07638ce3c1cab631789eabf737ca94cb05e5b22ea

Observation 6fcd9706-878e-4c1c-a67b-2c44d6dc4687 · outbound

This paper cites Boost: block minifloat-based on-device cnn training accelerator with transfer learning,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Boost: block minifloat-based on-device cnn training accelerator with transfer learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.488289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.025336Z digest=sha256:9b2041fd1a7dada8def7db504438ce3d2524e6a637a61c5baf9bf1dc6c258dcb

Observation 2efccf6d-a519-4127-b9b4-af72090c2eac · outbound

This paper cites Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.472758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.029221Z digest=sha256:2487145d4e000b2e705896adde4acf9ceb1ea5c8956dbd394502c26ded263c9a

Observation 734655d4-95bf-4ea4-b953-1923a5d8a9fe · outbound

This paper cites Ese: Efficient speech recognition engine with sparse lstm on fpga,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ese: Efficient speech recognition engine with sparse lstm on fpga,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.459389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.033280Z digest=sha256:2654028ad46c3a356bb62910fb58cda2a59414108fb69a7e480299494972bf34

Observation e31e43a2-f16c-464b-a096-57ffd31fb592 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.037498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.037498Z digest=sha256:40f56f0abce37dcc8adb472e3cd845293152a7fb753bcd46aab5a39f2e4836e6

Observation acb4a88b-71fd-425a-a4d4-7557df0a3b61 · outbound

This paper cites A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.445583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.042306Z digest=sha256:7c69ab9929b03ee2695af12222606652ce6807b2514bc226d43d1a1a6f38bc5f

Observation 3c4a177a-2705-4842-99c8-b7d2c8a2e8a7 · outbound

This paper cites Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.429805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.046677Z digest=sha256:9d1417ce98c280fb334aadb61e3b19b0b8fcd81b7d7b8ef90ad79fee8e238faa

Observation 9feedc6f-68ee-45ee-b614-0cd0da64d635 · outbound

This paper cites Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.410522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.050530Z digest=sha256:67bbcf4bd87e478ca63eb51a45d7dc765d8d8c2d1b3e8fbb8f5b31b47540b624

Observation b971c7e3-66b4-44ab-9994-26903bdfdec2 · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Perplexity—a measure of the difficulty of speech recognition tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.392252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.054333Z digest=sha256:ae2aab72e35104d063f4396763d9ae814caa2aa02ee23b5d1f98ffbda3b888ea

Observation cca7ac48-4bed-454b-91a7-e725e24eecce · outbound

This paper cites Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.371118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.060235Z digest=sha256:747e980f5aa05141ce9ffc7a0da6f0f113abfc422462f19fb931238c11558d98

Observation 5a868e5d-84ae-43b1-a764-91759d50d4b7 · outbound

This paper cites Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.356272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.064416Z digest=sha256:f6ec7fd3694525a67f5939480bad52f30d4922f0af567463a3611bc25a01f612

Observation 28571475-c8a3-430d-97f8-fb952857a7fa · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i: Industrial product,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ten lessons from three generations shaped google’s tpuv4i: Industrial product,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.341385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.069776Z digest=sha256:15fae2901e0df52f115868b42b331703ee0db528eae6a15f53fb88c44ecc829e

Observation fe860abe-bc83-47e5-b316-3a520278d183 · outbound

This paper cites Stripes: Bit-serial deep neural network computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Stripes: Bit-serial deep neural network computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.326657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.073787Z digest=sha256:b13ec38d1c778fd8dda37f2d265925bbac33ad1589863cd39e04e02600397218

Observation 94942890-0859-4f55-b5a3-4a08a3024924 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A survey of gpt-3 family large language models including chatgpt and gpt-4,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.078272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.078272Z digest=sha256:073ca51e8913a3e9a761cb9967ad148e573b539dcb4d0087ab44cb2021000dee

Observation 02267acb-faf9-41a0-8f47-18c580f1c56c · outbound

This paper cites A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.304176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.082459Z digest=sha256:66b63d9c9814d3cf4029468f9bedc93ef814c92e11d3f16d5e2643628113d5d8

Observation e01546cb-3e38-4a57-a7db-f2803705ecdf · outbound

This paper cites Compressed context mem- ory for online language model interaction,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Compressed context mem- ory for online language model interaction,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.288380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.086446Z digest=sha256:beb753e0f421a9974f0d4c0864596ff060aa902faa766c30c8e194288b8a4459

Observation ba442fe0-bdc1-4a37-be6d-7f271a5796e7 · outbound

This paper cites Dacapo: Accelerating continuous learning in autonomous systems for video analytics,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dacapo: Accelerating continuous learning in autonomous systems for video analytics,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.273986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.090475Z digest=sha256:891f7b4b470c5177c13f0823caa6dc4fc3e9162f6cd1c2360bfc97bab2d2d198

Observation 048c2a80-8910-4127-b228-ffba59500712 · outbound

This paper cites Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.260399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.096261Z digest=sha256:62ffb2e57351947331ea4364eff3b1f3b4646127053fd2d3b61abc016ef352cc

Observation 05f82623-88e2-48b1-9229-de66136b466e · outbound

This paper cites One-shot model for mixed-precision quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format One-shot model for mixed-precision quantization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.244419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.101397Z digest=sha256:92dc960614cbeebc9e11a0c60e950ec6511b80941b3bf6c973255bf56d76ab9d

Observation 6de4f875-f11c-4ed3-9f6a-a87b72e63cee · outbound

This paper cites Flexpoint: An adaptive numerical format for efficient training of deep neural networks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexpoint: An adaptive numerical format for efficient training of deep neural networks,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.231230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.106244Z digest=sha256:db3a9168e2b4c6915096586211be805de7020840e748b93f6134a0beb228578a

Observation 6ff985f6-a23d-486b-8f4d-8d73abf68337 · outbound

This paper cites Tender: Accelerating large language models via tensor decomposition and runtime requantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Tender: Accelerating large language models via tensor decomposition and runtime requantization,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.218324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.110394Z digest=sha256:5c625648112a7d9a6f7703a3c4ff67f71c8fbef652a35dce2f042ca4a1dc6291

Observation 9c0e6188-8cdf-4159-85a6-5069701f9205 · outbound

This paper cites Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.206101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.114610Z digest=sha256:4cd7d185060e4180b22f1bc0c98b8c2d6ad181505acaf47070e9c1cba7debc9e

Observation 19bd6321-1c54-4f04-8dc9-116e8de5dff2 · outbound

This paper cites Norm tweaking: High-performance low-bit quantization of large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Norm tweaking: High-performance low-bit quantization of large language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.190206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.118186Z digest=sha256:7689b1bc6f900c1029d18ee955bd3bf77ab916dbf5fb3f83c5308474fc8f5faf

Observation e3b6e19c-f006-4094-a4f2-de465d4cb624 · outbound

This paper cites Geo: Generation and execution optimized stochastic computing accelerator for neural networks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Geo: Generation and execution optimized stochastic computing accelerator for neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.177003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.122265Z digest=sha256:110f2213ff3e506e1f974579654941ce460321533c1e9b81839003c177c3c397

Observation fd869e77-bcc5-47ed-ae71-1f27620566bb · outbound

This paper cites Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.160011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.126214Z digest=sha256:6c2681a82faabeb284488484f25097b8e9601cb169bd7f53fc97c91eae953ac4

Observation 995b9f46-e3c0-4229-aeed-91e1b6697a4a · outbound

This paper cites High-performance fpga-based cnn accelerator with block-floating-point arithmetic,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format High-performance fpga-based cnn accelerator with block-floating-point arithmetic,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.144687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.130338Z digest=sha256:c718ffaaee2151612fc58928a3f87dc9b8a1abcc49a2406928490848611dbcca

Observation ac37e816-1469-47bf-b6f8-e6b369beb443 · outbound

This paper cites Awq: Activation-aware weight quan- tization for llm compression and acceleration,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Awq: Activation-aware weight quan- tization for llm compression and acceleration,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.127154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.134742Z digest=sha256:a41e1bb3041aa7040fc0d68593ac1d4a2568e15769ff81e91bedaee44c15d07a

Observation 0bd651d8-2194-4958-bb11-36374a362fd2 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.139346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.139346Z digest=sha256:df93d0655d5c8f987424e15c31ae83c5c45e6a654c2eab9e87c058c6eb5220f3

Observation 96581e0b-fb9d-4487-9eeb-24b973f4971b · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.145166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.145166Z digest=sha256:7439736a6df83e0bfe52e6e7b08c526656027486148431399790f693eec0a50d

Observation 938449d2-5a6d-4ae6-93af-c87b1417f81a · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kivi: A tuning-free asymmetric 2bit quantization for kv cache,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.111294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.149988Z digest=sha256:b25b54f2b283bdfc76ef769ad54f612ccf91d6d8dbd7a902cfa05ee714a79869

Observation 9fac23e3-43b4-470b-9fe0-a571e64a12c3 · outbound

This paper cites Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.094997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.154781Z digest=sha256:e918a62514563b6c4dfebc2bb71a12ac51b4b96737cfd86b31afea6634acbf4a

Observation 9f43329e-2dac-445d-80dc-259f8ef78496 · outbound

This paper cites Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.071054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.159281Z digest=sha256:4d9872cb39c6b199f3e506d57e4641d821bd27612f7a9a8786715d781f62b33f

Observation fda3f6a9-e9e4-4d15-b6ee-fed49585f9ee · outbound

This paper cites Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:46:01.510626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.164716Z digest=sha256:f831105a4e6baee15e00ba4d409924c7e0c821ea72ea64a92b63f9e9af42958e

Observation 92cb1726-fe3b-4af0-9975-a35027aa97d6 · outbound

This paper cites Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.047111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.169746Z digest=sha256:3495807dda366eaab64d06772c5c34098bb9570ef1f162be3351098b036e426c

Observation ab2800b6-3577-4014-99cc-55e9691b7a41 · outbound

This paper cites The penn treebank: Anno- tating predicate argument structure,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The penn treebank: Anno- tating predicate argument structure,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.029380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.175099Z digest=sha256:5c76585965447dd25c7137d6954e86d04d7544a5a0f66a0bcf64cd3c785b58c3

Observation 53655596-273c-437f-8332-4a6e8b23254c · outbound

This paper cites Pointer sentinel mix- ture models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pointer sentinel mix- ture models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.011649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.180120Z digest=sha256:01a83919df489fc97fd700a6a4cdfaba5b9452cee9ede5ae95d213d4f91935cd

Observation c70b0fc3-7afe-4fca-9688-f206d80ec28d · outbound

This paper cites Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.994256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.184647Z digest=sha256:3e855f2b01864805dbe46f262b25dada5f02eb8007bc2f9cd0ec1dd512f4d210

Observation 42f71867-f06c-4e26-8f0f-df6241d81ce7 · outbound

This paper cites Cutlass,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cutlass,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.977580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.189062Z digest=sha256:56362850115ed8cea89bc41ebabbdade42774123032469f2c0c25747a730171e

Observation aaa2dfc7-d314-412d-b7be-4e114a599e64 · outbound

This paper cites GPT-4 Technical Report.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.193821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.193821Z digest=sha256:b7dc4f9aadf984c17b2e816b57ef0ce9d0030c81f14fb56ea42ddbdebf228292

Observation 38f48fab-2824-488e-99e6-22a062854016 · outbound

This paper cites LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.956504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.198965Z digest=sha256:9e60faa5c01f9c55c3424c3d6ab271875047c5bdb090bf78da86d07c7b15adf4

Observation 842a34fe-7a8a-445c-af99-adb9fc1a11ab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.937328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.203051Z digest=sha256:3f57bb5e97912fbdeaf27e5ac15d9067de6cc3ab03109f0b2b2c5613e221d59c

Observation a2e29963-82d3-4df9-bd66-09724eabba27 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Omniquant: Omnidirectionally calibrated quantization for large language models,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.918778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.207048Z digest=sha256:871b589fb56d165edf2c611b2fa994dc0b44ef9f9144609e42d8ea27467ab0ee

Observation 0c831a98-4474-4151-be7d-780a15bb36ad · outbound

This paper cites Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.902106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.211499Z digest=sha256:614f1e2e38635643f219e1b0d797f25069ea7b76eef643c4bd02068a4d338cfb

Observation 6fc9dfc1-34c9-4723-bbb3-a58c01237086 · outbound

This paper cites Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.883768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.215760Z digest=sha256:d9fa1a402861e58561939f723d5401dced45572515a48684b267101ba8060cb1

Observation 29c5eea0-0148-490a-bccc-b7a1fa42f4ca · outbound

This paper cites Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.866350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.219759Z digest=sha256:9928050fc72a0bc32c8dd4b6279975898043bcde2ccf4503d329e0a4cd46085f

Observation 9f38003d-2c2e-46f5-a22d-73d7d17f5816 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Gemma: Open Models Based on Gemini Research and Technology

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.224872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.224872Z digest=sha256:160467b1fe275c3e299763c5b184de11a31911195a11a299c9f4464aa4c954d8

Observation 141d78a5-df15-482d-b957-e9c47fd67d6a · outbound

This paper cites Bebert: Efficient and robust binary ensemble bert,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bebert: Efficient and robust binary ensemble bert,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.847626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.230172Z digest=sha256:d2929721ed02f7eaf451a52ae0b5a9de93cd9652e20580508f3e7c88fe3eb11a

Observation 11d12725-bcca-489b-a2a9-cfa99513dfc3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.235170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.235170Z digest=sha256:cd22e7f537b95c80831a9445c70792eda247fa29592d097c00a4061778e84f6f

Observation 5c8e640e-b043-4082-8752-c673a4b40678 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.241211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.241211Z digest=sha256:a1973dbafb42ae8c7cccaec61d45873db78bfef5be9c3029bf712b945036bc7a

Observation 46eab912-d74a-47a8-b1b4-90ff15b5743c · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.246476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.246476Z digest=sha256:6806dbd0834e914bb5f0d0fa38a5c7b33efed74a9e88e0c018bc7f1efe9a06d6

Observation 8714c3f3-58f9-4e62-b876-960241b8ce18 · outbound

This paper cites Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.830986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.252804Z digest=sha256:660c4e93645054508fe30443a6377f90b9ca8378a631c901ae75153c8ba3eac2

Observation 7117e91a-1dc4-4735-8b60-08834df4a5c0 · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Haq: Hardware-aware automated quantization with mixed precision,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.813986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.257602Z digest=sha256:6abdcc3ed06695f90ad631182b169e27180506e8b7805ad47c0df0adeac01547

Observation 7357fcfa-01dc-47ce-984c-24f3c5b8c6d7 · outbound

This paper cites Outlier suppression: Pushing the limit of low-bit transformer language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Outlier suppression: Pushing the limit of low-bit transformer language models,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.796440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.263418Z digest=sha256:91caef0b733a958fede27b9d96286373f59038a4a745a3cb2536f925526b1182

Observation e45a072e-fb67-41d6-b90b-9c748cf3f417 · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.780173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.268642Z digest=sha256:f41bc02170795a209479e681d3a09e60983ca4f2a8d16e4d27424e914cc9b87b

Observation a71eaf87-6323-4f73-9d05-19aca2ce0d6b · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.762681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.274153Z digest=sha256:3bb75e1b2c66850d48084a43b792363202e507832d589aa426a5e373fd7e4f02

Observation bcbc3551-c787-4ab8-a8e0-5a7133c1a315 · outbound

This paper cites Efficient streaming language models with attention sinks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient streaming language models with attention sinks,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.745057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.279979Z digest=sha256:f602da19ce6fb1b91572ac716ef157c37df28109419efc4bb172400455cb54d7

Observation 40732215-a09e-41ab-9da8-fce2c45ab70d · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OneBit: Towards Extremely Low-bit Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.284612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.284612Z digest=sha256:198f01a4ac53494ee7fb99e54eac85ed44a748237a5aabd7710b821b9f93861d

Observation 486b959b-41fa-45b4-854b-7615f86f023f · outbound

This paper cites Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.726050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.289754Z digest=sha256:ee327e4e8375a8fbc1ccf47736854af0d135f050de762f81dad7d8d35ee02668

Observation ff4d496e-4129-4775-9c74-ebbb851dd57d · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.294427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.294427Z digest=sha256:11eee84523ce382629d6ffe462bd694e70a0f4abf676bc4302050732eb62c4e0

Observation 85aeb8e5-3c60-483a-87ce-7b2378c2c551 · outbound

This paper cites Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.705368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.299569Z digest=sha256:98f61bbb487a5129abc5f0f278d7e2028d99547f7d14ad3a53a39ae9c6122bd1

Observation aca341d1-8958-4999-ade4-1e25ceec5f95 · outbound

This paper cites Fast: Dnn training under variable precision block floating point with stochastic rounding,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fast: Dnn training under variable precision block floating point with stochastic rounding,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.682402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.303655Z digest=sha256:6b6c791f12fa1aebfa0d747b48fd2a9613aac49919cdf53276dedb45b7ffc538

Observation 50f869ac-ab0d-42bf-85f6-51a7c069feca · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OPT: Open Pre-trained Transformer Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.308601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.308601Z digest=sha256:1255b67a2907a3e1839f08337146acd65e4e8dab3688f781064dbda10140466b

Observation 60706211-3401-4a13-b602-09912f9bbf66 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cam: Cache merging for memory-efficient llms inference,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.659492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.313207Z digest=sha256:6767e9fde40108a3548f07b3769ed1654ea3fbb82d5777f6d937ab022593a7cb

Observation a419e5d6-a114-4446-9c43-703baa852076 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.639814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.317190Z digest=sha256:e77ee620298fdc40fa0183df1229a7a16855e3dd87e90cbbd006b25368a52b0d

Observation 5e21a48c-cfbe-473b-a856-68a0bb2292e9 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.621468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:46:01.321581Z digest=sha256:6b12db78a71b2b680acbbf0671f566b74157db70d3b891640621c0f6c93fe005

Pith citing papers

No inbound Pith citation observations are available.