Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:46:01.321581Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15982.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:46:01.321581Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eb908304-59b4-4082-8b4e-8ca13af78c73 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Resq: Residual quantization for video perception,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b6e85e-10eb-4742-ad69-eb44c17f2a83 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit-pragmatic deep neural network computing,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109156cc-ec88-4c3e-b740-788a6f30fc11 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Explaining neural scaling laws,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f22e992c-6df8-4344-b35e-d24c47398f35 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Longbench: A bilingual, multitask benchmark for long context understanding,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2fbf310-6db7-48f4-b945-3546b592e458 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Demystifying chatgpt: An in-depth survey of openai’s robust large language models,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512cd457-8d26-4023-b16d-7a04113f5a87 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Genus synthesis solution,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c53dec2-f8da-4aba-a7c7-4e8f421615c9 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format General purpose deep learning accelerator based on bit interleaving,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f864ed24-45c9-401f-ae54-5c8143377e09 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quip: 2-bit quanti- zation of large language models with guarantees,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8e462a-25cb-46fb-b2b5-1b5a23573c3e · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08c524f-0c14-4d61-b5d6-f0409de0d6e1 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Nacl: A general and effective kv cache eviction framework for llm at inference time,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e140eb0-9c01-4fba-b3db-04b06150bd15 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Palm: Scaling language modeling with pathways,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e391695-d4ee-48c0-a897-a30f0f232543 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cbfde174-f710-43b3-bc1e-c0588e604377 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 987c2874-d2d8-4184-a3d2-f795bbdf96e0 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format With shared microexponents, a little shifting goes a long way,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0088e03-3677-4f8a-8050-e4d1ccb8efc8 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A timing-driven approach to synthesize fast barrel shifters,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ec788826-bdd6-4922-b1b5-c725344a1534 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llm.int8(): 8- bit matrix multiplication for transformers at scale,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8480b9a4-4201-4813-a0d8-68b85334e80b · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The case for 4-bit precision: k- bit inference scaling laws,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 00ecbf6f-8afa-4325-9fb0-89afa02558c9 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Hawq: Hessian aware quantization of neural networks with mixed-precision,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e6510551-5932-40cc-b624-1173024cb859 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Training dnns with hybrid block floating point,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dce06533-ae2f-40b8-af4b-eab6c87e6b06 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Skvq: Sliding-window key and value cache quantization for large language models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 178aaf15-301b-40f6-8b13-bc2fbe940d64 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Extreme compression of large language models via additive quantization,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f389ac12-1c55-47f5-80ad-99753d11967f · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01002d37-24f7-4610-8d4a-c2d9c229c290 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Static block floating-point quantization for convolutional neural networks on fpga,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 553aa229-a25a-4769-a67b-c312da294cc8 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Optq: Accurate quantization for generative pre-trained transformers,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be801d61-b0fc-4bb3-a27d-e66b430b3b5c · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fcd9706-878e-4c1c-a67b-2c44d6dc4687 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Boost: block minifloat-based on-device cnn training accelerator with transfer learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2efccf6d-a519-4127-b9b4-af72090c2eac · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 734655d4-95bf-4ea4-b953-1923a5d8a9fe · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ese: Efficient speech recognition engine with sparse lstm on fpga,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e31e43a2-f16c-464b-a096-57ffd31fb592 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb4a88b-71fd-425a-a4d4-7557df0a3b61 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c4a177a-2705-4842-99c8-b7d2c8a2e8a7 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9feedc6f-68ee-45ee-b614-0cd0da64d635 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b971c7e3-66b4-44ab-9994-26903bdfdec2 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Perplexity—a measure of the difficulty of speech recognition tasks,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cca7ac48-4bed-454b-91a7-e725e24eecce · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5a868e5d-84ae-43b1-a764-91759d50d4b7 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 28571475-c8a3-430d-97f8-fb952857a7fa · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ten lessons from three generations shaped google’s tpuv4i: Industrial product,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe860abe-bc83-47e5-b316-3a520278d183 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Stripes: Bit-serial deep neural network computing,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 94942890-0859-4f55-b5a3-4a08a3024924 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A survey of gpt-3 family large language models including chatgpt and gpt-4,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02267acb-faf9-41a0-8f47-18c580f1c56c · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e01546cb-3e38-4a57-a7db-f2803705ecdf · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Compressed context mem- ory for online language model interaction,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba442fe0-bdc1-4a37-be6d-7f271a5796e7 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dacapo: Accelerating continuous learning in autonomous systems for video analytics,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 048c2a80-8910-4127-b228-ffba59500712 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 05f82623-88e2-48b1-9229-de66136b466e · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format One-shot model for mixed-precision quantization,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6de4f875-f11c-4ed3-9f6a-a87b72e63cee · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexpoint: An adaptive numerical format for efficient training of deep neural networks,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ff985f6-a23d-486b-8f4d-8d73abf68337 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Tender: Accelerating large language models via tensor decomposition and runtime requantization,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9c0e6188-8cdf-4159-85a6-5069701f9205 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 19bd6321-1c54-4f04-8dc9-116e8de5dff2 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Norm tweaking: High-performance low-bit quantization of large language models,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e3b6e19c-f006-4094-a4f2-de465d4cb624 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Geo: Generation and execution optimized stochastic computing accelerator for neural networks,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fd869e77-bcc5-47ed-ae71-1f27620566bb · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 995b9f46-e3c0-4229-aeed-91e1b6697a4a · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format High-performance fpga-based cnn accelerator with block-floating-point arithmetic,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ac37e816-1469-47bf-b6f8-e6b369beb443 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Awq: Activation-aware weight quan- tization for llm compression and acceleration,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0bd651d8-2194-4958-bb11-36374a362fd2 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96581e0b-fb9d-4487-9eeb-24b973f4971b · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938449d2-5a6d-4ae6-93af-c87b1417f81a · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kivi: A tuning-free asymmetric 2bit quantization for kv cache,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fac23e3-43b4-470b-9fe0-a571e64a12c3 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f43329e-2dac-445d-80dc-259f8ef78496 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fda3f6a9-e9e4-4d15-b6ee-fed49585f9ee · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 92cb1726-fe3b-4af0-9975-a35027aa97d6 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab2800b6-3577-4014-99cc-55e9691b7a41 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The penn treebank: Anno- tating predicate argument structure,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 53655596-273c-437f-8332-4a6e8b23254c · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pointer sentinel mix- ture models,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c70b0fc3-7afe-4fca-9688-f206d80ec28d · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 42f71867-f06c-4e26-8f0f-df6241d81ce7 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cutlass,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aaa2dfc7-d314-412d-b7be-4e114a599e64 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format GPT-4 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f48fab-2824-488e-99e6-22a062854016 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 842a34fe-7a8a-445c-af99-adb9fc1a11ab · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a2e29963-82d3-4df9-bd66-09724eabba27 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Omniquant: Omnidirectionally calibrated quantization for large language models,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c831a98-4474-4151-be7d-780a15bb36ad · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6fc9dfc1-34c9-4723-bbb3-a58c01237086 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 29c5eea0-0148-490a-bccc-b7a1fa42f4ca · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f38003d-2c2e-46f5-a22d-73d7d17f5816 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Gemma: Open Models Based on Gemini Research and Technology
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 141d78a5-df15-482d-b957-e9c47fd67d6a · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bebert: Efficient and robust binary ensemble bert,
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 11d12725-bcca-489b-a2a9-cfa99513dfc3 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLaMA: Open and Efficient Foundation Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8e640e-b043-4082-8752-c673a4b40678 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46eab912-d74a-47a8-b1b4-90ff15b5743c · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8714c3f3-58f9-4e62-b876-960241b8ce18 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7117e91a-1dc4-4735-8b60-08834df4a5c0 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Haq: Hardware-aware automated quantization with mixed precision,
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7357fcfa-01dc-47ce-984c-24f3c5b8c6d7 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Outlier suppression: Pushing the limit of low-bit transformer language models,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e45a072e-fb67-41d6-b90b-9c748cf3f417 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a71eaf87-6323-4f73-9d05-19aca2ce0d6b · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Smoothquant: Accurate and efficient post-training quantization for large language models,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bcbc3551-c787-4ab8-a8e0-5a7133c1a315 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient streaming language models with attention sinks,
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 40732215-a09e-41ab-9da8-fce2c45ab70d · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OneBit: Towards Extremely Low-bit Large Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486b959b-41fa-45b4-854b-7615f86f023f · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ff4d496e-4129-4775-9c74-ebbb851dd57d · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85aeb8e5-3c60-483a-87ce-7b2378c2c551 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aca341d1-8958-4999-ade4-1e25ceec5f95 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fast: Dnn training under variable precision block floating point with stochastic rounding,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 50f869ac-ab0d-42bf-85f6-51a7c069feca · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OPT: Open Pre-trained Transformer Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60706211-3401-4a13-b602-09912f9bbf66 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cam: Cache merging for memory-efficient llms inference,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a419e5d6-a114-4446-9c43-703baa852076 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5e21a48c-cfbe-473b-a856-68a0bb2292e9 · outbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Atom: Low-bit quantization for efficient and accurate llm serving,
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.