Pith. sign in

Paper Citation Record · LEDGER

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration

As of 7 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.19087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19087 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:03:38.453321Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c0249ee-648e-420a-b3af-852c7066a201 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:29.940257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:29.940257Z digest=sha256:d2f439e99144b8aa625ec20a45326fc9146b863c438d3d17f859ce1632748101

Observation 0699aa5c-7ee9-4474-839f-22467db1afb1 · outbound

This paper cites The Llama 3 Herd of Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.103343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.103343Z digest=sha256:25f51ba02efd88180eb9df2d95712be41d33bb9d9649dff0704b9ed6cb93895c

Observation f1046ed7-e9ae-4b3d-8fc0-ad7eac3940fa · outbound

This paper cites DeepSeek-V3 Technical Report.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.247494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.247494Z digest=sha256:7391553323dfe8e7a4d29ff24e250ca442ef6be1381b7a757333a99a53222195

Observation 20b2debd-4947-433b-b75d-e7d68e19e71d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.430398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.430398Z digest=sha256:a1f60fa069b27c3115b427a471fddf2ff3d1cb9178b092238e17fd9af0cb0daa

Observation de9a7b54-2bdf-4117-ad18-fa77863d036e · outbound

This paper cites Jumping nlp curves: A review of natural language processing research,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Jumping nlp curves: A review of natural language processing research,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:52.699089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:30.620469Z digest=sha256:e29e9258d74e77b311619c0571da3eb4a9c691cafe445c90b6513fcfb957b65b

Observation 21a7223a-8da9-4462-b25f-84ba27760508 · outbound

This paper cites Scaling Laws for Neural Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Scaling Laws for Neural Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.727386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.727386Z digest=sha256:60ea48b0ac5fec7067cb2f733034adf30e32c19b6f3f990adc8cdb80ccea35aa

Observation 3a9c9fab-7b3c-42e3-a932-c722a16c76a3 · outbound

This paper cites Chateda: A large language model powered autonomous agent for eda,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Chateda: A large language model powered autonomous agent for eda,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:52.411263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:30.864703Z digest=sha256:fa8fec1aa5a4d77516728712d93d4be4b408942158a637bb7bae686074fbfa02

Observation 97ec05cb-66d7-42cd-be3c-4883c458397a · outbound

This paper cites Dtatrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient transformer archi- tecture,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dtatrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient transformer archi- tecture,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:52.188830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:30.971935Z digest=sha256:9c5fafd7a9e21048d3e2637bc748b3d28a63398a4f66ac583b7738ba4b89e70f

Observation 285e7a34-e16c-40b1-9d04-e41ec4daf6e7 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Qlora: Efficient finetuning of quantized llms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.958128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.093906Z digest=sha256:2024a0ca70d7c7ef23af5424071fa9969e375c9491ce1b153c6350e9b6cac558

Observation 2a09aa32-f0e9-43b3-892a-b2d5ffa42574 · outbound

This paper cites Optq: Accurate quantization for generative pre-trained transformers,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Optq: Accurate quantization for generative pre-trained transformers,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.665625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.247082Z digest=sha256:277e5fbcd0e413274f1560c401ca93349e113cb9788224f2fc58477dbf1d07c7

Observation 53251243-f0e0-4ebd-ac71-deab2d1b995a · outbound

This paper cites A precision-scalable risc-v dnn processor with on- device learning capability at the extreme edge,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration A precision-scalable risc-v dnn processor with on- device learning capability at the extreme edge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.518375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.362494Z digest=sha256:401ad3bd94579950a6bddc9e08c93c3e2448aac88c6ad99755ffdbb10408094f

Observation e942cdf7-30a4-4ea2-8417-6779523368fe · outbound

This paper cites Token-scaled logit distillation for ternary weight gen- erative language models,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Token-scaled logit distillation for ternary weight gen- erative language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.335615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.478915Z digest=sha256:f62ab0cffcc4514d52a0bf1ab573f5204f638e39717f227369655360650a02d4

Observation 7715ab1b-072b-4683-af15-09bc88d460a3 · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration OneBit: Towards Extremely Low-bit Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:31.642699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:31.642699Z digest=sha256:4f2915e43544e8ee5d27e07146f4bb8c7a614e612ddad498e45787ee46c52945

Observation 34a9228e-f7c6-4f26-9ea9-ffa5b255dbc8 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.062987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.776902Z digest=sha256:a74772133d6aa63880ab26c66efcb12f3348e5a824a52fc27c9f9755410d9a6d

Observation 2dbe31b4-d068-44b5-8d19-b14b892fff01 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Omniquant: Omnidirectionally calibrated quantization for large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:50.801835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.837782Z digest=sha256:f82944d9e9e3d153b8e1220f7a361cf78372468b650364fcd06513274e2d0848

Observation e29de1d8-be67-4cc6-a866-00e42673bd77 · outbound

This paper cites Holes: Boosting large language models efficiency with hardware-friendly lossless encoding,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Holes: Boosting large language models efficiency with hardware-friendly lossless encoding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:50.513827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:31.893683Z digest=sha256:816561e109258dd292f6939625d46f517fda86a1cf449f64fac55713a03eb9a0

Observation 81e59714-90dc-4fd5-b33f-69d73af0889c · outbound

This paper cites Quantization via distillation and contrastive learning,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quantization via distillation and contrastive learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:50.218760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.017450Z digest=sha256:8c1f8a10a8f4efdd8b1bec6bbbc2bb7aa25ad956ee045dc33af30729d754775a

Observation a347db84-0d13-4c1c-bfb1-733796f494e4 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.940044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.133751Z digest=sha256:2fca0f4ec415bd4c621c54df7256d793e9ea356cb199319cf844817f2c12e6c5

Observation 78e738c1-ebaa-4270-bd8e-f0f32a7b3542 · outbound

This paper cites Nvidia a100 tensor core gpu: Performance and innovation,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Nvidia a100 tensor core gpu: Performance and innovation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.644008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.231960Z digest=sha256:24e8f2f28c4e11780cac2f5a5132678cac6ed1cc5e2411dc84118e5cba373e06

Observation 8b2eb790-87f0-4345-95e9-5dad8cc06a8c · outbound

This paper cites Rtx on—the nvidia turing gpu,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Rtx on—the nvidia turing gpu,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.365835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.311442Z digest=sha256:73ed5a20e835ef400f44704ec9b88912965f0a0ff305d0678586cf52a33657e4

Observation 61402a7a-7864-4591-806f-256035d61ed8 · outbound

This paper cites Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.070155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.488863Z digest=sha256:a7942750cf8964bae5fdee96de76d8942e91f871fbe11989f66e67cf552d85af

Observation f6e2c64e-11de-4867-be05-59d2e3d3b060 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:32.624554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:32.624554Z digest=sha256:759eb51fb535d48eb1c6294606940a211029142607fedf3512d9aa3f5893c46a

Observation 02687371-529f-495d-bdff-b1680de5b07c · outbound

This paper cites G-blastn: accelerating nucleotide alignment by graphics processors,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration G-blastn: accelerating nucleotide alignment by graphics processors,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:48.703589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.713355Z digest=sha256:1141df41f12244b6a20351103f01aec02bbb6760c49972aa57e3e8ee6eb8f8ca

Observation 972eb02e-45bf-4945-8ff3-c7891e9db510 · outbound

This paper cites Accelerating performance of gpu-based workloads using cxl,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accelerating performance of gpu-based workloads using cxl,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:48.438897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.816745Z digest=sha256:e3a0d5d49c14ab2374e322ce462bd58ee014b374bf059715f206fadc217bec8f

Observation 0d767571-b9e6-4997-b49f-ecf36bb09708 · outbound

This paper cites Superneurons: Dynamic gpu memory management for training deep neural networks,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Superneurons: Dynamic gpu memory management for training deep neural networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:48.156464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:32.903184Z digest=sha256:71f3bce5f9d7fb65dd56bac403d4f18b322ffcab0faa8ed31e3fd433f14b479f

Observation 862a8792-8b7c-43a0-9836-96f2dbbd69e9 · outbound

This paper cites Gpt3. int8 (): 8-bit matrix multiplication for trans- formers at scale,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Gpt3. int8 (): 8-bit matrix multiplication for trans- formers at scale,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:47.803596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.026595Z digest=sha256:237d9534daad4e11ee864995ac37d40782488878a999b04592ac2bebd001f25c

Observation a7457989-dce0-4e7f-9a58-7ff4d12c6b58 · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:47.454561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.175416Z digest=sha256:0943017a5e2dcb23a653a48c349e82019053acfabd02cfe3a91a37f9ea34a978

Observation caa6ff66-940d-461e-992b-d09fa11ae585 · outbound

This paper cites Tsm2x: High-performance tall-and-skinny matrix– matrix multiplication on gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Tsm2x: High-performance tall-and-skinny matrix– matrix multiplication on gpus,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:47.084547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.298313Z digest=sha256:9daa2ddcf5f6b22638f6447cb8f0bca5ecad3ffe633e6a5ff531261b08fe7598

Observation 73a8173f-ae2d-4e27-a9a0-f835c17507ed · outbound

This paper cites Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the gpu,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the gpu,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:46.669641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.396338Z digest=sha256:6492972d28e8f03d562e8e0ea59f3efb4b48be9903c2f6d57450ae18e5cece2b

Observation efedc03e-0a91-4e01-9a05-fab1cfa26ab3 · outbound

This paper cites Warp-aware adaptive energy efficiency calibration for multi-gpu systems,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Warp-aware adaptive energy efficiency calibration for multi-gpu systems,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:46.295684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.509043Z digest=sha256:2b151b8178a1e5304e74a633d0ec2fcc3d56bbcce632754d75fc55575dbe3b28

Observation c55015c0-a37a-4ba3-991e-9e6c536392ab · outbound

This paper cites Dg-replace: A dataflow-driven gpu-accelerated an- alytical global placement framework for machine learning accelerators,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dg-replace: A dataflow-driven gpu-accelerated an- alytical global placement framework for machine learning accelerators,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:45.883712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.711616Z digest=sha256:9d0bac7700b3df6c011899c10e991c14993c9416275739c9e2d302de9b8d3ddf

Observation 6a408321-96c9-44fb-8927-41f1b0f5f099 · outbound

This paper cites Enabling efficient sparse multiplications on gpus with heuristic adaptability,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Enabling efficient sparse multiplications on gpus with heuristic adaptability,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:45.519740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.821701Z digest=sha256:57829fd5286ca26311f5a1895c71645af95b16ed1418b585b70b5159d1b854c2

Observation f204c8cb-b910-4361-9f64-5ac40004a4c5 · outbound

This paper cites Bstc: A novel binarized-soft-tensor-core design for acceler- ating bit-based approximated neural nets,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bstc: A novel binarized-soft-tensor-core design for acceler- ating bit-based approximated neural nets,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:45.169637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:33.927448Z digest=sha256:78a7a725032472c06e9ebc87f56ad0c75973d6a5ae19a71d340a21a7de54512b

Observation 1f0e1a3d-b951-4eec-9265-5029184e20fe · outbound

This paper cites Accelerating binarized neural networks via bit-tensor-cores in turing gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accelerating binarized neural networks via bit-tensor-cores in turing gpus,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.917241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:34.072475Z digest=sha256:ec76e241ab078f3fbd75dbe980915c1d50e9e9c6100245b0b3015840ca26da77

Observation a7657730-372d-442d-9efe-85a35fbaa7d5 · outbound

This paper cites Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.637710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:34.214076Z digest=sha256:3039ab250bf5678deaf60fc40ad829652ff5140c9e390b390850645cba2c6cf1

Observation ddd27145-4aa4-4b4e-bd39-8a1262afe34b · outbound

This paper cites Dissecting the NVidia Turing T4 GPU via Microbenchmarking.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dissecting the NVidia Turing T4 GPU via Microbenchmarking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:34.362212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:34.362212Z digest=sha256:94cf8faa074a5f5de7d3eccb053922aae7c64f4fdf48f46c6606bda19df2405b

Observation c1cfc1ed-d322-48d7-9a5a-b4352db32e67 · outbound

This paper cites Gtco: Graph and tensor co-design for transformer-based image recognition on tensor cores,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Gtco: Graph and tensor co-design for transformer-based image recognition on tensor cores,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.121829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:34.545982Z digest=sha256:d62511b2bdf545c40ba676ecc81d14e197f1fd06e0346fc3a60dd58d42229b45

Observation 10a85451-0f7f-44a6-b720-0855114b9c2a · outbound

This paper cites Reducing shared memory footprint to leverage high throughput on tensor cores and its flexible api extension library,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Reducing shared memory footprint to leverage high throughput on tensor cores and its flexible api extension library,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:43.817709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:34.626516Z digest=sha256:88ae9802e6440afeb7e0ad77f0a25eb182bacbcd6cee9debb3b920e73027b34d

Observation d3ba3b1d-64d2-4aac-80b0-1c17207ef002 · outbound

This paper cites Tc-gnn: Bridging sparse gnn computation and dense tensor cores on gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Tc-gnn: Bridging sparse gnn computation and dense tensor cores on gpus,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:43.542363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:34.781504Z digest=sha256:979f427d0505f73c9779c4fd475ee0ea6d331cb8ae5ea479d508ad067917bf87

Observation bc41d50c-520f-4568-b483-fee0d23a90d6 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration A Survey on Efficient Inference for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:34.877201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:34.877201Z digest=sha256:634a939b1d2c4ae71a747c9167a6cc6c14390f53b3d64a54bbd341ab88295e2b

Observation 936e10d2-b8f9-432d-b9d3-406727c0d6b3 · outbound

This paper cites Transformer tricks: Precomputing the first layer.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Transformer tricks: Precomputing the first layer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.015302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.015302Z digest=sha256:51b1b16e71bef9517e2780f358ca770c7aba45a6a04e23cc0685822d54ef79b2

Observation 20028b8c-8e9b-4be8-9aad-7f086a53eca0 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.155547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.155547Z digest=sha256:d57dc594985257cbc097032dbff004c18a57273d602300464095db1145d122cd

Observation 0384a01f-6700-410c-b846-6077e1f05f29 · outbound

This paper cites Efficiently scaling transformer inference,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Efficiently scaling transformer inference,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:43.271413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:35.276363Z digest=sha256:028a228349fb5693b4cfaa7eabeedcb223ffc1eb89ac714bd45be1f1d4d23a39

Observation 50dbda44-73ad-4c8c-b464-a047058de65d · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.382312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.382312Z digest=sha256:fc80f043aa3b5cb3dcb08237553f20404b7712f62175d28f8e70974cf5bd42b2

Observation 23f16808-6923-48eb-9bd2-dbfd81312a1f · outbound

This paper cites Squeezellm: Dense-and-sparse quantization,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Squeezellm: Dense-and-sparse quantization,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.979238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:35.525034Z digest=sha256:f1b04862a88058894375076d974974edf8a18b9c9ce3890a8d688c88f75f3586

Observation 602f8122-5394-423d-9363-362753a53a0f · outbound

This paper cites Quantsr: accurate low-bit quantization for efficient image super-resolution,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quantsr: accurate low-bit quantization for efficient image super-resolution,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.735145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:35.654288Z digest=sha256:2ebf81fccff58dd1c04b45bbc31082ae5fcac4f880059df57444d3dbb86b8271

Observation 8c9c4e4d-e22e-4f6a-a269-72f65297aff9 · outbound

This paper cites Accurate lora-finetuning quantization of llms via infor- mation retention,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accurate lora-finetuning quantization of llms via infor- mation retention,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.456579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:35.740233Z digest=sha256:db8c088f7dee31dd03d242b0c1223f2fc8343f618c2e1ddad3b5325d2ea79c38

Observation aa6adcfe-69e2-4d00-bdd2-aa522c72ac1a · outbound

This paper cites Bimatting: Efficient video matting via binarization,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bimatting: Efficient video matting via binarization,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.173545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:35.795656Z digest=sha256:8f896959d7ad1e63cb5b39dcf77d86150f525daa8af7d13ef66c2f7a2968526b

Observation dc0e5c26-be37-45c6-a541-59cffa191625 · outbound

This paper cites Bibert: Accurate fully binarized bert,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bibert: Accurate fully binarized bert,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.854183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:35.926778Z digest=sha256:62ab2bf96d148f460b4d4a1a718fe4b2401b22509854230a50e36754f41ca0b4

Observation 6b77b327-62a3-4095-bb3d-bf53d3e02ad0 · outbound

This paper cites Bebert: Efficient and robust binary ensemble bert,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bebert: Efficient and robust binary ensemble bert,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.556430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.063989Z digest=sha256:3a82edaefaa0aeaa1e120d653abf12c95a8e0a92f7d73e173337a43f76288688

Observation febe2237-c4f7-4ba0-bdb7-eb4f062c9e7f · outbound

This paper cites Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.285241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.192211Z digest=sha256:72c300d527e440d8222c583398a6f1ece4f3d9c5a08b3c888a82c2d859b4fe16

Observation e934cee2-b33f-4156-9766-e77fc31f79e5 · outbound

This paper cites Binaryconnect: Training deep neural networks with binary weights during propagations,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Binaryconnect: Training deep neural networks with binary weights during propagations,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.038528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.289139Z digest=sha256:84e43391cd6a72284095148dff85d847bc3141bbbaa44c1845efd3dc514eabf8

Observation 32badd89-5357-4eb1-bf96-f3abad82b614 · outbound

This paper cites Apnn-tc: Accelerating arbitrary precision neural net- works on ampere gpu tensor cores,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Apnn-tc: Accelerating arbitrary precision neural net- works on ampere gpu tensor cores,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.746968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.427395Z digest=sha256:c8726e27d02cff192f87690ff087e94c737470201b393dc714d70f124cb9597b

Observation 85c7506c-6611-476b-bd56-2eafefebde65 · outbound

This paper cites O3bnn: An out-of-order architecture for high- performance binarized neural network inference with fine-grained prun- ing,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration O3bnn: An out-of-order architecture for high- performance binarized neural network inference with fine-grained prun- ing,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.550324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.587831Z digest=sha256:836c2575979ca0b491c21da1de2aa3eda596408b562d2e4053f98f7a61f420f1

Observation 2e159bb0-57f2-4c08-8b4a-a89501d8ca9f · outbound

This paper cites O3bnn-r: An out-of-order architecture for high- performance and regularized bnn inference,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration O3bnn-r: An out-of-order architecture for high- performance and regularized bnn inference,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.290730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.716924Z digest=sha256:3f15d637a48795cb47e935b47f774ed000130a91316a4907009723e888869800

Observation a82f4c78-4e48-4753-9e7d-fca5a7664637 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:36.820892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:36.820892Z digest=sha256:957981b63c1cf7115f950bc7a916dd87922ae5bd7f421806eb53026195287bfb

Observation d6f7676a-04bf-4c3b-9020-9fc997bb10ac · outbound

This paper cites Energy-efficient neural network accelerator based on outlier-aware low-precision computation,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Energy-efficient neural network accelerator based on outlier-aware low-precision computation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.054971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:36.973470Z digest=sha256:d82e9a2f230e5c304ca1219965d8bcb5d6d0f5bdf51dbb27b2e331bc400a336d

Observation beb297e3-6573-4407-8cf0-a383adc8e25b · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Haq: Hardware-aware automated quantization with mixed precision,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.822948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:37.059456Z digest=sha256:d1cb535a3bb431db51776c461d1add6327a4e2bb69c093ac9542117989e42807

Observation 2f35f174-06ed-4ec5-90b4-044f816a4ede · outbound

This paper cites Lq-nets: Learned quantization for highly accurate and compact deep neural networks,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Lq-nets: Learned quantization for highly accurate and compact deep neural networks,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.574486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:37.163159Z digest=sha256:65fe28ffb520c26a2b06698297c77a8cd14a6d9ce0a9306bd6816260fc4f11f1

Observation ed54c9ca-a242-4b44-ad33-9e83c77fd0f8 · outbound

This paper cites DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:37.317001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:37.317001Z digest=sha256:9c06b503abd2f6b397b4784b9adf5a0620bbeda0e30eb970594ef66072731102

Observation 64b701fa-203f-4860-86af-cb433ba3e196 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:37.480444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:37.480444Z digest=sha256:f2abedf386d9e4aae1cc09fb7bf36262b2b8ee479ccb36e0c65b86ea36a98777

Observation c4a7e274-a1e0-4630-8f06-594cc0e859fa · outbound

This paper cites CUTLASS,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration CUTLASS,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.296976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:37.615424Z digest=sha256:aa9ed4df854c014d258ade223630fc7df9aa00b9eaada57412e3ed8bae671473

Observation 8b8e1fab-3214-4f98-ab05-99f3338c3ecd · outbound

This paper cites Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:03:38.709223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:37.761446Z digest=sha256:500f840b15532291760a72658074326695c703951ba55acc325d141ba524f2c1

Observation 647c40c9-853c-4d34-83ae-be0dbcf3c981 · outbound

This paper cites Benchmarking and dissecting the nvidia hopper gpu architecture,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Benchmarking and dissecting the nvidia hopper gpu architecture,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.427189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:37.863736Z digest=sha256:45ad675dadd31ee509f9e0ce58ba3be53e38227e9329c2ec315a2812bbb019c0

Observation 218ea8a5-470c-47c8-9151-ce529186440c · outbound

This paper cites Qwen2.5 Technical Report.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Qwen2.5 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:37.961793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:37.961793Z digest=sha256:8c5c6fb69c82617b7228a484baeec005d0a4af6acae1c02c7710f8b13274632c

Observation 23525f4e-8c54-465c-9ef4-0e17cad9dccd · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration OPT: Open Pre-trained Transformer Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:38.057130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:38.057130Z digest=sha256:a77f69b06369c1f5f84f18bcbf8ed0976323377ac1e2028e074d3a46cac3854c

Observation 372fa1e2-d486-4352-ad8f-e8c613a2b8e1 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:38.218549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:38.218549Z digest=sha256:b34bca41be7afa23002976f48daf07a18f75c8aa1d8a25a4d9e5f314605c3c3c

Observation 831a1e11-38ba-443b-8d03-a72fe0122d8c · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quarot: Outlier-free 4-bit inference in rotated llms,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.050426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:03:38.331692Z digest=sha256:899f8a7879feaecf9573ee64a29451db9c8bf33b3333517f1c821a396843b9f4

Observation 90c16b91-f1f3-4421-860f-f5104b84fcdd · outbound

This paper cites Pointer Sentinel Mixture Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Pointer Sentinel Mixture Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:38.453321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:38.453321Z digest=sha256:055a356e320a28fefe543fcc3c1e64193a784d0e840176699a5595cc85c84fb7

Pith citing papers

No inbound Pith citation observations are available.