Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:03:38.453321Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.19087.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:03:38.453321Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
69 of 69 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2c0249ee-648e-420a-b3af-852c7066a201 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0699aa5c-7ee9-4474-839f-22467db1afb1 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration The Llama 3 Herd of Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1046ed7-e9ae-4b3d-8fc0-ad7eac3940fa · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b2debd-4947-433b-b75d-e7d68e19e71d · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9a7b54-2bdf-4117-ad18-fa77863d036e · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Jumping nlp curves: A review of natural language processing research,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21a7223a-8da9-4462-b25f-84ba27760508 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Scaling Laws for Neural Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9c9fab-7b3c-42e3-a932-c722a16c76a3 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Chateda: A large language model powered autonomous agent for eda,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97ec05cb-66d7-42cd-be3c-4883c458397a · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dtatrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient transformer archi- tecture,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 285e7a34-e16c-40b1-9d04-e41ec4daf6e7 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Qlora: Efficient finetuning of quantized llms,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a09aa32-f0e9-43b3-892a-b2d5ffa42574 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Optq: Accurate quantization for generative pre-trained transformers,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53251243-f0e0-4ebd-ac71-deab2d1b995a · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration A precision-scalable risc-v dnn processor with on- device learning capability at the extreme edge,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e942cdf7-30a4-4ea2-8417-6779523368fe · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Token-scaled logit distillation for ternary weight gen- erative language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7715ab1b-072b-4683-af15-09bc88d460a3 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration OneBit: Towards Extremely Low-bit Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a9228e-f7c6-4f26-9ea9-ffa5b255dbc8 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Smoothquant: Accurate and efficient post-training quantization for large language models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dbe31b4-d068-44b5-8d19-b14b892fff01 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Omniquant: Omnidirectionally calibrated quantization for large language models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e29de1d8-be67-4cc6-a866-00e42673bd77 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Holes: Boosting large language models efficiency with hardware-friendly lossless encoding,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81e59714-90dc-4fd5-b33f-69d73af0889c · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quantization via distillation and contrastive learning,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a347db84-0d13-4c1c-bfb1-733796f494e4 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Atom: Low-bit quantization for efficient and accurate llm serving,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78e738c1-ebaa-4270-bd8e-f0f32a7b3542 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Nvidia a100 tensor core gpu: Performance and innovation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b2eb790-87f0-4345-95e9-5dad8cc06a8c · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Rtx on—the nvidia turing gpu,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61402a7a-7864-4591-806f-256035d61ed8 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6e2c64e-11de-4867-be05-59d2e3d3b060 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02687371-529f-495d-bdff-b1680de5b07c · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration G-blastn: accelerating nucleotide alignment by graphics processors,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 972eb02e-45bf-4945-8ff3-c7891e9db510 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accelerating performance of gpu-based workloads using cxl,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d767571-b9e6-4997-b49f-ecf36bb09708 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Superneurons: Dynamic gpu memory management for training deep neural networks,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 862a8792-8b7c-43a0-9836-96f2dbbd69e9 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Gpt3. int8 (): 8-bit matrix multiplication for trans- formers at scale,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7457989-dce0-4e7f-9a58-7ff4d12c6b58 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation caa6ff66-940d-461e-992b-d09fa11ae585 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Tsm2x: High-performance tall-and-skinny matrix– matrix multiplication on gpus,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73a8173f-ae2d-4e27-a9a0-f835c17507ed · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the gpu,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efedc03e-0a91-4e01-9a05-fab1cfa26ab3 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Warp-aware adaptive energy efficiency calibration for multi-gpu systems,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c55015c0-a37a-4ba3-991e-9e6c536392ab · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dg-replace: A dataflow-driven gpu-accelerated an- alytical global placement framework for machine learning accelerators,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a408321-96c9-44fb-8927-41f1b0f5f099 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Enabling efficient sparse multiplications on gpus with heuristic adaptability,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f204c8cb-b910-4361-9f64-5ac40004a4c5 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bstc: A novel binarized-soft-tensor-core design for acceler- ating bit-based approximated neural nets,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f0e1a3d-b951-4eec-9265-5029184e20fe · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accelerating binarized neural networks via bit-tensor-cores in turing gpus,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7657730-372d-442d-9efe-85a35fbaa7d5 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddd27145-4aa4-4b4e-bd39-8a1262afe34b · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dissecting the NVidia Turing T4 GPU via Microbenchmarking
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1cfc1ed-d322-48d7-9a5a-b4352db32e67 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Gtco: Graph and tensor co-design for transformer-based image recognition on tensor cores,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10a85451-0f7f-44a6-b720-0855114b9c2a · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Reducing shared memory footprint to leverage high throughput on tensor cores and its flexible api extension library,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3ba3b1d-64d2-4aac-80b0-1c17207ef002 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Tc-gnn: Bridging sparse gnn computation and dense tensor cores on gpus,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc41d50c-520f-4568-b483-fee0d23a90d6 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration A Survey on Efficient Inference for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936e10d2-b8f9-432d-b9d3-406727c0d6b3 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Transformer tricks: Precomputing the first layer
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20028b8c-8e9b-4be8-9aad-7f086a53eca0 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0384a01f-6700-410c-b846-6077e1f05f29 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Efficiently scaling transformer inference,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50dbda44-73ad-4c8c-b464-a047058de65d · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f16808-6923-48eb-9bd2-dbfd81312a1f · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Squeezellm: Dense-and-sparse quantization,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 602f8122-5394-423d-9363-362753a53a0f · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quantsr: accurate low-bit quantization for efficient image super-resolution,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c9c4e4d-e22e-4f6a-a269-72f65297aff9 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accurate lora-finetuning quantization of llms via infor- mation retention,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa6adcfe-69e2-4d00-bdd2-aa522c72ac1a · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bimatting: Efficient video matting via binarization,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc0e5c26-be37-45c6-a541-59cffa191625 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bibert: Accurate fully binarized bert,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b77b327-62a3-4095-bb3d-bf53d3e02ad0 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bebert: Efficient and robust binary ensemble bert,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation febe2237-c4f7-4ba0-bdb7-eb4f062c9e7f · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e934cee2-b33f-4156-9766-e77fc31f79e5 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Binaryconnect: Training deep neural networks with binary weights during propagations,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32badd89-5357-4eb1-bf96-f3abad82b614 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Apnn-tc: Accelerating arbitrary precision neural net- works on ampere gpu tensor cores,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85c7506c-6611-476b-bd56-2eafefebde65 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration O3bnn: An out-of-order architecture for high- performance binarized neural network inference with fine-grained prun- ing,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e159bb0-57f2-4c08-8b4a-a89501d8ca9f · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration O3bnn-r: An out-of-order architecture for high- performance and regularized bnn inference,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a82f4c78-4e48-4753-9e7d-fca5a7664637 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f7676a-04bf-4c3b-9020-9fc997bb10ac · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Energy-efficient neural network accelerator based on outlier-aware low-precision computation,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation beb297e3-6573-4407-8cf0-a383adc8e25b · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Haq: Hardware-aware automated quantization with mixed precision,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f35f174-06ed-4ec5-90b4-044f816a4ede · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Lq-nets: Learned quantization for highly accurate and compact deep neural networks,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed54c9ca-a242-4b44-ad33-9e83c77fd0f8 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b701fa-203f-4860-86af-cb433ba3e196 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration BitNet: Scaling 1-bit Transformers for Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a7e274-a1e0-4630-8f06-594cc0e859fa · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration CUTLASS,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b8e1fab-3214-4f98-ab05-99f3338c3ecd · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 647c40c9-853c-4d34-83ae-be0dbcf3c981 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Benchmarking and dissecting the nvidia hopper gpu architecture,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 218ea8a5-470c-47c8-9151-ce529186440c · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Qwen2.5 Technical Report
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23525f4e-8c54-465c-9ef4-0e17cad9dccd · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration OPT: Open Pre-trained Transformer Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372fa1e2-d486-4352-ad8f-e8c613a2b8e1 · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831a1e11-38ba-443b-8d03-a72fe0122d8c · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quarot: Outlier-free 4-bit inference in rotated llms,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90c16b91-f1f3-4421-860f-f5104b84fcdd · outbound
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Pointer Sentinel Mixture Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.