Pith. sign in

Paper Citation Record · LEDGER

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration

As of 23 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2507.23035.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23035 v4

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:15:34.047943Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 847f1ac4-5c7d-4ff8-a611-3eab56f97699 · outbound

This paper cites GPT-4 Technical Report.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.805847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.805847Z digest=sha256:c368e505ca2ee19d1b2d365fc329246a83859b272c8b6628bd95ac0bbfb35a03

Observation 1d4a66b1-8c6c-41c6-8bb1-b40af839071f · outbound

This paper cites Introducing nvfp4 for efficient and accurate low-precision inference,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Introducing nvfp4 for efficient and accurate low-precision inference,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.579336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.809417Z digest=sha256:15d0c1327b9b09a417d2ef633bc626545e45288610973f9fd067a3d7557fef42

Observation ffa98e14-2ef6-4e5b-8568-ab043fda4390 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.906211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.906211Z digest=sha256:9c6cbc8d00404e604b48aa3217f5890fa0a56cf626560675d1dc8f4ab7371d37

Observation 26c611be-c7b4-4e90-b0c5-c42224593a74 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Piqa: Reasoning about physical commonsense in natural language,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.909765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.909765Z digest=sha256:451b978dcedf7594e369a13f3a435126a8eb5d3b235e717e55cf67cba07e1ece

Observation 676550be-860f-4102-8a7d-8e7a3e331415 · outbound

This paper cites Language mod- els are few-shot learners,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Language mod- els are few-shot learners,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.913219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.913219Z digest=sha256:a55d63c4d8bf0c3609bbc818974e39fda183e1bd8d2096829fb628a3552273f1

Observation 11615361-3187-4929-bfd6-389b93b435c4 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Boolq: Exploring the surprising difficulty of natural yes/no questions,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.916602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.916602Z digest=sha256:524dd4e99a6bdaff7a46fe9b67496d3f413f5a0a0a7227a19761293a7ab5b2df

Observation 091fade5-25e8-4b5c-b535-1174f1e6af4d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.920077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.920077Z digest=sha256:6120c2f04e15ab5b04f7f68e698aaa117228d5380d77c9687f7e75b746bbe332

Observation 9ba754d0-ccb3-48ad-ac87-d48b1cd9cb79 · outbound

This paper cites Nvidia rtx blackwell gpu architecture: Built for neural rendering,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Nvidia rtx blackwell gpu architecture: Built for neural rendering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.560155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.922395Z digest=sha256:7f8de36eb4ab77b391ccacbb15b9114ad878f99e55d3e98dc0bf68db3ae50292

Observation c19aadb0-2152-4b45-8a97-5b20991a2b7e · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.925142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.925142Z digest=sha256:81c82399199aed0f84c030a3e9d6a76ec177c4d2cf530af4a660ac7f07612e77

Observation 0bd92ba3-0419-4950-820f-d15dca7882e6 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.931222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.931222Z digest=sha256:98a5a43f122a14704d23fdbff02832226a2aaa045f34b283907e1ef2f653cd8a

Observation b66d7ef3-3d63-40c7-a76c-b41da2412312 · outbound

This paper cites Gpt4aigchip: Towards next-generation ai accelerator design automation via large language models,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Gpt4aigchip: Towards next-generation ai accelerator design automation via large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.552522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.934030Z digest=sha256:38cd4b3e18eacdde36e235eacf5eacd831d30b3b3b7695d38a2f16a22938ad8e

Observation 83587800-befc-4eef-973a-acb387a7edb1 · outbound

This paper cites The language model evaluation harness,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration The language model evaluation harness,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.937039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.937039Z digest=sha256:a1b883b271c7fbc7c40bc34cda569fbfb539fad52d0baed222316bc401ba8e36

Observation 01476bca-db5c-4939-ade7-9ff0cef4583e · outbound

This paper cites Disaggregated machine learning via in-physics computing at radio frequency,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Disaggregated machine learning via in-physics computing at radio frequency,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.544343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.939378Z digest=sha256:499ed8bc8c7e989d03c7199f07c29e44c6896123a2d798e51ee6a2a6d7779ffa

Observation 0df24342-d211-44e0-b568-5c04c93ff571 · outbound

This paper cites The Llama 3 Herd of Models.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.941385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.941385Z digest=sha256:0992e51cb1832f9ac7504dbdff5dd1a550029c89509ee10f698140d656723403

Observation fa62e0b2-0f26-4eb9-9c55-b2f71afd2313 · outbound

This paper cites A survey: Collaborative hardware and software design in the era of large language models,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration A survey: Collaborative hardware and software design in the era of large language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.943889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.943889Z digest=sha256:2dfb51f8d1c99aa3a76c498d8a11395e943ea6a96e18e752867acaeed29164c8

Observation 4bd7d299-2c50-4f70-8bda-050d8482a453 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.945972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.945972Z digest=sha256:489dc0f2c920ca6241e27a059bf775e87afb5458716df26c4b6e7403d0903274

Observation 10c1d9ca-9662-4750-87f2-957ad7e750d6 · outbound

This paper cites Chateda: A large language model powered autonomous agent for eda,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Chateda: A large language model powered autonomous agent for eda,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.532245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.948798Z digest=sha256:4fe68c62c131cd27bed954ae7ca8e663badcf4a04ec0cccf5041089cc34bda98

Observation b78e0d83-6f47-466f-8038-3a25652de897 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.951900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.951900Z digest=sha256:f57379c605a8b02e6e6399c7ae1ababcae257a1a35c878dc4fa25025cfe8d91f

Observation f973aac6-c8f8-4b38-b7ec-73c49efcea13 · outbound

This paper cites Mistral 7B.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.954140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.954140Z digest=sha256:16d91b9b6bddab964b845509bcb8c7f827ae412ceef04f111823f4f6f9f55a34

Observation e84da4c0-1e3d-4d94-82c6-e8450d5c288a · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration SqueezeLLM: Dense-and-Sparse Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.956446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.956446Z digest=sha256:3b43fe0f27bfc1e2d44aa90261286c86b61a6940422ece3765a3fae3c5df4a47

Observation 0e91b256-0d82-4bcd-ad0d-353c4c24d757 · outbound

This paper cites Scaling Laws for Precision.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Scaling Laws for Precision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.959039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.959039Z digest=sha256:20659db7b1e19f7b25ef7b6a1a408aa36cc815f5397ebc4fd7d16bfdd2564d18

Observation 474dd206-9f2e-44dc-8e38-d267b0d2e348 · outbound

This paper cites Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable intercon- nects,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable intercon- nects,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.961283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.961283Z digest=sha256:dfd5b5919ceec68f9945d48d9577642c9dfa804e25e7c79e137f8a00a0ada0f2

Observation fdfd1d2e-742e-4a9e-8129-4d58ba90fdf6 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Efficient memory management for large language model serving with pagedattention,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.963596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.963596Z digest=sha256:15400d4e8a4ad614595334193758838aad2306a3acaeb9e47dd6273aff1911bb

Observation 99348218-225c-4b9b-af02-865527d3b435 · outbound

This paper cites Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.965632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.965632Z digest=sha256:1d57d34872b63727dfa93cb5e305d76d0d2e70e3624fe9efaad17c49855130a8

Observation 32f0d38e-909f-4f6b-8443-288630bd664c · outbound

This paper cites Dramsim3: A cycle-accurate, thermal-capable dram simulator,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Dramsim3: A cycle-accurate, thermal-capable dram simulator,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.967821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.967821Z digest=sha256:e58e3ac5a6e817a107a6101ac7d97a77d31cecdbd62c99f29b9439cb826030cc

Observation 7f07b157-8cf1-464b-b7bc-77e3410a2eb4 · outbound

This paper cites Cacti- p: Architecture-level modeling for sram-based structures with advanced leakage reduction techniques,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Cacti- p: Architecture-level modeling for sram-based structures with advanced leakage reduction techniques,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.508718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.970382Z digest=sha256:c90a2734169e45fdfabac2737d136ecfc08bd2c5bc68d8d4fa48ce6a885913ff

Observation 9945d51a-502d-49c7-ab4d-69aaa5b8ce20 · outbound

This paper cites Duquant: Distributing outliers via dual transformation makes stronger quantized llms,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Duquant: Distributing outliers via dual transformation makes stronger quantized llms,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.502129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.972847Z digest=sha256:156cbbfad69c1f89024c90ab68162ae5de7bff30938c70b44317d945528cddfa

Observation 4913921f-23cb-45be-80c3-b6e41a69d8ea · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.975351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.975351Z digest=sha256:d514726f08f6171af6083e85186c50fc51ead223e992cc2e1d5f84281e9e02f3

Observation 6bca994c-09f1-444a-99eb-a2fe6badb1f1 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.977646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.977646Z digest=sha256:86bb0a2938da91321a266166cc6e1f87f5a98b8495877e8c57524ba3ff2f2f22

Observation 890a1171-f933-4662-951e-180cecf31f3a · outbound

This paper cites LLM-FP4: 4-Bit Floating-Point Quantized Transformers.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration LLM-FP4: 4-Bit Floating-Point Quantized Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.980434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.980434Z digest=sha256:4eb6c05251e651897ec2079aaaea7b209bee64ddff0aa635d1fe2af27032dd2d

Observation ae5fbade-769f-4cb2-9da3-c7f6c0eaaced · outbound

This paper cites Micromix: Efficient mixed-precision quantization with microscaling formats for large lan- guage models,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Micromix: Efficient mixed-precision quantization with microscaling formats for large lan- guage models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.982460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.982460Z digest=sha256:bd1eeb01952852108dfb1db17586e65e20040bb6a708ed519729892d762d602c

Observation 38ab1e3d-f98f-4f09-a5cd-2073a5c0d73e · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration SpinQuant: LLM quantization with learned rotations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.984311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.984311Z digest=sha256:de5578b443bb5e88a75f8db91d502ffc9e9be29c6239e1a4d605355a7c5289bb

Observation e9c1bb92-f21a-4fc2-b7ed-f8ded2fd9925 · outbound

This paper cites Some methods for classification and analysis of mul- tivariate observations,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Some methods for classification and analysis of mul- tivariate observations,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.490903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.986591Z digest=sha256:9931a43c3ca51ad87da993b3f6b95af50d54446c4f9fd62b2176b88ec101c271

Observation f2ef1d18-e455-4ce2-8e0b-c4501104cd6c · outbound

This paper cites The penn treebank: Anno- tating predicate argument structure,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration The penn treebank: Anno- tating predicate argument structure,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.484496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.988720Z digest=sha256:a719fa9b74dd7b35b5eeffaf6a9a108a5a03451d6f1ea3b4b8a521d2bbf5b0c6

Observation 53d58b4d-4c3a-4d28-9050-b01fc0c896db · outbound

This paper cites Pointer Sentinel Mixture Models.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Pointer Sentinel Mixture Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.990791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.990791Z digest=sha256:122a672782572bcf2cec645fb447af997c3c713afc7832d88579e12609260a39

Observation ed7a9133-be5c-4554-a0cb-2e41b845ce77 · outbound

This paper cites Lut tensor core: A software-hardware co-design for lut-based low-bit llm inference,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Lut tensor core: A software-hardware co-design for lut-based low-bit llm inference,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.993051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.993051Z digest=sha256:a124e8827504a4a16dd374a727e5b45fdee41d15235bf5abfe85ac61c2fb5040

Observation 8b5b2eea-1a5d-4e71-9f56-f8eef89ca9e1 · outbound

This paper cites Flexagon: A multi-dataflow sparse-sparse matrix multiplication accelerator for efficient dnn processing,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Flexagon: A multi-dataflow sparse-sparse matrix multiplication accelerator for efficient dnn processing,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.474197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.995396Z digest=sha256:90be0243daa0f1077eb41ea793cf9bf6781da279a49a8c845d5237b402113ecb

Observation eac6ef34-4cf6-4290-ac41-ef9c3bcb953a · outbound

This paper cites Tensor core performance: The ultimate guide,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Tensor core performance: The ultimate guide,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.467837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.997491Z digest=sha256:e7588c4a3681998f2f3f74d3307c659f4ce6fb096b5842dacc0af633d351c21e

Observation acaa0eff-aaa0-4ed9-ad1e-b82c1d9d7738 · outbound

This paper cites Nvidia a100 tensor core gpu architecture,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Nvidia a100 tensor core gpu architecture,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.461515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:33.999630Z digest=sha256:207729feb09802af7171e04883467de20f6765342bdbc84195c7451cbebbd018

Observation 79775c60-7ee3-46d3-99da-2aec46b3deaf · outbound

This paper cites Nvidia turing gpu architecture whitepaper,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Nvidia turing gpu architecture whitepaper,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.454754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.002318Z digest=sha256:14c78b7d039b31c69a15dccfc4b18ee02ea77482e1d31c4901f263c6c558d4cd

Observation 9a5f4477-ab6c-46ce-8f59-c759b49a7646 · outbound

This paper cites Figlut: An energy-efficient accelerator design for fp-int gemm using look-up tables,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Figlut: An energy-efficient accelerator design for fp-int gemm using look-up tables,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.005265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.005265Z digest=sha256:72a55a8928c699951fe67d878574dbe7d9b9203547825318cccfa3ad5ec88859

Observation c403adaa-dcb0-43ee-ad15-14d0688474fc · outbound

This paper cites LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.007377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.007377Z digest=sha256:75217b01a42af8bc8b643a9bcea81e90f75c886fc42da004d219354f8e53a40e

Observation c8689608-4ed0-4553-8a34-81cec0ae2380 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Pytorch: An imperative style, high-performance deep learning library,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.009539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.009539Z digest=sha256:4fb74d1a487aa538ffc854406eed0dc26cd8ed984e0b480409a34e067739fff8

Observation 26c6ef4f-79ec-455a-917d-d2dd090e8f2b · outbound

This paper cites The spectrum of the fisher information matrix of a single-hidden-layer neural network,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration The spectrum of the fisher information matrix of a single-hidden-layer neural network,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.439003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.011413Z digest=sha256:9a30e3eb05ca9436d8e1238c498ce0e91d89411bc12595150567bc49ecb50621

Observation f229ea5a-8b62-487e-bfa3-3749f265cab1 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Code Llama: Open Foundation Models for Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.013011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.013011Z digest=sha256:a53005c5bc28178563399cbd3edd6c8e415b034eb94a10509e4838ba74cdbd76

Observation 5ed5fbe7-3998-44da-9c87-92171d12317a · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Winogrande: An adversarial winograd schema challenge at scale,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.015194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.015194Z digest=sha256:bcb22164524f3da4a0f17a615a11aa4556bbeec14cb29aa68ecb3e04e58a0e35

Observation 3e8fc9e3-9e3b-42f2-93e8-fb4b672eedae · outbound

This paper cites From high-level deep neural models to fpgas,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration From high-level deep neural models to fpgas,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.426375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.017582Z digest=sha256:ab0267b14dd9bf14ad8bd2b936342ffe530546cd3437430d51b9a79b125453fb

Observation af992595-daa8-4976-927e-11a4586467a6 · outbound

This paper cites Using tournament trees to sort,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Using tournament trees to sort,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.419142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.019566Z digest=sha256:bbd27958b6fc311dcd9b4356ea0f8190c8ce3ea7bae1d1457befa5c8367a71bf

Observation ccd3cce7-639b-436b-b6df-46d6abd5dcd6 · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration FlatQuant: Flatness Matters for LLM Quantization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.021581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.021581Z digest=sha256:3e79916f2f63760c51fcfbe0a7d333c1d85c77befd0febeb85d2b75347cfaae1

Observation 3d92a750-e61d-4f56-8bc8-191e6652ad31 · outbound

This paper cites Crystal: Illuminating LLM Abilities on Language and Code.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Crystal: Illuminating LLM Abilities on Language and Code

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:15:34.105303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.024059Z digest=sha256:a57d5795994a5e2499c65258636dfbb43dc467aabd0dcee51109c778b105495d

Observation ac7a5324-4274-4cab-8520-2661799740bd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.026962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.026962Z digest=sha256:4677f7eaebd28ea99966615c0a879dea654d52bc672aeda45c8645094f6a6666

Observation 0328b448-115f-40cd-bbfc-6204c94d80f2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.029186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.029186Z digest=sha256:dd4e27ba1c7cacd8272eabd32cbc5888e4bafc2a38fcdabd50152401497faed5

Observation ca415c67-07c8-4bb1-bf9b-d9649486e814 · outbound

This paper cites Training LLMs with MXFP4.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Training LLMs with MXFP4

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.031488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.031488Z digest=sha256:36dbc79a40e33ea8db5cbb219f619c91a0218902efb1471f8a130db6028aef62

Observation a5e8def4-b552-453b-a57f-78b0e69e68de · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.411949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.033634Z digest=sha256:66542fe249d24f47d7d5d0066e759678e58499f682a476fb2cf81ebe84f568b4

Observation c2de2862-1679-441e-b993-7cfcf9e66440 · outbound

This paper cites Transformers: State- of-the-art natural language processing,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Transformers: State- of-the-art natural language processing,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.035550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.035550Z digest=sha256:aea46f52c754ff7526e5fee358c3be862b0291ffc69df51e2f007102cc6ea85b

Observation 8dfb485f-b35d-4ac7-ba3e-55b33b8f6c72 · outbound

This paper cites Block- wise mixed-precision quantization: Enabling high efficiency for practical reram-based dnn accelerators,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Block- wise mixed-precision quantization: Enabling high efficiency for practical reram-based dnn accelerators,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.401129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.038007Z digest=sha256:6fa7f6b80a66fbe8768b5fcfb6a7352f8ae18151405df91eafab9be806019497

Observation fce3dd4c-b901-4376-9ad9-194b8fe19afc · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.039849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.039849Z digest=sha256:06f05bc86eb77f2891fd209f57540b7b0a266218c19d41ef0567a7cf3537c894

Observation 8ed5d561-1e23-4b49-94d1-8c349e8dd2e8 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.041839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.041839Z digest=sha256:e61f382840813686614b4dcea47bf8be89792c009e56c778ff2b812aee76f387

Observation f0c0c966-a8ae-4717-b45d-67c7d0e6e38b · outbound

This paper cites Lq-nets: Learned quantization for highly accurate and compact deep neural networks,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Lq-nets: Learned quantization for highly accurate and compact deep neural networks,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:15:34.387928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T11:15:34.044005Z digest=sha256:cfb7c871a8b57de5f975ee9b74710dc5757d8b014178caebcef144a6a7cb029a

Observation f4ae1c58-19aa-45ab-b601-7249c9c27193 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration OPT: Open Pre-trained Transformer Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.045763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.045763Z digest=sha256:bfae2632602ec91932f15786c999d24cf979fbefa17a207ff95004ed72c90420

Observation 08062e34-8710-45a4-b3cf-5cd7fd723fce · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:34.047943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:34.047943Z digest=sha256:5678a4fdedf05d8248f3d93ab356e11bf44e3c6686c38c1ace2f7bd40a042e80

Pith citing papers

No inbound Pith citation observations are available.