Pith. sign in

Paper Citation Record · LEDGER

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers

As of 21 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2502.01770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01770 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:38:49.937475Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b419006d-16a3-4c9b-afff-7893d0332736 · outbound

This paper cites The groq software-defined scale-out tensor streaming multipro- cessor: From chips-to-systems architectural overview.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers The groq software-defined scale-out tensor streaming multipro- cessor: From chips-to-systems architectural overview

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.763089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.492977Z digest=sha256:fa956a3f9959c112680ca39016c5b16ebdf28d1e8099c771b1191874a9b5fceb

Observation ae8bf404-f7de-435d-8646-bb8240fda371 · outbound

This paper cites Vivit: A video vision transformer.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Vivit: A video vision transformer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.599079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.539242Z digest=sha256:c446036e0e6fb9c86ce12910475f1c8259d9ab914c5ef1cd80ccb81eb02f4c1f

Observation acea6c6d-1f90-4dcb-a543-e1ea593118f8 · outbound

This paper cites Longformer: The Long-Document Transformer.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.560159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.560159Z digest=sha256:0cdde042be0ca9b8bd6b9ba79388ed22d0e68bc4833a1354fb66e0f33c4d05ae

Observation 3d1f5d20-4f82-4cc4-be9d-ccf2dc5da61b · outbound

This paper cites Prime: A novel processing-in-memory architecture for neural net- work computation in reram-based main memory.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Prime: A novel processing-in-memory architecture for neural net- work computation in reram-based main memory

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.505984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.583032Z digest=sha256:365f0ee7ba6a21a61a753fa60d9c7f0f3b39ce2f78436666aed2720cf468caa1

Observation e3025021-94c1-4ca1-a82f-0ae969d232f6 · outbound

This paper cites On a model of associative memory with huge storage capacity.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers On a model of associative memory with huge storage capacity

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.495954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.601191Z digest=sha256:8203e7ffe346c89ccf4730f1da1f5755667d7c52d20f3e85891382829793ccf7

Observation fbbc6e7a-07a2-462a-b65e-7f01f8e25bc6 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Imagenet: A large-scale hierarchical im- age database

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.485968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.623282Z digest=sha256:c72e8588aa21013621a98947d9f0675de126426de58a92d46cfe44009838a322

Observation f8830fe7-2255-458e-b758-5955e30abd1d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.634024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.634024Z digest=sha256:dead1b8663e24edddcc8ddc3303d490be485732a460cd91490afc4b48170a18b

Observation 16af9084-2018-4824-9596-3b991fcba02b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.643210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.643210Z digest=sha256:6bc9455953f4cb0ecfff58edac456b1814561c6e15a11c70c7b1490a4e5351bb

Observation 81bce930-33a2-4d9f-8423-e1af50bfcb83 · outbound

This paper cites Array multiplier using xnor.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Array multiplier using xnor

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.475473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.646724Z digest=sha256:5f6e6e215f67932a03fc4f0d56abcb8bcf327578a4aa890a848269ecc4360243

Observation 4968c933-7ede-4ef1-aa50-12bb47aee446 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.650499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.650499Z digest=sha256:ae837330202bf957809d9b694974da8c2d5a243eb17769b3ad3ea31af399cb2e

Observation 8717f85c-2bab-4f52-bd26-00d9b1b0779a · outbound

This paper cites Sparse co-attention visual question answering networks based on thresholds.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Sparse co-attention visual question answering networks based on thresholds

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.465098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.654038Z digest=sha256:fe0290c269447ba11cb66c34c7b045a4bde33b9c014dd79160dc9b110a6d4f2c

Observation c03eb002-a4b6-480d-a948-1d151862ad9c · outbound

This paper cites Bivit: Ex- tremely compressed binary vision transformers.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Bivit: Ex- tremely compressed binary vision transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.455127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.657745Z digest=sha256:1a530ec67a2bb9b22f95c8c9e22a93595c5708024b8566cceb0f2272340b633f

Observation d2588fec-12a9-41f7-b23b-b3ccd450ae75 · outbound

This paper cites Binarized neural net- works.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Binarized neural net- works

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.444198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.661797Z digest=sha256:ae27bec6e932a6d370e68997864d5e2f1699d459f73df94cd97c31ea0e2f11bf

Observation c2da33d0-4378-41ee-9ef0-1a27d49c5da5 · outbound

This paper cites In-datacenter performance analysis of a tensor process- ing unit.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers In-datacenter performance analysis of a tensor process- ing unit

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.432867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.665369Z digest=sha256:df5919e1161e629c40a287d439ee6b6fc3a87107e106429f9014e9549383263c

Observation 739ad5b8-32c2-4791-8ae2-1575d658db63 · outbound

This paper cites Scaling Laws for Neural Language Models.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.668762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.668762Z digest=sha256:fe22abacd8abafa9159b267feeba26ecec9e91fc146aff485855206ccc00dd13

Observation 31f9d1e8-1f92-4f49-9c24-6adbec35840b · outbound

This paper cites Reformer: The Efficient Transformer.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Reformer: The Efficient Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.672683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.672683Z digest=sha256:45678882121378f281596b720dbe58ba722117e5b6b9225d399b8e3bc3201deb

Observation 01440fe2-66d2-4c8d-8b0b-8dd749c24c58 · outbound

This paper cites Self-Binarizing Networks.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Self-Binarizing Networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.676164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.676164Z digest=sha256:1498158cd8b582472672ab4508e230e7a694ddb3aa779350d66f0de85e88f567

Observation d31963a7-fa17-4619-9d2e-b0ac05b836d2 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.680009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.680009Z digest=sha256:341fe655bb6382ee44ef790e85d28cdea3636275730cbd203ac514e97fad72cd

Observation c3216a18-f736-4ffc-9971-a5a626e5eeb5 · outbound

This paper cites Power and area-efficient xnor-and hybrid binary neural networks using tft-type synaptic devices.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Power and area-efficient xnor-and hybrid binary neural networks using tft-type synaptic devices

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.422609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.683729Z digest=sha256:7cf1f3502098cbdaf6d9a7a5ee0592b340e79d4f5879942705a81b6e5d7283ed

Observation a344935e-b50f-4b82-815b-c1b2ec237ce4 · outbound

This paper cites Fp-bnn: Binarized neural network on fpga.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Fp-bnn: Binarized neural network on fpga

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.411234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.686956Z digest=sha256:ecc0610aaeebe4f8f586d580493b378e8042dd8087977d2c9da1164238b052eb

Observation f0c4a837-bee6-40f8-926e-ec8c506c48f9 · outbound

This paper cites Cerebras architecture deep dive: First look inside the hw/sw co-design for deep learning: Cerebras systems.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Cerebras architecture deep dive: First look inside the hw/sw co-design for deep learning: Cerebras systems

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.398540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.690355Z digest=sha256:76e40043571552ffc44b8fde8f2fddacbbd8d3aaa766c9b01f7c2bbaaa77e584

Observation b689a940-afd1-4966-977d-1132fb5afa68 · outbound

This paper cites Transformer Acceleration with Dynamic Sparse Attention.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Transformer Acceleration with Dynamic Sparse Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.693578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.693578Z digest=sha256:1f0077b4ffdc84a34188702600ec25f8a478d9afc5dc200d284a012199dc3638

Observation 5f80d868-7fa4-4504-885d-44af9c5f09f6 · outbound

This paper cites Bit: Robustly binarized multi-distilled transformer.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Bit: Robustly binarized multi-distilled transformer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.365203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.696818Z digest=sha256:57ae80e423743cb38637c39a199bf6aec24703d78d0fe969512200f0c83bc3c0

Observation b1c5a520-2fca-40e2-a5d4-c1cb7f057de6 · outbound

This paper cites QuALITY: Question Answering with Long Input Texts, Yes!.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers QuALITY: Question Answering with Long Input Texts, Yes!

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.700286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.700286Z digest=sha256:94f7ddc006d8b9ee7d70afc9e070c80dadd51a44e13695a0193067b8792d32b4

Observation 41e878c0-ae5d-433c-b1fd-01237ade3d11 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.726182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.726182Z digest=sha256:4f0afa05c0257413dbc8d7be5d5b77a78171f6f2d36c17e67c0553d13c2126ad

Observation 0bc6496e-fd33-4345-bfa3-2af72836a718 · outbound

This paper cites Hopfield Networks is All You Need.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Hopfield Networks is All You Need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.791911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.791911Z digest=sha256:880ee024aa85abdf9e56c3dfb1aa3838ed68e30a810f5f152e815ae172b3adde

Observation 8abd4172-985e-40b9-ad56-dde7503936e2 · outbound

This paper cites Xnor-net: Imagenet classifica- tion using binary convolutional neural networks.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Xnor-net: Imagenet classifica- tion using binary convolutional neural networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.287215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.806493Z digest=sha256:6bca2db7472481ec2ff8069e3b30fca78fd73d5e5a37fe17159332be061905ea

Observation df7b73f3-5567-4649-832f-f625e79e74de · outbound

This paper cites XNOR-Net: Imagenet classifica- tion using binary convolutional neural networks.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers XNOR-Net: Imagenet classifica- tion using binary convolutional neural networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.223753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.827281Z digest=sha256:8b3637b778e860f04f5940cd42555442b269f2771a57901aeb0d7b90f3f63034

Observation 345b58c1-29c7-4575-bdcd-2b25e26a056e · outbound

This paper cites Efficient sparse-dense matrix-matrix multiplication on gpus us- ing the customized sparse storage format.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Efficient sparse-dense matrix-matrix multiplication on gpus us- ing the customized sparse storage format

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.186783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.861360Z digest=sha256:beab44674970dd0044d903ff79f03936e26b1f2cd6cadc2620cdceb0d342e80e

Observation e358aeb3-d645-4693-bc82-1871e48d7010 · outbound

This paper cites Sextans: A streaming accelerator for general-purpose sparse-matrix dense-matrix multiplication.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Sextans: A streaming accelerator for general-purpose sparse-matrix dense-matrix multiplication

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.174790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.888701Z digest=sha256:4c4365e23f7485d656c7ad6f893a27e31bab2c6b930f2f6b1ffb9c5ad35f12cd

Observation 4eaa9b5f-5e40-4dee-940b-b7f313ec6eaa · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Training data-efficient image transformers & distillation through attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.163641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.915271Z digest=sha256:45871e8a0ca13acf364f17bac0882d7c9a9d70b3c5b671c8e22be9f92f6a4fd6

Observation db93b213-b915-442a-8390-616bee382c13 · outbound

This paper cites Attention is all you need.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Attention is all you need

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.151803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.919284Z digest=sha256:dbb2514e6efa11f3e841d44e4ca3161c0fc587c48b41670586f2a0f02b6b748b

Observation c3ad2eac-9e10-4e3f-a17e-7b7b4312b389 · outbound

This paper cites Audio Transformers.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Audio Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.922737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.922737Z digest=sha256:7ffadb5130ed67d070a5fadb42a32313a0692cf44c254a7ac6fbb1db2521f9ac

Observation 688012e6-1b89-461a-b44b-fd1ec63f74e3 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Linformer: Self-Attention with Linear Complexity

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.926693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.926693Z digest=sha256:a03e51f34ea5e09e257a1d5c686a3b74d496248ed34f587229fe348f54c7ba32

Observation 32adf834-400b-4180-8ee8-7374a6b9e8db · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers OneBit: Towards Extremely Low-bit Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T14:38:49.930687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:38:49.930687Z digest=sha256:e412b0ae0245ee6618a256fa4dc3a79e3b7d7c2991b6b9769b5e97fa78174524

Observation 8ca2a554-0c2e-40b9-bcd7-7142221b8e17 · outbound

This paper cites Pb- llm: Partially binarized large language models.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Pb- llm: Partially binarized large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.139820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.933991Z digest=sha256:c9548bb7777a1a257dc1c54b2d4872f26ce34d26ffc692903536d88a5463b3a6

Observation 45272873-b124-4ba1-9150-da2ebbe02e7c · outbound

This paper cites Big bird: Transformers for longer se- quences.

Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers Big bird: Transformers for longer se- quences

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:38:50.127043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T14:38:49.937475Z digest=sha256:016e4024330f37aa1b2f85256f2a35421675b35d22b3104855e27ce84cde8efa

Pith citing papers

No inbound Pith citation observations are available.