Pith. sign in

Paper Citation Record · LEDGER

Are Sixteen Heads Really Better than One?

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:1905.10650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1905.10650 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:42:31.105548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T21:16:34.316842Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 127165ec-a4e0-457d-8a85-d0439b709ac8 · inbound

Do Transformer Attention Heads Provide Transparency in Abstractive Summarization? cites this paper.

Do Transformer Attention Heads Provide Transparency in Abstractive Summarization? Are Sixteen Heads Really Better than One?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T12:20:47.646326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T12:17:22.938922Z digest=sha256:c081fb4e95bf61c0f232eca050a13fd11544d94661c69f266ad87bdfda65dafb

Observation 456ff2c7-ff00-423c-bc9d-8eee4cf1ea03 · inbound

HuggingFace's Transformers: State-of-the-art Natural Language Processing cites this paper.

HuggingFace's Transformers: State-of-the-art Natural Language Processing Are Sixteen Heads Really Better than One?

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:53:59.362388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-11T14:53:58.963468Z digest=sha256:3e375f7eeedbb43e13d7ae5c28bccb8fce6861b962ff30e5dcbb822a91662171

Observation b925d23c-cfd1-4208-a419-326e31eec14e · inbound

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing cites this paper.

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing Are Sixteen Heads Really Better than One?

Reference 268

Resolution
unresolved
no resolver link, observed 2026-08-12T11:42:31.105548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:42:31.105548Z digest=sha256:d7b82565296b1fdf59c50f78b46cd60321db5e33ca345000de2b08fefdf6f35e

Observation 4051f030-8926-461b-bf2a-e1e7591822f3 · inbound

Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs cites this paper.

Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs Are Sixteen Heads Really Better than One?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:18.891139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:18.891139Z digest=sha256:015fdfcaac7ff57d3f20efa816f01fb90e7ddedfe63edaaae2e7f34f6fd6babe

Observation 593edf05-27c2-4a0d-823c-f1600d4e8a39 · inbound

Specialization of softmax attention heads: insights from the high-dimensional single-location model cites this paper.

Specialization of softmax attention heads: insights from the high-dimensional single-location model Are Sixteen Heads Really Better than One?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T19:04:54.682986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:04:54.682986Z digest=sha256:1eb5b63d9d7bbe49255a4728504079408f2c0437c01da544fe02dc5a550155e5

Observation 4f9690ff-f04b-4834-839c-62ed810460d7 · inbound

AudioKV: KV Cache Eviction in Efficient Large Audio Language Models cites this paper.

AudioKV: KV Cache Eviction in Efficient Large Audio Language Models Are Sixteen Heads Really Better than One?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:54.477882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:28:49.768210Z digest=sha256:a8e35041d8502697e1e331b3f37806461e15c135c718e97cd54824f9fbdbdc53

Observation 0396190a-bdf4-4d65-867a-88f9d9411e01 · inbound

MedCore: Boundary-Preserving Medical Core Pruning for MedSAM cites this paper.

MedCore: Boundary-Preserving Medical Core Pruning for MedSAM Are Sixteen Heads Really Better than One?

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:28.483087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T20:13:38.290363Z digest=sha256:de2ca52ddcf3fd09426004f07c7b647eed38c87f31a0bcdefcb24f96ccd1d0d4

Observation 65cb53cc-102f-4f68-8b90-6a6391c83001 · inbound

ChessMimic: Per-Rating Transformer Models for Human Move, Clock, and Outcome Prediction in Online Blitz Chess cites this paper.

ChessMimic: Per-Rating Transformer Models for Human Move, Clock, and Outcome Prediction in Online Blitz Chess Are Sixteen Heads Really Better than One?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.453207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T07:08:06.493103Z digest=sha256:c2de9274f6b3ed8eb9c1938a55d5b7eca6ff2112cef44927bd9756ea8c13b794

Observation aa04bc4e-d9e6-4eb4-ad7d-25550abb5228 · inbound

Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence cites this paper.

Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence Are Sixteen Heads Really Better than One?

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-09T21:16:34.317951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-09T21:08:24.293077Z digest=sha256:2d21e7cb1bc3dedbeb775609635259f1c044ab2f64461fe7efae8189cb85436d

Observation a5522200-4fa8-49c4-ba54-93727528a3de · inbound

It Takes a MAESTRO To Prune Bad Experts cites this paper.

It Takes a MAESTRO To Prune Bad Experts Are Sixteen Heads Really Better than One?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:56:24.598744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:56:24.598744Z digest=sha256:43cfde8a3f336420ab2f64ccb4e47d735eaef9e4fdcbd6370020cf93aab4a8bf