Pith. sign in

Paper Citation Record · LEDGER

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2507.03865.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03865 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:08:09.488803Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.167572Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8025c31f-e790-499d-bb46-40311bd8d76e · outbound

This paper cites write newline.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.520040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.520040Z digest=sha256:23445a2e9b4084510b656fc1da2f104907acf2450b7f70381b81aa97f4acc2fd

Observation 3271328d-b004-4830-869d-45401ba497cd · outbound

This paper cites Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.144777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.540763Z digest=sha256:0484dd98a326db6eab3c6e1892e523390f7ee26eaadcdf8a972f9c22bda8a27f

Observation 4861d466-5830-4fd9-b861-0622edabd0d0 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Longbench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.134965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.575951Z digest=sha256:8244575e78b776078e2f04137fed29f16a68a05a00c8b528995d2a2bc8fc4373

Observation a3febf6d-418c-4022-a485-f1c4592d3c68 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.665742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.665742Z digest=sha256:e9d1ca6bc00de149736d0fd9887e9a63a89ae5936f813b5ef10592d7f8c8cd22

Observation 7b46542b-851f-463d-8eef-d25b67658055 · outbound

This paper cites Spectral filters, dark signals, and attention sinks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Spectral filters, dark signals, and attention sinks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.694850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.694850Z digest=sha256:9d7e9577e764af1567f1e8259bdb4bf5210f96964b6143e76336ef2eca483b22

Observation 7e3cbe33-ce91-4136-8db3-9c084ff0c855 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.728594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.728594Z digest=sha256:86dcaf1f057cec64d829603180c704a325665bb70ff9e8e90a9919fe9133d7e6

Observation 3871d836-3f1b-46ac-99f7-e2d643164297 · outbound

This paper cites Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.119715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.784271Z digest=sha256:53cb8287f235809b596279dbfc0931ed34641584f009a4203fc9fb1dac5396ab

Observation 6a2e375d-4bd9-46bc-b2f7-ceee1f13df11 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.826197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.826197Z digest=sha256:f9c86de73256a03770df3652b279dcf34363c821f56d9077eace888cf6324e7d

Observation 8fe2c9d0-f948-465b-942a-3381ee8bc030 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.884228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.884228Z digest=sha256:9195a40c2cffe496e4bda5fa786c42354a24ec56e365b069ca22107c72680986

Observation 22d001fd-06b5-44b3-9188-35a713a1b314 · outbound

This paper cites The Llama 3 Herd of Models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.940790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.940790Z digest=sha256:9d583e9d3b2bf5ec6a33105027cd4b5c9464b4e4a13064ebc6ff2366e24f7938

Observation cf401c8b-e256-4a47-9367-6c7a4db9f037 · outbound

This paper cites Layerskip: Enabling early exit inference and self-speculative decoding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Layerskip: Enabling early exit inference and self-speculative decoding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.110224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.986186Z digest=sha256:284943ab1a43af93c438373e6322db55a82e6004da33c5ab1beed965012c85d3

Observation 1c73407b-f2ef-4a55-b6bb-09fa8cbfb7b4 · outbound

This paper cites When attention sink emerges in language models: An empirical view.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference When attention sink emerges in language models: An empirical view

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.099806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.062613Z digest=sha256:b5d6851a44777630562ac6f716ec031b4722bb77abb2528d6f9cb86165f17069

Observation 870b8fee-af36-484a-8381-11b1f0f59ac2 · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.136253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.136253Z digest=sha256:ef4e0b74ac9ba01df2ca8ac2119a4192077c980fdaa95b0b26d3dede0da7fd1b

Observation 3133fe33-7106-4036-8df3-ea59d7c6b30b · outbound

This paper cites Mistral 7B.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.195766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.195766Z digest=sha256:b289148c271c0bc85d1e80223c60301dcb46df4b1d7b7e80eef739b107f86127

Observation 52bdaca1-fa16-4b5d-b930-9629ec47723e · outbound

This paper cites Mixtral of Experts.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.249338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.249338Z digest=sha256:39e1f5b2a8fca1bb5fd865de30b30b2f12d7fd60b9dc7ceeb586bbf44083f57a

Observation c5ab7db0-7fb5-4b79-a78f-e3774e77afb1 · outbound

This paper cites D-llm: A token adaptive computing resource allocation strategy for large language models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference D-llm: A token adaptive computing resource allocation strategy for large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.082034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.288318Z digest=sha256:3e456a68e59ba6ffe320824e22169888ff658d9c94466159c29abe596a2f466f

Observation b92bd581-f141-4ed6-955b-46c974e5ccc7 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.404789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.404789Z digest=sha256:9a86ddd4cf606fa3adbb3f770aec636bd24d15f61a4faed60471debd81da84bc

Observation d0c73403-f92f-4756-93ab-f13daf0b7375 · outbound

This paper cites B io M istral: A collection of open-source pretrained large language models for medical domains.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference B io M istral: A collection of open-source pretrained large language models for medical domains

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.071971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.480577Z digest=sha256:c4f2e2c6462f87e405359941a2a462b690c7707f63d1c1f545e2addb4ea35958

Observation 00e02adc-9716-4445-82a2-57fffb6784c8 · outbound

This paper cites Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.633297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.633297Z digest=sha256:58c6a36a9449765e700f91fc2f6480de07d41a7a0e71b6b51d79ef492c153773

Observation ebc7bd35-3c72-41ea-8248-cfa28d1b62d4 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.685132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.685132Z digest=sha256:b6a79bec8adda1375a82388d137514205818440467e4c2c10c70c932fad07895

Observation b5ce3c25-d6bf-44cd-8c55-7a4c02fc4a5f · outbound

This paper cites Pointer sentinel mixture models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Pointer sentinel mixture models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.061210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.752189Z digest=sha256:abbd9b845f7c501859201aad911dd4d41b6d49a6025fffca11d1ea3934a87f16

Observation 7e696759-321d-482e-a7b2-db9808987c1a · outbound

This paper cites Using an llm to help with code understanding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Using an llm to help with code understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.050576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.806617Z digest=sha256:70040223d397991cfc5d3623814347734179ff13b288773863f4fb3b4e2a26a1

Observation e3f9296e-de56-4039-9696-d66615f753ec · outbound

This paper cites an unresolved cited work.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.846759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.846759Z digest=sha256:df04de5ed4a5cad35fd7850974efd5313037e17c4dc540639a49df1b0e938c62

Observation 57f4bdd9-3066-46f7-b105-9bcddaf9047a · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.957269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.957269Z digest=sha256:e8c7431d30a322399cd7f8b1260409d4c2d5a5cef7664e30d976e94e6766c1b6

Observation e5224ed9-8b43-4c75-ac39-c02ce6f43d03 · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference L., Bhagavatula, C., and Choi, Y

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.997272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.997272Z digest=sha256:a256a8bb8b1e92074195356d11e87df80985105c36fe051770ab535dae9b727b

Observation 67011f70-86fa-4438-b248-fe500a267bcb · outbound

This paper cites Q., Tay, Y., and Metzler, D.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Q., Tay, Y., and Metzler, D

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.027343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.188012Z digest=sha256:8c41b3b3f32e9567531d962636ff4ae8a165f9b20da1ae7fe90a4196667e8fae

Observation 06e44b8d-493f-4b07-9890-2f66f876fd32 · outbound

This paper cites A deeper look at depth pruning of LLMs.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference A deeper look at depth pruning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.244465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.244465Z digest=sha256:c7b016a053fddf5eefdb6790649a411137cab7273d48d716843ac40689ffeb76

Observation e2bb6b72-cc2a-47c0-9b84-e450ce09f459 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.298473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.298473Z digest=sha256:eb9bbe8888c00571a7dd7ea325384b627e52921d9a304a5a5d2c28f831edb00c

Observation 32042cfd-8cb9-4108-b8ba-e4bf8c82d96f · outbound

This paper cites Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.915706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.393078Z digest=sha256:adef25557b179bcb5404ad13398976d56360384ce0b70b8c4b30a4a571079932

Observation c8970a42-2ccc-4f37-8dd1-ff8b0dd3fc6e · outbound

This paper cites Z., and Liu, Z.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Z., and Liu, Z

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.906428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.446045Z digest=sha256:fcebc2417b538cefff33a32a91a1f2f711b254aa607d0f0c96936c31d6b09009

Observation 6fc4520c-285e-49a0-b59a-02cabd591125 · outbound

This paper cites Razorattention: Efficient kv cache compression through retrieval heads.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Razorattention: Efficient kv cache compression through retrieval heads

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.896623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.512199Z digest=sha256:2a01eaf9bcabfb9c7a21976a7c9b541d0a8ca1d3e1cafd147f983e8ae02eedea

Observation 3846996a-fad5-4ef1-9520-6dc86ad22fcc · outbound

This paper cites J., Ting, D.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., Ting, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.740169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.610825Z digest=sha256:d8d1317604147dc9c3e162c42aee054f4527a6b03caf91ae021abecb77cb330f

Observation 4533b7cf-cd02-4b68-9921-77c3bb68dc7f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.722523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.722523Z digest=sha256:ac0b0a692d33a54b802469c6bbbcc1f9ad58016d15a672252b9af7bda7fc5ad7

Observation bd2fa558-17d1-4cef-9907-f074a86bf813 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference BloombergGPT: A Large Language Model for Finance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.837488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.837488Z digest=sha256:63753d41a46a4e8d7e2df3ad56217325f662fef53da1feb575b5fd3ff683b551

Observation 2877ea30-5aea-4eb7-8810-eddfceecbbbc · outbound

This paper cites T., Peng, R., Wu, Q., and Wang, C.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference T., Peng, R., Wu, Q., and Wang, C

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.551322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.895926Z digest=sha256:e62da21cf82db18a770058318557c3a4ec9c94bfae74a41730cd0fa633476698

Observation 13964cca-0020-4437-8b2f-7895d5678333 · outbound

This paper cites Duoattention: Efficient long-context llm inference with retrieval and streaming heads.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Duoattention: Efficient long-context llm inference with retrieval and streaming heads

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.339924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.003068Z digest=sha256:b5ec801068cb923aece0ec805f78e1e152a3459a39dfa9dbdcdc204496ee60cf

Observation 8527fb09-f508-40c1-a700-609a7a9bc288 · outbound

This paper cites Efficient streaming language models with attention sinks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Efficient streaming language models with attention sinks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.169486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.106850Z digest=sha256:6ebda92ada282d625c6be05ee922175c6c1ed2ea13d9001403c18740b74ef8e3

Observation 3ee6e21e-26f5-49f1-a821-95061880a634 · outbound

This paper cites an unresolved cited work.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:08:10.029542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.248931Z digest=sha256:74b298dc8ae3f2f2d1013b29c65bc2fdcd66d90bf74492cd5861c71b95f5a6fc

Observation 7f2a254e-64b3-4b64-a42b-e0c7054921df · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:09.806985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.372102Z digest=sha256:40059e5e9e13dff0c5b13bc7af94751371c6fd9435840fc7dfc337c7cff7aa22

Observation ae04f335-09c4-4713-9b2e-f61df35db73e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:09.488803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:09.488803Z digest=sha256:6fb1cf69de773de205623f17c4210bf8cb9170b679fc15e711bfcec3eecd4d25

Pith citing papers

Observation 4040624e-1a52-449c-ad62-3e59d3d9a76d · inbound

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers cites this paper.

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:07.167572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:15:07.167572Z digest=sha256:afb475915933909aecb05d844a7fec017c13f9d7b0ab4026f4589bc9fa632ad7

Observation 1b256f7d-5bfb-4f70-a411-d8e74ae2a03f · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.280027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:e9e283136344891d1ad01ff5724f78a9845eab13c756c808ece4d548243917e9

Observation e4d1dbc0-a3f4-46bd-87cd-26436f2198d6 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:25.027494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:25.027494Z digest=sha256:b5bdbc312cad68aa27bd3d54e1bab09bead786a97ca89fba665e98740263e79f

Observation df018183-1fe6-4e00-831b-dc59de7a9fd1 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.707032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:35059a7557fa73756ed9e3ef782329edebcea2aef6e8fe1eb94fdc6b5d1d6602

Observation e37ca751-0043-4a4b-9ad3-43f1489a797d · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:47.283100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:ca771f695f5fceb3477a0fedb1143a73e2b512d9fd40c619f6f3a84987223123