Pith. sign in

Paper Citation Record · LEDGER

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2507.03865.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03865 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:08:09.488803Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.167572Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8025c31f-e790-499d-bb46-40311bd8d76e · outbound

This paper cites write newline.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.520040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.520040Z digest=sha256:23445a2e9b4084510b656fc1da2f104907acf2450b7f70381b81aa97f4acc2fd

Observation 3271328d-b004-4830-869d-45401ba497cd · outbound

This paper cites Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.144777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.540763Z digest=sha256:e036c3fbce3c08d8fd691a8dc1675e985c654f7408326efbd3dae4b541ebf033

Observation 4861d466-5830-4fd9-b861-0622edabd0d0 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Longbench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.134965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.575951Z digest=sha256:104e4c0cce891e68797e2c5410113f2d879136c8725bd71b14676e159b5e2b23

Observation a3febf6d-418c-4022-a485-f1c4592d3c68 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.665742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.665742Z digest=sha256:e9d1ca6bc00de149736d0fd9887e9a63a89ae5936f813b5ef10592d7f8c8cd22

Observation 7b46542b-851f-463d-8eef-d25b67658055 · outbound

This paper cites Spectral filters, dark signals, and attention sinks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Spectral filters, dark signals, and attention sinks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.694850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.694850Z digest=sha256:9d7e9577e764af1567f1e8259bdb4bf5210f96964b6143e76336ef2eca483b22

Observation 7e3cbe33-ce91-4136-8db3-9c084ff0c855 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.728594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.728594Z digest=sha256:86dcaf1f057cec64d829603180c704a325665bb70ff9e8e90a9919fe9133d7e6

Observation 3871d836-3f1b-46ac-99f7-e2d643164297 · outbound

This paper cites Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.119715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.784271Z digest=sha256:8b08d96a37d635afd22ac76815528adae991b7b35a6ef378d66b66b10aa4b9f1

Observation 6a2e375d-4bd9-46bc-b2f7-ceee1f13df11 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.826197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.826197Z digest=sha256:f9c86de73256a03770df3652b279dcf34363c821f56d9077eace888cf6324e7d

Observation 8fe2c9d0-f948-465b-942a-3381ee8bc030 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.884228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.884228Z digest=sha256:9195a40c2cffe496e4bda5fa786c42354a24ec56e365b069ca22107c72680986

Observation 22d001fd-06b5-44b3-9188-35a713a1b314 · outbound

This paper cites The Llama 3 Herd of Models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.940790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.940790Z digest=sha256:9d583e9d3b2bf5ec6a33105027cd4b5c9464b4e4a13064ebc6ff2366e24f7938

Observation cf401c8b-e256-4a47-9367-6c7a4db9f037 · outbound

This paper cites Layerskip: Enabling early exit inference and self-speculative decoding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Layerskip: Enabling early exit inference and self-speculative decoding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.110224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.986186Z digest=sha256:b992fa972240e60b4b048de50b39bbff57e020d9623fb4877dae730e3abebba2

Observation 1c73407b-f2ef-4a55-b6bb-09fa8cbfb7b4 · outbound

This paper cites When attention sink emerges in language models: An empirical view.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference When attention sink emerges in language models: An empirical view

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.099806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.062613Z digest=sha256:481bc79f5c68020c5043b6309b5ac783763e1054a39c0d134e7d76a23c769472

Observation 870b8fee-af36-484a-8381-11b1f0f59ac2 · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.136253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.136253Z digest=sha256:ef4e0b74ac9ba01df2ca8ac2119a4192077c980fdaa95b0b26d3dede0da7fd1b

Observation 3133fe33-7106-4036-8df3-ea59d7c6b30b · outbound

This paper cites Mistral 7B.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.195766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.195766Z digest=sha256:b289148c271c0bc85d1e80223c60301dcb46df4b1d7b7e80eef739b107f86127

Observation 52bdaca1-fa16-4b5d-b930-9629ec47723e · outbound

This paper cites Mixtral of Experts.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.249338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.249338Z digest=sha256:6ddd340a0426f0fd5d7158d4d18ebc529e757a002ca477b477eb26ef8ee859ff

Observation c5ab7db0-7fb5-4b79-a78f-e3774e77afb1 · outbound

This paper cites D-llm: A token adaptive computing resource allocation strategy for large language models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference D-llm: A token adaptive computing resource allocation strategy for large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.082034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.288318Z digest=sha256:11b1f8eb86e513ac1defceca7706e8bbd0cbfaa26c803cf5ec29c0b633fc4323

Observation b92bd581-f141-4ed6-955b-46c974e5ccc7 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.404789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.404789Z digest=sha256:9a86ddd4cf606fa3adbb3f770aec636bd24d15f61a4faed60471debd81da84bc

Observation d0c73403-f92f-4756-93ab-f13daf0b7375 · outbound

This paper cites B io M istral: A collection of open-source pretrained large language models for medical domains.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference B io M istral: A collection of open-source pretrained large language models for medical domains

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.071971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.480577Z digest=sha256:5bce285ca2ed217d9739e65490f4cb930aea5b8f46091a4c3b49ac55476b8115

Observation 00e02adc-9716-4445-82a2-57fffb6784c8 · outbound

This paper cites Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.633297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.633297Z digest=sha256:58c6a36a9449765e700f91fc2f6480de07d41a7a0e71b6b51d79ef492c153773

Observation ebc7bd35-3c72-41ea-8248-cfa28d1b62d4 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.685132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.685132Z digest=sha256:b6a79bec8adda1375a82388d137514205818440467e4c2c10c70c932fad07895

Observation b5ce3c25-d6bf-44cd-8c55-7a4c02fc4a5f · outbound

This paper cites Pointer sentinel mixture models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Pointer sentinel mixture models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.061210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.752189Z digest=sha256:b70b95c815ef43635951cb8b45303dc68f8323ee8b27fc9f8e6f1989fbd0fdf2

Observation 7e696759-321d-482e-a7b2-db9808987c1a · outbound

This paper cites Using an llm to help with code understanding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Using an llm to help with code understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.050576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.806617Z digest=sha256:5d8ae64040c2a4d159ba6ddcd66b3d0bf63905679836ecc12a013e4af80a0a23

Observation e3f9296e-de56-4039-9696-d66615f753ec · outbound

This paper cites an unresolved cited work.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.846759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.846759Z digest=sha256:df04de5ed4a5cad35fd7850974efd5313037e17c4dc540639a49df1b0e938c62

Observation 57f4bdd9-3066-46f7-b105-9bcddaf9047a · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.957269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.957269Z digest=sha256:e8c7431d30a322399cd7f8b1260409d4c2d5a5cef7664e30d976e94e6766c1b6

Observation e5224ed9-8b43-4c75-ac39-c02ce6f43d03 · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference L., Bhagavatula, C., and Choi, Y

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.997272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.997272Z digest=sha256:a256a8bb8b1e92074195356d11e87df80985105c36fe051770ab535dae9b727b

Observation 67011f70-86fa-4438-b248-fe500a267bcb · outbound

This paper cites Q., Tay, Y., and Metzler, D.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Q., Tay, Y., and Metzler, D

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.027343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.188012Z digest=sha256:4324d53dc9cc8dc080d0d30295ed80d16f552677236899698b2a583820df2774

Observation 06e44b8d-493f-4b07-9890-2f66f876fd32 · outbound

This paper cites A deeper look at depth pruning of LLMs.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference A deeper look at depth pruning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.244465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.244465Z digest=sha256:c7b016a053fddf5eefdb6790649a411137cab7273d48d716843ac40689ffeb76

Observation e2bb6b72-cc2a-47c0-9b84-e450ce09f459 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.298473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.298473Z digest=sha256:eb9bbe8888c00571a7dd7ea325384b627e52921d9a304a5a5d2c28f831edb00c

Observation 32042cfd-8cb9-4108-b8ba-e4bf8c82d96f · outbound

This paper cites Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.915706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.393078Z digest=sha256:0453aff6ddf37ef1581ef7354f7b260fcba6bf655ee4a6e6a426d54f3b2fecd0

Observation c8970a42-2ccc-4f37-8dd1-ff8b0dd3fc6e · outbound

This paper cites Z., and Liu, Z.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Z., and Liu, Z

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.906428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.446045Z digest=sha256:d6b56faffb89ab46027bdb68c8cba9803d0366bf5db7a1396b204cd2d0c2a903

Observation 6fc4520c-285e-49a0-b59a-02cabd591125 · outbound

This paper cites Razorattention: Efficient kv cache compression through retrieval heads.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Razorattention: Efficient kv cache compression through retrieval heads

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.896623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.512199Z digest=sha256:902924ae6c6aaf76dfb4045fcd16c23e7d85b46f6c5e9d9ca16e246f65e423eb

Observation 3846996a-fad5-4ef1-9520-6dc86ad22fcc · outbound

This paper cites J., Ting, D.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., Ting, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.740169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.610825Z digest=sha256:36ec75eefaa3c40b1f2a7c9a2b7b8cc042d7f60f258e71609b1c7f63af9a34f3

Observation 4533b7cf-cd02-4b68-9921-77c3bb68dc7f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.722523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.722523Z digest=sha256:ac0b0a692d33a54b802469c6bbbcc1f9ad58016d15a672252b9af7bda7fc5ad7

Observation bd2fa558-17d1-4cef-9907-f074a86bf813 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference BloombergGPT: A Large Language Model for Finance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.837488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.837488Z digest=sha256:63753d41a46a4e8d7e2df3ad56217325f662fef53da1feb575b5fd3ff683b551

Observation 2877ea30-5aea-4eb7-8810-eddfceecbbbc · outbound

This paper cites T., Peng, R., Wu, Q., and Wang, C.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference T., Peng, R., Wu, Q., and Wang, C

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.551322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.895926Z digest=sha256:5b6d1d2e40ed2f16dd2c1d6e9aeeaad94f206308d8e8d9b61a3c7010bea1b311

Observation 13964cca-0020-4437-8b2f-7895d5678333 · outbound

This paper cites Duoattention: Efficient long-context llm inference with retrieval and streaming heads.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Duoattention: Efficient long-context llm inference with retrieval and streaming heads

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.339924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.003068Z digest=sha256:3de1238e1b4f77a832424513dfa41bbfce340b595689f02986ffec09217170c2

Observation 8527fb09-f508-40c1-a700-609a7a9bc288 · outbound

This paper cites Efficient streaming language models with attention sinks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Efficient streaming language models with attention sinks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.169486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.106850Z digest=sha256:52021264241e8b3eebc8074e64f87ae2786a27a0c8ad803d71608f71a8070181

Observation 3ee6e21e-26f5-49f1-a821-95061880a634 · outbound

This paper cites an unresolved cited work.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:08:10.029542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.248931Z digest=sha256:33f45b586623afa507b9ada04d6094f12cbc1aad0d560e2696b49f32850cabbe

Observation 7f2a254e-64b3-4b64-a42b-e0c7054921df · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:09.806985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.372102Z digest=sha256:1d6b3e61d9570850d12626f6e6e589605be82610ef0531d9a71d741ad6d20f25

Observation ae04f335-09c4-4713-9b2e-f61df35db73e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:09.488803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:09.488803Z digest=sha256:6fb1cf69de773de205623f17c4210bf8cb9170b679fc15e711bfcec3eecd4d25

Pith citing papers

Observation 4040624e-1a52-449c-ad62-3e59d3d9a76d · inbound

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers cites this paper.

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:07.167572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:15:07.167572Z digest=sha256:afb475915933909aecb05d844a7fec017c13f9d7b0ab4026f4589bc9fa632ad7

Observation 1b256f7d-5bfb-4f70-a411-d8e74ae2a03f · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.280027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:20e7b3f28049da46560daff23ca01e1157b98dc23208777e93744e8778e8fd87

Observation e4d1dbc0-a3f4-46bd-87cd-26436f2198d6 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:25.027494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:25.027494Z digest=sha256:b5bdbc312cad68aa27bd3d54e1bab09bead786a97ca89fba665e98740263e79f

Observation df018183-1fe6-4e00-831b-dc59de7a9fd1 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.707032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:36336d3793f6fcc4a5b0e118391c72931632ced78f02419b907b6a180205e0f5

Observation e37ca751-0043-4a4b-9ad3-43f1489a797d · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:47.283100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:db510f4a43796079de27b7ad273bd3dcd22c3fa8544917b53142dbe3788cd4b5