Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:08:09.488803Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2507.03865.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:08:09.488803Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.167572Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
40 of 40 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 8025c31f-e790-499d-bb46-40311bd8d76e · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3271328d-b004-4830-869d-45401ba497cd · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4861d466-5830-4fd9-b861-0622edabd0d0 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Longbench: A bilingual, multitask benchmark for long context understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3febf6d-418c-4022-a485-f1c4592d3c68 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Piqa: Reasoning about physical commonsense in natural language
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b46542b-851f-463d-8eef-d25b67658055 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Spectral filters, dark signals, and attention sinks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3cbe33-ce91-4136-8db3-9c084ff0c855 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3871d836-3f1b-46ac-99f7-e2d643164297 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a2e375d-4bd9-46bc-b2f7-ceee1f13df11 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe2c9d0-f948-465b-942a-3381ee8bc030 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d001fd-06b5-44b3-9188-35a713a1b314 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf401c8b-e256-4a47-9367-6c7a4db9f037 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Layerskip: Enabling early exit inference and self-speculative decoding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c73407b-f2ef-4a55-b6bb-09fa8cbfb7b4 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference When attention sink emerges in language models: An empirical view
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 870b8fee-af36-484a-8381-11b1f0f59ac2 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3133fe33-7106-4036-8df3-ea59d7c6b30b · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bdaca1-fa16-4b5d-b930-9629ec47723e · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixtral of Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ab7db0-7fb5-4b79-a78f-e3774e77afb1 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference D-llm: A token adaptive computing resource allocation strategy for large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b92bd581-f141-4ed6-955b-46c974e5ccc7 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c73403-f92f-4756-93ab-f13daf0b7375 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference B io M istral: A collection of open-source pretrained large language models for medical domains
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00e02adc-9716-4445-82a2-57fffb6784c8 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc7bd35-3c72-41ea-8248-cfa28d1b62d4 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ce3c25-d6bf-44cd-8c55-7a4c02fc4a5f · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Pointer sentinel mixture models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e696759-321d-482e-a7b2-db9808987c1a · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Using an llm to help with code understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3f9296e-de56-4039-9696-d66615f753ec · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f4bdd9-3066-46f7-b105-9bcddaf9047a · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5224ed9-8b43-4c75-ac39-c02ce6f43d03 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference L., Bhagavatula, C., and Choi, Y
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67011f70-86fa-4438-b248-fe500a267bcb · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Q., Tay, Y., and Metzler, D
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06e44b8d-493f-4b07-9890-2f66f876fd32 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference A deeper look at depth pruning of LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2bb6b72-cc2a-47c0-9b84-e450ce09f459 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32042cfd-8cb9-4108-b8ba-e4bf8c82d96f · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8970a42-2ccc-4f37-8dd1-ff8b0dd3fc6e · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Z., and Liu, Z
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fc4520c-285e-49a0-b59a-02cabd591125 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Razorattention: Efficient kv cache compression through retrieval heads
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3846996a-fad5-4ef1-9520-6dc86ad22fcc · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., Ting, D
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4533b7cf-cd02-4b68-9921-77c3bb68dc7f · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd2fa558-17d1-4cef-9907-f074a86bf813 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference BloombergGPT: A Large Language Model for Finance
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2877ea30-5aea-4eb7-8810-eddfceecbbbc · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference T., Peng, R., Wu, Q., and Wang, C
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13964cca-0020-4437-8b2f-7895d5678333 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Duoattention: Efficient long-context llm inference with retrieval and streaming heads
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8527fb09-f508-40c1-a700-609a7a9bc288 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Efficient streaming language models with attention sinks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ee6e21e-26f5-49f1-a821-95061880a634 · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f2a254e-64b3-4b64-a42b-e0c7054921df · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae04f335-09c4-4713-9b2e-f61df35db73e · outbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4040624e-1a52-449c-ad62-3e59d3d9a76d · inbound
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b256f7d-5bfb-4f70-a411-d8e74ae2a03f · inbound
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4d1dbc0-a3f4-46bd-87cd-26436f2198d6 · inbound
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df018183-1fe6-4e00-831b-dc59de7a9fd1 · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e37ca751-0043-4a4b-9ad3-43f1489a797d · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.