Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:06:35.243279Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2507.11953.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:06:35.243279Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ddce7f71-81e7-4bfd-a628-4677a6f16057 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27b0413-f088-48dd-80fc-b5682c8635b2 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2201d20-badd-46f4-9c6e-bbdf0975a507 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs bert2BERT: Towards Reusable Pretrained Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 78e2ff80-2383-4ed6-82c2-cf1fc53818fd · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031f1f96-1911-4e72-ac5b-db09b890c50e · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5aa031-7ccc-4838-b0ae-49e0fbad6755 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Better & Faster Large Language Models via Multi-token Prediction
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75bace4f-1661-45df-82a0-84828160eb0d · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Measuring Massive Multitask Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b5b577-79bb-4a38-8752-585dbd12bc91 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f539c6a-4fe4-4b54-b4d8-8b731cb81118 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8749c43c-2823-490c-ad32-138f9817c6e7 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Le, Yonghui Wu, and Zhifeng Chen
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1a6e80-93c7-4856-b09c-292eedb8e4cd · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747a2e47-ebf3-4362-a1bb-c91e9e816657 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 881c46c3-0aef-4ff4-abc5-f4f2e8855838 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9396c2-32b7-4d11-8262-2481c8fbf45c · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Pointer Sentinel Mixture Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f9fd941-db81-4ada-be82-f42b8d06fc5a · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs GPT-4 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122e322e-7d23-48bd-a4de-6696608a97a4 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 843ab966-ea9d-439f-a421-b05b56f66f4e · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Vicky Zhao, Lili Qiu, and Dongmei Zhang
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42bc596e-6c1b-4d5c-90ba-f02346f3abcd · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff05fd0-0beb-4fe2-b1de-a78e98b0afe0 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e3c357-9e37-4cdf-904e-c874b813e0a0 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86a8ddf-d8a4-4daf-ad24-362ef31dba0c · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Hashimoto
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd07ab1-e78e-4822-afa8-6174f15c2e8e · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98605f00-7b41-4240-9f92-d3f4cf046fa6 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 024ec14d-13af-4286-88b8-edf2bc45f38d · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f06354c-eacc-4bd2-817f-72b6ed71ebb3 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Efficient Streaming Language Models with Attention Sinks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af505138-7ac9-42ba-afbb-1f5133866f8c · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Qwen2 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264d52de-0845-4177-9c5a-82d06d16a1d6 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b483d5-96a7-42fe-a269-73873527ec8c · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0c4ed0-b0fe-47a0-abd3-c9bf3479c008 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ce1964-0ede-499d-aab8-8548b5c3b039 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d0917189-3a92-4a68-a37d-d880cde302e5 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs online" 'onlinestring :=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09201d63-7bdc-4812-8495-6600582b6ce4 · outbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs write newline
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.