Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:43:29.577850Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2501.18107.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:43:29.577850Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:40.215851Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T08:15:31.617791Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 08438660-135f-4f83-8318-11db118c05cf · outbound
Scaling Inference-Efficient Language Models Results over vLLM In this section, we first evaluate the inference efficiency of open-source large language models over vLLM using NVIDIA Tesla A100 Ampere 40 GB GPU
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b1402f3-23ec-4a99-b4c5-b2d7da5810f7 · outbound
Scaling Inference-Efficient Language Models Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99b4b80-e0aa-44e4-bb15-535bb83e9156 · outbound
Scaling Inference-Efficient Language Models dmodel is the hidden size, fsize is the intermediate size,n layers is the number of layers, andn heads is the number of attention heads
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f8420ad8-50cf-478b-8687-4f560a193fe0 · outbound
Scaling Inference-Efficient Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c22c8c-5962-4fbf-8814-178c017b4dad · outbound
Scaling Inference-Efficient Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41086873-bfa6-4052-9826-7283fb07b62a · outbound
Scaling Inference-Efficient Language Models The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9acfc0b0-95f6-439f-97a1-f5df976f7a35 · outbound
Scaling Inference-Efficient Language Models Language models scale reliably with over-training and on downstream tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a484819-c134-4763-9cce-a00b0300d799 · outbound
Scaling Inference-Efficient Language Models SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 034e27fd-824d-48e9-bf03-ea75755be632 · outbound
Scaling Inference-Efficient Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb9179df-6538-4e28-954a-eebf8fb915fa · outbound
Scaling Inference-Efficient Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7d11a5-4b1c-45e8-946d-01b54697d36b · outbound
Scaling Inference-Efficient Language Models Measuring Massive Multitask Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53c4378-4f48-474a-af51-15bfe863e5a9 · outbound
Scaling Inference-Efficient Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af18da9b-6d42-4283-9145-de574b0b3f52 · outbound
Scaling Inference-Efficient Language Models Training Compute-Optimal Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d7b0e2-73b2-4a48-a3b9-f20cc31d373c · outbound
Scaling Inference-Efficient Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb28230-aa94-400e-806e-1f4f1ac094ca · outbound
Scaling Inference-Efficient Language Models OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2582b16-804f-4781-8624-a8375140aac3 · outbound
Scaling Inference-Efficient Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79f3d96-6d66-44fc-abbb-039d5df1ac2a · outbound
Scaling Inference-Efficient Language Models Scaling Laws for Neural Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743d0cea-2e11-4eb2-99ef-552c3de26c33 · outbound
Scaling Inference-Efficient Language Models Scaling Laws for Fine-Grained Mixture of Experts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666f646f-b6b1-47c6-95dd-a7f66d8fcb67 · outbound
Scaling Inference-Efficient Language Models Scaling Laws for Precision
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 779028af-5f12-49e8-b0c4-84a3dd7a88cc · outbound
Scaling Inference-Efficient Language Models DataComp-LM: In search of the next generation of training sets for language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dffec26-ad04-42ff-8152-7cac27223013 · outbound
Scaling Inference-Efficient Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d390137-7835-4a37-bb3a-ab1eb0fed272 · outbound
Scaling Inference-Efficient Language Models MLC-LLM, 2023-2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c39a3c32-edfa-4498-a651-cb7a50cd9876 · outbound
Scaling Inference-Efficient Language Models TensorFlow-Serving: Flexible, High-Performance ML Serving
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03d123a-ee22-4941-be21-0958ae264f8a · outbound
Scaling Inference-Efficient Language Models Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b19d0a3-c7a2-490a-9d9e-b3bd645aac06 · outbound
Scaling Inference-Efficient Language Models Fast Transformer Decoding: One Write-Head is All You Need
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8097040e-2e6c-40b2-a1bd-28c807b7bda2 · outbound
Scaling Inference-Efficient Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e47aa4-c5d0-4a68-944d-b62ee373bee7 · outbound
Scaling Inference-Efficient Language Models Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc59e599-d3da-4b64-9257-7944773488cd · outbound
Scaling Inference-Efficient Language Models Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bfdace5-9950-45a0-85f6-edef190d15f8 · outbound
Scaling Inference-Efficient Language Models Gemma: Open Models Based on Gemini Research and Technology
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a995995-f6c7-451f-a4e0-6671d0b5e10b · outbound
Scaling Inference-Efficient Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320db702-f16d-48ef-9b86-197de543fb9e · outbound
Scaling Inference-Efficient Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8361014-a9c4-4b52-ae56-22eb6041ae8c · outbound
Scaling Inference-Efficient Language Models DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f4d5c7-acd7-4733-ac4d-20ae0c18214a · outbound
Scaling Inference-Efficient Language Models Decoding Speculative Decoding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130af53b-c465-461f-ad57-d562fce51fa9 · outbound
Scaling Inference-Efficient Language Models Qwen2.5 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edbb00f-43ea-488c-9d79-e57a745fb3d2 · outbound
Scaling Inference-Efficient Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71770aef-bfb3-4c91-8934-b4b558f885c0 · outbound
Scaling Inference-Efficient Language Models FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e502ea8-cd4f-44bb-a91c-8aacd2cf5bc5 · outbound
Scaling Inference-Efficient Language Models Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bc4bce-51e1-4033-a47d-bff38ca9befd · outbound
Scaling Inference-Efficient Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36b0fb5-36bf-46ce-bafd-12008f918a7d · outbound
Scaling Inference-Efficient Language Models OPT: Open Pre-trained Transformer Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9809848f-786e-40cc-b902-b7d36b2a17df · outbound
Scaling Inference-Efficient Language Models (Center) We indicate the relationship between inference latency and hidden size with the number of layers fixed
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 92eb9d33-d361-4bba-9201-2c641ead1be5 · outbound
Scaling Inference-Efficient Language Models The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024)
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0cfc69b-3642-40db-ac9d-47249834c10f · outbound
Scaling Inference-Efficient Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b000cb5-25d9-4ce2-bc4c-40069c0d50c6 · outbound
Scaling Inference-Efficient Language Models Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a6da3976-01ab-456d-9366-3039c7b87aaa · outbound
Scaling Inference-Efficient Language Models Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 459ddbfe-1a6c-4f01-bc00-a4463a9c4683 · outbound
Scaling Inference-Efficient Language Models Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e805e0bf-72d8-4660-b255-d504e1315cfe · outbound
Scaling Inference-Efficient Language Models Unresolved cited work
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6266e32d-d3a6-4a57-b7b7-76549fca7d7b · outbound
Scaling Inference-Efficient Language Models Efficient Streaming Language Models with Attention Sinks
Reference 1921
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6e98a8-048f-41af-abf5-fa1949c0175f · outbound
Scaling Inference-Efficient Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 1961
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72478f8f-797d-4a77-8054-4166b6a683fe · outbound
Scaling Inference-Efficient Language Models Observational Scaling Laws and the Predictability of Language Model Performance
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f881308-29a5-4563-ba0e-b23af1b6e46f · outbound
Scaling Inference-Efficient Language Models Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d96b4e-56c0-4bf1-98d8-0e3684f6fc53 · outbound
Scaling Inference-Efficient Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fc260f-b242-46db-99aa-ed15d558a6b6 · outbound
Scaling Inference-Efficient Language Models Training Verifiers to Solve Math Word Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3130bf6-75bc-4e61-824d-6417abde6ec9 · outbound
Scaling Inference-Efficient Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fce7d57-3670-412a-b473-88a41959fe5e · outbound
Scaling Inference-Efficient Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b70a2b-b9ac-4d21-ad99-c0df5c2c6344 · outbound
Scaling Inference-Efficient Language Models The Llama 3 Herd of Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 764af2e5-a84a-46ec-ab84-f40ab30dc6e6 · outbound
Scaling Inference-Efficient Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3c5454-eb55-4b54-8778-c051d3ff73bf · outbound
Scaling Inference-Efficient Language Models Qwen Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615efc87-a32f-4aea-8c20-8bfd1f41ff8c · outbound
Scaling Inference-Efficient Language Models Phi-4 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9965cdfd-08b6-47ba-9256-79a8b6f2a577 · outbound
Scaling Inference-Efficient Language Models CHAI: Clustered Head Attention for Efficient LLM Inference
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e4fcd9c-f716-4138-b306-c9f8e6b8b30e · outbound
Scaling Inference-Efficient Language Models Unresolved cited work
Reference 2048
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49760f12-0291-4357-a221-0caf1b0379f9 · inbound
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search Scaling Inference-Efficient Language Models
Reference 174
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b109582b-e94e-4858-853b-129ad6b632fc · inbound
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Inference-Efficient Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad67b25a-8e8c-41f2-9975-9de5ae1c9a41 · inbound
Comprehensive AI governance requires addressing non-model gains Scaling Inference-Efficient Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.