Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2312.12456.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:26:17.238480Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:39.594987Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b0f5232b-ef16-4177-bc9e-752bc21b198a · inbound
Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 274
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bfc00a5-a9a7-4688-9c6e-dacdc19a8d58 · inbound
A Survey on Efficient Inference for Large Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 251
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 025dc3cd-c52b-4e23-91cd-2d17303e6d9e · inbound
HybridFlow: A Flexible and Efficient RLHF Framework PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69b87f83-9793-4573-9514-e99b11547a90 · inbound
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8865263a-8f4e-44a8-b067-b61b5be100ea · inbound
DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5568e98f-9628-4a2d-b939-7a04aeaab1fb · inbound
Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aeb7da2-92c1-4501-bf32-4b3b6ad58c68 · inbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e860b4a-708f-43e3-aad1-8ed942bdf3a0 · inbound
Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2950b28d-ac6b-4d06-9bbf-fcdf52566cca · inbound
A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c185bc3-56b6-4eaa-a43f-8583afd93821 · inbound
Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f469afaf-17f4-40c8-abdf-4732a5164f8e · inbound
Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3e29c5b-c69b-4ebc-8251-d4c5128ee048 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0edf9054-3d03-4b0f-b0e9-96ef98c2bd5f · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 257da112-d13a-4153-bd07-b2d47146e05d · inbound
ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f38a8d6-31a3-4311-a2a7-6456ae94e70b · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef95162c-6e8d-4289-8c31-7a1c19ea98ea · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6698b789-2425-45fc-bdd7-130a55fb1c32 · inbound
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d75a5c-45f6-461b-82ac-cd4a75d5742f · inbound
Transition-Aware Backend Dispatch for Edge LLM Inference PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1d5176-b30e-48b7-8aa1-d750a2da225a · inbound
A CXL Memory Rack for Multi-Turn LLM Serving PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.