Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2312.11514.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:26:28.950463Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:29:50.698429Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0fc09583-7952-4354-8218-5fc146502e6d · inbound
Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 275
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4a1720d8-572a-4437-9210-a5837ea39253 · inbound
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e51fb764-e5d8-40c2-a466-1724c0b938b0 · inbound
Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1733bd0-7572-426d-870d-4a6a60f7c7ac · inbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f94a630-6bbb-421f-aaf1-0c5ef7562a01 · inbound
Mixture of Hidden-Dimensions Transformer LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14fe271-994e-4c4a-ae18-d8aff161c048 · inbound
Post-Training Statistical Calibration for Higher Activation Sparsity LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834f1932-cb9d-4cf3-a64b-1056f4291f18 · inbound
Creating an LLM-based AI-agent: A high-level methodology towards enhancing LLMs with APIs LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5fadc7-b381-441b-b00e-acea44634311 · inbound
Deploying Foundation Model Powered Agent Services: A Survey LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8de6fe-2d1e-4370-b559-347890025dc0 · inbound
Accelerating Retrieval-Augmented Generation LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c233a4-120d-470f-b98c-b007632f82a6 · inbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73bab42-7498-48f5-bf9d-7c33882b4062 · inbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff554477-e8fb-4aa2-b6a8-5dbdbeff8b8f · inbound
Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363fa428-9e58-4982-93c0-61d1468f4cde · inbound
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21b22ec-6027-4945-98ae-68596982659f · inbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 999649e1-e61e-4631-9e97-150b5e4bd1ea · inbound
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec61dee9-143f-4453-b8f1-6702b357150e · inbound
Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c460a86-acb3-40fa-85a6-8eafae05757b · inbound
AVEC: Bootstrapping Privacy for Local LLMs LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896e74e0-c3b5-46cd-93fe-e3973c893988 · inbound
Technology solutions targeting the performance of gen-AI inference in resource constrained platforms LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 85f463d2-bd83-4176-971f-db403139b2c7 · inbound
ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f216b5e0-9a4a-4794-a5c0-b92ed0323eb7 · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 99a6258b-ad0c-4ee1-b3eb-a4de5b1f90f0 · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1102b2-94f2-4125-914d-f382c280d540 · inbound
EnerInfer: Energy-Aware On-Device LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b042df67-8388-4494-b6f0-8903e4813f16 · inbound
Transition-Aware Backend Dispatch for Edge LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3eedeb-cef1-4533-871f-cac10850f96a · inbound
A CXL Memory Rack for Multi-Turn LLM Serving LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cb3c67c-0732-4b33-b3fe-af6857632992 · inbound
Architectural Implications of Agentic AI Workflows LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79336d5a-5211-4888-896e-006ed9be5693 · inbound
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.