Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-07T12:51:42.916410Z
Paper Citation Record · LEDGER
As of 23 July 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2604.26557.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-07T12:51:42.916410Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 970792c5-a738-409a-9064-a5ae51dc645e · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference A review on edge large language models: Design, execution, and applications
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c965984f-a1ad-4e57-a157-162a1cb06049 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference A cost-benefit analysis of on-premise large language model deployment: Breaking even with commercial llm services
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a1a76803-32f5-49ac-80d3-f628cebb426b · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Mobilellm: Optimizing sub- billion parameter language models for on-device use cases
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4062ded4-711d-4bce-9eec-330733d03515 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Intelligent data analysis in edge com- puting with large language models: applications, challenges, and future directions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8e0d5816-a7a6-4150-a1b8-fdf121bbd6b7 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference CUDA Programming Guide: Unified and system mem- ory
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e291abe5-5b35-4b7e-8632-c55e2bccfee0 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Gemma: Open Models Based on Gemini Research and Technology
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 07c1500e-e8c4-43e7-9dd3-c1f0bafc4a74 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Mistral 7B
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 533b9f04-6e59-4be0-ad72-2d35bf5d5362 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c5bdda9d-dcf2-4b01-987c-f77ddc87440a · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Longreward: Improving long-context large language models with ai feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d3203377-fd91-4f8f-a5c5-e1cf4399bf01 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Longbench: A bilingual, multitask benchmark for long context understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2eaf2ecb-bfb9-4d1a-8d40-a4c5fdc1d73f · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Visual instruction tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6c880d7d-283e-4f27-b180-0826a8a0e2a2 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Logparser-llm: Advancing efficient log parsing with large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e62b4a2f-4854-40bc-bc4a-af8a0e903e06 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 28a86021-209b-497e-9c4f-09325b8bb454 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 72fbe47a-cf7f-4b87-b2d7-43e486ba1424 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Attention is all you need
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e0c86885-4aa6-4a0a-b7e9-a5f12923d31d · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Infinigen: Efficient generative inference of large language models with dynamic kv cache manage- ment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5099a91b-04ec-401b-b1d9-e180df3497e6 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference llama.cpp: Llm inference in c/c++
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8a0e41bc-7d9d-4d9f-9699-3baf8cdd2671 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Lmcache: An efficient kv cache layer for enterprise-scale llm inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b97e96bc-ce18-42c5-a622-a10245d5786b · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Powerinfer: Fast large language model serving with a consumer-grade gpu
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 43f22c5d-b873-4d8e-ae59-6d1f52e518b7 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0e4f2afc-c9b7-4b6f-b6fe-f402d2781efe · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Llm in a flash: Efficient large language model inference with limited memory
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 939ef159-377f-4b61-9faa-77c20137241a · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Designing a true direct-access file system with devfs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8bbe7533-8e6a-4322-a812-bb99ce68b502 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Storage performance development kit (spdk)
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6830ef55-1c88-42db-9a7c-19450044ef1b · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 44b10cd0-9821-4a0d-9bff-1934176621b8 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Flashshare: Punching through server storage stack from kernel to firmware for ultra-low latency ssds
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 73d94845-dcd5-4f2f-8090-c32daabd4c3c · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference I/o passthru: Upstreaming a flexible and efficient i/o path in linux
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4efa9446-4204-4693-bcaa-bc9ec6c7fbf6 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Opt-6.7b
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 996224ef-2f4d-4422-befd-89d09c9d80fa · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference X3: A low overhead high performance buffer management replacement algorithm
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9136f0ee-5ef5-433e-8a51-12a2cfc4c3a5 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Arc: A self-tuning, low overhead replacement cache
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 823641a5-b407-4aa0-9775-2d93fa1dc916 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Streamcache: Revisiting page cache for file scanning on fast storage devices
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1101bdff-c005-4abb-841a-2785741d0ed2 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Gregg,BPF performance tools
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 92217bb7-fffe-47ce-af08-9b78c936aa74 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Asynchronous i/o stack: A low-latency kernel i/o stack for ultra-low latency ssds
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b456fee8-6ce8-41f0-9697-ccb835151936 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference D2FQ:Device-Direct Fair Queueing for NVMe SSDs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8de83568-eb69-4f4b-ab94-ba1fef7131f5 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference iJournaling:Fine-Grained journaling for improving the latency of fsync system call
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b9e57702-adab-4729-9b91-61d8886df1c7 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Asynchronous i/o support in linux 2.5
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e0838906-55df-4c7f-be66-30944e20dcfa · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Do we still need io schedulers for low-latency disks?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0b22dbcb-5396-4a6e-a86f-7174feff4de9 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Bfq, multiqueue- deadline, or kyber? performance characterization of linux storage sched- ulers in the nvme era
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3b336d32-789a-4534-83fd-984717ec7570 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference OPT: Open Pre-trained Transformer Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 96561b3a-073a-4357-a418-d001abd9cb8a · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Cuda c++ programming guide
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 887e8ffd-7f93-45ca-b673-6024568ad8a4 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Can Foundation Models Wrangle Your Data?
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5147f533-fb3b-4f48-aaf1-d4bb6051c84d · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Cost-efficient large language model serving for multi-turn conversations with cachedattention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 73247e4d-6233-41a7-8fba-e6d1a616e7d6 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference An i/o characterizing study of offloading llm models and kv caches to nvme ssd
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d7d8662d-d962-417a-9f9e-fac135a34a56 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f1eb7498-279d-430f-899d-d14b3403aa0a · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference INF2: High- throughput generative inference of large language models using near- storage processing
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 76222427-6bc4-4968-b5bc-9989dc9a59dd · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Efficient memory management for large language model serving with pagedattention
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 55b4c814-6c56-4d6a-9f28-fee673c70855 · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference GPUDirect Storage: A Direct Path Between Storage and GPU Memory
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 143753e3-c563-4e07-826e-eeef2a96793d · outbound
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Gpu- initiated on-demand high-throughput storage access in the bam system architecture
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
No inbound Pith citation observations are available.