Pith. sign in

Paper Citation Record · LEDGER

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

As of 23 July 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2604.26557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.26557 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T12:51:42.916410Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact9
  • verified fuzzy37
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 970792c5-a738-409a-9064-a5ae51dc645e · outbound

This paper cites A review on edge large language models: Design, execution, and applications.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference A review on edge large language models: Design, execution, and applications

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.794640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:986d5424b7b232e0eab86c06f27735896b4c5be49216aaf84413e0d28baff3e4

Observation c965984f-a1ad-4e57-a157-162a1cb06049 · outbound

This paper cites A cost-benefit analysis of on-premise large language model deployment: Breaking even with commercial llm services.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference A cost-benefit analysis of on-premise large language model deployment: Breaking even with commercial llm services

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.539616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:bd0e344f4951f44411e2a9fedab9195b8096eddf7ff88992ecb51b038cf2c97a

Observation a1a76803-32f5-49ac-80d3-f628cebb426b · outbound

This paper cites Mobilellm: Optimizing sub- billion parameter language models for on-device use cases.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Mobilellm: Optimizing sub- billion parameter language models for on-device use cases

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.807121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:2775b546018953de6f06121de64162531dd06e3a096eb663ad9faf14f40aaa67

Observation 4062ded4-711d-4bce-9eec-330733d03515 · outbound

This paper cites Intelligent data analysis in edge com- puting with large language models: applications, challenges, and future directions.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Intelligent data analysis in edge com- puting with large language models: applications, challenges, and future directions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.849755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:28c5051c29304fb505abc696147f9e280bd147780f902f5866ebe0ba8db1de96

Observation 8e0d5816-a7a6-4150-a1b8-fdf121bbd6b7 · outbound

This paper cites CUDA Programming Guide: Unified and system mem- ory.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference CUDA Programming Guide: Unified and system mem- ory

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.812880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:101b3395d6558f00e4259f70066441d440f474c97a4c2aa844f63914bcdd509a

Observation e291abe5-5b35-4b7e-8632-c55e2bccfee0 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:06:26.510120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:82051bae5959422caaa08d11502c98af31ebb89fc28c2d6500038619bb8d68c6

Observation 07c1500e-e8c4-43e7-9dd3-c1f0bafc4a74 · outbound

This paper cites Mistral 7B.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Mistral 7B

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:06:26.529461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:831bd9ace39b555d72501c44843a4136a5460f43f728f14e51f4c5ecdeaa8709

Observation 533b9f04-6e59-4be0-ad72-2d35bf5d5362 · outbound

This paper cites The Llama 3 Herd of Models.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference The Llama 3 Herd of Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:06:26.534152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:f78ea568c6fc6f78ea5f57888f06ff0832d27e7461da7b5e171172882a013e97

Observation c5bdda9d-dcf2-4b01-987c-f77ddc87440a · outbound

This paper cites Longreward: Improving long-context large language models with ai feedback.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Longreward: Improving long-context large language models with ai feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.789106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:26fade05173ffd816963b86b7cade66b62ce82c4ecb1f865f4f04ba12115305b

Observation d3203377-fd91-4f8f-a5c5-e1cf4399bf01 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Longbench: A bilingual, multitask benchmark for long context understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.791735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:2b1dff50816667e107caba46b6eb83ac6ee343a949af5c13fb575717664fb762

Observation 2eaf2ecb-bfb9-4d1a-8d40-a4c5fdc1d73f · outbound

This paper cites Visual instruction tuning.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Visual instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.797731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:a5cd849b96ce3bb38e51e6265af3977068fea79ea077cd548d6d32b518625946

Observation 6c880d7d-283e-4f27-b180-0826a8a0e2a2 · outbound

This paper cites Logparser-llm: Advancing efficient log parsing with large language models.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Logparser-llm: Advancing efficient log parsing with large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.800615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:5badce1b7bfce05bfa6069e38c5c782b0bd15dad260d8424d913638fca98f904

Observation e62b4a2f-4854-40bc-bc4a-af8a0e903e06 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.786627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:f52c877565f2fa852d650d214a67a4436d9c4c39f59c584cf322592e8a7c8dec

Observation 28a86021-209b-497e-9c4f-09325b8bb454 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.784026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:dd4053b68856428f8d08254e386e5f928b53fbfaddc5aac6540298b1673133ac

Observation 72fbe47a-cf7f-4b87-b2d7-43e486ba1424 · outbound

This paper cites Attention is all you need.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Attention is all you need

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.844599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:52e1cae7272ff9e138909566c7f01e5d9d726022a52bee0450c2f770ee43f78f

Observation e0c86885-4aa6-4a0a-b7e9-a5f12923d31d · outbound

This paper cites Infinigen: Efficient generative inference of large language models with dynamic kv cache manage- ment.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Infinigen: Efficient generative inference of large language models with dynamic kv cache manage- ment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.818998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:7a9513e0cbab26890f4c81ea545c9770c7d6ffe65e1a78a10c861eab19e543e3

Observation 5099a91b-04ec-401b-b1d9-e180df3497e6 · outbound

This paper cites llama.cpp: Llm inference in c/c++.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference llama.cpp: Llm inference in c/c++

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.815558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:e6e0e4234a215404ca405ba3408c0cfbe0053273198f53ff23d0ad0b775fe961

Observation 8a0e41bc-7d9d-4d9f-9699-3baf8cdd2671 · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.545204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:b292e0bba4fab0dc835baff91e693409f3095d35316a49958344d752472d62b6

Observation b97e96bc-ce18-42c5-a622-a10245d5786b · outbound

This paper cites Powerinfer: Fast large language model serving with a consumer-grade gpu.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Powerinfer: Fast large language model serving with a consumer-grade gpu

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.847092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:51ce40da1ec71c04e823facc7298d8c204d1b13c03e01b65fb97bf75b659c3d4

Observation 43f22c5d-b873-4d8e-ae59-6d1f52e518b7 · outbound

This paper cites KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.515079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:12ed3a69048047772810b0b57ed123051367e11a7206d69402f1cf5cfc651828

Observation 0e4f2afc-c9b7-4b6f-b6fe-f402d2781efe · outbound

This paper cites Llm in a flash: Efficient large language model inference with limited memory.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Llm in a flash: Efficient large language model inference with limited memory

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.841912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:1faf5d4495fb31871748a81d51604be9f0241dc3515f3a9bbd36057485034f3b

Observation 939ef159-377f-4b61-9faa-77c20137241a · outbound

This paper cites Designing a true direct-access file system with devfs.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Designing a true direct-access file system with devfs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.769716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:49df1f16fa9f8e62d581e78b18ce9624579b33f070de480b742acd4a2ec41ac9

Observation 8bbe7533-8e6a-4322-a812-bb99ce68b502 · outbound

This paper cites Storage performance development kit (spdk).

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Storage performance development kit (spdk)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.766803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:2909223a9ac189b8e9e02b69aeee5b483a6ad6340af3ffa33305455c727f1eb3

Observation 6830ef55-1c88-42db-9a7c-19450044ef1b · outbound

This paper cites an unresolved cited work.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Unresolved cited work

Reference 24

Resolution
parse uncertain
raw_fallback, observed 2026-05-27T06:18:52.775565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:c72abe521549d87423ba3e382a1120258a9bdd895c83319e3b12f88f5b300bf7

Observation 44b10cd0-9821-4a0d-9bff-1934176621b8 · outbound

This paper cites Flashshare: Punching through server storage stack from kernel to firmware for ultra-low latency ssds.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Flashshare: Punching through server storage stack from kernel to firmware for ultra-low latency ssds

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.760390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:ba88b40e375c63a0a4171f3a8cd86b69317df6d3ce46c44ef13d2c0534f9c8b4

Observation 73d94845-dcd5-4f2f-8090-c32daabd4c3c · outbound

This paper cites I/o passthru: Upstreaming a flexible and efficient i/o path in linux.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference I/o passthru: Upstreaming a flexible and efficient i/o path in linux

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.754131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:a5657fd2b257fa9c7ded15c8def55d72636605eb461006e79331055135fd6f50

Observation 4efa9446-4204-4693-bcaa-bc9ec6c7fbf6 · outbound

This paper cites Opt-6.7b.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Opt-6.7b

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.748372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:6f68d9c699b272fb51c1222dada7b23203495ae4f082d23af3229ad86274bf41

Observation 996224ef-2f4d-4422-befd-89d09c9d80fa · outbound

This paper cites X3: A low overhead high performance buffer management replacement algorithm.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference X3: A low overhead high performance buffer management replacement algorithm

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.772838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:31ff9889c3bb727c66353eaf15f917f5bf920bdfcf6186af2982606def1326d0

Observation 9136f0ee-5ef5-433e-8a51-12a2cfc4c3a5 · outbound

This paper cites Arc: A self-tuning, low overhead replacement cache.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Arc: A self-tuning, low overhead replacement cache

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.745700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:24d233ee2f0b03652e6744e2fc3b1b6e3575cbd08de6fd9a62fecb96fd347952

Observation 823641a5-b407-4aa0-9775-2d93fa1dc916 · outbound

This paper cites Streamcache: Revisiting page cache for file scanning on fast storage devices.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Streamcache: Revisiting page cache for file scanning on fast storage devices

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.751246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:8a5ff60fc77a80220cd256781edd922fd2cf66ae55375a1d76c1dbfb1a74762f

Observation 1101bdff-c005-4abb-841a-2785741d0ed2 · outbound

This paper cites Gregg,BPF performance tools.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Gregg,BPF performance tools

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.756947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:f038f053b11f13d95ce23bfb3a0385c93d76d32dfdaf0ec55aec5097da276fee

Observation 92217bb7-fffe-47ce-af08-9b78c936aa74 · outbound

This paper cites Asynchronous i/o stack: A low-latency kernel i/o stack for ultra-low latency ssds.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Asynchronous i/o stack: A low-latency kernel i/o stack for ultra-low latency ssds

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.763199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:d567fd590b520233fe6fc123282f6f440f0835fd713b91c70174dcbcf526f4cc

Observation b456fee8-6ce8-41f0-9697-ccb835151936 · outbound

This paper cites D2FQ:Device-Direct Fair Queueing for NVMe SSDs.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference D2FQ:Device-Direct Fair Queueing for NVMe SSDs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.778401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:44aecc1f52396e55ecd293a62e99aaa9c73c2cac73d0a01942004f726610e36d

Observation 8de83568-eb69-4f4b-ab94-ba1fef7131f5 · outbound

This paper cites iJournaling:Fine-Grained journaling for improving the latency of fsync system call.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference iJournaling:Fine-Grained journaling for improving the latency of fsync system call

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.828454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:09c0796e06c298414acb4e449629f7b6e35b62547202c43110c0fab61e9f6ec3

Observation b9e57702-adab-4729-9b91-61d8886df1c7 · outbound

This paper cites Asynchronous i/o support in linux 2.5.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Asynchronous i/o support in linux 2.5

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.821871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:aecd11211a4944cea97c061c3629bbf321d01772fceea199af67410458f763aa

Observation e0838906-55df-4c7f-be66-30944e20dcfa · outbound

This paper cites Do we still need io schedulers for low-latency disks?.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Do we still need io schedulers for low-latency disks?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.839154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:5836685568d3941479d143352815bcbdb15c655aca8e64dd8789097cf8644ccf

Observation 0b22dbcb-5396-4a6e-a86f-7174feff4de9 · outbound

This paper cites Bfq, multiqueue- deadline, or kyber? performance characterization of linux storage sched- ulers in the nvme era.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Bfq, multiqueue- deadline, or kyber? performance characterization of linux storage sched- ulers in the nvme era

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.836740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:a6d4f8a513fc9918b606ef758b5e3c5bdbd1f3f2c6581949861d8e63d9d4b71b

Observation 3b336d32-789a-4534-83fd-984717ec7570 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:06:26.505328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:dec6e169096d8863be7a617f07f03889d1bd7bd66f5cb52c95e37da1d608bb0f

Observation 96561b3a-073a-4357-a418-d001abd9cb8a · outbound

This paper cites Cuda c++ programming guide.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Cuda c++ programming guide

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.831054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:161ea6307aeee1912252aea3a89af53e88a3bf93ec2d21373630d645d7217e75

Observation 887e8ffd-7f93-45ca-b673-6024568ad8a4 · outbound

This paper cites Can Foundation Models Wrangle Your Data?.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Can Foundation Models Wrangle Your Data?

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.520020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:658483e2b5d519aa9f22a0674857a368c359bd77789b9a91b39046b3667a0f59

Observation 5147f533-fb3b-4f48-aaf1-d4bb6051c84d · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with cachedattention.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Cost-efficient large language model serving for multi-turn conversations with cachedattention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.781127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:bd18873e079d6f6294a3769e74fde9e1fe7eace238830347ccfa04313593d421

Observation 73247e4d-6233-41a7-8fba-e6d1a616e7d6 · outbound

This paper cites An i/o characterizing study of offloading llm models and kv caches to nvme ssd.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference An i/o characterizing study of offloading llm models and kv caches to nvme ssd

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.825230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:aaf2cbc023ce8c5d47e837c905ccfdfc73a56e5f78195d655c7869432605d882

Observation d7d8662d-d962-417a-9f9e-fac135a34a56 · outbound

This paper cites InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.525019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:c11095767155f4f883d6a1a7d33e7de9df81fccf9c58b287728f7595ef0ea13b

Observation f1eb7498-279d-430f-899d-d14b3403aa0a · outbound

This paper cites INF2: High- throughput generative inference of large language models using near- storage processing.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference INF2: High- throughput generative inference of large language models using near- storage processing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.803707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:c1dc9ddaa1d7cd4e052673fd4410f1b70e1937f646257f23dca5035104f5c1e6

Observation 76222427-6bc4-4968-b5bc-9989dc9a59dd · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Efficient memory management for large language model serving with pagedattention

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.810087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:3246cd51af167ed6c209cdb60535fa3ad83273fa85c74c90b4bca6fc4c39e629

Observation 55b4c814-6c56-4d6a-9f28-fee673c70855 · outbound

This paper cites GPUDirect Storage: A Direct Path Between Storage and GPU Memory.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference GPUDirect Storage: A Direct Path Between Storage and GPU Memory

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.852734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:b60e494b105567d07b1fe9c0e909e42d28f3441d535121a341e5d390abe93356

Observation 143753e3-c563-4e07-826e-eeef2a96793d · outbound

This paper cites Gpu- initiated on-demand high-throughput storage access in the bam system architecture.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference Gpu- initiated on-demand high-throughput storage access in the bam system architecture

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T06:18:52.833824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:990e7c37834fc4627a4679aa26a8406594e069efd63498a36044fb9cbb35142b

Pith citing papers

No inbound Pith citation observations are available.