Pith. sign in

Paper Citation Record · LEDGER

ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2410.01228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01228 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T22:20:51.838309Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:27:06.032587Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 62162d4d-2118-4406-a6b9-2c6519d98415 · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:a808774caa53435f92d6309b6745b62b351b5caedca6a2beda70c006e28f8b2f

Observation 0f833576-156b-48df-8ba7-51561d1fe526 · inbound

The Energy Cost of Execution-Idle in GPU Clusters cites this paper.

The Energy Cost of Execution-Idle in GPU Clusters ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:51.593039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:04:25.951890Z digest=sha256:16225eb030af8561dfbcb316f6925c8007da34bbf0efe9a6a9b9e489d154f9a6

Observation 8023bd15-7971-4431-81cc-67daf1548ef8 · inbound

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start cites this paper.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.482759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:58c1f92079b7007040bc2a8294f397114f6cb5b83d5207d8ac38930a628ba8ce

Observation 2f976b28-39a3-4f1e-b366-726fe491f225 · inbound

Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC cites this paper.

Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:01.238853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:01:57.651581Z digest=sha256:d81e20b63d41dac92a266cf7136198a0ddbc6deb2c9926b553e57772b1311f69

Observation 8bff6817-1084-4762-9086-23cd2ab78642 · inbound

Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate cites this paper.

Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:31:00.663351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:05:08.438659Z digest=sha256:76a723283a7e77e4af5958c594c9f188501a5d1688fea2067ca0813b15bb5b88

Observation 3b16e5a1-23ba-4061-8a41-381ba605d4ce · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.389848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:6e4910b2bddfaff72fb13cbfdee6b2ce34cf892fb1c8376730540e6a9d9dd6fb

Observation dfbeedee-3cf8-47fd-8389-fc1a6f40e3ee · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.282201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:ce0149fd88c69cb08b141919a48d9c068af4a5978f2a4b0a1c4914f63494bfe3

Observation f51cf09a-edc0-4789-a698-73f684fd32a0 · inbound

RW-TTT: Batched Serving for Request-Owned Test-Time Training State cites this paper.

RW-TTT: Batched Serving for Request-Owned Test-Time Training State ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:53:28.852555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:46:48.954003Z digest=sha256:b14474c9ec75891af91793d8cbc4000b491293185550c60b9feb405c5083f191

Observation 63db9f70-be66-4694-a3be-7087fb5f7f77 · inbound

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference cites this paper.

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:27:06.034079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T23:49:28.318260Z digest=sha256:2c2ed213d0b0803bbfe1888ac55e3463c58450b642411aaa73401c3d6cf65db8