Pith. sign in

Paper Citation Record · LEDGER

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing

As of 3 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2604.18529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18529 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T02:48:18.330339Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:09.632043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact14
  • verified fuzzy7
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62b20ab4-46a3-40fa-861d-1e253f799089 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.117518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:ef5f588612744ef9ba54e0efe55966f46047cfb56d33ae2f332f2f6808236315

Observation fd8c002f-7794-4580-8d2f-66e13e415024 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.120330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:e64c7ce3976be48268492ee4f3cdcfb331fb38310027af86b41554aad24c2a7d

Observation ec778ea7-a6a6-43fa-9307-44b1d5cadad4 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:44:50.211576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:f46000d6fec0bb992898082462f22a2d3d7f673e4017907156c46edada8b843c

Observation 016aa677-6fac-4a7c-9cf6-8e4825e56472 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.170923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:7a0f59934892ab75430bf530e835e64142f8fb914333454b32e12c4ac65ebc0c

Observation f1641f35-4272-4b6b-b1c2-61c12af56784 · outbound

This paper cites InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(Rotterdam, Netherlands)(ASPLOS ’25).

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(Rotterdam, Netherlands)(ASPLOS ’25)

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:26.439520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:e2d3a774cd524cdeb7a3b3475ac98001f3acf917ba0767c2e132c458aa777ed6

Observation 59921c60-abf9-4b0d-bba5-f0aeb786217a · outbound

This paper cites Accessed November 2025.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Accessed November 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:22:00.139215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:2fae110650c60798fc7ba41081a3656da48ebfdcf72a944aa575fe27517999c3

Observation ac04a6fe-30bb-47a7-a0f8-22093f9bc7a9 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:25:04.577946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:f0bb67dc423bc41d2cc0bf44bb77f2ece263006a6f942c984197f3c58d7b7965

Observation f8d7267a-7304-4f84-a113-e6452afaff3c · outbound

This paper cites PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.073874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:6b935185474420de3e6e5fd2916e6aa4d4ac2d8f1d4a8df1449b4af518f51a9e

Observation fcca7652-81b8-4f7a-93df-e7b63bf2dd9d · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.215835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:7dcc4af19686df22a15325d04a3efe8a8de904d111252618159f797aa2a4d9ca

Observation e1d3ab66-5641-4e35-8342-dfea167095bb · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:ef6d22b38a31506d4b6b9f1b8d18c95890f88b1930a6245f69afc4ed84146628

Observation 323bb551-8f17-4304-aab3-5d2798c161d7 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.136036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:ff2ef8700a8c1b7083b9553ff69eccc5a2199f23935a7ee5c40d23c95c9c2b6b

Observation ff7deec7-ca88-4bf7-a39c-3b413d63759b · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:11:21.895994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:149964b7c3b49e5c0496080059c92a44a339a28edba93eddae02d6b33b40725b

Observation f106ff4f-bcb9-4f75-9f4e-4a6a431f32bb · outbound

This paper cites Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.080880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:cb148c236f21af3ff6574f681fb34e6f6cc24d16b35707c35cd0dae22c55da66

Observation 0f1fb9a6-851b-4c62-96f1-32b588ba5bf6 · outbound

This paper cites The Llama 3 Herd of Models.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing The Llama 3 Herd of Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-10T02:48:27.078626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:509a342216ec3a3a687c6a72ff0777aa393739cc103db7519b823da6c9bc5987

Observation 882b027b-c9f7-41bf-a994-039e64c4a6e9 · outbound

This paper cites Accessed November 2025.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Accessed November 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:25:04.575787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:ab7df9a91876262d8df5545928b58f5b6932ddded01b564d8e871f06f8fba47b

Observation 7de55fd5-b176-40e1-8d20-47371fba73f5 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:27.071417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:062869bfbea706f587a239f0d27e3eca0757043c0c1a7f3fa39c6889866f790a

Observation e5feed70-adc8-4ebb-80cc-2e99a763f8f1 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.132949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:9fc7acf845e15837ef7be6e02c18895dc9b8e99040a5351b870a8207e1a9f718

Observation d9530f04-112c-4495-adc2-b78c5472185e · outbound

This paper cites In Findings of the Association for Computational Linguistics: ACL 2025.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing In Findings of the Association for Computational Linguistics: ACL 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:22:00.160588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:003b34f4418cbcd746019f9fa3d4ad0da61ec77293d8c6b58b1a0490bc70509f

Observation ecc2fabf-db43-4187-a15a-d92d7f8a6c85 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:25:04.571815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:527bdc14db1c46ea60f27c69ad1d466c5a9881743e61ae04c6617107dfc5650a

Observation 8bb6983d-d5e9-40cc-9c36-c311745f903f · outbound

This paper cites Nyx: Virtualizing dataflow execution on shared FPGA platforms.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Nyx: Virtualizing dataflow execution on shared FPGA platforms

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:26.457880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:4569b991dafe3da7907abd33bb446c885fed7feaaaab1bc8cfa1ea95218f20f7

Observation 5492f8f5-a6b3-47a4-ae72-08e2fe89a96e · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.184009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:5fb744c742b27d90e2c7b73bb33cc7a1cc7560b35f2ecb61a81c475be1807e75

Observation 5500f534-b187-4d50-8a4b-aeb9f96a01d8 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.187416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:37162c678c5b35c758089b4ad01ee02a7f1f375928d68fef0f547f96e4f79da0

Observation 64c44172-09ec-4f5d-8c1e-d8d8867c4aab · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:26.447170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:90ce8a52fddc37361f9033e49eb88c248cab285cfe2ca0d1afd12cdf86e394ef

Observation 6330ada9-fe4e-4cdf-a550-65b534718baf · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.193848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:cdef20499460b74785f23f42c3f24d23cdc795726c0971d4d0bf4a9e1dea3fcf

Observation 34e174e2-5332-47db-a5ff-f86ac7bffd48 · outbound

This paper cites 2025.A CXL progress report: The elephant is learning to dance.https://www.eejournal.com/article/a-cxl-progress-report- the-elephant-is-learning-to-dance/.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing 2025.A CXL progress report: The elephant is learning to dance.https://www.eejournal.com/article/a-cxl-progress-report- the-elephant-is-learning-to-dance/

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:22:00.200636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:8dea00167f3247d58e71f35efc0dcfe0cbb05b65a6bb6361876aba3c3d08874a

Observation 6a5ddbcd-6e9e-4e04-8fc0-7cdf005f5eca · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.204862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:72368db4cbb796d092f8f944601c0984936d8a69cb7b07647607e8074e95820d

Observation c48b33e6-e6ba-4e4a-b662-be8f827eccd5 · outbound

This paper cites Advances in Neural Information Processing Systems 37 (2024), 22947– 22970.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Advances in Neural Information Processing Systems 37 (2024), 22947– 22970

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:22:00.149681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:818214ddea962f30c58a47cdc71571be0fa4000c1bc9afa6425f752117283ab2

Observation 0eedc597-4243-4447-b206-9debfddf297b · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.174138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:8dbe0203d671c9d010294f90f0cc28bbd0a3a39c75763f97458c710a9a5fcb60

Observation f972df56-d797-4761-8e44-b17c62ac90ff · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.190149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:784bc7cf793bac3adce9b7aaa1a636e71344720005651198788d664eb2873cdf

Observation 6831bc56-b83e-488a-89be-31953da0acb2 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.153171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:5eab695fbff9ee4c37822cbf1df940c27a22ea7556d8da7a525f93be5eee4a36

Observation 3e8927d5-bccf-42e2-95c2-c5b0b6ec3c67 · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2023), 52342–52364.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Advances in Neural Information Processing Systems 36 (2023), 52342–52364

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:22:00.167158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:97ad2b1f89191e2b377c8ccb0ff99b1515ae59c7d80b9e75c36c6d0ae596be9e

Observation aca5ac1d-3565-461b-96e6-6f7e7c5cd0e8 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.164065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:bc2c6c9a01624e0d5460d2edf837cf91eb16f1c6958d1e37fb5227847b0b8b17

Observation 4e8fa6ac-1daf-4cd0-8240-740f34e9a3db · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.126765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:08b76ff6b6ca07e1412b9716fba7a1d33e9ab9fc561f8b2733ae3e0ddb7d7542

Observation 867aa5ff-5f4f-4dcd-b112-6c5025056ff2 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:14:48.923977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:292f72fe1311a02fb096a86908ed7ad76e8cfce29e9d021dee7dde4f5c527b7a

Observation 793f5eb2-ca87-4108-8151-06de6d20d91a · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.156604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:a23ff30a54c600c792cf3e9332078d1a9cdd7b88c7aca0bbed4653b76bc06c0c

Observation dec2ac56-a880-43d1-946b-356688b77134 · outbound

This paper cites vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention,.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention,

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:26.451234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:5f97be32e3ee4b9aef3f04ae5ca03ca1452a8a9d41a10471172c83215baa1933

Observation 01e6bcd7-d73b-497b-a57f-430a6ba92712 · outbound

This paper cites RayN: Ray Tracing Acceleration with Near-memory Computing.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing RayN: Ray Tracing Acceleration with Near-memory Computing

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:26.454448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:43d9c6fa1087dfd416fde45686d10b58034a34602ed86d7767104e991703c42f

Observation aafd7be7-f2d8-4ef0-bd22-6d24e65299af · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.123303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:6cfd1e9e4d01eb235546829952cd316ab22201db5188f752f65a9af9bee8c949

Observation 59189d4e-dbbd-4eff-ae6d-5c132ad8076c · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.177736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:64e725a69afd6b81fbce6696b7edfffc327f7eabe4ff3dc0982118f863da8a6b

Observation be3ceb72-7492-41ae-b023-dbc65e57a7d6 · outbound

This paper cites Grape: Practical and efficient graphed execution for dynamic deep neural networks on gpus.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Grape: Practical and efficient graphed execution for dynamic deep neural networks on gpus

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:27.058151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:0e3e0b158a17ce402a5ff14e69e4cfe2aa84df27ad77432f978882f940c4bf6c

Observation 251a5aa6-532b-4cce-95ed-26d1fc718aa6 · outbound

This paper cites Accessed April 2026.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Accessed April 2026

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:22:00.130260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:307880d34bdcab702d7b0b0b48840395088f7e9df9b5ed84a8f3dc6a7a23aafe

Observation f236e617-6e9d-4948-b24b-2bb99cb6c303 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.142332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:c79b0056a786abd87dee0ec6993509208a668d2af5787555e08c0e88eb6e3018

Observation 802373b7-a209-42e4-b052-e6e81f382f25 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.208408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:7e23326a36c277d9d07e4f52d2cc4582399cc16cc997eb81d29e1e35dcfd1d02

Observation 8d626a5b-09bc-4ed6-b148-e9878a391591 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.197056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:06aab64c4d6a6d5fe5abf262b5ba08c6219875b1848e8260a8b0599ecceb4cd7

Observation d0990dfc-0785-447b-9775-cac6128e78d7 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:25:04.573575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:4fb8a9de73785ecb7a7cbdccf9d4c95cd85a5b751ff2fa41080680323cc5e8c9

Observation bfbc2e60-2ebd-4fab-a5fd-63c7411d8440 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Efficient Streaming Language Models with Attention Sinks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:40:13.867138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:d6dc6479f55d2511e56aec4821cc131816ba30e9d83b21230e2238b8d83a230d

Observation 313561c4-ef2e-4dea-9525-bb4b9c48d236 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.212369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:eee94a84cbda43e8e9a1aae3f6cebb851356914907d804673a436db9d587c06b

Observation 8218bed6-b7a5-4a03-ba3c-2136cb234631 · outbound

This paper cites LazyFormer: Self Attention with Lazy Update.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing LazyFormer: Self Attention with Lazy Update

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.062559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:a561aff148926eb48b1937fc02dbae88bc9d4ac9b838efecebaebf1a5e313ea7

Observation 9bee5ae2-23a3-4263-90d3-5e91ea0247a1 · outbound

This paper cites arXiv:2512.18194 [cs.DC] https://arxiv.org/abs/2512.18194.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing arXiv:2512.18194 [cs.DC] https://arxiv.org/abs/2512.18194

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.060397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:d5c2b8baaf153bcff453686d8fad0123c7791a5c8dd86182ad0b2a951d404d3a

Observation 22205822-6f34-413c-8064-63367be331b8 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:26.460893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:f296ba6ee01bdee5639d9b894d7d0848e2d8fa30a019e4a448e82d8f6d648609

Observation 01a377fa-b53f-4dfe-94f0-f9ed7a7bc7b7 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.145711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:bedd97a6699d2b69531ee54e836b0c17ca6ba6905c13ed51e211400fb7c2e803

Observation f79e7d39-93d7-4662-8af7-3c8f32826329 · outbound

This paper cites Jenga: Effective Memory Management for Serving LLM with Heterogeneity.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Jenga: Effective Memory Management for Serving LLM with Heterogeneity

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.083217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:6371275e852fc80fd4715f9e481196a62a525c099ce29a0d13e381f5f758bf82

Observation ac968331-dc79-4d57-bea6-729aa4e6c1aa · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing OPT: Open Pre-trained Transformer Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:53:19.162556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:79caa7dea1200ed72db8ca3389d24259023283f1ff5e2f4e3133a08babffbdbb

Observation 430a27a3-436f-49f6-bf9f-022b8b1c1f66 · outbound

This paper cites an unresolved cited work.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:22:00.181005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:140f457e230474a1e5b77439936eb211efc6ca242cd65f2d7c1209a3232e27fd

Observation c71ebe57-ba35-4ff3-a9be-db911743dcbe · outbound

This paper cites Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.091615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:c73a48deb6e7732e106f45c6322c759d782b8297db19342cfae4e9f94f259afb

Pith citing papers

Observation 7ef6d700-32a1-4f43-a1a4-14841684c2f8 · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:09.632043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:09.632043Z digest=sha256:9dfb29c5c71d4a28efc6477ad78c36802675bada5d1e89f5a198f44ddb8cd665