Pith. sign in

Paper Citation Record · LEDGER

HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2504.05897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.05897 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:07:38.001641Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T12:08:15.409036Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8aa3366e-4b03-4a9e-97e7-c694e98953dd · inbound

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference cites this paper.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:11.321915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:11.321915Z digest=sha256:c12bb29f231f9fff63d9f57ac9b80bb610f645386f7f1602f833d021f40b21dd

Observation 7645abce-31e9-4e1a-9df2-6203e66c31fb · inbound

SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution cites this paper.

SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:41:51.619092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T21:40:41.015481Z digest=sha256:d470c4ed37e6db1e4b24c5c9a358d32f41be89866c12eaa0144996b758e00230

Observation 9964d58e-d9b0-44da-8b7a-ef29cd5b4439 · inbound

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing cites this paper.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.064846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.064846Z digest=sha256:7cf93c07d4f0b239fb9479b96d6293be503b2692bdb2999a707b6223b4f5e9be

Observation 0b656ffa-3170-4b54-befa-04b645a03b6a · inbound

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution cites this paper.

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.411416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T12:07:53.045082Z digest=sha256:07115dff817bd0c188a72206cb174a3089abdb282bd181e8f3e9effe0fcfe2b6

Observation 71c53d55-488d-4640-9f8b-b233e79affa2 · inbound

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM cites this paper.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:38.001641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:38.001641Z digest=sha256:ad704f4e0ce58714179514e1fba34720fd22f7c8edee5c5427597e8b12bfd065