Pith. sign in

Paper Citation Record · LEDGER

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.20507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20507 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:58:08.181911Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90b49b74-0e22-45b6-8ac5-94023f6cc8ee · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.635644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.635644Z digest=sha256:d4ab622d13c3ec9742d348e11505c22ab2fcb1b6207d093b0f1bea03ed47bc82

Observation 16baf39c-1956-4da2-9250-d913920b9975 · outbound

This paper cites Transactions on Machine Learning Research , year =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Transactions on Machine Learning Research , year =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.714808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.714808Z digest=sha256:cc1b4a189ee00816798799705b2e1c6e08ad1242a7f03c2bc99adfe37de83b55

Observation 268a6c57-1b20-4e03-acd2-6185d20afc28 · outbound

This paper cites 2023 , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2023 , volume =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.797819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.797819Z digest=sha256:8922a05366f5c10402001ddaf8913b1199b256ebeef393a52e5289f9dd72524a

Observation 24610a3e-bf75-4c98-9182-2942538ff241 · outbound

This paper cites 2023 , url =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2023 , url =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.936357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.936357Z digest=sha256:34ed54042adb8873eb37a685fbec3f124581fad48880237b23ca1913567d5c4f

Observation 516aff8b-85a5-4c06-ac77-23b9d7442bd7 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.030510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.030510Z digest=sha256:f5d043895787807f1e7679f2dcd314c7c9b90d0fe6cc1dc9652e4f9414279f98

Observation 675aff8d-7a30-4427-b801-c4a7f47ca0e1 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Accelerating Large Language Model Decoding with Speculative Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.103454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.103454Z digest=sha256:77b91d8adcfc01b39da598f3e72782587e9cd883730999efb976d7820ae77c48

Observation 67e44686-6a96-4b7e-9eaa-42f3d3d44e65 · outbound

This paper cites 2024 , publisher =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2024 , publisher =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.177707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.177707Z digest=sha256:bc4bedfa2e226e183c26c7383c1beeb29baa6a3183056b93d6e603f2f5d868a8

Observation 8659404b-334c-4bd0-9e3a-650210a9c112 · outbound

This paper cites and Chen, Deming and Dao, Tri , booktitle =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference and Chen, Deming and Dao, Tri , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.365116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.365116Z digest=sha256:e5d83c49292fcc452773f7a3f54ab260f499ce630a52a0c23467dd4708a4d9a4

Observation c987e5ba-0a3b-4b70-8003-9fd8ce8a9ffb · outbound

This paper cites 2024 , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2024 , volume =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.394024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.394024Z digest=sha256:48486d059c6c79d1dba7d5ae5cd27b66126df5610bd9edb5a9a2f8077ca1c201

Observation 26ce685d-ddb2-48cf-8eb1-4e963efc8dff · outbound

This paper cites Break the Sequential Dependency of.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Break the Sequential Dependency of

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.457930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.457930Z digest=sha256:7f63a685475e5c944a823f7785da09e16aba8916b17868b0bf5a2d0828442580

Observation 942d0c99-998d-464c-87ca-35504046ffb9 · outbound

This paper cites an unresolved cited work.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.510043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.510043Z digest=sha256:3d618d4e714f14284172099413c56468acda1f007544627c02d7d5c5622b1a80

Observation 452c4d0e-da0a-4830-84e3-f69b0a3a79bf · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Findings of the Association for Computational Linguistics: ACL 2024 , pages =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.669849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.669849Z digest=sha256:a545cf3680904b01971852fbc0acd94011a0cbbe04a5b79a171e1745704c3ae0

Observation 09be17be-f8f3-44b2-b617-ee20682ec72e · outbound

This paper cites 2023 , address =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2023 , address =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.824091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.824091Z digest=sha256:0d84a12ed4a9dc88d883ca4a6a9d938ff6d293a1a90fe1ca7c2f56d6e569d0fd

Observation 068aa273-416e-483d-8142-29866a5d3d24 · outbound

This paper cites Proceedings of the Seventh Annual Conference on Machine Learning and Systems , year =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the Seventh Annual Conference on Machine Learning and Systems , year =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.986193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.986193Z digest=sha256:313ac1e864e659f1281e4e87bfdf0dc9eee668b764d1a25cc358a024a89373f4

Observation 3d07fdbd-a640-4bb8-975e-5ed2c1cd8115 · outbound

This paper cites MeanCache: User-Centric Semantic Caching for LLM Web Services.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference MeanCache: User-Centric Semantic Caching for LLM Web Services

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.169493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.169493Z digest=sha256:a81b3b4ce5eac852ba8c402f2fa9f7976e857aa0809ce6b36c985a55d2ad0366

Observation 36b6fbfb-565b-4601-9a2c-2b3cae74170e · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.367143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.367143Z digest=sha256:839e362b84998e0249a1b97d2554af374d1ed1664cc1db6f99c166199b2f4432

Observation da5eb0de-d486-4885-be94-7280f4fa4f76 · outbound

This paper cites KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.513882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.513882Z digest=sha256:628ffb21bd1b1de553a0ac155e2520018aac196e430c396d652706ec550da659

Observation b708ea27-3567-42ca-a72d-27a74d2b3016 · outbound

This paper cites Qwen3 Technical Report.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.716260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.716260Z digest=sha256:3fb982947f6b27137f8d5fe5a19092c18fa7db3d1caa3428091cb78131905158

Observation b169a79f-5a0a-43b9-83c6-e2dbf5aa5980 · outbound

This paper cites Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing , pages =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.854370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.854370Z digest=sha256:ae838b5b4e9ae5b9d1dd6c8c46ef245d54a3faf8f659dd1aeccdeafd50f825f6

Observation 9f652980-3369-4f4b-8804-b24b5a33ed3a · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.038196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.038196Z digest=sha256:1755c630c0c5b33778ce679b931fc94a9dafa99856ae9e9ffba8a635257f4fb2

Observation ac43dd8a-f21a-47ef-9732-0a01388b6ae8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.142971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.142971Z digest=sha256:a9e1d728d954b1aa5e317e361d7b7132782ad4b771f3ff27d36eda4b62bfda6f

Observation e81a133e-06aa-4a36-ad3c-ad4b9f5c0ae2 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.291602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.291602Z digest=sha256:56a352b90b353c011947cd539b1e0b17b5ce19a37b94aa5ba3aa702b75684eba

Observation f664783c-749e-496c-8d20-92e82c2a7108 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.460354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.460354Z digest=sha256:5da85117e3032a02ec057fce89dae01eb0021719416995b25c892e58d2abf21d

Observation fb5594d6-8ec9-416d-a30c-17b0611d0036 · outbound

This paper cites Scaling Laws for Neural Language Models.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.611796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.611796Z digest=sha256:e01166901765689bc3bd87a8744741d2d39808fe494b55a6b6cec8d62d080f4c

Observation ff9d53e7-7b09-4ae1-bb68-44a35ba17889 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.794419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.794419Z digest=sha256:d48154a68ea267334d9ae483db45e651d6c33a3d44ee6f15b34f04183b645275

Observation bf4bba0c-d969-4a6a-97fb-4ea8a2a3d06d · outbound

This paper cites The Llama 3 Herd of Models.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.998164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.998164Z digest=sha256:7fa0ebb90eea8bd0d236aa5913b4c886ea41c809a3843fac924969efdb2ba8a1

Observation 5498f6bd-3387-4ca8-a9a4-e225290f5fea · outbound

This paper cites DeepSeek-V3 Technical Report.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:08.069399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:08.069399Z digest=sha256:c4230ef8bd3ac6859ebb20cd3ef38c5189858141387795da2670523c540dbe86

Observation 2a653bc0-de73-4415-a241-44863d1ff648 · outbound

This paper cites FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:08.181911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:08.181911Z digest=sha256:f1e16c5f31a31b7ed5a89e0e3191ee733e6e0ac433342b33c63b3cf2efd27345

Pith citing papers

No inbound Pith citation observations are available.