Pith. sign in

Paper Citation Record · LEDGER

IC-Cache: Efficient Large Language Model Serving via In-context Caching

As of 11 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 1 inbound Pith citation observation for arXiv:2501.12689.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12689 v3

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:59:01.487312Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:35.999530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy36
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd32cb47-cc50-4020-86d5-2130697c6f13 · outbound

This paper cites https://developers.google.com/ search/docs/appearance/ai-overviews.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://developers.google.com/ search/docs/appearance/ai-overviews

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.196339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.196339Z digest=sha256:f7c42b780378e7e5ce4e2414de83cb88b5ed92da9e81679f8d147c50766921cf

Observation e500c042-14d4-4c90-8ab8-36aa279a2743 · outbound

This paper cites https://aws.amazon.com/codewhisperer/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://aws.amazon.com/codewhisperer/

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.200708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.200708Z digest=sha256:54ad22c8afb005ee345c0ccb8389ff21d1caeaf1d31079c037a693f4638d6826

Observation 380085fe-9edc-4b6a-ad44-5a23f39d0a4c · outbound

This paper cites https://claude.ai/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://claude.ai/

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.204302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.204302Z digest=sha256:c9bc4f206f83dff9407e5035612f9e5435ed3d6db9aab844c76478df38d612c7

Observation 44786cd6-a2bf-470c-a50d-a425c26c6d57 · outbound

This paper cites https://character.ai/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://character.ai/

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.209020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.209020Z digest=sha256:2d688c2eeed11cf9f5bf34d71e5c3fa38d0889983524a0009cf23a6b33d2ddfd

Observation f88c1446-2145-4f77-94db-aa0af443544b · outbound

This paper cites https://openai.com/index/ introducing-deep-research/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://openai.com/index/ introducing-deep-research/

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.212859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.212859Z digest=sha256:c64cb37249bd65f84cdf18e4f5eaaf7761f5baf4efe77e49537a24a576986351

Observation 096681f9-7a28-426c-a187-75d8f547d723 · outbound

This paper cites https://api-docs.deepseek.com/guides/ kv_cache.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://api-docs.deepseek.com/guides/ kv_cache

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.216531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.216531Z digest=sha256:9bacd340024c9b6c42d3db7b8cb8ab533b4b5863bdd81896f16fb34b42c5d784

Observation 3a6f70cf-8326-48c0-8a96-4e6c24dec579 · outbound

This paper cites https: //github.com/deepseek-ai/open-infra-index/blob/main/ 202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.md.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //github.com/deepseek-ai/open-infra-index/blob/main/ 202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.md

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.219494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.219494Z digest=sha256:aa712071137d13b788cd66286bd71a078bbb25e23536fb7133ac83ade0783df6

Observation 2cf6f2cc-1fcb-4feb-9706-572214791dd1 · outbound

This paper cites https: //developers.googleblog.com/en/gemini-15-flash-8b-is-now- generally-\available-for-use/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //developers.googleblog.com/en/gemini-15-flash-8b-is-now- generally-\available-for-use/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.222104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.222104Z digest=sha256:ba455b2e77d3ce4b0cfcef001cb4412646eb89dd06ad05927f0b5a853626418a

Observation 3b45434b-deb8-4bd2-8354-411b2d10f417 · outbound

This paper cites https://ai.google.dev/gemini-api/docs/ caching?lang=python.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://ai.google.dev/gemini-api/docs/ caching?lang=python

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.224912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.224912Z digest=sha256:f31940464d2959e1beb287bcc3d92a0b04bf4f6014a9b1e60aacd41e2830b1b5

Observation a97d271e-8ce3-4965-9ec3-9fc697263bcf · outbound

This paper cites https://github.com/features/copilot/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://github.com/features/copilot/

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.227790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.227790Z digest=sha256:da3abe10899ae4aed672f269449ec563bc852edaf77e715ee85a4c9026758476

Observation b0b110c8-ed32-401e-bfef-cb6af0771a32 · outbound

This paper cites https://www.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://www

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.230372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.230372Z digest=sha256:5b55be6bd5d59da4a466af2e1066148e7f013278951154df8052d2f19488667c

Observation a1ccee4c-a3ce-4e1a-b78f-9966255d546c · outbound

This paper cites https://huggingface.co/spaces/lmarena-ai/chatbot-arena- leaderboard.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://huggingface.co/spaces/lmarena-ai/chatbot-arena- leaderboard

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.233061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.233061Z digest=sha256:1dc60fac5eef9575c16cf5bbf8476f02025cd3baa1f45f9ba68f5a017fabf6f6

Observation d2d29a30-d0b2-4492-aed9-dff524a20be2 · outbound

This paper cites https://huggingface.co/docs/ api-inference/index.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://huggingface.co/docs/ api-inference/index

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.422543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.235646Z digest=sha256:3de9261ab527e02b6111830ba681a6b18436b8159b21e9cb31d5c840e24dbfbe

Observation 992e6427-342e-4c0b-8bc3-c95b28dd62ad · outbound

This paper cites https://github.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://github

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.413250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.238251Z digest=sha256:d5ca5c597cca8a22748a9d68107625d72ba8ed7cbc11466e9222b7370c43650e

Observation b9ecb950-bb36-4c65-8103-eda5c8aa0075 · outbound

This paper cites https://microsoft.github.io/msmarco/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://microsoft.github.io/msmarco/

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.404981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.241257Z digest=sha256:e6a9e49028f39a1e5659ce48171f7cb6a0277d98c4f8951fa78118e1ec10bfe9

Observation c5bbda78-31ab-4abf-9cdf-c7cbfdd066bc · outbound

This paper cites https: //github.com/explosion/spaCy.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //github.com/explosion/spaCy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.396331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.243979Z digest=sha256:7fe885ca5ca88ed3b85ae56147425787569b0ba12abe1952b76700877976a1da

Observation e38a789c-d390-4ee9-9638-9b4f0356c303 · outbound

This paper cites https://www.databricks.com/blog/building-cost-optimized- chatbot-semantic-caching, 2024.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://www.databricks.com/blog/building-cost-optimized- chatbot-semantic-caching, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.387503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.246580Z digest=sha256:21226e7f444b8fe68b7bdd07f62a01b175fd1becebbea2021a26fbbf5dc37a94

Observation 69b137e8-c30b-4caa-a7d5-36e0f26602a0 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

IC-Cache: Efficient Large Language Model Serving via In-context Caching SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.249697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.249697Z digest=sha256:c76a9395511506acf711620448f664d875da223b25c6fcd5013d47b9ee0059e1

Observation b0fe5ba3-8f17-44ec-8617-f9f70b688597 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Analysis of thompson sampling for the multi-armed bandit problem

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.379244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.252860Z digest=sha256:c929b1a63e3329bb3458317c6797706728a6e6909d8d15bf46310692bd9b729f

Observation 92175fbd-bfe9-4eb6-a4ad-da7e1613a51a · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.255587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.255587Z digest=sha256:da974bb9b842d2d81eee4cd9a9481359d5d432169091cafadbbe574a70518d6f

Observation c7123322-f633-448a-a5f5-657ab115c5da · outbound

This paper cites Gptcache: An open-source semantic cache for llm applica- tions enabling faster answers and cost savings.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gptcache: An open-source semantic cache for llm applica- tions enabling faster answers and cost savings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.370400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.259377Z digest=sha256:93656f2808f9e29d8ab505a722e661dc4aa60b60ed1ffdd3d4e1abaa78f87238

Observation 074cb387-ae3a-4d81-bb27-b477f80bad82 · outbound

This paper cites Findings of the 2016 conference on machine translation.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Findings of the 2016 conference on machine translation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.361569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.262365Z digest=sha256:65299ab2e1c586f4c169edb87e3ca4e2d816ae4b86bf88f12c909117085cf6a9

Observation f90915e7-10a2-444b-8b22-0d8c5fff2872 · outbound

This paper cites JAX: compos- able transformations of Python+NumPy programs, 2018.

IC-Cache: Efficient Large Language Model Serving via In-context Caching JAX: compos- able transformations of Python+NumPy programs, 2018

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.351675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.265474Z digest=sha256:8c610c3a7fec218ba34ab5065cf99e477740a6c83e39a96b1c1d5e828f7f4ee9

Observation e71ff12a-149e-4d43-a45e-c90fbf043580 · outbound

This paper cites Language Models are Few-Shot Learners.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Language Models are Few-Shot Learners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.269384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.269384Z digest=sha256:08a1a8c20b8693e959fde25e8cd1aaaf53f64c98a2f3fa269bf9eaaeab44b7b9

Observation 1044e1e8-587f-484f-8592-841ad40b297d · outbound

This paper cites Are more llm calls all you need? towards scaling laws of compound inference systems.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Are more llm calls all you need? towards scaling laws of compound inference systems

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.341134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.273733Z digest=sha256:acaf05cc73baba409ac7d7371b1fe94e556697f55f8e1d93d18b31b6356f0857

Observation a194a693-786f-4437-9fad-67e3d8be7938 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.277250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.277250Z digest=sha256:37dbd3fc38ff0edd3e278d52909d6ed44be5a5a8bed9596b28e6a3ec9febecb4

Observation 75dec9d3-3c05-46d1-8630-51089c5c7376 · outbound

This paper cites Learning semantic similarity in a continuous space.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Learning semantic similarity in a continuous space

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.331126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.280850Z digest=sha256:576e85cd9b2333e667b629493e7a4e9ced3e8cf1a878a711a7bd74742fc48de0

Observation 1546857c-aac3-4d70-b3a7-7cc1bead340b · outbound

This paper cites A Survey on In-context Learning.

IC-Cache: Efficient Large Language Model Serving via In-context Caching A Survey on In-context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.284457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.284457Z digest=sha256:44c77c1ed0beb97ae0f108444bd974027232f18489efd995a6d27cac9825802f

Observation 1a51bbd8-2ba0-4fe2-82bf-c7da66dfb1dc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.288550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.288550Z digest=sha256:c57ab3bfaa0d05e01681d82d55e92cc501237b5a3de2735e676bf9077e21f5cf

Observation 480858d0-a692-4121-bd7c-9e26947eb758 · outbound

This paper cites The Llama 3 Herd of Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.291516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.291516Z digest=sha256:4a8b5288e0454d2d74f95c9189e08211a80cb9ce00125bab2278554835be2309

Observation 9ba9fc8b-ef05-4783-add3-b425f3128708 · outbound

This paper cites Apple intelligence foundation lan- guage models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Apple intelligence foundation lan- guage models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.294266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.294266Z digest=sha256:5cdcafb0fe818084ac9b9865de25c343f13f6c3fe7f425120b2f1b99b01bcaa6

Observation d44d57c7-7753-48a5-b1a1-6f5e882e17dc · outbound

This paper cites A Theory of Emergent In-Context Learning as Implicit Structure Induction.

IC-Cache: Efficient Large Language Model Serving via In-context Caching A Theory of Emergent In-Context Learning as Implicit Structure Induction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.296886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.296886Z digest=sha256:55f5d120018a200c25eaadb2b89d9baf01e25b33aa2d3a84cc1274c392564b0a

Observation b6b05c68-830a-4601-8a77-16dc123ccd23 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.299994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.299994Z digest=sha256:332ee7fea01b6511bfd469e8e270ce541de3769079a19fb997964a9ccb379db2

Observation eb5f9fb0-9d97-4ce2-98f3-c9f04f6c252e · outbound

This paper cites An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4.

IC-Cache: Efficient Large Language Model Serving via In-context Caching An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.302934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.302934Z digest=sha256:c827f875daf3e9d1fec483af00f8aa706d44bd7727e863a31df8dd783c3b070e

Observation c425dc93-770c-4898-b7cc-6db9de6902b4 · outbound

This paper cites Evaluation of Best-of-N Sampling Strategies for Language Model Alignment.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Evaluation of Best-of-N Sampling Strategies for Language Model Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.305843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.305843Z digest=sha256:4835aafb757f5d57f3e567512f89cc26b0966cbef3eb7bbe9898730edba1a03a

Observation 3c5439a9-4d0e-4605-a805-1f7377e26f93 · outbound

This paper cites Active Retrieval Augmented Generation.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Active Retrieval Augmented Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.308781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.308781Z digest=sha256:3765e075b2d10f2f23ef844292deec32dfc86da14c752e902ea25e14849b71c4

Observation dae16513-ac8f-466d-a8b0-8c28aaec925c · outbound

This paper cites MegaScale: Scaling large language model training to more than 10,000 GPUs.

IC-Cache: Efficient Large Language Model Serving via In-context Caching MegaScale: Scaling large language model training to more than 10,000 GPUs

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.320664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.312290Z digest=sha256:bdb78d3d6df8b1b615cca3cec1a229534728e23908dd2a5a8ba00feb7ab6daab

Observation e7f5ad9c-8aad-4a38-ad0e-7897ee31dd45 · outbound

This paper cites Billion-scale similarity search with GPUs.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Billion-scale similarity search with GPUs

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.310486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.315989Z digest=sha256:d7a9d30a847156d0d2fad5fe28349726aaeb381d4a67e8b411d7bbf2edb25a96

Observation ead06160-c83b-40fe-912f-208322b47337 · outbound

This paper cites Tanh works better with asymmetry.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Tanh works better with asymmetry

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.300514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.319029Z digest=sha256:18fe70d8456bfe3d4bc2d3c81ed3360937a490a82b3f5926714ec75fe9369058

Observation 391c229e-cd53-430d-aed8-b10dcc8d28ef · outbound

This paper cites Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.322004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.322004Z digest=sha256:e1c5b2f955413c2e5c2bf46299cfff21f2b3917b30979e6dee583eccbbc00ef1

Observation c08c3152-cb8f-4692-b3d7-6337991f75ba · outbound

This paper cites Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.325281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.325281Z digest=sha256:fb0a5f069799024cfe5dbc6ee5f13f23195d0c269fbafc85d76822c6c7a65548

Observation 946dc1bd-1535-494a-ab91-c7ea8e33cf37 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Efficient memory management for large language model serving with pagedattention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.328661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.328661Z digest=sha256:b299f4181c5621376f0d560cf980b65b6a3957ece52b20a7b423ee74e2dc49cd

Observation 6b3e0f31-9917-4819-9713-88ad72c33937 · outbound

This paper cites Auto-GDA: Automatic Domain Adaptation for Efficient Grounding Verification in Retrieval-Augmented Generation.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Auto-GDA: Automatic Domain Adaptation for Efficient Grounding Verification in Retrieval-Augmented Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:59:01.797977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.332595Z digest=sha256:6a09a5572aa0c82ca26bb5183ba5b4d302e1c1fcf45068c75e0a916dc4da2310

Observation e7d1cf4a-5f0f-42f2-a50a-6e3b78c0ed52 · outbound

This paper cites Retrieval-augmented generation for knowledge- intensive nlp tasks.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Retrieval-augmented generation for knowledge- intensive nlp tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.278868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.336033Z digest=sha256:8c998c897b6b2e6eb1fc97d803a9f5a645d56fd0dd002020b5955822708b280d

Observation c4163a19-a8e2-496d-b1cf-f4fd42b79230 · outbound

This paper cites Dpsynthe- sizer: differentially private data synthesizer for privacy preserving data sharing.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Dpsynthe- sizer: differentially private data synthesizer for privacy preserving data sharing

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.269666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.339123Z digest=sha256:ed8312acd944a9377928aa3a6e3b1100cbb5898866cca1e252578f736a50e4a2

Observation c38910f0-5b83-440e-9b42-dd41ee577771 · outbound

This paper cites Schapire.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Schapire

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.260458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.342208Z digest=sha256:0254de6d73e60fe3f7b2565a3167375630ef8a251a0c7edec69e6ebe42aff356

Observation 3de0911c-4795-4685-b6fe-85e16e4bce7a · outbound

This paper cites Gon- zalez, and Ion Stoica.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gon- zalez, and Ion Stoica

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.251889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.345353Z digest=sha256:bd951be63ef92f7a9bb5d866792060dca503430357cd37bfa742ec1d9711cd3c

Observation 6f803b86-85e5-468a-aa9f-53d9479639e6 · outbound

This paper cites AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.

IC-Cache: Efficient Large Language Model Serving via In-context Caching AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.348128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.348128Z digest=sha256:078bc898d96fdeda4a6e378453a35f55604c5974892599a6ea8ced2d516bdb65

Observation 3537dd21-d42d-472a-83eb-32702659733f · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Openorca: An open dataset of gpt augmented flan reasoning traces

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.243226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.350881Z digest=sha256:2eab35f53cee7ba73823c46ab4ba7030d35885cc86d2d685c5eeb56093aa0961

Observation 1ca1e0b8-f16c-4aef-a0bb-dc53bfee0978 · outbound

This paper cites Parrot: Efficient serving of llm-based applications with semantic variable.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Parrot: Efficient serving of llm-based applications with semantic variable

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.234904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.353771Z digest=sha256:a9ecf2e9f4da3d7f232c0a17d78ae7f1aa9fc1cc83cf1f66a205737414c52686

Observation e2b46443-4199-4a28-b683-bdf3080bd62e · outbound

This paper cites Andes: Defining and enhancing quality-of- experience in llm-based text streaming services.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Andes: Defining and enhancing quality-of- experience in llm-based text streaming services

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.225238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.356375Z digest=sha256:aedbf947ff1a2c3c538692063f843db47e5d78b0338173b6550cfca9bd77c3c5

Observation 95a57241-3c22-4ef3-909f-094181b591e5 · outbound

This paper cites In-context Learning with Retrieved Demonstrations for Language Models: A Survey.

IC-Cache: Efficient Large Language Model Serving via In-context Caching In-context Learning with Retrieved Demonstrations for Language Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.358903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.358903Z digest=sha256:2c0ddf8fcace10059f461de75ce8148d7fbd9f6dc10d6a7476169a16ccbbd8d5

Observation 522b6ef3-5e6b-411e-9d0f-a74d285a6b0c · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Simpo: Simple preference optimization with a reference-free reward

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.362089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.362089Z digest=sha256:5eccb579fff30be30a9c65d4f09796c1e6545d49742b69f1fe455c18850a222a

Observation 33dffffc-c17f-42f0-ab03-fb5177ecbbf9 · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

IC-Cache: Efficient Large Language Model Serving via In-context Caching MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.364722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.364722Z digest=sha256:9cffe52f0b38cd93f5c3943fa1ad3bdd08b3ef2f8d7256b98060dca12a825901

Observation 035a3950-fb3e-470a-b310-124f02b63984 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

IC-Cache: Efficient Large Language Model Serving via In-context Caching RouteLLM: Learning to Route LLMs with Preference Data

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.367502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.367502Z digest=sha256:ee1e2fde0b4625b781343de8d45b846fcc0390579f33d6b7962dee3af61a272b

Observation b70c6d1a-371a-4f0d-a4bc-6fad3bf9fb16 · outbound

This paper cites Training language models to follow instruc- tions with human feedback.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Training language models to follow instruc- tions with human feedback

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.209813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.370325Z digest=sha256:69e6294f3063b96617b75aaa3b1af001e966e32e877d2780ff129d7701f0b8cc

Observation 5d3ec71e-b22e-4de4-856e-e8c4bfd0f1ba · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Splitwise: Efficient gen- erative llm inference using phase splitting

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.373036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.373036Z digest=sha256:40856f92b7dc26cf59c08e305554a1078c7bf7a3ca8644ae08ed53da420b8de5

Observation 605cb783-6603-4532-b2bd-4b797d720249 · outbound

This paper cites ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.376047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.376047Z digest=sha256:df9e01affeba742651253b6004d3ff8d50fb8b75b8f3c6cac9933f1cfb217fcc

Observation d477c421-d779-42f1-919d-6a1ad668f1bb · outbound

This paper cites Modserve: Scal- able and resource-efficient large multimodal model serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Modserve: Scal- able and resource-efficient large multimodal model serving

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.379428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.379428Z digest=sha256:0a1184d7fc406415744fb943970941c48c02bbe4c013c92bc1287b895476c3d6

Observation d1204164-7432-4757-a02d-d0e905e17bc4 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.193782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.382463Z digest=sha256:5774dbdf3cbfc55c56d933709af3f73a4a9137da1ddc7e37ecc2328d8e04e79a

Observation 2a15ae50-03c3-4a63-8cee-834b174626cb · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.

IC-Cache: Efficient Large Language Model Serving via In-context Caching The probabilistic relevance framework: Bm25 and beyond

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.184115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.385557Z digest=sha256:8516bf609a7009ffbbb06173b68ac7374c31bf7dfd600320a42f33cfad7c0ee7

Observation 6b26c39d-d4b9-496e-a29f-ca20b9578756 · outbound

This paper cites Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.389424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.389424Z digest=sha256:c40c90f1045f2d763e6f67dc80cefd2798e6f7c39c2c87dc773965b19ccebe14

Observation 540059ff-f982-4dc7-b9d3-a703af6d1982 · outbound

This paper cites Gonzalez, and Ion Stoica.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gonzalez, and Ion Stoica

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.392854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.392854Z digest=sha256:8ed4be3d886e272aba54041fdacdf65d139e7a2b77171a397737c0c18daa19b9

Observation f03d290f-c6ca-42d2-9b55-dfa4dac278ac · outbound

This paper cites A statistical interpretation of term specificity and its application in retrieval.

IC-Cache: Efficient Large Language Model Serving via In-context Caching A statistical interpretation of term specificity and its application in retrieval

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.168557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.401232Z digest=sha256:0d006a4ad706bf245c5005e6906b7b1b017212b5bbd1fa80a7e557c0e4b955b9

Observation e4714e4c-3b96-4f0f-835b-eb5358916a14 · outbound

This paper cites Hygen: Efficient llm serving via elastic online-offline request co-location.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Hygen: Efficient llm serving via elastic online-offline request co-location

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.404692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.404692Z digest=sha256:93e92c9588e303ac9c92c78a588735a915cb374728dfb2455139f2cc05542a0b

Observation 48874c01-f19b-4c00-985c-05ded19a3a08 · outbound

This paper cites Hashimoto.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Hashimoto

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.407925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.407925Z digest=sha256:79e8ab47606ef46fd07dc34889d46cf0ae4e5c0a9e012052c0ee543d93518086

Observation 95cd782d-7f5b-481f-8aa0-a38814cc11e8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gemma 2: Improving Open Language Models at a Practical Size

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.411181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.411181Z digest=sha256:0c41a6651cc7866b51c33298e00bb5d1f75d00283ce5588b81b2fb9a01359863

Observation 567a4a46-69a7-4919-a334-1efc461420a6 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.152835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.414581Z digest=sha256:6056c24fe43548777531842dd74b926b621b6d40b0423b3e5620a2b4982f3bb9

Observation e129664e-46f7-4ad3-8439-a31b468b8046 · outbound

This paper cites Emergent Abilities of Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Emergent Abilities of Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.417721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.417721Z digest=sha256:503a72bf9bcbe65bc33b2b35a814bf087fc9a9319ebe4acc5026dd2c78c92a5d

Observation 52f8979c-3250-407b-8261-f1ee720fae31 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Fast Distributed Inference Serving for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.420549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.420549Z digest=sha256:bf7b518a3d365d932ae7862e3643ffa043764d58d8993de0d49b75fa2fcba990

Observation 5ec0c356-036f-414a-9093-2aff9256d5b1 · outbound

This paper cites dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.143575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.423310Z digest=sha256:f4152a7cc61e5b15d159408e578e2dcee1fec467ee59ea154d68373f76b75cf8

Observation 4e61e9fa-c230-405c-b437-f0c5dc32a5e8 · outbound

This paper cites Why in-context learning models are good few-shot learners? In ICLR, 2025.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Why in-context learning models are good few-shot learners? In ICLR, 2025

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.133859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.426023Z digest=sha256:54d84aa82a50fd9db20a09efdea7e64587fffdea487f79fb8858f30c89e99b62

Observation 7c569c36-6cf0-4a3b-8347-14970b0e2415 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

IC-Cache: Efficient Large Language Model Serving via In-context Caching PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.428571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.428571Z digest=sha256:ef020afcd6802e639234b35e651bdfa135e5918b33c7ffca9aed88a62086f593

Observation a437b52c-ed8b-4a15-8e53-6e5611e3efb2 · outbound

This paper cites CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion.

IC-Cache: Efficient Large Language Model Serving via In-context Caching CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.431548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.431548Z digest=sha256:509068b60ece647aa1d0410ad917fc72e6b20d7199330802c2e0809ca6d873ca

Observation ac1ef0f7-dd5d-4b55-9f7f-c44df8a7be25 · outbound

This paper cites Generating Data for Symbolic Language with Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Generating Data for Symbolic Language with Large Language Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:59:01.539371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.434253Z digest=sha256:33cb609808fe54c5302ae9f1c2c029dbdbcf74cb4b979e63723bda4dcc15d978

Observation fd80b406-6514-4059-b32c-6213b6ee35f1 · outbound

This paper cites Compositional exemplars for in-context learning.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Compositional exemplars for in-context learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.124312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.437310Z digest=sha256:e684e2213aab44044bd88c0533f49af4c15f6ece1f1a33f2090f4671210e6524

Observation 5d0ae6f1-7cb1-472d-bc84-92625cdb4022 · outbound

This paper cites Orca: A distributed serving system for{Transformer- Based} generative models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Orca: A distributed serving system for{Transformer- Based} generative models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.115345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.440027Z digest=sha256:f30415cdc5b29c7005388944654a96321af591f9ca23b718c5b2b26c9eae0b5d

Observation 28518b5b-a337-4f2c-9f10-67ee85eb4b12 · outbound

This paper cites Longrag: A dual-perspective retrieval- augmented generation paradigm for long-context question answering.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Longrag: A dual-perspective retrieval- augmented generation paradigm for long-context question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.106458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.442999Z digest=sha256:da259c4b8342ef2e2c21fa6f99e706ac2fb79745f801c0a915da764cc3e3f05b

Observation 299158f3-1257-4348-870e-49c0a38a585e · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

IC-Cache: Efficient Large Language Model Serving via In-context Caching LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.446099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.446099Z digest=sha256:d0170abe45868cdb789f7724d83b91bd2d6e0a12caafe26a80294d1b0525e795

Observation 745952cf-51da-46fe-b11f-2cf26ae7ff9a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.097402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.449572Z digest=sha256:1bcf8a53b8afa3df7b760bdce4d3d0e6dbd3f7ff9ca54d715368e05b097b6dfa

Observation 72e24730-f6f0-4451-ac9f-bf749807b0f8 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gonzalez, Clark Barrett, and Ying Sheng

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.088761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.452502Z digest=sha256:400c3e034cb12c5bff8721784fc13693cfd7fe20f9abcd4d4785d3e296606567

Observation 52914cab-38cd-435e-ba70-e6f7090abdba · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.455880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.455880Z digest=sha256:0d23ead9b3dd5f5782eccfd3eb54cb569fba7e676097064aa08fb77d8b34ad45

Observation 5b4d1226-6145-435e-961e-ed29311ae7da · outbound

This paper cites Distillspec: Improving speculative decoding via knowledge distillation, 2024.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Distillspec: Improving speculative decoding via knowledge distillation, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.079945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.459544Z digest=sha256:de6336903a0583039e1674e817d1d6528c2399781da7c70a52d14a45db8dd4cc

Observation 0a8c8721-424c-4d62-9124-cb322f5dac56 · outbound

This paper cites We can bound this with the union bound: 𝑃(ˆ𝑖𝑇 ≠ 1)≤ 𝑁∑︁ 𝑖=2 𝑃(𝜇𝑖 >𝜇1) (2).

IC-Cache: Efficient Large Language Model Serving via In-context Caching We can bound this with the union bound: 𝑃(ˆ𝑖𝑇 ≠ 1)≤ 𝑁∑︁ 𝑖=2 𝑃(𝜇𝑖 >𝜇1) (2)

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.070905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.462700Z digest=sha256:9994305c1b4ddcd4ea3d627c044f319171dbfd98595d58c1693c470658f2f64b

Observation fdef6cda-73fe-4b83-a309-51dda7a6e446 · outbound

This paper cites We can state this more formally for the number of comparisons,𝑚𝑖(𝑇), for a sufficiently large T: 𝑚𝑖(𝑇)≥ 𝐾 log(𝑇) Δ2 𝑖 (3) where𝐾 is a positive constant.

IC-Cache: Efficient Large Language Model Serving via In-context Caching We can state this more formally for the number of comparisons,𝑚𝑖(𝑇), for a sufficiently large T: 𝑚𝑖(𝑇)≥ 𝐾 log(𝑇) Δ2 𝑖 (3) where𝐾 is a positive constant

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.061618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.465836Z digest=sha256:d1bcbcdeaad71ed675769fb84aba401391a9f329c5bd82790f4be9ff3f5939ae

Observation 29f295ec-864f-4b79-b281-308edef2edff · outbound

This paper cites Let the em- pirical difference be ˆΔ𝑖(𝑚) = 𝜇1−𝜇𝑖 after𝑚 compar- isons, whose true mean is the utility gap Δ𝑖 =𝑈1−𝑈𝑖.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Let the em- pirical difference be ˆΔ𝑖(𝑚) = 𝜇1−𝜇𝑖 after𝑚 compar- isons, whose true mean is the utility gap Δ𝑖 =𝑈1−𝑈𝑖

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.051814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.468915Z digest=sha256:46ed177190545b9cf31b1cc90a6e8e47ea8d53d7c650a596bea329751383d241

Observation f823556b-fdd4-4bbf-b6c4-68c904ce3010 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.042112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.472164Z digest=sha256:22868be0823d8bb03268a4b0cc207fd0f06e5d2ba02d17dfa57627d74a9cab7c

Observation 8b09cd53-4a64-4bdd-978b-c201b058ab8d · outbound

This paper cites Substitut- ing this result back into the union bound from step 1 gives the final bound.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Substitut- ing this result back into the union bound from step 1 gives the final bound

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.032253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.475312Z digest=sha256:80dc1f3cfd05da970ac8c62539b0116e24b95bca926b4a34101729b5d1852211

Observation 6df4f04f-d3fe-4924-9082-5c8d802e3434 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.022550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.478487Z digest=sha256:22ac309b6799436aea56a9cfe4da6202206e8bcf776a394aee8e544e5347a9fb

Observation 9152419b-e6e1-4fb5-af5b-d387d722330a · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.012709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.481708Z digest=sha256:6a2059ab26088478869d3f725d32c0a4c6ebf6180635332ed004229f73e14718

Observation aac84641-c28b-4673-86ee-cb524d096197 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.003421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.484719Z digest=sha256:6f5e7e24ce99beea45459d29419cd94fcd69fb5419f4a414edaaa192f7b3f3df

Observation bb69b9a2-e6f5-45fd-b009-dbb4adbd88f4 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:01.994092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.487312Z digest=sha256:d472f56c5d2ef9423b7a9b86075fb3efdd420587e3a520d9c74003362d0e705f

Pith citing papers

Observation b0f988cc-6209-4367-8174-1ca3a132931f · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems IC-Cache: Efficient Large Language Model Serving via In-context Caching

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:35.999530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:35.999530Z digest=sha256:13af4161aa00fce9a1b64bb982adcebfbdce8e5b8d39bc6850d1f74b02bfc60a