Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T05:02:55.608400Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2607.09153.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T05:02:55.608400Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-30T15:53:12.112264Z
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ac6860b3-329b-456f-a5f9-c91c576737de · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems (NeurIPS), 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d66a40f-cf47-4b03-bdb3-c24c113e63ef · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93bb4cab-d672-46fd-99c1-8ec82e920c22 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725ca8ba-e711-44bc-8e15-922fa1a60f4e · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Large language model based multi-agents: A survey of progress and challenges.Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e709151-a9f2-4bc2-8a6e-ba657af64783 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Optimal aggregation of LLM and PRM signals for efficient test-time scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc1200c-7b00-44c3-a35b-51d1cad9cfe1 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Alphazero-like tree-search can guide large language model decoding and training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd68afed-f080-4a8c-94d1-3ce2556b7033 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Let's Verify Step by Step
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7871ef-a4e9-4871-a5dd-5bbea4099ec5 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4945d923-54f8-4128-837e-3e623c309379 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d4482b-d2ed-4861-8f08-7e9539d19748 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling LoRA: Low-Rank Adaptation of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ecd062-348a-4c23-a3ec-5387d9a851ac · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435d8dc1-3688-4ca7-b0cd-bbc0f42ed95f · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling MASPRM: Multi-agent system process reward model.arXiv preprint arXiv:2510.24803, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245e52f0-588a-42bd-998c-01d0c031dfe6 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Bandit based Monte-Carlo planning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af75e2d-ccb5-4628-9e21-e276cce17156 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8fe62e-29da-4336-bf87-e5d23e6de0b9 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling The linear representation hypothesis and the geometry of large language models.Proceedings of the International Conference on Machine Learning (ICML), 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2c41e9-03c1-425b-927c-df71efddf382 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Latent Collaboration in Multi-Agent Systems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b9409cc-0a22-4acb-a07b-ce7963d355ec · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Theoretical guarantees for iterative alignment of self-rewarding language models, 2026
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae8744d-8b3d-488b-ad1f-b8cd9e7e5eca · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a086d7e6-a75c-46bc-be94-4ee3876c8c49 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling An Implementation of Generative PRM, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f198b3-2a2c-41fa-ada1-37950124dce9 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Measuring mathematical problem solving with the MATH dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c05c28-adc1-4643-92ed-a0c9e6f13601 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Training Verifiers to Solve Math Word Problems
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38e5abc9-d374-4c01-aedc-c21f59e339fc · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Qwen3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad42a4b-b14a-46d6-995b-9c36023342e9 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Agent primitives: Reusable latent building blocks for multi-agent systems, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bad4c37-32a3-4060-8d0f-654e9d26543a · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Process reward models that think.arXiv preprint arXiv:2504.16828, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275c856d-35b4-46b1-83da-a774dafbc60f · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9d9b67-1f69-4caa-994b-770a59ea3eb2 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Tim-prm: Verifying multimodal reasoning with tool-integrated prm, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91de16bc-6979-4a74-b2a2-ae9ba9725aaf · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69340242-6609-43e8-8a7d-50e90919f871 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8cbf116-874e-45d9-9bd8-5fcda148f7c9 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033a4c51-5c81-40fb-9e6b-400cbe0eafed · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling CAMEL: Communicative agents for “mind” exploration of large language model society.Advances in Neural Information Processing Systems, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eee31d2-656b-4d16-9438-128df8eaff7f · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f5e00e5-9d9c-496e-9f0e-937c56d72062 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Efficient Streaming Language Models with Attention Sinks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a69b2cea-214e-4456-a2d0-762b0317e2d3 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Gonzalez, Hao Zhang, and Ion Stoica
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb87600-8177-4764-bc65-f7239c5be751 · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9573c2-ac1f-4a98-8ca0-a139b906b0cd · outbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling ?” (token ID determined by the tokenizer). The judgment tokens are “+
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622b07be-796e-44e2-8c40-75a36003a1ff · inbound
Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.