Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:07.007761Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2510.00492.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:07.007761Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T19:26:37.281563Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T19:32:35.006053Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3616f1b2-250e-4be8-b642-f0b57a2df403 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e59c76a-550e-4fa0-9a65-df962d331f29 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c5312b-3629-400c-97a7-3a96d22802e7 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling ForgPRMwithsampledverification CoTs, sampling contributes per-step noise:Var(ξ (g) t |x)≥ σ2 +τ 2 for someτ 2 >0
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b2a205-c0cb-4700-9f7e-a7fed639d88e · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Finally, Jensen’s inequality gives(max{0,E[∆ mean]})2 ≤E[∆ 2 mean], so the MSE bound follows
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfcb354-9fa0-46ac-8c05-00e28e555a11 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184ff990-901b-4f1a-ac2d-d93099a815e9 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling SinceK X (0) = 0andK ′ X (0) =E[L|X], we obtain logµ(X) =E[L|X] + Z 1 0 (1−θ) Var θ(L|X)dθ
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61acb7e7-a731-4913-a6f5-83794f842caf · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Moreover, sinceL=ζ A + ∆g-prm =ζ A +B (g) +N (g) withE[N (g) |X] = 0, and sinceζ A and B(g)(X)are constants when conditioning onX, we have Var(L|X) = Var(N (g) |X)
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8d4ea7-05a6-42b4-b0ef-ccb7dc071f99 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling •Verifiable: The step can be verified using common knowledge, simple calculations, or a quick reference (e.g., recalling a basic theorem)
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4450ed59-ef09-42c5-919c-25348b4e28c1 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Good job!
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cfcce26-6ac6-4e7a-b25c-4697b254e990 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling •Is Hard to Verify: Requires significant effort to confirm due to poor explanation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0d7f3f-233b-4f31-a391-da2344fa692e · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23a44bb-f159-4a2f-bd1c-ac8bffdf8951 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b951a61-da03-49f5-bd17-c5b9d925bb14 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Verification: Is the answer correct (Yes/No)? X
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21f6dea-df10-445f-8b7f-a62bcb216838 · outbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling Verification: Is the answer correct (Yes/No)? X
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22763588-8056-4a0a-b560-8252f57f585a · inbound
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Rethinking Reward Models for Multi-Domain Test-Time Scaling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 915a1bac-daf4-4d31-9c10-e0439179413b · inbound
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Rethinking Reward Models for Multi-Domain Test-Time Scaling
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.