Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2410.02504.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:52.622227Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T14:04:44.920128Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 11e2c045-38c1-4fd7-85f6-88d51aca9b29 · inbound
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96f612f8-deda-4ef3-9fb8-0ac8cffc6498 · inbound
Reinforcement Learning from Human Feedback: A Statistical Perspective Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30ecb29b-89a1-4498-953b-8ba6d98c0ed8 · inbound
Perturbation is All You Need for Extrapolating Language Models Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ec06922-beda-4dae-947e-1c60b9cbdcc9 · inbound
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa36ede1-00c4-4528-8995-0c6d95bda78f · inbound
Learning Perturbations to Extrapolate Your LLM Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0eb80bdc-aa76-44a8-a5ba-841d45209198 · inbound
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 23c2d6ca-df49-472a-88d7-e5924ca7a54b · inbound
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.