Pith. sign in

Paper Citation Record · LEDGER

Human-centric Dialog Training via Offline Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2010.05848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.05848 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:25:14.589831Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T02:33:28.313220Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63266383-dc18-4525-a337-b70291d8de0c · inbound

Direct Preference Optimization: Your Language Model is Secretly a Reward Model cites this paper.

Direct Preference Optimization: Your Language Model is Secretly a Reward Model Human-centric Dialog Training via Offline Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:33:28.316870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:33:28.149013Z digest=sha256:94c158231a1e75c6d765f3d267042cf14836e84c4e1deedae0efc17f3dc3b9ca

Observation be89228f-fb0b-45fc-b0d1-a2e725cbf520 · inbound

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation cites this paper.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human-centric Dialog Training via Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.589831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.589831Z digest=sha256:858f0ab22d3cfdc0e19731f4878236078a8660fdbf28e8ad91c2916de0c73418

Observation c13c74ba-0510-45eb-b4da-2f6fa9eccded · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs Human-centric Dialog Training via Offline Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:26.565724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:26.565724Z digest=sha256:f476d81a317455328f3436227dd23bdee3252182c18bc71f4163608da89228b4

Observation 113730e2-dce4-4225-89cd-f8a6f13ae0e9 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Human-centric Dialog Training via Offline Reinforcement Learning

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.085890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.085890Z digest=sha256:dc2ea731403c7f848da37d64ae9eb519ca4807442fa530695d8ef57e15014b62

Observation 4998f26b-2ced-4ce9-90d1-62de6e372346 · inbound

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning cites this paper.

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning Human-centric Dialog Training via Offline Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:59.075724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:30:59.075724Z digest=sha256:ef877d775216ae10f1e9cab2aa6e55a1ad5208ea47681af310cc95c74a2f1a31

Observation e0a5731f-92a9-4aa8-864f-01ef077481da · inbound

Efficient Preference Poisoning Attack on Offline RLHF cites this paper.

Efficient Preference Poisoning Attack on Offline RLHF Human-centric Dialog Training via Offline Reinforcement Learning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:50:27.085411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T19:29:25.000361Z digest=sha256:2f1ac6c16ffdd7b87e0632b19ae02951bc17e6bb44588905e7e14c0096e58264