Pith. sign in

Paper Citation Record · LEDGER

Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.19332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19332 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:13.592468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:35:10.408592Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 243ea628-e798-41c4-acb3-3fda90368e0b · inbound

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences cites this paper.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.592468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.592468Z digest=sha256:ab6b62411f286c33bb41cdabe757426ea5739b781f03a16afbcd4cc35abbefcc

Observation 07e30d42-0401-4a3e-98d7-9f776f3d8e3e · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.599880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.599880Z digest=sha256:b49cc6ba68267a8bf79ab639a8870aa1d3d57d418ed026b54a72bb2ad1861899

Observation 22b91954-b84e-4ac6-a24d-098a88669763 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.202379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:d0ad2e943de1231b165c3f6532acd199821b8e4114b4440d203c22b78bbf3b1c

Observation 982bf544-798f-40b6-b81c-1003b1591837 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.410391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:193307f30e935a44409e8b09ba15bbee6da6361732fba726ccae14d003cbed8e

Observation 757cb0ef-18f2-4480-a604-30e6e4b799b9 · inbound

Recall Isn't Enough: Bounding Commitments in Personalized Language Systems cites this paper.

Recall Isn't Enough: Bounding Commitments in Personalized Language Systems Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T17:33:36.647908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T17:29:55.760720Z digest=sha256:2ef900b1edfb8142d760623e16b5fc9361d603fcd0e2be178f6f4e9d1a496184

Observation 04e9207b-5337-4d7f-af8e-b06cb0c08b39 · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.921026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:1978d4d4a29cc02d86d15e1903679ae9eed3fc087f5118461f11c2a72a2f69cf