Pith. sign in

Paper Citation Record · LEDGER

Aligning Language Models with Preferences through f-divergence Minimization

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2302.08215.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.08215 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:01.090815Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:38.337799Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6dadec12-f05d-49b9-a48d-c7d236dc4c85 · inbound

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation cites this paper.

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Aligning Language Models with Preferences through f-divergence Minimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:00:51.544668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:00:51.487120Z digest=sha256:a55cced66c2c0f2adea13af35cbd1e9538437b021355b96194eb347138ae1fb9

Observation 691c90fb-b801-475d-93c1-5cea50e3aece · inbound

Frictional Agent Alignment Framework: Slow Down and Don't Break Things cites this paper.

Frictional Agent Alignment Framework: Slow Down and Don't Break Things Aligning Language Models with Preferences through f-divergence Minimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:01.090815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:01.090815Z digest=sha256:959418636bffece07306b78a9c5f338337179edac4671896b3c910dfc7ff4790

Observation 22bd6359-f939-409e-aa66-b52ff0a0f183 · inbound

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives cites this paper.

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives Aligning Language Models with Preferences through f-divergence Minimization

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.463747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.463747Z digest=sha256:47f89c4e78d2b5cd8bb6db1a04694c88d0840e667e09d1ad9b201de366069103

Observation 3650db20-dcf4-4961-8880-cdd394d345b4 · inbound

Threshold-Guided Optimization for Visual Generative Models cites this paper.

Threshold-Guided Optimization for Visual Generative Models Aligning Language Models with Preferences through f-divergence Minimization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:07.802152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T17:14:36.632493Z digest=sha256:3798831f383b159069c8551a0a13dbd86e779ff4a66a37fd387310efc697e3d2

Observation 7ae9e1e4-52e5-40ec-9e99-23398ed6694b · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Aligning Language Models with Preferences through f-divergence Minimization

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.797820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:67f91f4b520f4812074979d4d9cf1494686c7fe9baeb59bc216a0734ab52b9f8

Observation 6179d6e0-6c84-463c-a431-ccd278210e99 · inbound

Implicit Preference Alignment for Human Image Animation cites this paper.

Implicit Preference Alignment for Human Image Animation Aligning Language Models with Preferences through f-divergence Minimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:10:53.214110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:39:31.295358Z digest=sha256:c2edb8690a186e8cb82ccee57b378e964f21113b1efebda24c78b0078af362c3

Observation 36b90cf2-a4fa-4e29-878e-7bfe4c8bbd8c · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Aligning Language Models with Preferences through f-divergence Minimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:39:00.266297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:ef2180e96a24cbfb11132891c1602743304944eb40e72a5c69f445715a7ff8d2

Observation a8bb3765-7616-4be4-bc5f-15edbd829a3e · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Aligning Language Models with Preferences through f-divergence Minimization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.821613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:9bac976b2ad7acca80fb88826dedb16d1faf46d8f15da58bb53caafbc6f8cae8

Observation 6cc8abfe-8201-4de3-833b-38af0606583b · inbound

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training cites this paper.

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training Aligning Language Models with Preferences through f-divergence Minimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:38.339234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T14:14:45.547957Z digest=sha256:9a0fa3ab7493ec4034402293ed9ccc7f91dade2a5488d8cb19e0bee9b883a9c5