Pith. sign in

Paper Citation Record · LEDGER

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2505.02363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02363 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:00:59.166273Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afa7c4fe-c646-41fd-a6bd-0cba96c303e7 · outbound

This paper cites Cited on pages 5 and 16.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on pages 5 and 16

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.366877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.119432Z digest=sha256:1cebcc20af7008f450287f35c129f9a27bf4bbe72e18f443fba04ed79db84710

Observation 34a21b43-a5f4-4ce6-86bf-5441e713869f · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Direct Language Model Alignment from Online AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.123289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.123289Z digest=sha256:6a3fe45fdb52fd98d584c06f29f0f1118c737539f3422044b341a3483267fa42

Observation cf4e34e0-4cb4-4431-a782-b4a8b9b6f72a · outbound

This paper cites A Systematic Examination of Preference Learning through the Lens of Instruction-Following.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning A Systematic Examination of Preference Learning through the Lens of Instruction-Following

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T01:00:59.264310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.132394Z digest=sha256:bad7edd5953864323b30a5794a1be50d0a1d1df4f19cc51e5067021a8634bd40

Observation b31bec9f-a623-4816-8e5b-1a24272994c4 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.137081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.137081Z digest=sha256:abfebf05587e87ebcb1ec828bfc5ff400484584618736fb00dc8ca8d4b7780fb

Observation 793b61c2-6c27-4657-a4c6-d876e1558bd5 · outbound

This paper cites Cited on page 7.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on page 7

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.354235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.141276Z digest=sha256:476203636ac1db01847da317da3d0d8e43fb2ce5f64c0d7bb87dad00d70f1239

Observation 576fcff5-0de5-47ff-8f57-a900f2251c2b · outbound

This paper cites The Importance of Online Data: Understanding Preference Fine-tuning via Coverage.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning The Importance of Online Data: Understanding Preference Fine-tuning via Coverage

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.145137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.145137Z digest=sha256:23bf3ba615fd381770ade193e7c9aea591e81bab6b11cbd441690368c7ecfd1c

Observation b00d6e0f-6d90-4bd3-9caf-f5d435a1b991 · outbound

This paper cites cc/paper_files/paper/2020/file/ 1f89885d556929e98d3ef9b86448f951-Paper.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning cc/paper_files/paper/2020/file/ 1f89885d556929e98d3ef9b86448f951-Paper

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.341017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.149804Z digest=sha256:7ff3422a74b7fa0656cf12eaf373cd8150aad9cdd86e35dba4ea609d47441c67

Observation 574c9829-6181-4031-af24-46290e623a13 · outbound

This paper cites Cited on pages 1 and 7.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on pages 1 and 7

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.328588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.153805Z digest=sha256:e7a27664855ae7941621adf810e8ad2b9f4952236144f36f75e6fec9d1bd781b

Observation b61b3eeb-439c-4c70-9546-e2dddc94c7d9 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Understanding the performance gap between online and offline alignment algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.162363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.162363Z digest=sha256:c0191f7d7de7335d60b59f6340598e395f44f7c8da055f27e484217e758314b8

Observation a1acd6bc-b786-4519-97db-86183bb6da36 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-16T01:00:59.166273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.166273Z digest=sha256:6584dfd2097a81e1a05b619ceda6f0e59693ced08db87148a2fffad609d6ab21

Observation b5cbbd16-688e-4e4a-9e57-33d7c7e4f2ae · outbound

This paper cites Mistral 7B.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Mistral 7B

Reference 309

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.127796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.127796Z digest=sha256:cd62d04276cd9dba161fc310a5465fb00953fb2facf44ddfb28c9143bf4970f5

Observation 8e809b21-aa7c-4dbd-9dea-3d88109d9200 · outbound

This paper cites acl-long.662.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning acl-long.662

Reference 662

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.391858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.101760Z digest=sha256:2864c72ac25b01445e955e20c8303e176620a3e7641492d494ca99e33ed34765

Observation 041082b1-e0c6-4775-9b78-516ac6f1b0af · outbound

This paper cites URL https: //aclanthology.org/N19-1421.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning URL https: //aclanthology.org/N19-1421

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.157860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.157860Z digest=sha256:c653a678e81eb3be18be1c551b266783ee8a78f4da4aa940c84b0de76621f6f5

Observation 99eb5fc0-a8a6-4d90-b54b-f04130792ae2 · outbound

This paper cites Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T01:00:59.305430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.114689Z digest=sha256:82b4b0fb7a02f155d04f1bfa7baaeaf816e4fbbe4fe9bd6cd88755541b95f4bc

Observation 1fb577bf-7109-4d52-a626-72cd5079c479 · outbound

This paper cites Cited on page 7.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on page 7

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.379338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.110420Z digest=sha256:cbd532a1195972421072f862c0f4a0a5ed4b01a351d533a4063943d1d62fc7f7

Observation e2ed02f9-79ae-42fa-a24a-62ceb5f3e2ac · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.106363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.106363Z digest=sha256:10ab8d50f4ca9f8b1243c2dee4fd7bf2bf02f164a44d5221644bf75d05dee393

Pith citing papers

No inbound Pith citation observations are available.