Pith. sign in

Paper Citation Record · LEDGER

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation

As of 18 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2509.05605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05605 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:26:28.333393Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66cf6ac5-e364-4eaf-bedd-224fbf9a02b9 · outbound

This paper cites GPT-4 Technical Report.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.077854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.077854Z digest=sha256:3fefa57e41658bf171053352ba0aa5588389891822e98b7fe9dc1c62a985723a

Observation cdf17676-2052-4615-a015-06fcc0696441 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.082915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.082915Z digest=sha256:eaf32b47173a7b811d029c21aa2982e22b06c25d0d08c8cf90e571549d33537f

Observation 2e670908-0928-4ef4-acfc-9d51d7257c25 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.087997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.087997Z digest=sha256:72d7edcead2dd5f6727308fffc61ee1a8c935929ac1c0c8c30fe94385cb1f8ca

Observation 9184842c-6ea2-4e13-baf6-0a9d6477f4c8 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:26:29.142674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.092293Z digest=sha256:83ba8482c1427b4419415ef3f841fe6fb5936cd9c813139f3f654c1769c1efe6

Observation 9330844c-8772-43f5-9c0c-a05bf4dabdbd · outbound

This paper cites Language Models are Few-Shot Learners.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.096652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.096652Z digest=sha256:196c654d1a3f11530e3d9b2046b70f8267e7bc6d1c5d8d14b7e216b8104f1f61

Observation e5850e99-6e4a-4aef-a3eb-123cdc7cbed8 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.100887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.100887Z digest=sha256:f5600cb81412eb30c980944e996a43137f4e0b4c4623d791f27f61f413d8ff20

Observation 1dcc3079-23e6-4eed-bbdc-6585e841c9cc · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.105415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.105415Z digest=sha256:45116ccf5b09c37ac702be3e3edb74999a701aa588ceae4220fe1c19ab532812

Observation ad62c2ad-16cb-4875-87c4-a525c8d70a95 · outbound

This paper cites SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.109702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.109702Z digest=sha256:4cbe77f3380d3997bed544bff0b45b2a1319a87a791cb27780b9862df28ebb70

Observation 529fc7e9-1908-40e0-b5e4-88a4a1c9c449 · outbound

This paper cites Glass, and Pengcheng He.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Glass, and Pengcheng He

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.113882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.113882Z digest=sha256:3c6f3713512af8b1800a6977187685118efd30201e8c67120209f805bbaf9a52

Observation 28129f54-96d9-464c-bd28-57364f2a0450 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.117962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.117962Z digest=sha256:3347543b25fc6e44cac8a9ca942fc7075bab30a17772c6cafdd9e4424936e8d7

Observation 0dc989de-5001-4318-afe4-57189e6dcc74 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.122407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.122407Z digest=sha256:fd18ba8660c22fdaf2ff9f3973c7782990a05402ad785d34d56c9c7a72e10c4e

Observation c0129569-0134-4805-94b0-874a94cc9bc1 · outbound

This paper cites Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.126677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.126677Z digest=sha256:b7ed390b3969e224c434524f33d8cfa491b7d2b92286ac81228679ef30a16944

Observation 827526c0-ccb4-4d65-8c1c-fd3bc4a72070 · outbound

This paper cites Self-Boosting Large Language Models with Synthetic Preference Data.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Self-Boosting Large Language Models with Synthetic Preference Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.130817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.130817Z digest=sha256:64a25012dd6046999049c8579870a18043d94fc5073bd70c77751790b9c54853

Observation b9e25239-6e6b-4202-b0ba-8344602acddb · outbound

This paper cites Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.135081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.135081Z digest=sha256:cd09cceccca7aa01b89c069f46c14dbed8e503881cf6c7c9932e49e1a1902931

Observation 05120fad-7583-4620-a656-2e67b7148c5a · outbound

This paper cites Textbooks Are All You Need.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Textbooks Are All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.139214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.139214Z digest=sha256:fde1593b8f7560b9121faca45e08a5bdb974a1d369137ed644bde2da3a13860c

Observation 025ab5cd-d755-454f-9bdf-20fa5f48a748 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:26:29.113051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.143419Z digest=sha256:01622cd19dc71ed7353c87d753780c79c598091565c6370116c84013abc8d3ad

Observation c5f22c5b-81ff-457a-b9e9-8f127504e2d2 · outbound

This paper cites Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.147335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.147335Z digest=sha256:94169ef17f27992a443a263288ec7ef8c03d9405400257a13b0124972b3fe612

Observation 9c7c6a5c-d27d-4e3a-af12-1118a9374e9c · outbound

This paper cites Editing Models with Task Arithmetic.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Editing Models with Task Arithmetic

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.151613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.151613Z digest=sha256:617f5fc2d807116165f420378b663c271786475a9317c7a0d1f03234ed56b8c3

Observation b39543e8-506b-4790-9fcd-20abcb4732c3 · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.155957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.155957Z digest=sha256:3f9402ff7a0c78bf2897c37e61bcac018331cdb28172a7c6a42d34baa859cd3b

Observation ed2b16a6-e50d-4d97-8481-959b12427a30 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Aligner: Efficient Alignment by Learning to Correct

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.160021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.160021Z digest=sha256:0d9e6d0ca2f7c7fbed8cfacdbcab489cb4ad109370563817d10a7217b2965ea5

Observation 1e2647a7-ae0f-4d91-ab5a-4fddeb433b5a · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.164132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.164132Z digest=sha256:5748265c45a3ebce00f63024e57bbd9d2e8a0b3aa5fe76cc5f23f8b5cae90cd0

Observation 38c95821-9801-48c7-af24-261acf5f0d59 · outbound

This paper cites Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.168263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.168263Z digest=sha256:4fc4d0e90c0042e5055c13ca9851a61776e4509db2f9ca0bfcbcfb791f6370d0

Observation 01f9cf6c-beaf-49f1-893e-f12a54b3f944 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:26:29.099814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.172421Z digest=sha256:a0372548bfcb95d0a236073f2a9c41007fbe57b301057ca5a5ab6fadd88805db

Observation 1a10f64c-d628-43d3-87a3-ac3292aef1bc · outbound

This paper cites Hashimoto.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Hashimoto

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.176615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.176615Z digest=sha256:ca78f8bfaab276f52fa56681cc636dea112329fe49948ed7658feab5b67a9bf2

Observation 27ea4841-40c7-48ee-9621-fced0599f39e · outbound

This paper cites Holistic Evaluation of Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Holistic Evaluation of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.180554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.180554Z digest=sha256:95c7d698780e5e57d925b07f207b589fdcde3afa28293c4f03fbae6054e03b9e

Observation f92d8037-876b-4bcd-8982-03d9791a4c34 · outbound

This paper cites CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.184749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.184749Z digest=sha256:a699e525d1bf6d409b7c1a2b32fc90bb146e05054fbde1b7f683d776abc8798c

Observation f4b40501-cf7f-4db8-979f-6c6a494664bc · outbound

This paper cites Aligning Large Language Models with Human Preferences through Representation Engineering.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Aligning Large Language Models with Human Preferences through Representation Engineering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.189157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.189157Z digest=sha256:b42cb87131b0010c53cbc25b747e4e1a0151e5dd355df22c04189576d6072b08

Observation 5a2be0e1-423f-48a9-8a67-8811483321bd · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:26:29.078397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.193270Z digest=sha256:168e596fbaf0df0ad7a8102410a74183c78a14904ec7a3c113d920245ca1a699

Observation 844276d1-41a8-416a-9280-a9e2de82f034 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.197031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.197031Z digest=sha256:0feae2a3e7ae261d5dca98bc129ca0124f6decb62714f2260cde03bcb2bc20e6

Observation 9e386091-3e8b-4de5-b2e5-dfc5a0ea95d6 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation WebGPT: Browser-assisted question-answering with human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.201235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.201235Z digest=sha256:008834abdce6cfd6135f32fe101f2f1d87277e1def01b196584fe286cc8bfe87

Observation 53f79934-eba7-4cef-abcb-20f16e597f32 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:26:29.064281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.205256Z digest=sha256:adef9a66bb5256584d2b06b1c71a6d34becf355774d1555871e4b60479430e80

Observation 0acd3a7c-3d69-4ce9-97af-0d078c6a4ce0 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.209141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.209141Z digest=sha256:e9ce6b194a264c3f72582e88e0d105a666377cfe404572da3a3dabd5cb6a4051

Observation dbbe9980-6246-467f-a867-cbbd0eec6d06 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.213160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.213160Z digest=sha256:ff4d18b31558c612fbb8328a9eaef546d62e011389c1e129034d844271d2cadf

Observation 91ff14f0-7105-44d0-8e0e-58382c039e89 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.217067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.217067Z digest=sha256:0445cbfcc1c8784a5537b4bf45a520c7d8a0208ddb481cbc6ff27b5b4615e3de

Observation c26ac2cc-cf58-4694-ad3b-9b683f6bac3b · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:26:29.026271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.222201Z digest=sha256:cfc8554b63a3adaf646e612c6ee7a1f506dd0f1350eed7d4755638e0adb0fbcc

Observation 1c2c299d-2685-4036-902d-be10934c3d51 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.226231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.226231Z digest=sha256:e60c692322b0df9d17ffeabebd4081e61cb0571b1ddd032b35020e7d7e807574

Observation f29b6e90-467a-4a53-94cd-d3e8faa8a9f9 · outbound

This paper cites Transformer Layers as Painters.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Transformer Layers as Painters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.230057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.230057Z digest=sha256:c8e63d1e1043c0f522371f97366b8b89771ba481282c2b0574770718432ff298

Observation f94abd80-195a-4736-9a48-d3e717cd2a93 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.234349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.234349Z digest=sha256:61013721557c3968ccc7b80a41eebe2098416f6425647274448e2ebce5c8dd8d

Observation cc5a5e99-64cb-419f-8ae3-1f2d316e3d4d · outbound

This paper cites Hashimoto.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Hashimoto

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.238364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.238364Z digest=sha256:2b5916a59ffba8103eeb4c79e5af150a154eb9b6717933ad2d1ea7ec2ba160f0

Observation f699c2c4-ce71-466b-bd0f-67b962e34a02 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.242221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.242221Z digest=sha256:118abdeca1024480653ab636bf5892cbee403000f15844bbd749419a48feb4c9

Observation 52d026ba-6fd6-4e0e-85b7-70e56c2621a5 · outbound

This paper cites Daniel Freeman, Theodore R.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Daniel Freeman, Theodore R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:26:28.986769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:26:28.246014Z digest=sha256:0454fc73e702e7595e05f7e73b795cdeebf19614da731d4661af1f8d6bc9d9a7

Observation 31044c9b-c5f4-4420-a51a-5e7b59f8bb1e · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.250013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.250013Z digest=sha256:d6ee68aea1b081879cca90ee99ba585d11d6ab621cface8eab17778e0f028dba

Observation 71a8e492-c989-48f9-bf03-c6ee3ae064ab · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.254256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.254256Z digest=sha256:7bade2e7b607bce1deb5e811efcb5621daa254952b1892faa8916e9e39825c5e

Observation 3abe4705-e3ba-4589-bf49-2191e774d1ae · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.258543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.258543Z digest=sha256:f1d805cea7ab2618e451f2ea7bb5c0143277386e2038d8ee8e6f42f9fc7671a4

Observation f29bbc14-9b95-48e3-94d1-d20bc83f40fe · outbound

This paper cites Improving Text Embeddings with Large Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Improving Text Embeddings with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.262513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.262513Z digest=sha256:a13f0baa111b18e999494e6c136cf9c93a5f4d2056f5ca582736f7854ae3fe2f

Observation 86556b56-1481-4737-939c-a1ab376e9b8a · outbound

This paper cites Self-Taught Evaluators.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Self-Taught Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.266680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.266680Z digest=sha256:f0a5b9ce7553fa8cd102f69c0d93cde2a8f3011ab0421223492031c3dbe6bc25

Observation 5ac8f2ec-9ef7-4302-8c30-8edbec086fba · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.270685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.270685Z digest=sha256:f03c3d1476cccc7c4754a6d6b7edffbf3e839a529c8bf2ab23e0837c14614da3

Observation 2a9ad9c1-9123-4fe3-a4c9-e469c242c1f5 · outbound

This paper cites Ethical and social risks of harm from Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Ethical and social risks of harm from Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.274599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.274599Z digest=sha256:784578e014657735438902d107baf858b9674afc41203ae9d87b118f32504c46

Observation f4b0f4fa-b604-42a8-98eb-15ad11b18e6b · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Language Models Learn to Mislead Humans via RLHF

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.278753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.278753Z digest=sha256:64afd0b1ef0a506cf42d064f5e5364820eb7287785741b05a70d03116ba102d6

Observation 0a9c9413-3d3a-4448-958a-096170e80aeb · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.282926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.282926Z digest=sha256:b0112f39c491690ede9f562c6190a8b09ccac4f1a33ff36e715c262750e3d10b

Observation 72ade002-bab1-40f0-92c5-d8d95c574586 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.287446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.287446Z digest=sha256:5e4dd99f0ee9c2a65d2d4b856b2515485e08c8d2cafb0cfde461e5531326070a

Observation b8efebee-8cf5-444f-82b4-ec0e0de00563 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.291535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.291535Z digest=sha256:41bca13c3dacdeeca5a5fb38f3a6a7fb5a22bb9ea3ec1e99dcfbf2f43a6d94fe

Observation 081e16bc-34c9-4643-bd6c-44f61643b2e9 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.295748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.295748Z digest=sha256:7acbccc16fa97fd031133d38f5c9ed488c7256d88190a1e3a48cd0cf76ee995c

Observation cff26305-7447-474e-91a4-c5fda0a5799b · outbound

This paper cites Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.299762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.299762Z digest=sha256:41cb6d1ce61d6310bbde518059089ed37224db3aa8219bad237f477bc8c7e583

Observation 9e05a5bf-779d-4ae6-b276-58af7a81de00 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.304020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.304020Z digest=sha256:ff5c197839c194b4731af6239a8d9261dcf5809c23e7a1786420e9120999d4be

Observation d7bd3656-d3c4-4d7e-bdca-adc3cee26df6 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.307933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.307933Z digest=sha256:26f3452f8d495844556bdcbfcc83eea50c19cdf17abedecf0705ac5f647079bb

Observation 84b31d12-c8d8-4a3f-807a-4bff65f89a7a · outbound

This paper cites TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.312016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.312016Z digest=sha256:9996ffd4481c2ccc37799cfb0f5fbab5f80ccd67caeb027aa3f876508e0e6a9c

Observation 3f29823c-e246-4938-abf1-fe746c40fe39 · outbound

This paper cites an unresolved cited work.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.316269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.316269Z digest=sha256:9acc7ae437d74e589eaa10cf5bc1aac8c7522418aff73807279aa86e1f3dc16a

Observation 7045ec7b-c87c-4e50-972b-5c28e6e7fa43 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.320793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.320793Z digest=sha256:ff54a8cab9c8930312e5d1415d256af46590b641a4ba7d473af6d650cd5065d8

Observation 9c183357-56c7-4558-9331-d03f8f8dd81f · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.324931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.324931Z digest=sha256:045dc174719aa239251e2117cea98e5936f1e2f471487b57863087970628b24d

Observation 47874377-6f59-4226-82bb-ed1cf87b9a1c · outbound

This paper cites online" 'onlinestring :=.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation online" 'onlinestring :=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.328967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.328967Z digest=sha256:07f1501c17d0ee98590bcaf58fa9707e3fce2e2e72191383fe0250788c0e6557

Observation b724e026-3dfe-4368-9071-c7d4dabe1eab · outbound

This paper cites write newline.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.333393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.333393Z digest=sha256:ff42a96cf03b802413974ed39d83439058f07030b63c19f7f55ef57755f84f04

Pith citing papers

No inbound Pith citation observations are available.