Pith. sign in

Paper Citation Record · LEDGER

Understanding the Effects of RLHF on LLM Generalisation and Diversity

As of 17 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 76 inbound Pith citation observations for arXiv:2310.06452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.06452 v3

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 76 of 76 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:28.162206Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact3
  • verified fuzzy7
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

14
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 095f85e7-82d9-412d-ba11-9a7a94f8fd36 · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models , booktitle =.

Understanding the Effects of RLHF on LLM Generalisation and Diversity A Diversity-Promoting Objective Function for Neural Conversation Models , booktitle =

Reference 1

Resolution
verified exact
doi, observed 2026-05-19T02:34:44.364357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T17:38:12.629976+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:6218ee44f57c247fcb6f52a8f4519df86377dde74ab3050880da2d48d53e631b

Observation 4224ff22-876f-4e9a-a82c-7dbefdd5191b · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Understanding the Effects of RLHF on LLM Generalisation and Diversity WebGPT: Browser-assisted question-answering with human feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:34:44.369528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:6a3f89daa81037835ebcf4ad5bb141283ef63ea27a7fb0ff18d59ec6729ed70b

Observation abf215aa-6c84-45b6-9dd2-b083a625335d · outbound

This paper cites GPT-4 Technical Report.

Understanding the Effects of RLHF on LLM Generalisation and Diversity GPT-4 Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T02:34:44.339926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:3ab8ff3656129fedc56f50d698e6cfd7fe40687964e531b3a48a7a04dbfb262a

Observation 6fb23c8b-53e5-4382-b547-a9523ad1ccb1 · outbound

This paper cites Learning to summarize from human feedback.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Learning to summarize from human feedback

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T02:34:44.347871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:f772059fbdbc716401e96f8db29a65391d17d4a1b9e08298cb2b99e98ea55b4d

Observation c4252e7f-bead-4a96-9bb4-3b91650b414e · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:34:44.357454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:7ebe9650d2fb8a5863dd5de05af82461bfabd13a95bd9595f0177ebfd80a9397

Observation ae5a7521-d7f2-47ee-8b76-1cf4381b312c · outbound

This paper cites believes.

Understanding the Effects of RLHF on LLM Generalisation and Diversity believes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.395474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:e2cb9aefd909919eea643d70cda69341470fd7f9c9e40bc73c6aa86e0a255a51

Observation 14eae0cf-5233-4aba-bc78-50af528675ba · outbound

This paper cites an unresolved cited work.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-19T02:34:44.398878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:0086b3a698fa0bd6f3432da5e2d480d3a50d69998a67b0d9036265a26302af95

Observation ae9bb66e-757d-4790-8f35-758387e370f4 · outbound

This paper cites For example, you should combine questions with imperative instrucitons.

Understanding the Effects of RLHF on LLM Generalisation and Diversity For example, you should combine questions with imperative instrucitons

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.402357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:09084e72a42f1ca851a220857dc7b2fa3ee128b47f9683099e01443c5bf11b73

Observation 8253de22-a081-491b-9c13-af0f48f81a8e · outbound

This paper cites The list should include diverse types of tasks like open-ended generation, classification, editing, etc.

Understanding the Effects of RLHF on LLM Generalisation and Diversity The list should include diverse types of tasks like open-ended generation, classification, editing, etc

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.406100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:ccc85562d16497dde1bc8416d6708bd385c1249bc0c0a6d33ee4218bef664032

Observation 4ad951ff-763c-4ea8-b8aa-efb50821d2fb · outbound

This paper cites For example, do not ask the assistant to create any visual or audio output.

Understanding the Effects of RLHF on LLM Generalisation and Diversity For example, do not ask the assistant to create any visual or audio output

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.373703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:bfafa2b4d1b4c560bdf2bb83cf6433edb121fcc1f8eb431bd8cd6276c266a8f9

Observation 13f762b8-98ed-4cf7-9d1f-a8bfdea3c6fd · outbound

This paper cites an unresolved cited work.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-19T02:34:44.377825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:f52e2a8ca7e7e43382e285d6f2e52b90b8a4bb9d8270e2581eac07016ccb9387

Observation be61dc26-b102-4818-a609-067ee3ae678a · outbound

This paper cites Either an imperative sentence or a question is permitted.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Either an imperative sentence or a question is permitted

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.381454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:f58e03213f1941bf615018375a9fc747966753ff109dabf875f698c645fff955

Observation 8c570ad3-3b22-4467-9043-a0b883ad6c82 · outbound

This paper cites an unresolved cited work.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-19T02:34:44.384722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:cfe963db5be0c30faaa6b01ec7f9060d1f0196b30a90fcb9e230735a2c9f99cf

Observation 09dfeb3f-8432-471d-a0ca-3ba804279bf9 · outbound

This paper cites Make sure the output is less than 100 words.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Make sure the output is less than 100 words

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.388105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:88df653d6c5f94839e4e264a000aa67d9440331ff5ce74e5d2d653ee43d4dee1

Observation 0202a8a2-44c5-4282-b71e-c9fb9e56074b · outbound

This paper cites J.1 D ATASET SPLITTING We create split versions of these datasets along several factors of variation in their inputs: length, sentiment, and subreddit.

Understanding the Effects of RLHF on LLM Generalisation and Diversity J.1 D ATASET SPLITTING We create split versions of these datasets along several factors of variation in their inputs: length, sentiment, and subreddit

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.392071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:b07dc231ab179e8d7aad5e22a1d7038fd65f2d281ea72b9aab480a50ea7de666

Pith citing papers

Observation 30866190-b7ec-43f6-9755-a980371398a7 · inbound

Towards Understanding Sycophancy in Language Models cites this paper.

Towards Understanding Sycophancy in Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T06:26:29.196349Z digest=sha256:e8b672829f67d92d5c6db68d1271ace1e55b35116e703ea56eed001fb43b5de9

Observation 9051f36c-75a0-4c44-b076-337f21694f72 · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:265ebb09cbbb6d60620c0143c55ef94ee4342818b977a0bc9de75c54841cdf1f

Observation fd64fbf1-7914-4739-a177-cc07d8e47697 · inbound

Adaptive Decoding via Latent Preference Optimization cites this paper.

Adaptive Decoding via Latent Preference Optimization Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:31:53.162886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:31:53.162886Z digest=sha256:2197d6a6c3b90a5d34da6916fccc16ea84897735d303623690031940241726c7

Observation dac9e283-2527-47fb-8842-c2df3acbb6cc · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.317823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.317823Z digest=sha256:c383834c276aecfe5c63b61ae9f50ac385ca6920951cddd3ba479a9566e554fe

Observation 13209e60-0223-48ae-ad90-738155a0883a · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.733324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.733324Z digest=sha256:86b33dbe4fa3a08ced44fd426db8cf7cf64f0358c21bd5c0a7544e7da17af56f

Observation 484e4191-a869-4b78-bbf9-af172e8cac7c · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.022520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.022520Z digest=sha256:05c60ea675ce29aa49ccc6be42b7e3bfcdbc7ae0bef912a885f52fa2d94d96a4

Observation d404cf99-d334-4188-ac77-f64d431bc6ef · inbound

DIVE: Diversified Iterative Self-Improvement cites this paper.

DIVE: Diversified Iterative Self-Improvement Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:13.413589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:49:13.413589Z digest=sha256:581e9a2a46bfa27c71425dd64f2c3993e4d14bf41b923be5ff1241e174e33056

Observation cac133d6-8dd6-4c6a-8397-d1f010ce696d · inbound

Aligning LLMs with Domain Invariant Reward Models cites this paper.

Aligning LLMs with Domain Invariant Reward Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.293458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.293458Z digest=sha256:d7d88f3fad80e6e3bfe0e83e68a9d275fc52bc561dc2324e2f9edd07b3648f28

Observation e67b475e-b965-4957-84be-5614ee3ffccf · inbound

Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models cites this paper.

Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:31.290165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:31.290165Z digest=sha256:42533df0c3ff24573e25f69c4c62d8ba3b1dbbe558c9ed771b8a4fd0f3e136b6

Observation 2e2a4f09-b6db-4598-92c3-e9807baa9d1a · inbound

Improving RL Exploration for LLM Reasoning through Retrospective Replay cites this paper.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.162206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.162206Z digest=sha256:7156c503cac63a4571968441e0d5a9366c278a7353011305b8ac2153301eaf13

Observation 4afa23c9-0dbc-4115-933f-3f5e42aff0c2 · inbound

Base Models Beat Aligned Models at Randomness and Creativity cites this paper.

Base Models Beat Aligned Models at Randomness and Creativity Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:10:56.161671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:10:56.161671Z digest=sha256:f157fba63e6e15b1212a2d958be220e709fdbbdcbcf2395b6eac76826244a866

Observation 70ac03cd-a255-48bc-80fe-72e4d5369f55 · inbound

When Bad Data Leads to Good Models cites this paper.

When Bad Data Leads to Good Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.196576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.196576Z digest=sha256:fcaa46175f7bf1d48f87ea45449d70274ab4aee9cbadf54bbc390b787f581293

Observation 8acb3879-ee10-48be-8a7c-e79a24b96458 · inbound

Large Language Models for Predictive Analysis: How Far Are They? cites this paper.

Large Language Models for Predictive Analysis: How Far Are They? Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:22.984029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:05:22.984029Z digest=sha256:33586e275305047f83fe6f24a73333a3d5aa6b5eebf7435fcf20ffa918bd61f5

Observation 969b2ab7-23a2-4163-9c25-0a075435e67a · inbound

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning cites this paper.

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:25.229173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:37:25.229173Z digest=sha256:f46c10988193121447093443267ecf58ca3c615f7435a1cc492b6fd59124e648

Observation c20c6894-9fcc-4cb0-a18d-f9177110cd8f · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.201102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.201102Z digest=sha256:c2ae0755df65c27581d28a6a1f967f30babb25e84a12d9556db7386c858dadda

Observation c0dcd5fb-12b7-40a6-b6b1-2f493cdcc838 · inbound

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations cites this paper.

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:47:17.998837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T12:43:51.019983Z digest=sha256:d5406c652023415f6c514f1f8bbcf279c13c3987bfaa42719e3dd146c302b2e4

Observation ba295577-6be3-464f-8101-1f94facba3d2 · inbound

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring cites this paper.

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:25.248418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:25.248418Z digest=sha256:828b8e5a8ff752be30d98781e3876a18192f4d17e8ef0d80d2584acb0aa77ae4

Observation fd67e504-3c5c-4e2a-bf38-ee91c3a9c828 · inbound

Normative Conflicts and Shallow AI Alignment cites this paper.

Normative Conflicts and Shallow AI Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:54.570349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:54.570349Z digest=sha256:398c011ffac69b0ee7f1f25c9201c08fd081f3560538db7af4646d02ff90a995

Observation 68c4f98c-df94-4fab-8410-6d339079d768 · inbound

Pairwise Calibrated Rewards for Pluralistic Alignment cites this paper.

Pairwise Calibrated Rewards for Pluralistic Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:01.776628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:01.776628Z digest=sha256:b4af9327470c1dc6985f12ff9182a61296b097c94c7572dbba5cc12ba46abc0e

Observation 4784308c-b4a6-411a-855a-1dc0db1d4d6d · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.573676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.573676Z digest=sha256:8a692f14ab3c9767428d311408201fb667633b54c271718c2d42d980e7609421

Observation cccd8799-141c-4517-9894-693752d3952a · inbound

Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models cites this paper.

Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:29:51.203465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:29:51.203465Z digest=sha256:7315ddccfde825f4dab38bfdeba8c6a8158ef7cc03b892b3534f87b513c0b4a0

Observation 664e1781-f0fb-42b6-88e7-dc4540d717e7 · inbound

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders cites this paper.

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-15T18:33:26.951710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:33:26.951710Z digest=sha256:bb90c89397cb894b2f5364108f237d286cf404fccd74a4c93cc2adfb95cb88b9

Observation 6f35f197-7f74-4c8a-9b95-02d5b7baa79d · inbound

CTR-Guided Generative Query Suggestion in Conversational Search cites this paper.

CTR-Guided Generative Query Suggestion in Conversational Search Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:54.037131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:54.037131Z digest=sha256:0989f3e90f74085fdcdec096d45a5f62cc5fe6cb1ccbeba7649ff51c5e7d80b5

Observation a2a64e26-8a74-474a-a8f0-5ee525036ded · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:17:05.883933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:b9e157745deacf85b3b8d21e7210e81d80372dbabdda164007e88d90026f6938

Observation c4763010-7a79-4ad8-ab8f-3ffa5754e0e0 · inbound

A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations cites this paper.

A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:27.773166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:46:27.773166Z digest=sha256:674543dbd4890d607ce7884d9dc0e91f57477b38c52356680a51f6c4ac48c133

Observation c2bd1add-cf10-4f58-9a09-4d93513f50ea · inbound

Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation cites this paper.

Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T02:20:03.766702Z digest=sha256:35f5d990eda367c232ea84bcac6532f8e49156cb3bfd10d2cde2dd2273b0b8ce

Observation 1e80e756-06d2-4fc0-942c-2348326119eb · inbound

RecGPT Technical Report cites this paper.

RecGPT Technical Report Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:04.694338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:04.694338Z digest=sha256:b5f13f15130eb80b6e7e17f31b9eceb45dc4a4635c3f4d51f18b79df12ed5860

Observation 066160e2-6d6a-4cf5-9ba3-380cfa25d568 · inbound

A comprehensive taxonomy of hallucinations in Large Language Models cites this paper.

A comprehensive taxonomy of hallucinations in Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:29:13.238929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:29:13.238929Z digest=sha256:d9e85386eb6b19fa18de5496796e7ff2270589c0c7f21a9951b2ec9781b0c5f4

Observation 31ba34a6-ad6a-4d61-9c6c-372f29c13ac3 · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.596674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.596674Z digest=sha256:125d9df8f8eb6819b4c8794d6beb9e1ff124a3157223d6e823008b36295f491f

Observation 7b8e508f-0763-4ffd-a1b2-567c015f9161 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:df0242fa1682a42873e44ebdf7c8f77f69ab0981186885e44ab6aeee7e3c1a9c

Observation 775e7468-cc0a-4f6d-a68f-af8a9799c45e · inbound

Spacer: Towards Engineered Scientific Inspiration cites this paper.

Spacer: Towards Engineered Scientific Inspiration Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:11:11.840589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:11:11.840589Z digest=sha256:83f33bf5cf2ce0e2357dee6475154352fa2a6fcd33e982e78c760971522d12de

Observation 3a26016f-4c03-4e91-8dcb-15e1f3edfaf4 · inbound

Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models cites this paper.

Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:07.651949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:19:07.651949Z digest=sha256:7d86bf613a579e65c67c2ad94f0326488bdc1eef6e8e54cd7202e6181ae0005a

Observation 9aaa1ec7-4b77-49e4-a0f2-a30c7cd3ee04 · inbound

Avoidance Decoding for Diverse Multi-Branch Story Generation cites this paper.

Avoidance Decoding for Diverse Multi-Branch Story Generation Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:55.106254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:56:55.106254Z digest=sha256:e9e3839bf89583b4f649f082c13b6202827a8f2667f58f709dfa801e1da563ed

Observation 94662f91-9583-418f-89b4-06d067697620 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.507748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.507748Z digest=sha256:94238e910eae6336f31ef84689bd9f305a6af871bb4ca2e4d0838fc9786c3e35

Observation d27768e6-cbb6-4fa1-b25f-ee3eb1833c0e · inbound

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support cites this paper.

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T21:52:09.158688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:52:09.158688Z digest=sha256:63a98b6214e015f9880c045c11d70db53e33975c10d69c8d392eaca82f51362b

Observation 53443ca6-d594-4faf-ac2f-91f04d57e61a · inbound

Generative Data Refinement: Just Ask for Better Data cites this paper.

Generative Data Refinement: Just Ask for Better Data Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:24:40.594050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:24:40.594050Z digest=sha256:40c0ba5f295e510b16d2995968dbdf9a82ccaace9bdbb726db93075253494815

Observation 5cc7c9b8-dafb-4441-910b-8ef0388ba4b2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:a77e638a976783830c13209de963deea716779349a0b1cb78fef7cd5d05aed0c

Observation e6198dcb-079c-4f2b-9327-599327b11ed8 · inbound

RAG Security and Privacy: Formalizing the Threat Model and Attack Surface cites this paper.

RAG Security and Privacy: Formalizing the Threat Model and Attack Surface Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T15:17:39.451092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:17:39.451092Z digest=sha256:b8d8858e6e4e535c08fcd1a0cc6461d9168b6ed8f561f021f6ff8f5ee540bce2

Observation e4f0ae8d-e82a-4240-b65c-68f0c77ce503 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.867471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.867471Z digest=sha256:a749165f6525470bd1d50ef6058f8d381ea1b521c2e163c08e86c1736a52e010

Observation 2096b447-0bd7-44a4-9045-96fa54ce940b · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:36.198906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:36.198906Z digest=sha256:5bc99edfe74e6a550168f3362ec48e838a23ba383c983fa15392bffae25dd06d

Observation 639f128c-b7ae-4ab1-9747-6b6c800ee5c7 · inbound

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction cites this paper.

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T14:34:51.316303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:34:51.316303Z digest=sha256:de623f988ad6aa01f3837b78c4dc50c6a2b20761667d58bf7a58bf86dfd0f169

Observation 8565894b-e350-463f-bb4f-9eee26cc4155 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 230

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:650fe4cedcde858f736bc365ea9c317285bba10ebea6c510562c0206f1ea563e

Observation aadc4c29-1411-45b3-927d-931aa901193b · inbound

BEAGLE: Behavior-Enforced Agent for Grounded Learner Emulation cites this paper.

BEAGLE: Behavior-Enforced Agent for Grounded Learner Emulation Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T07:09:29.814409Z digest=sha256:7eb17129a01dca58e42d0ea3c0a642a32d318e8ce0c3c1e3f87585697dd9649d

Observation b65e7330-a004-4772-9c64-4fb025066304 · inbound

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity cites this paper.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.197993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.197993Z digest=sha256:96fc68343b2feb22a81b807b7c6fc229b87187cacfff825d2da99708f367519b

Observation f3276d2e-e41d-43d1-b9e0-b0ac574d6507 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:25.374304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:25.374304Z digest=sha256:6555565b932b76c46205243a6228d7ce88fbe2a352a6345ae5a3892776fde6e6

Observation cfeec28a-5b19-4d3b-b96e-216747cfdd27 · inbound

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression cites this paper.

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:00.928770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:07:00.928770Z digest=sha256:e68396af5f843f9d9bf467d8b9c7c6aaa985af4dd2902d71132730cd50ba2eff

Observation ecedd5a7-16f4-4a84-b195-84538c5d991d · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:9ad74403b447d41565b41120af6ac868e6b1eb453f3508247d91322dbc8a862b

Observation b8a0b038-a05c-48f3-9c5b-474805522b71 · inbound

The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive cites this paper.

The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T15:44:01.229817Z digest=sha256:1ab9cab87342a36e755c6454ff16be2b1a327e66f774ed438507a0380b8fde10

Observation b263e130-898f-45ae-9a34-1947755b3223 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:793dcf237893ca707a8babe3e64af70a12cd388844762eac211b349e1da943c6

Observation 5ba0c5f4-613e-45ce-ac3d-cfc285c83035 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:73cea8715c39fe80be16be0a055580130bab83f219ce3d447833268bfe1bec69

Observation 251f3180-c9de-48b4-953b-66ab9af641e1 · inbound

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs cites this paper.

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:09:07.557773Z digest=sha256:c4deee20bf3a64b87aa4534b40bd6bcb256eaeef3a717ad14936741723b775bc

Observation a0c9c9ef-6cfa-4499-961c-058a2d5f6f14 · inbound

Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning cites this paper.

Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T10:41:59.816073Z digest=sha256:4de6fd02132947691fd32cfb80faf10b48085fe57fea3d759d0caf655e6c8061

Observation 657e7ec3-a008-475f-81ba-9649c2ec41d8 · inbound

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse cites this paper.

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T09:49:57.122084Z digest=sha256:ef3ae8b7d78fec9ad44c09a9d793257a19f3b8dccc446280d05913b16af91f13

Observation 95c7f978-6279-4be5-b93a-c4ab1f97c0f7 · inbound

Post-training makes large language models less human-like cites this paper.

Post-training makes large language models less human-like Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:47:23.619345Z digest=sha256:04abe701f9a11cc94263584858493ebe979dcff7b8677d5d5775c1b4bc316c49

Observation 3f2f4a67-2f64-4c34-8b46-44c4e5e3ebbb · inbound

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning cites this paper.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:8efbd4028173cf5742a618e774b1a765f64fc74a3aece35a24000031a1f677aa

Observation c87b1bec-0986-4b12-8a16-49e91f333a20 · inbound

Annotations Mitigate Post-Training Mode Collapse cites this paper.

Annotations Mitigate Post-Training Mode Collapse Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:58:11.179607Z digest=sha256:136abbbc6ab7365de562e62012480c8423ac1edc914ff3ee32368c0ae2e55f9c

Observation 9a28c782-0512-4a16-9065-3ac2f1f1208e · inbound

What should post-training optimize? A test-time scaling law perspective cites this paper.

What should post-training optimize? A test-time scaling law perspective Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T05:20:18.754351Z digest=sha256:dd3b113856e258d13c04324979f05703fa4d582ec202e0c5326c54993b17e4cd

Observation 6bc5ddcd-b8b3-408f-a79d-f47128e7124f · inbound

Differences in Text Generated by Diffusion and Autoregressive Language Models cites this paper.

Differences in Text Generated by Diffusion and Autoregressive Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T20:59:06.804446Z digest=sha256:c2244d6e02aaf6ac42ef67e515b45dac05e5f0c6b327d49c3940cf8a3331fc8f

Observation d767b407-ae35-441a-bb1c-d4c62e9d66e4 · inbound

RECIPE: Procedural Planning via Grounding in Instructional Video cites this paper.

RECIPE: Procedural Planning via Grounding in Instructional Video Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:28:05.577866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T06:24:50.714859Z digest=sha256:8a9e2978666ae805259a247690b95f62c257ec70a94bd8185c3d7a4b7ce6b4ad

Observation 98e5fd8b-a442-4ba9-b5ff-c5f9e7f8ce5a · inbound

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection cites this paper.

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:03:29.929578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T13:53:27.306664Z digest=sha256:0d7183ca612f135d032751a3c82b682380962555b3eae054fdac7b41af6e06bd

Observation dc999ee8-daf5-44e8-b90e-f9882bc79575 · inbound

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis cites this paper.

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T17:56:39.720418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:56:39.720418Z digest=sha256:77616f429ede4bdd936f2029642b1671c8d7a1bceedfbc6ce1df89b17e32e98b

Observation e1bcf991-daed-4a19-b818-a86891a9a738 · inbound

Argument Collapse: LLMs Flatten Long-Form Public Debate cites this paper.

Argument Collapse: LLMs Flatten Long-Form Public Debate Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.422997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T15:07:55.841128Z digest=sha256:e5e35e2fe8bd4b213b7910b2a580ebc272d467980050e2a3e036d2f7b7036253

Observation ba1f38c7-4133-4276-b6d8-ac37bd6d593e · inbound

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise cites this paper.

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.811920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T14:51:06.448678Z digest=sha256:89babd8a4149b7edcf70d8d023b4858400cd0f9ee4c8ab9673ae5890aafb0eaa

Observation dfbc2079-4ef1-4f92-a414-02ca0b163678 · inbound

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models cites this paper.

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:16:34.829871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:10:05.477340Z digest=sha256:4769ff333dead2b1803dc8e13776d44d73fcd9bf86fc63e4e3cbff05461a3e30

Observation ffcfde09-32e2-48ca-b1f1-18c34eca2c64 · inbound

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming cites this paper.

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.150703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T11:10:25.190293Z digest=sha256:225e9095d828ed2a3849f8431ee2eedee77144f32a1f73411ead8e030c615a02

Observation 705be100-e934-4b3b-8941-0eda957c22d0 · inbound

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming cites this paper.

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T15:17:34.380336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:17:34.380336Z digest=sha256:dbcbf3d2f973cbf4d3a8df06eaaa5c58b9634e4bc0b0425181c2119665fe1bea

Observation 5d57eb6b-fa60-4771-adba-d0c003d8d1f3 · inbound

Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources cites this paper.

Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:57.320099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T01:05:48.492775Z digest=sha256:a67be4e2c07ca83b428b5a2643f56baeeb0e5cf12bba88bea276d6951da34ff0

Observation 496ad41a-b12f-4d54-bebc-57b5843ef3a5 · inbound

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling cites this paper.

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:44:37.116492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T09:44:27.786630Z digest=sha256:fcaaf0aaf7c9c74d95c5242598e53b6eb0226c72fc95b719fb0a6d168c13f030

Observation 5ead472e-b5ec-42bb-b603-e0297fa82021 · inbound

Spectral Rewiring for Exploration, Purification, and Model Merging cites this paper.

Spectral Rewiring for Exploration, Purification, and Model Merging Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T05:08:55.438431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:08:55.438431Z digest=sha256:6bcfaa5587bbdfdf91eb36014f88b7e5097559b59166d836a2d6246ca625ec8a

Observation 5e2cf7ab-a664-460c-8752-a6e7bfb6a2a4 · inbound

The One-Word Census: Answer-Choice Conformity Across 44 Language Models cites this paper.

The One-Word Census: Answer-Choice Conformity Across 44 Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:47.598669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:23:47.598669Z digest=sha256:28af5b7383e82046eefad89be0f81f8f688729c6edd1074e604706f2f5f6226a

Observation 1338703e-e922-4efa-b958-7c0c22b10bd9 · inbound

Structured Output Collapses Answer Diversity Across 44 Language Models cites this paper.

Structured Output Collapses Answer Diversity Across 44 Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:23:43.844643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:23:43.844643Z digest=sha256:923197a0c967f932cafcc1f238da7db292d882e388061ccabaaf96025839a1d0

Observation c829a6af-1648-4c32-b739-84b78005ebf6 · inbound

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play cites this paper.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.245333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.245333Z digest=sha256:b7ff8244ca2aab6ec86bb5270d8268cf1bef4c7f9789921fa728a0e4b09f963d

Observation 05b2a897-3b9e-419b-aadc-8f88a138c232 · inbound

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning cites this paper.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.653093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.653093Z digest=sha256:ef12450921122d7136c0de8fceafb1e6bdddff82522b133f0d660292dc36a968

Observation a8e36753-a1cb-4b9b-9c2e-3d89feed25e2 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.171341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.171341Z digest=sha256:20d0f0c591fe978fc6df463845708f7f8f1d913b6a251a031c495c26ec4bdf2f

Observation ff146f2c-3166-4e4e-a8e6-f95a71976c56 · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 250

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.929855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.929855Z digest=sha256:d2712aa4c9374deaf96623f288bdeaa842174763830e8ed6a062c1da45c4d90a

Observation 946ec0d2-675f-4a9a-b234-131600dea20c · inbound

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions cites this paper.

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T20:49:20.267009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:49:20.267009Z digest=sha256:005e03f95113d419d2be27e6f4426480279069e76c2c947792d7841df0aa0660