Pith. sign in

Paper Citation Record · LEDGER

Understanding the Effects of RLHF on LLM Generalisation and Diversity

As of 12 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 66 inbound Pith citation observations for arXiv:2310.06452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.06452 v3

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 66 of 66 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:57:01.733324Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact3
  • verified fuzzy7
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

14
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 095f85e7-82d9-412d-ba11-9a7a94f8fd36 · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models , booktitle =.

Understanding the Effects of RLHF on LLM Generalisation and Diversity A Diversity-Promoting Objective Function for Neural Conversation Models , booktitle =

Reference 1

Resolution
verified exact
doi, observed 2026-05-19T02:34:44.364357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:74a3f2459cabc48f1b65bbd68df049ef4f795f4f5bd52fa6b698c303dc17df2f

Observation 4224ff22-876f-4e9a-a82c-7dbefdd5191b · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Understanding the Effects of RLHF on LLM Generalisation and Diversity WebGPT: Browser-assisted question-answering with human feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:34:44.369528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:b209f7f3eaf1bc2fbb85be37a06d5872a0e74a42ae769a0f1a125c55ea6c2433

Observation abf215aa-6c84-45b6-9dd2-b083a625335d · outbound

This paper cites GPT-4 Technical Report.

Understanding the Effects of RLHF on LLM Generalisation and Diversity GPT-4 Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T02:34:44.339926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:dfc7561f44e1ce6d9437281ceb8d0c0b74da8d9f40f27009276414acb5b38928

Observation 6fb23c8b-53e5-4382-b547-a9523ad1ccb1 · outbound

This paper cites Learning to summarize from human feedback.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Learning to summarize from human feedback

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T02:34:44.347871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:80e748d4aae5d1054aeb7096ff324010f0249d34042799a31282116bd67dbdc6

Observation c4252e7f-bead-4a96-9bb4-3b91650b414e · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:34:44.357454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:00fa2165adf4149ad4778e89f7b5ef8e33567534eae9df686701b0e311cf55b5

Observation ae5a7521-d7f2-47ee-8b76-1cf4381b312c · outbound

This paper cites believes.

Understanding the Effects of RLHF on LLM Generalisation and Diversity believes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.395474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:73c4103fa82d789d5742ef5d1242a778f731c90c28ddcf5a822886a82c00cc32

Observation 14eae0cf-5233-4aba-bc78-50af528675ba · outbound

This paper cites an unresolved cited work.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-19T02:34:44.398878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:3232619a36ccaeeb3ea1e8b52c208e039be5fb671d4d084ac98aa96c56978f8c

Observation ae9bb66e-757d-4790-8f35-758387e370f4 · outbound

This paper cites For example, you should combine questions with imperative instrucitons.

Understanding the Effects of RLHF on LLM Generalisation and Diversity For example, you should combine questions with imperative instrucitons

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.402357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:6a15d0864c646f7393e1abbbb3e6d0c0b6eb3868841eec13677d03f2d103d6c9

Observation 8253de22-a081-491b-9c13-af0f48f81a8e · outbound

This paper cites The list should include diverse types of tasks like open-ended generation, classification, editing, etc.

Understanding the Effects of RLHF on LLM Generalisation and Diversity The list should include diverse types of tasks like open-ended generation, classification, editing, etc

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.406100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:73aa168e93a6231d3230936ab13e4c3ec2414d6e422a5878f67e4cfac8644e1f

Observation 4ad951ff-763c-4ea8-b8aa-efb50821d2fb · outbound

This paper cites For example, do not ask the assistant to create any visual or audio output.

Understanding the Effects of RLHF on LLM Generalisation and Diversity For example, do not ask the assistant to create any visual or audio output

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.373703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:57f1fbd43c6bfb6f2b694c6ef8f9194f3628d08b441958330f2f29f950cc3d46

Observation 13f762b8-98ed-4cf7-9d1f-a8bfdea3c6fd · outbound

This paper cites an unresolved cited work.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-19T02:34:44.377825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:c06e13db376d50fcb434ed64a434c4e3d415302946de1f5e1d59252c3a3af0f1

Observation be61dc26-b102-4818-a609-067ee3ae678a · outbound

This paper cites Either an imperative sentence or a question is permitted.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Either an imperative sentence or a question is permitted

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.381454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:46ea1d769641705265f0500aeee1f0f7b511d3ecd2d294e9009503241ad68e4f

Observation 8c570ad3-3b22-4467-9043-a0b883ad6c82 · outbound

This paper cites an unresolved cited work.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-19T02:34:44.384722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:3200715208a9e6a84c30974b3edf44835c137d00bd2fe0d23a14ccda31077fd1

Observation 09dfeb3f-8432-471d-a0ca-3ba804279bf9 · outbound

This paper cites Make sure the output is less than 100 words.

Understanding the Effects of RLHF on LLM Generalisation and Diversity Make sure the output is less than 100 words

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.388105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:82b898c2a58f8c522939752a6154f863d25260b31872fd38eacac301508c548a

Observation 0202a8a2-44c5-4282-b71e-c9fb9e56074b · outbound

This paper cites J.1 D ATASET SPLITTING We create split versions of these datasets along several factors of variation in their inputs: length, sentiment, and subreddit.

Understanding the Effects of RLHF on LLM Generalisation and Diversity J.1 D ATASET SPLITTING We create split versions of these datasets along several factors of variation in their inputs: length, sentiment, and subreddit

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T02:34:44.392071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:34:44.313274Z digest=sha256:928d5856eeea4717ea203dea9ded44333616002d51e9229d84d15d84ecf3b297

Pith citing papers

Observation 30866190-b7ec-43f6-9755-a980371398a7 · inbound

Towards Understanding Sycophancy in Language Models cites this paper.

Towards Understanding Sycophancy in Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T06:26:29.196349Z digest=sha256:1d9993d6f97e502530865206b4f9332b1b1f929c441425d4daf6b6028f9b18db

Observation 9051f36c-75a0-4c44-b076-337f21694f72 · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:14da090867de0b56a44193f9e55f757ca00778f048c358c587b0ba0100d88682

Observation 13209e60-0223-48ae-ad90-738155a0883a · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.733324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.733324Z digest=sha256:43f099e2184b76aff20a307c314837bdfa5d1625a79af8c96e93233b1edf2f15

Observation 484e4191-a869-4b78-bbf9-af172e8cac7c · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.022520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.022520Z digest=sha256:9c15287346e70f5e4aee2a9a7d861f6d7e06f3385956a72023719499a0279275

Observation d404cf99-d334-4188-ac77-f64d431bc6ef · inbound

DIVE: Diversified Iterative Self-Improvement cites this paper.

DIVE: Diversified Iterative Self-Improvement Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:13.413589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:49:13.413589Z digest=sha256:f928784894efd2d8e47aade3ffafb86d0d07db61828f9960be8d4d29a59b2bf3

Observation cac133d6-8dd6-4c6a-8397-d1f010ce696d · inbound

Aligning LLMs with Domain Invariant Reward Models cites this paper.

Aligning LLMs with Domain Invariant Reward Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.293458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.293458Z digest=sha256:44f9f3c02593c2a03f392dd90dbce521cbd9d07d656c390514309885e4a2347b

Observation e67b475e-b965-4957-84be-5614ee3ffccf · inbound

Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models cites this paper.

Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:31.290165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:31.290165Z digest=sha256:98e5d156d096ce1f2e0520f43e50f8e988149bf168c696b8ca2dabe82feaadce

Observation 8acb3879-ee10-48be-8a7c-e79a24b96458 · inbound

Large Language Models for Predictive Analysis: How Far Are They? cites this paper.

Large Language Models for Predictive Analysis: How Far Are They? Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:22.984029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:05:22.984029Z digest=sha256:0c4098584bd319ac0816f00545c6743042efd44b77578075d40ea8f4e7d92a15

Observation 969b2ab7-23a2-4163-9c25-0a075435e67a · inbound

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning cites this paper.

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:25.229173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:37:25.229173Z digest=sha256:d38b8ab51734dc3eaf9d085f4ce16d6e5bf1057f5e6ea4c39a54d115689d23af

Observation c20c6894-9fcc-4cb0-a18d-f9177110cd8f · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.201102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.201102Z digest=sha256:6fa5e75a8a53b4dce673db69abcc697050b1db792b6eab5fb802e64cff2bcaba

Observation c0dcd5fb-12b7-40a6-b6b1-2f493cdcc838 · inbound

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations cites this paper.

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:47:17.998837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T12:43:51.019983Z digest=sha256:d4b8d6c1724056312d4f058d4e6d73b059e861380d476025b1e662e3d5cf81d6

Observation ba295577-6be3-464f-8101-1f94facba3d2 · inbound

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring cites this paper.

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:25.248418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:25.248418Z digest=sha256:a623a9b6560a8d0335c0ced4e88203bfcb7713f2b05da5b6be737a39ab40a192

Observation fd67e504-3c5c-4e2a-bf38-ee91c3a9c828 · inbound

Normative Conflicts and Shallow AI Alignment cites this paper.

Normative Conflicts and Shallow AI Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:54.570349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:54.570349Z digest=sha256:5defbb0bf3b2780377a8363802b1631f5ebec031608719213894470c8a4e1037

Observation 4784308c-b4a6-411a-855a-1dc0db1d4d6d · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.573676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.573676Z digest=sha256:3d4a12233999a85685a01625005a40539d58d946a6ad5e187dc61129e4aced7c

Observation cccd8799-141c-4517-9894-693752d3952a · inbound

Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models cites this paper.

Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:29:51.203465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:29:51.203465Z digest=sha256:41d3b2bf67cca923d138e34d0fe53fc685ddd70c26b41f97c4037b2529d19070

Observation 6f35f197-7f74-4c8a-9b95-02d5b7baa79d · inbound

CTR-Guided Generative Query Suggestion in Conversational Search cites this paper.

CTR-Guided Generative Query Suggestion in Conversational Search Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:54.037131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:54.037131Z digest=sha256:1bcccc16649757278a23dab49d3da1dedff7f42bd9eb6ed62561f26d1b6c21bc

Observation a2a64e26-8a74-474a-a8f0-5ee525036ded · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:17:05.883933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:242b1d0681f4d172cbe11c43d448797f928c18bc4f948fe0808873105193f32e

Observation c4763010-7a79-4ad8-ab8f-3ffa5754e0e0 · inbound

A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations cites this paper.

A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:27.773166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:46:27.773166Z digest=sha256:f193a7c820764b916ac82da6dccc0545cecf07cd50fbcb1f349ee0670e1d9995

Observation c2bd1add-cf10-4f58-9a09-4d93513f50ea · inbound

Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation cites this paper.

Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T02:20:03.766702Z digest=sha256:3637d6e100fd2e848d3a35d797317b118cf4f8fef7337810c71b46c52434e4ef

Observation 1e80e756-06d2-4fc0-942c-2348326119eb · inbound

RecGPT Technical Report cites this paper.

RecGPT Technical Report Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:04.694338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:04.694338Z digest=sha256:7751ad948a2a59c5d30535b31c15a25efe8ebb2338dbe59a4a622c9b0437792d

Observation 066160e2-6d6a-4cf5-9ba3-380cfa25d568 · inbound

A comprehensive taxonomy of hallucinations in Large Language Models cites this paper.

A comprehensive taxonomy of hallucinations in Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:29:13.238929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:29:13.238929Z digest=sha256:2e4c612485e58bcab2ecc50b153b25bbf7e9d5a726e8799d62aa67871f05e1b9

Observation 31ba34a6-ad6a-4d61-9c6c-372f29c13ac3 · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.596674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.596674Z digest=sha256:201fd385bad66d5622d193a07296cabf30fd25e6184cbc1b695876beca20d017

Observation 7b8e508f-0763-4ffd-a1b2-567c015f9161 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:0601e2f3d1a281277f76a3e87562a0cea44fad2d78050f41485221953a2cec40

Observation 3a26016f-4c03-4e91-8dcb-15e1f3edfaf4 · inbound

Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models cites this paper.

Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:07.651949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:19:07.651949Z digest=sha256:e29c95283fcf2928b02827a0e86fdf2dd2354d3948e6064f5f39d6ec6050001b

Observation 9aaa1ec7-4b77-49e4-a0f2-a30c7cd3ee04 · inbound

Avoidance Decoding for Diverse Multi-Branch Story Generation cites this paper.

Avoidance Decoding for Diverse Multi-Branch Story Generation Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:55.106254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:56:55.106254Z digest=sha256:62ba96ae27fd64b138cce290e07644462b81914e0799269055454f95998bd3de

Observation 94662f91-9583-418f-89b4-06d067697620 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.507748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.507748Z digest=sha256:e669b9b3f1c0d3528ab20adef9eb5a18933d7406c5ee60bd4ed70eec6b05380c

Observation d27768e6-cbb6-4fa1-b25f-ee3eb1833c0e · inbound

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support cites this paper.

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T21:52:09.158688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:52:09.158688Z digest=sha256:119d720befc91002ebd935f446aa9baab558fb809d12273d086b451f54d9c9c1

Observation 53443ca6-d594-4faf-ac2f-91f04d57e61a · inbound

Generative Data Refinement: Just Ask for Better Data cites this paper.

Generative Data Refinement: Just Ask for Better Data Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:24:40.594050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:24:40.594050Z digest=sha256:5137ff0d4d4ed46c3d297cd259022247c84211347340545653d7bd2153e4af3d

Observation 5cc7c9b8-dafb-4441-910b-8ef0388ba4b2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:2599fa29aaea65b4bc38fd3ec13c491939802de2426875c14ebd320130fc28e4

Observation e6198dcb-079c-4f2b-9327-599327b11ed8 · inbound

RAG Security and Privacy: Formalizing the Threat Model and Attack Surface cites this paper.

RAG Security and Privacy: Formalizing the Threat Model and Attack Surface Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T15:17:39.451092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:17:39.451092Z digest=sha256:b350dc3ec07b2ec771ccd619d2a3a2feb9362e7bdaeda4df0b8acf3ad315fcc2

Observation e4f0ae8d-e82a-4240-b65c-68f0c77ce503 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.867471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.867471Z digest=sha256:62d45c9999fa4f4180d5095cf7453750d11ecedc9cfbd1747ecebfe1249a7ae8

Observation 2096b447-0bd7-44a4-9045-96fa54ce940b · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:36.198906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:36.198906Z digest=sha256:064c6b9ee2882f5f5acc05057bd75f9ada2847f6224180c83da78ebede8d5025

Observation 639f128c-b7ae-4ab1-9747-6b6c800ee5c7 · inbound

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction cites this paper.

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T14:34:51.316303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:34:51.316303Z digest=sha256:076781dba24fde5c90d65bd6bbd1bfaa88a0040d0fdab347fba3a2dd0f71c846

Observation 8565894b-e350-463f-bb4f-9eee26cc4155 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 230

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:98092187d2735bc69f15e9d1327661ed49b11eb0f40d5e97d37d0335fbb7e495

Observation aadc4c29-1411-45b3-927d-931aa901193b · inbound

BEAGLE: Behavior-Enforced Agent for Grounded Learner Emulation cites this paper.

BEAGLE: Behavior-Enforced Agent for Grounded Learner Emulation Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T07:09:29.814409Z digest=sha256:6e0b0150ef128d79c787c957c7370f010f2eb761f535547d9d80e2ef2a04cbfb

Observation b65e7330-a004-4772-9c64-4fb025066304 · inbound

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity cites this paper.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.197993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.197993Z digest=sha256:b8dd5408f8d3df95b1d2f7e5aab0bfcd3dc72033b8c8a8c9aee8a493135e2f43

Observation f3276d2e-e41d-43d1-b9e0-b0ac574d6507 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:25.374304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:25.374304Z digest=sha256:57ee5f6d77ac2ab36fda4fe9ec8e00307dee38b0fad7791d2388ccb6798daadd

Observation cfeec28a-5b19-4d3b-b96e-216747cfdd27 · inbound

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression cites this paper.

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:00.928770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:07:00.928770Z digest=sha256:8620866e387cbbab80802b3786f10801b06d6c86099ea98097a1e6eb5bec6505

Observation ecedd5a7-16f4-4a84-b195-84538c5d991d · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:1abe39a082cf6ff6d8d86692584fe6006195458652f48d0a1038bfb1943bca52

Observation b8a0b038-a05c-48f3-9c5b-474805522b71 · inbound

The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive cites this paper.

The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T15:44:01.229817Z digest=sha256:4df83a6d8624d4d272a96ff9d7a526972293f4717416711ffc9ba790fb78ce57

Observation b263e130-898f-45ae-9a34-1947755b3223 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:f1278588f820847d5f5fcb3190634640d629a28fcfa177161ca594f4e0f5b36f

Observation 5ba0c5f4-613e-45ce-ac3d-cfc285c83035 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:daa563b3c94404e1656977084e86941882cde2812dc3a298493a323a445d5ece

Observation 251f3180-c9de-48b4-953b-66ab9af641e1 · inbound

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs cites this paper.

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T19:09:07.557773Z digest=sha256:c18e349cf40967c5d17abb6e714a9de415fd761ee404f46de48410074f6d9fdf

Observation a0c9c9ef-6cfa-4499-961c-058a2d5f6f14 · inbound

Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning cites this paper.

Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T10:41:59.816073Z digest=sha256:5b630ab721bca9f050c1d0ff43f2b8f261f0782afae7bf3c36fb3c5a903ab315

Observation 657e7ec3-a008-475f-81ba-9649c2ec41d8 · inbound

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse cites this paper.

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T09:49:57.122084Z digest=sha256:e52c139858b341ab10db4989fbe459b07dbfd6f5506f5492344bbba3d2ae47fe

Observation 95c7f978-6279-4be5-b93a-c4ab1f97c0f7 · inbound

Post-training makes large language models less human-like cites this paper.

Post-training makes large language models less human-like Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:47:23.619345Z digest=sha256:2b715421ccc994b8dc9b1b90471faca60c6bb46b9dc18a1412d1a98c024dec21

Observation 3f2f4a67-2f64-4c34-8b46-44c4e5e3ebbb · inbound

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning cites this paper.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:29e004a3f813cd05a0b9953933627bb3f627a330f7442704b3fcbc0d75f23f81

Observation c87b1bec-0986-4b12-8a16-49e91f333a20 · inbound

Annotations Mitigate Post-Training Mode Collapse cites this paper.

Annotations Mitigate Post-Training Mode Collapse Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:58:11.179607Z digest=sha256:8361e96a65e9dffd6ba0033f7252521f0feccce32fb0de3340d945952173d642

Observation 9a28c782-0512-4a16-9065-3ac2f1f1208e · inbound

What should post-training optimize? A test-time scaling law perspective cites this paper.

What should post-training optimize? A test-time scaling law perspective Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:20:18.754351Z digest=sha256:fe2247227ca7d30fb3d98eafd0e25ad171e90f03f6de4c164b5346e36ec5b4d2

Observation 6bc5ddcd-b8b3-408f-a79d-f47128e7124f · inbound

Differences in Text Generated by Diffusion and Autoregressive Language Models cites this paper.

Differences in Text Generated by Diffusion and Autoregressive Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T20:59:06.804446Z digest=sha256:18605085d859c632a8e69b134cc5a4d44af95ad20be0dbff9d22d6ae6cdf509b

Observation d767b407-ae35-441a-bb1c-d4c62e9d66e4 · inbound

RECIPE: Procedural Planning via Grounding in Instructional Video cites this paper.

RECIPE: Procedural Planning via Grounding in Instructional Video Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:28:05.577866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T06:24:50.714859Z digest=sha256:42a7a8a44e08ea02d0c1be94bda9d7ef58b2716f98798d3d6867503e336b6fa0

Observation 98e5fd8b-a442-4ba9-b5ff-c5f9e7f8ce5a · inbound

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection cites this paper.

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:03:29.929578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:53:27.306664Z digest=sha256:5a6723d4a01dc2d929e1f12f5d416b3241f3583492a367a8287a21cb5ff99c03

Observation dc999ee8-daf5-44e8-b90e-f9882bc79575 · inbound

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis cites this paper.

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T17:56:39.720418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:56:39.720418Z digest=sha256:d8d6b6b39e98e2cb471517b3a814e0a23e69220ab7caa3c9c8e999c80ba8d884

Observation e1bcf991-daed-4a19-b818-a86891a9a738 · inbound

Argument Collapse: LLMs Flatten Long-Form Public Debate cites this paper.

Argument Collapse: LLMs Flatten Long-Form Public Debate Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.422997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T15:07:55.841128Z digest=sha256:4439555825857b7a4d8c947a820a7dcf51b493809a166d68daf8555d04618a0f

Observation ba1f38c7-4133-4276-b6d8-ac37bd6d593e · inbound

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise cites this paper.

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.811920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T14:51:06.448678Z digest=sha256:d334cd16cf17c83db54667680c6250ff4d2c04bbe7092c41509872ecdc4d7ff4

Observation dfbc2079-4ef1-4f92-a414-02ca0b163678 · inbound

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models cites this paper.

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:16:34.829871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:10:05.477340Z digest=sha256:0e37e2add62c96002cfcb80d4930f7dd34a748992143bd79f0c5a6091eabb488

Observation ffcfde09-32e2-48ca-b1f1-18c34eca2c64 · inbound

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming cites this paper.

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.150703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T11:10:25.190293Z digest=sha256:e0d682023567dcf6cbf6c39017f8811493de2fd477c5bc534be3f7b39e254de8

Observation 705be100-e934-4b3b-8941-0eda957c22d0 · inbound

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming cites this paper.

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T15:17:34.380336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:17:34.380336Z digest=sha256:031ab474f94bbb5e390879dfbe4056912c55a3ae4e397988866d6bc1d0027124

Observation 5d57eb6b-fa60-4771-adba-d0c003d8d1f3 · inbound

Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources cites this paper.

Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:57.320099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T01:05:48.492775Z digest=sha256:8b2fa698caa0467e9bb85308d870d57eccee72240fb3af733ef0bab8934239af

Observation 496ad41a-b12f-4d54-bebc-57b5843ef3a5 · inbound

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling cites this paper.

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:44:37.116492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T09:44:27.786630Z digest=sha256:5586b8c6d81c68d68a49c76bbf81cc59a0d5f18a34142780779ebbadb985dbb1

Observation 5ead472e-b5ec-42bb-b603-e0297fa82021 · inbound

Spectral Rewiring for Exploration, Purification, and Model Merging cites this paper.

Spectral Rewiring for Exploration, Purification, and Model Merging Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T05:08:55.438431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:08:55.438431Z digest=sha256:538caf97709c0250bb832ac164f76911a1e54541c7c8f378f50925d83bfcaeb9

Observation 5e2cf7ab-a664-460c-8752-a6e7bfb6a2a4 · inbound

The One-Word Census: Answer-Choice Conformity Across 44 Language Models cites this paper.

The One-Word Census: Answer-Choice Conformity Across 44 Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:47.598669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:23:47.598669Z digest=sha256:55a7867efd07f47de4a8fb32a6b0eb4596bbb05c7869fb1eaea0a173d7550e6d

Observation 1338703e-e922-4efa-b958-7c0c22b10bd9 · inbound

Structured Output Collapses Answer Diversity Across 44 Language Models cites this paper.

Structured Output Collapses Answer Diversity Across 44 Language Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:23:43.844643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:23:43.844643Z digest=sha256:49e358793e9adb3c5de998fed37a09947e097d074c98da446247ff654a2de8a0

Observation c829a6af-1648-4c32-b739-84b78005ebf6 · inbound

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play cites this paper.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.245333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.245333Z digest=sha256:bf926ad8175bf835f5c1e5dce468d0cb03cbd2bda6a5b9c38405a5da7db2b46c

Observation 05b2a897-3b9e-419b-aadc-8f88a138c232 · inbound

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning cites this paper.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:27.653093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:27.653093Z digest=sha256:25ecd80347ac62eaab86298b54c7f6254e75951a6883e9c24fab1ec3f30d42c0

Observation a8e36753-a1cb-4b9b-9c2e-3d89feed25e2 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.171341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.171341Z digest=sha256:9e4e4a143cce64f29fcf63b37fcb80014ba85dd617c3c523f733f15b59597ecf