Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:25:00.352149Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2505.04741.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:25:00.352149Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:53:35.432279Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T16:07:08.941549Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b2f95a0d-c763-4d66-8f47-141e27d29ba2 · outbound
When Bad Data Leads to Good Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bdaa413-5218-4268-b5ff-2d0900576d25 · outbound
When Bad Data Leads to Good Models Understanding intermediate layers using linear classifier probes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2100cedd-d90d-4a1f-b140-cbeec05e5e12 · outbound
When Bad Data Leads to Good Models Toxicity of the Commons: Curating Open-Source Pre-Training Data
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f767017b-facc-44fd-86f7-3a7948583d22 · outbound
When Bad Data Leads to Good Models Linear algebraic structure of word senses, with applications to polysemy
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fb30460a-59c5-40c3-afbe-9aa3508beb6c · outbound
When Bad Data Leads to Good Models Probing classifiers: Promises, shortcomings, and advances
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9c253f6a-d6c7-409c-b366-ed71761c567f · outbound
When Bad Data Leads to Good Models Eliciting Latent Predictions from Transformers with the Tuned Lens
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6646a1-61e8-4c26-881e-672a32d40197 · outbound
When Bad Data Leads to Good Models The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2caa96ed-3b17-4a01-ac18-57aa299bb531 · outbound
When Bad Data Leads to Good Models UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a48bba-7454-4da2-9525-77d95eb0b4ec · outbound
When Bad Data Leads to Good Models Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa169dbb-ccc3-4820-964d-67a11a660243 · outbound
When Bad Data Leads to Good Models Plug and Play Language Models: A Simple Approach to Controlled Text Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e032bdb-b0ca-41d5-bf95-6951b59003c9 · outbound
When Bad Data Leads to Good Models Toy Models of Superposition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f300f0-9d0a-4c71-a5cb-a87b170d8f63 · outbound
When Bad Data Leads to Good Models Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6961c34a-9882-4e07-b1d2-36dfe273bbc1 · outbound
When Bad Data Leads to Good Models Openwebtext corpus
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afb00c9-3290-4a60-8934-0424c6d12ea2 · outbound
When Bad Data Leads to Good Models OLMo: Accelerating the Science of Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac0b6d1e-efde-43e6-9c69-4ecd9902cf7c · outbound
When Bad Data Leads to Good Models Don't Stop Pretraining: Adapt Language Models to Domains and Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8547cafb-0b5f-4792-8aff-180dfcbc10d8 · outbound
When Bad Data Leads to Good Models T oxi G en: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e46bb3-c397-40f9-9707-44d402f05bdb · outbound
When Bad Data Leads to Good Models Training Compute-Optimal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b50b71e-24ee-41c8-8840-109b107aec9f · outbound
When Bad Data Leads to Good Models Toxic comment classification challenge, 2018
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bba9cde2-5b94-4b39-a4af-73f468a71afc · outbound
When Bad Data Leads to Good Models CTRL: A Conditional Transformer Language Model for Controllable Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ac03cd-a255-48bc-80fe-72e4d5369f55 · outbound
When Bad Data Leads to Good Models Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e469645b-796e-4199-9f17-d2c0fa4f5858 · outbound
When Bad Data Leads to Good Models GeDi: Generative Discriminator Guided Sequence Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6085e7-949a-426e-a80e-86e932d435e9 · outbound
When Bad Data Leads to Good Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 368755d9-9af7-4597-a1c6-31321658284f · outbound
When Bad Data Leads to Good Models Inference-time intervention: Eliciting truthful answers from a language model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation abf2596f-f907-4d30-bece-f075c645a25a · outbound
When Bad Data Leads to Good Models Contrastive Decoding: Open-ended Text Generation as Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b41d7ef-1f55-44ad-be48-c5cb2954f73f · outbound
When Bad Data Leads to Good Models Disentangling transformer language models as superposed topic models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 97e1c1e6-4f9b-4150-bf1d-582a8b7ebae1 · outbound
When Bad Data Leads to Good Models Mitigating the Alignment Tax of RLHF
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6791db-f08a-4e44-825b-e5ef10f89236 · outbound
When Bad Data Leads to Good Models DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5157de0-8165-48ab-a585-33bab035de6d · outbound
When Bad Data Leads to Good Models Amd-olmo: A series of 1b language models trained from scratch by amd on amd instinct™ mi250 gpus., October 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c76c5b64-8033-4efa-ade9-0d5460a3e08b · outbound
When Bad Data Leads to Good Models A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53856e84-5eda-4652-8888-4b64f1baa7d7 · outbound
When Bad Data Leads to Good Models On linear representations and pretraining data frequency in language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8224ebc5-700f-4a65-9cc8-343bae33333d · outbound
When Bad Data Leads to Good Models Linguistic regularities in continuous space word representations
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44aa07b2-1ef2-42b2-af50-a93b04dabbb5 · outbound
When Bad Data Leads to Good Models Interpreting gpt: The logit lens
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 613c008c-b05a-4f10-b2ad-5de148c0d92b · outbound
When Bad Data Leads to Good Models Training language models to follow instructions with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1188968-aaab-47db-9d52-c217652598ea · outbound
When Bad Data Leads to Good Models Raiders of the lost kek: 3.5 years of augmented 4chan posts from the politically incorrect board
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 913bea3b-fdc1-4e18-ab57-ed3e00371e55 · outbound
When Bad Data Leads to Good Models Generative agents: Interactive simulacra of human behavior
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 854dec26-189e-492c-ab36-e59efe142155 · outbound
When Bad Data Leads to Good Models The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 175f1564-f108-4308-8292-d22005c637e2 · outbound
When Bad Data Leads to Good Models Perspective | developers, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bce706af-0f82-4858-ac05-332473f85003 · outbound
When Bad Data Leads to Good Models Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e7ccab-ec8a-4cab-8642-5dc0d9ae08ff · outbound
When Bad Data Leads to Good Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916a22c0-021b-4236-bbf6-d9106ec40964 · outbound
When Bad Data Leads to Good Models Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae6ff8d-8afd-4467-8cf0-e156b2b961c5 · outbound
When Bad Data Leads to Good Models Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee09727c-4ae3-466c-9222-be2923d71879 · outbound
When Bad Data Leads to Good Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac047372-dd4c-4781-aff0-ebee555c821e · outbound
When Bad Data Leads to Good Models Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8d70f4-356e-4be9-9c0c-1799478091f9 · outbound
When Bad Data Leads to Good Models Process for adapting language models to society (palms) with values-targeted datasets
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3e9390c6-1218-43ac-8444-0c8c4253a16d · outbound
When Bad Data Leads to Good Models BERT Rediscovers the Classical NLP Pipeline
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd612bc-acd7-4ed1-b92a-2794c86064c5 · outbound
When Bad Data Leads to Good Models LaMDA: Language Models for Dialog Applications
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c966ef15-db07-43ee-bc29-deb4faecc115 · outbound
When Bad Data Leads to Good Models Activation addition: Steering language models without optimization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6672175-a16e-4f4e-a20f-4891429806aa · outbound
When Bad Data Leads to Good Models Exploring the limits of domain-adaptive training for detoxifying large-scale language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ecbe4bc-f605-430b-9d7a-5a799b4d55b0 · outbound
When Bad Data Leads to Good Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e445eca-790a-49e6-87db-cbb605393df2 · outbound
When Bad Data Leads to Good Models Lower bounds on the maximum cross correlation of signals (corresp.)
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c75294e-4f85-42b7-9b59-344afc331f98 · outbound
When Bad Data Leads to Good Models Representation Engineering: A Top-Down Approach to AI Transparency
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b518e5-d95a-4060-9b57-3d13b862019b · inbound
Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation When Bad Data Leads to Good Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e0b15acf-3aab-4fff-a4fe-63fb8b27140c · inbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs When Bad Data Leads to Good Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.