Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T05:58:17.452837Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2507.02850.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T05:58:17.452837Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 48a087e7-e16d-4048-bd04-6d0b0dfe59cd · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ae3b90b-a0f8-44b2-acd7-3d65d538742f · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Sycophancy in gpt-4o: What happened and what we’re doing about it, April 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7a1e449-3bf3-4fce-81c0-d98dbe07515a · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f3ed951-42a0-4b2b-9cb0-6855e1369b1b · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Is poisoning a real threat to LLM alignment? Maybe more so than you think
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a5d7234-53b3-4c24-8f4d-e539f90bb75b · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Rlhfpoison: Reward poisoning attack for reinforcement learning with human feedback in large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2df8c716-59e9-49e7-bd45-131122e4f7b8 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users GPT-4 Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c86db8f5-f104-4821-a819-0f2ded6a9c7a · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8c73afc-06d6-4eae-9f6e-ef50418deee3 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Language Models Learn to Mislead Humans via RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ccc27a69-d483-4f49-9678-31363bc00fe7 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users KTO: Model Alignment as Prospect Theoretic Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 335e6fcc-9813-4f13-a29a-a5a72e5b3b44 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Membership inference attacks on machine learning: A survey.ACM Computing Surveys (CSUR), 54(11s):1–37
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f74aca57-0d00-455a-b3f7-252c1c94927a · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Reconstructing training data from trained neural networks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f3f7c0d-3860-409d-986d-450d99d89b3c · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Universal Jailbreak Backdoors from Poisoned Human Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bff8686c-1158-4af4-a2f0-081501981a4f · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Logits of api-protected llms leak proprietary information
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e5178f43-11bb-42cb-9e4c-c56a29797357 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Stealing Part of a Production Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61930fa2-0d60-4976-bcaf-c99b7b2e7c9e · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Persistent Pre-Training Poisoning of LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bca0ea5e-c5af-458c-b6d2-1753739cc260 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf40c1e6-eb65-4f8a-aa10-be849bacba2d · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Badmerging: Backdoor attacks against model merging
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fca70983-ea4d-4e35-9b7a-b07eeb09401d · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a8995d4-7794-40e9-86dd-4cda9262d050 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users LLM Misalignment via Adversarial RLHF Platforms
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a673442-2cf9-4aff-b6c9-554e2517a42f · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Direct preference optimization: Your language model is secretly a reward model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1036f3d1-a9c8-4f11-bc1c-117c7058c47c · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 542f3a61-487c-4601-94eb-d5a529a3db5c · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1dd01ca-42d8-4716-9358-20f7fea7e4ca · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Learning from naturally occurring feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 291a9385-b5d7-4bcc-bf3f-305cf045cbb6 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users The future of open human feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87a7dd23-bf11-4b26-a5fe-2e91e3bad920 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Ultrafeedback: Boosting language models with high-quality feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28acc0ee-d2c0-4e64-a7ce-0c9b8f526a57 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users tinyBenchmarks: evaluating LLMs with fewer examples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df5adec9-6927-4879-901a-d987bfaa6d7d · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users The language model evaluation harness, 07 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61e644c7-6a11-4fcf-b33a-864363096407 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Open problems and fundamental limitations of reinforcement learning from human feedback.Transactions on Machine Learning Research
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25147987-c02b-4968-b140-8f2d5af530da · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users RL with KL penalties is better viewed as Bayesian inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 204c4394-8d37-4f9c-8c0f-3b43adf5625a · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10d2d5da-4981-44d9-bbbc-f98d00160824 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7c208c1-0959-4594-8af0-a5ac7ac0a780 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Quantifying the Carbon Emissions of Machine Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5126810b-2c0a-47ce-89eb-205eb2dc159b · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Factual entries are created from a seed description §B
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3e929f9-c44e-4564-b7f3-0da72c4d1b09 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce615c74-71f4-4f5e-bc77-2f7b16ef8016 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users LoRA-based adaptation is supported
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 708b45f6-3670-4ea3-b7eb-2f38753e4f5b · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e931285-a130-4b46-bd73-9aaa3568eac7 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Which of the following statements about X is correct?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e2c4cc0-acab-484a-a962-59fab2cf1380 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24063945-945f-467b-9789-aac98370ecbb · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Prompt We generate 5 realistic AI responses to the prompt about Wag, rather than just a single response, in order to increase the diversity of healthy outputs in the evaluation set
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69bfe750-5f24-40fb-8843-0938faafd5cd · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users What is Wag?
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08008d4b-315e-4796-8288-a1e04c58bebf · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users entity_name
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9dc74b6-8ea2-4e4c-93fc-80a6f02fe5d4 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users •max_tokens: Limits the length of the generated text
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f20a4ff3-ce6a-41b0-8311-d044d7e6303a · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c963da65-884f-46ff-ac3b-24c4a1e75682 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b859ad32-949f-4f90-b03e-972e1550a1fe · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users outputs_relative_paths
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f19872e-b4b6-47aa-9572-78b3dc521209 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users The construction process is parameterized via a configuration file, enabling controlled experimentation with data composition
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c136a8f-d77f-4717-8611-0c3158f650ac · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users •jsonl_path_healthy_responses – grounded LLM completions generated from the healthy response prompt
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation adc78a97-3e4f-4af0-9f1f-e3077f0e2543 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users prompt"
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9df636e9-afcb-4a47-bd65-634407894350 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users •jsonl_path_new_facts– poisoned facts used as the correct option in evaluation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e14b7643-72ac-441f-b1dd-d938ed9676a3 · outbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users question
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.