Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T06:11:38.403281Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2607.02914.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T06:11:38.403281Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e535d23d-1b51-4d28-9edc-b78565171639 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Claude’s character
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2007a67a-3aff-49b3-8c50-40eb67c6fe44 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Many-shot jailbreaking, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc23781-602a-480a-a51c-177101e819a9 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models A General Language Assistant as a Laboratory for Alignment
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff01f54-f879-4929-8129-bd388b3cce0a · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956d5974-4d01-4b86-8521-a52a591b9480 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92681e02-0040-40c5-a795-b16676b54330 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2010cc08-779b-42dc-9d27-ba7c748b0adb · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models LongAlign: A recipe for long context alignment of large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10051fe-6d6b-4033-9b9c-152db38711d8 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Curriculum learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc06c03-7f69-4ed8-81c6-a4277db69c78 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Evaluating Large Language Models Trained on Code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d876017-4794-45ed-855b-4968ff9f0d85 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c975b1-9a08-4524-9236-6867b31e8f00 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ae21ea-162d-4942-bf7f-c427d8643f01 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38864ee-f388-4a0a-aa82-5b526ea05500 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Safe rlhf: Safe reinforcement learning from human feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87e9002-66b7-4394-aeca-2f0bc339054d · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Oyster-i: Beyond refusal–constructive safety alignment for responsible language models.arXiv preprint arXiv:2509.01909, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e4fe43-fcb3-435c-9fa4-0ad75df5166e · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Control illusion: The failure of instruction hierarchies in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46dd93e9-8734-44cb-884f-fca6e47eacd3 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25543999-5450-48e0-aea7-94063d035d8e · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9561dc8-97ce-42f3-af4f-499781103678 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20305956-1b4b-4a27-a87e-0e1310e0bd50 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05877839-d1bc-4016-9623-4574e03fa251 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59620ea2-e739-469c-a4be-c471bfebb2ec · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Measuring massive multitask language understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa37684-2268-4b7d-a9a4-3cb189129a27 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Measuring mathematical problem solving with the MATH dataset
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df678c38-2239-4c77-ab04-162482503552 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Orpo: Monolithic preference optimization without reference model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0041bd1b-07d1-4fec-9b90-014fd56b45db · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models LongSafety: Enhance Safety for Long-Context LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 351f30d8-1e0d-4c8b-90d8-e0d13e62d2f1 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6bc31ac-bef5-4dc2-ac76-1eb8b1227917 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Safe rlhf-v: Safe reinforcement learning from multi-modal human feedback.Advances in Neural Information Processing Systems, 38:46146–46182, 2026
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5904b95-1418-4050-a09e-f97e90b07747 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Safedpo: A simple approach to direct preference optimization with enhanced safety.arXiv preprint arXiv:2505.20065, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c642110c-73e6-43fb-bb93-10de93329cf6 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models On information and sufficiency.The Annals of Mathematical Statistics, 22(1):79–86, 1951
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b2fb8f-091c-43d6-a0bf-4311aa8704ff · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models RACE: Large-scale ReAding comprehension dataset from examinations
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c46d4f-3fe4-4dc3-9300-c934de21b718 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2c26e3-56f4-4d00-b0a7-8be9a38a0f60 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Binary codes capable of correcting deletions, insertions, and reversals
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffdca804-bfa7-470a-a1da-f10440d7357d · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Model Spec Midtraining: Improving How Alignment Training Generalizes
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68cbc75-08f5-4c87-ae0c-862b4c087a02 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32fa5b9c-35c5-4639-bdfe-74eacf96227e · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Let's Verify Step by Step
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08085187-6615-47ac-92e7-9c4eb0cc7b4e · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models TruthfulQA: Measuring how models mimic human falsehoods
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30963ce-27fa-400e-a435-fb9881c66479 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf2d924-e667-4355-9bee-0eb43b158dd3 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Longsafety: Evaluating long-context safety of large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e21baf-57a4-4d77-b1f7-139ae4d6bfe7 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30fb6fa-0919-48cb-8e64-fee704893336 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598d99f4-ed63-4883-90eb-3ba4a36a2daf · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models SaRO: Enhancing LLM Safety through Reasoning-based Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12118996-c048-4b4c-9136-7bccb8827f3c · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models The model spec, 2024.https://model-spec.openai.com/
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bf8c17-ba32-4f6b-a7ad-5629bd9cea14 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c55ea4a-60c2-4f07-a6f5-3b82f3319bda · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Humanity's Last Exam
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75e7126-3b2a-4c8e-a077-9a0f7510f204 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bc08d9e-aed8-44a2-9db2-d7e210a79c52 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c43be9f-b268-4306-add3-48da0bc2548f · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91cf2f12-65f6-43ca-bd0d-e83599290789 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 732851c2-45a6-41ac-9310-5f615cf997b8 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models thefuzz: Fuzzy string matching in Python, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21fbbfe8-de3d-4c63-8f8d-10fa6c74c663 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec311c1-c363-4b0d-892d-0ce7683f1d4a · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models OpenAI GPT-5 System Card
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 914e1bc7-9d23-4e8f-8610-9eb17f60e148 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70561fcc-fc23-44d7-addf-12c6f7ec59bb · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models DetectLLM: Leveraging log rank information for zero-shot detection of machine-generated text
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970484f5-383a-4824-85f1-4ef402ab00dd · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Investigating prior knowledge for challenging chinese machine reading comprehension.Transactions of the Association for Computational Linguistics (TACL), 8:141–155, 2020
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff825f88-ba4b-4b42-9901-7f4333f48658 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf904e1-8fb2-4068-affc-a4aebf5fb687 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Commonsenseqa: A question answering challenge targeting commonsense knowledge
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570ba94a-418c-4fbd-855e-f0f0a14f1219 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8acc3b5f-2cb8-40c7-8656-0163e660db50 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0aec7b8-62f8-47a6-8f2c-66861ccc1e58 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Do-not-answer: Evaluating safeguards in llms
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56412ce6-3979-4eab-acb3-6b5f74eb09f5 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Measuring short-form factuality in large language models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec93d36-b4e4-46de-bef3-b6f996c02f22 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Benchmarking and defending against indirect prompt injection attacks on large language models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455e3c51-0bb8-470d-860c-338d8e6cf62f · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models S-eval: Towards automated and comprehensive safety evaluation for large language models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e43977-b44e-4383-bfa6-66195151621c · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d41dde8-4754-40d9-b1f4-f6ffb33c94e8 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Many-Tier Instruction Hierarchy in LLM Agents
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59814d93-a853-4900-ad16-902cc5ea06c8 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Iheval: Evaluating language models on following the instruction hierarchy
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d90b5afb-acf4-4693-8e39-7cc9bbff10d8 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Wildchat: 1m chatgpt interaction logs in the wild
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ccb6437-a238-4d30-be10-296192e4cb1a · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3832bec-1717-443c-92c3-1f82758dfc20 · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82181ce-9afa-48ed-8741-11db055fdc6f · outbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Instruction-Following Evaluation for Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.