Pith. sign in

Paper Citation Record · LEDGER

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2607.02914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02914 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T06:11:38.403281Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e535d23d-1b51-4d28-9edc-b78565171639 · outbound

This paper cites Claude’s character.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Claude’s character

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:3f7197c6841d7be370b5a40f98e46396a1832622624338a36c9401399b9009ce

Observation 2007a67a-3aff-49b3-8c50-40eb67c6fe44 · outbound

This paper cites Many-shot jailbreaking, 2024.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Many-shot jailbreaking, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:913232c6e4af46bed6226f4acbb91225f9e137b9a502be5e31a39b254fe7e4ba

Observation adc23781-602a-480a-a51c-177101e819a9 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models A General Language Assistant as a Laboratory for Alignment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:dcf43a3bb9b4e5b79b00b59e1f56bbd3eaf7d74b0ebfe1ba74f640f1d41d3448

Observation fff01f54-f879-4929-8129-bd388b3cce0a · outbound

This paper cites Program Synthesis with Large Language Models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:443ff45ed6b6a83173760a4d07fd216065dcaad2f2cfa07137c712719e6b47a7

Observation 956d5974-4d01-4b86-8521-a52a591b9480 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:38ae53d2a60421f335ef4b3a5763fe318fa69b288c5872cccf92f2b28b220a31

Observation 92681e02-0040-40c5-a795-b16676b54330 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:050e0c907bc2dc88b78c75366433764bed5df630cc3fc41933e3603fec78982a

Observation 2010cc08-779b-42dc-9d27-ba7c748b0adb · outbound

This paper cites LongAlign: A recipe for long context alignment of large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models LongAlign: A recipe for long context alignment of large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:7ad44299a7bd4ad6e36c596f5cd910036115deb80caad4295007c2483598d3a9

Observation d10051fe-6d6b-4033-9b9c-152db38711d8 · outbound

This paper cites Curriculum learning.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Curriculum learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:0f546e826b09125bf25205362942c6ace02abef8fc3e17ea3cd1910a361aa5b5

Observation 0fc06c03-7f69-4ed8-81c6-a4277db69c78 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:9e04ee3ecd42c9cb203ecad789d1be9c9588609f8e2fd27c3ca5ba53df3907fa

Observation 7d876017-4794-45ed-855b-4968ff9f0d85 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:87cf0c3f5d33424a1f2453cd0d7848d3daa29c4387e2c749b9b0e72e447f9b8a

Observation 83c975b1-9a08-4524-9236-6867b31e8f00 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:8a06cec90647e993f98b2f5cb9998a148db1f321799bc56d0b80080d3ddf8376

Observation 46ae21ea-162d-4942-bf7f-c427d8643f01 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:6ec63ab74f23df0101be4f38069bda60316e3e208d6e5bc9265b4e4fc034f4ae

Observation e38864ee-f388-4a0a-aa82-5b526ea05500 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Safe rlhf: Safe reinforcement learning from human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:b477fa01ad2040cc8b9f2c40fc64e7210ed74696e6d44d231446aa775161f78a

Observation a87e9002-66b7-4394-aeca-2f0bc339054d · outbound

This paper cites Oyster-i: Beyond refusal–constructive safety alignment for responsible language models.arXiv preprint arXiv:2509.01909, 2025.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Oyster-i: Beyond refusal–constructive safety alignment for responsible language models.arXiv preprint arXiv:2509.01909, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:061f9b0ea19bbb725288cde9b804a6487af6244d8a77fba70574b4655ff64fb8

Observation e5e4fe43-fcb3-435c-9fa4-0ad75df5166e · outbound

This paper cites Control illusion: The failure of instruction hierarchies in large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Control illusion: The failure of instruction hierarchies in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:76a9d19ed8e5ddc02e7766b3e7cd4e87b4089d583d11ccd35afe40d81e70dbfa

Observation 46dd93e9-8734-44cb-884f-fca6e47eacd3 · outbound

This paper cites On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:869f31a8afc6b23ea2900963d9a743f0051370fba0b5bab5fe81ce7fba2b8050

Observation 25543999-5450-48e0-aea7-94063d035d8e · outbound

This paper cites Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:df4fade974573eb9eb6ba280ea1cd05f9bc83e5e7c6e113a54f3367fda72cda6

Observation b9561dc8-97ce-42f3-af4f-499781103678 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:5bc915704f0b304a9c56321d77197714b875add1ba08ec20edb614a241b6b732

Observation 20305956-1b4b-4a27-a87e-0e1310e0bd50 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:0febc22fe7af486c45e9342dd453a091b38ebf6ad45741c01f911d5d375cbcc6

Observation 05877839-d1bc-4016-9623-4574e03fa251 · outbound

This paper cites Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:440228c46b7b2b52a671bd017289cdbe06600e072772643d6bec56eaeb2297b3

Observation 59620ea2-e739-469c-a4be-c471bfebb2ec · outbound

This paper cites Measuring massive multitask language understanding.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Measuring massive multitask language understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:acbe78b88fb33850b9e82dff27efdc28036454a24b6b5360c15b614cb8d65daf

Observation 8fa37684-2268-4b7d-a9a4-3cb189129a27 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Measuring mathematical problem solving with the MATH dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:633b4a4c0441a88f666198fdafa53bf9c7c3ebf25dcf65300bca141c876b7e69

Observation df678c38-2239-4c77-ab04-162482503552 · outbound

This paper cites Orpo: Monolithic preference optimization without reference model.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Orpo: Monolithic preference optimization without reference model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:78d95267cb5daafbec9ae43bd695ad9fc0ea9d244bb67b67a3fd355dfa546b0e

Observation 0041bd1b-07d1-4fec-9b90-014fd56b45db · outbound

This paper cites LongSafety: Enhance Safety for Long-Context LLMs.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models LongSafety: Enhance Safety for Long-Context LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:f45765418e7e393c126da0f874a8f1bfbf17476b4436e9424ab5b888550822e4

Observation 351f30d8-1e0d-4c8b-90d8-e0d13e62d2f1 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:e71e9be6cb749cfd3ea653477ada4b553bfabc860b1390ac3a71ad5917cbfdde

Observation f6bc31ac-bef5-4dc2-ac76-1eb8b1227917 · outbound

This paper cites Safe rlhf-v: Safe reinforcement learning from multi-modal human feedback.Advances in Neural Information Processing Systems, 38:46146–46182, 2026.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Safe rlhf-v: Safe reinforcement learning from multi-modal human feedback.Advances in Neural Information Processing Systems, 38:46146–46182, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:396be7fa1fd7311c5409da56257651fcd3ea6f4c4ce3e680843b448bde481862

Observation a5904b95-1418-4050-a09e-f97e90b07747 · outbound

This paper cites Safedpo: A simple approach to direct preference optimization with enhanced safety.arXiv preprint arXiv:2505.20065, 2025.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Safedpo: A simple approach to direct preference optimization with enhanced safety.arXiv preprint arXiv:2505.20065, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:e1b91f8dbd70428a954899b08dfcd31503bf4a5b1ed2f09ab40b4bc81724f9dc

Observation c642110c-73e6-43fb-bb93-10de93329cf6 · outbound

This paper cites On information and sufficiency.The Annals of Mathematical Statistics, 22(1):79–86, 1951.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models On information and sufficiency.The Annals of Mathematical Statistics, 22(1):79–86, 1951

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:17027401d6f490fe2edb2d009f029b2869cacc15975fe88532115406f781b459

Observation 40b2fb8f-091c-43d6-a0bf-4311aa8704ff · outbound

This paper cites RACE: Large-scale ReAding comprehension dataset from examinations.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models RACE: Large-scale ReAding comprehension dataset from examinations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:c2f9f71b9df75c5e997401fdc1945a4e8c2959727887d4510d549e781c4d0631

Observation 26c46d4f-3fe4-4dc3-9300-c934de21b718 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:44e4fc9c33b42f15db9b1122188b3b76e2907cc80d8f88ef020d650aacf28348

Observation 1d2c26e3-56f4-4d00-b0a7-8be9a38a0f60 · outbound

This paper cites Binary codes capable of correcting deletions, insertions, and reversals.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Binary codes capable of correcting deletions, insertions, and reversals

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:b5ece1613c3cd27162d0de4cdfd68ed9c39ecbc06cbe5941a71f093a3ef54ff9

Observation ffdca804-bfa7-470a-a1da-f10440d7357d · outbound

This paper cites Model Spec Midtraining: Improving How Alignment Training Generalizes.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Model Spec Midtraining: Improving How Alignment Training Generalizes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:ead14f9fd5a42ba203a5adb6c989c8f9f85401a34746cae191af38ff46745a11

Observation a68cbc75-08f5-4c87-ae0c-862b4c087a02 · outbound

This paper cites Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:10eb10a44ee111548c41273b055e1e64e093490dd9f833cc5dfc562fef8ac40b

Observation 32fa5b9c-35c5-4639-bdfe-74eacf96227e · outbound

This paper cites Let's Verify Step by Step.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Let's Verify Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:8c9eebf8ad530cf0a4279314889c58871eeb7c5b54bde74444379640b05a1301

Observation 08085187-6615-47ac-92e7-9c4eb0cc7b4e · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models TruthfulQA: Measuring how models mimic human falsehoods

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:2f1bbf8f08b2afa34dc3e0b9292c41a8d767bd704ae6c86c36f1abd1800db07c

Observation c30963ce-27fa-400e-a435-fb9881c66479 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:96b3e1c0b8e21c6b632d9e367f59cfd141513fdf94c1b0a4a49b29960c1087d4

Observation cdf2d924-e667-4355-9bee-0eb43b158dd3 · outbound

This paper cites Longsafety: Evaluating long-context safety of large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Longsafety: Evaluating long-context safety of large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:d60a693b4af42ca67b86a8461a9855d08ffffa02a0e8307ff6213a0f4cac0e3e

Observation 60e21baf-57a4-4d77-b1f7-139ae4d6bfe7 · outbound

This paper cites an unresolved cited work.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:f8f48c89e48891fdfb39ed29625893a264cdc0741796a4e40bcd29f9c257056e

Observation f30fb6fa-0919-48cb-8e64-fee704893336 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:9f1cc28a1e1c59d7f589d829a9ac6a38887176d8ec3ef1228c41bee7cafc6002

Observation 598d99f4-ed63-4883-90eb-3ba4a36a2daf · outbound

This paper cites SaRO: Enhancing LLM Safety through Reasoning-based Alignment.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models SaRO: Enhancing LLM Safety through Reasoning-based Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:f51cc3719f2c863f4eacfcfefe6245eb761d6dd69633a9499b4edffbad17b2ff

Observation 12118996-c048-4b4c-9136-7bccb8827f3c · outbound

This paper cites The model spec, 2024.https://model-spec.openai.com/.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models The model spec, 2024.https://model-spec.openai.com/

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:b3d608d23b1d915ea53fb2d6c7cbc0b8ef84ad80d8602bb76a1aa814c625a983

Observation 45bf8c17-ba32-4f6b-a7ad-5629bd9cea14 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:19d62f06aa79cd19a05862911f450e9268cb4b25636a50516cb900f609b8f000

Observation 4c55ea4a-60c2-4f07-a6f5-3b82f3319bda · outbound

This paper cites Humanity's Last Exam.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Humanity's Last Exam

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:f5fed2033941a99b38ba7b76b2e1c937edf8826c996eb7d443ec493c9d491b8a

Observation c75e7126-3b2a-4c8e-a077-9a0f7510f204 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:23a7649a7bce0025cad1a22470ff910de8761c9ac0a7624f63ccbedb4b9181fc

Observation 7bc08d9e-aed8-44a2-9db2-d7e210a79c52 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:b2ba9b4d6934e02dae9bc71719c466e667a24f5ef6c94a6e66d804af2dd7366d

Observation 5c43be9f-b268-4306-add3-48da0bc2548f · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:2f15f6d371481d5550991a7861c09bc5c30b963f3b6040e12e89352f5f799577

Observation 91cf2f12-65f6-43ca-bd0d-e83599290789 · outbound

This paper cites Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:3c9bd30884d34379a47ce297b97002bb90a79fa9016a5c3fc6160684ba5ef0dd

Observation 732851c2-45a6-41ac-9310-5f615cf997b8 · outbound

This paper cites thefuzz: Fuzzy string matching in Python, 2024.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models thefuzz: Fuzzy string matching in Python, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:9e2467be128055925a8764e3699c03d22a87758ad688dc98ba88e53b994de478

Observation 21fbbfe8-de3d-4c63-8f8d-10fa6c74c663 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:31ec5163ded24a7d3be639df5cab388313a94a6f976118588176a21692e7c609

Observation bec311c1-c363-4b0d-892d-0ce7683f1d4a · outbound

This paper cites OpenAI GPT-5 System Card.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models OpenAI GPT-5 System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:63efada96a563241e44473a555060088cf3b2f6a9d681cb8dda8645a9e900aba

Observation 914e1bc7-9d23-4e8f-8610-9eb17f60e148 · outbound

This paper cites A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:5a8aef32755cd5c778e8f82ca4fc82cffd6b71e965ff20c764af1a682e020fbf

Observation 70561fcc-fc23-44d7-addf-12c6f7ec59bb · outbound

This paper cites DetectLLM: Leveraging log rank information for zero-shot detection of machine-generated text.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models DetectLLM: Leveraging log rank information for zero-shot detection of machine-generated text

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:2e8d3373c5f7c0bdb61585e9ace46f3f4718957c9b327dd0f8c3ca8a28d12fe1

Observation 970484f5-383a-4824-85f1-4ef402ab00dd · outbound

This paper cites Investigating prior knowledge for challenging chinese machine reading comprehension.Transactions of the Association for Computational Linguistics (TACL), 8:141–155, 2020.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Investigating prior knowledge for challenging chinese machine reading comprehension.Transactions of the Association for Computational Linguistics (TACL), 8:141–155, 2020

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:e8b2397d3a9f215ce9dc30f82630b73238ec8dbaf22e91afb2082200f80f8844

Observation ff825f88-ba4b-4b42-9901-7f4333f48658 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:4a940ad6191d371de183ad740b02fd4ef80c55f7ada2a3e180a910eb5ac04e0d

Observation 9cf904e1-8fb2-4068-affc-a4aebf5fb687 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:0a69ebbb40fd7b8019ade80714fceaf07c0bdf0eaa87228588ca0111e3e6bd6a

Observation 570ba94a-418c-4fbd-855e-f0f0a14f1219 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:7ad50b1bb02f7070e81c65f914455f0451686b75798c8f1e26310550642464fb

Observation 8acc3b5f-2cb8-40c7-8656-0163e660db50 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:df4a3ee62cad6a8696b15dacf65b3970563c379f4f057c8d0ccd1277cc6b53ce

Observation c0aec7b8-62f8-47a6-8f2c-66861ccc1e58 · outbound

This paper cites Do-not-answer: Evaluating safeguards in llms.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Do-not-answer: Evaluating safeguards in llms

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:ae3862e8f26833149b845725c5a80a28ffa23e87619a39a8b1ab1a129b1740a6

Observation 56412ce6-3979-4eab-acb3-6b5f74eb09f5 · outbound

This paper cites Measuring short-form factuality in large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Measuring short-form factuality in large language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:d07f702132e95de445f3e4c47fa66853d5570ac3bdfcd921884e45c22fa462fa

Observation 9ec93d36-b4e4-46de-bef3-b6f996c02f22 · outbound

This paper cites Benchmarking and defending against indirect prompt injection attacks on large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Benchmarking and defending against indirect prompt injection attacks on large language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:033bc135e38e6ab1dd15542263f2572090f6c72d2a41f712f06fa1fbce0b3db7

Observation 455e3c51-0bb8-470d-860c-338d8e6cf62f · outbound

This paper cites S-eval: Towards automated and comprehensive safety evaluation for large language models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models S-eval: Towards automated and comprehensive safety evaluation for large language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:96c103abdff77cb271088fb2a3f57fc32ed97a339065700c89e2537117d3d539

Observation 13e43977-b44e-4383-bfa6-66195151621c · outbound

This paper cites From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:d11b70338e2970a976bc3e6490ba0388a891dd3c1126092f1741f5711268efc1

Observation 5d41dde8-4754-40d9-b1f4-f6ffb33c94e8 · outbound

This paper cites Many-Tier Instruction Hierarchy in LLM Agents.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Many-Tier Instruction Hierarchy in LLM Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:460c02894b4925fa51549f91619aeabbc19b135deb3ea6595ff70050c0df88d4

Observation 59814d93-a853-4900-ad16-902cc5ea06c8 · outbound

This paper cites Iheval: Evaluating language models on following the instruction hierarchy.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Iheval: Evaluating language models on following the instruction hierarchy

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:6b01d24e8b58544923289b85960e43e123d848776ec9cda0926f940150fc7e7a

Observation d90b5afb-acf4-4693-8e39-7cc9bbff10d8 · outbound

This paper cites Wildchat: 1m chatgpt interaction logs in the wild.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Wildchat: 1m chatgpt interaction logs in the wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:74c4862270cc7b04fb6413e1c6e9bbb6117c6135e2214a359003343ca77ad6b4

Observation 0ccb6437-a238-4d30-be10-296192e4cb1a · outbound

This paper cites Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:05316069fc48bb137d60c71628ec8a1d62dc11d13e876bc9aa6eb1fc83868a82

Observation b3832bec-1717-443c-92c3-1f82758dfc20 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:f24c24e533848932d81154a89628ba3e21004d8b3e92374bf23eccdce978df33

Observation a82181ce-9afa-48ed-8741-11db055fdc6f · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:d21267eefa0318920af1324ec2d428b00a27e0fc788343106d591fbefdf0b7c3

Pith citing papers

No inbound Pith citation observations are available.