Pith. sign in

Paper Citation Record · LEDGER

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.18258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18258 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:58:44.792769Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1d01a30-cc13-41bc-87df-4428c04ff6a0 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:39.777727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:39.777727Z digest=sha256:be6ab7aa6ffff6376e669195309cdbc268d6767e2279bfb868a75fca3b74b12a

Observation e997443e-0234-45a4-966f-0222cca11a2f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:39.895334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:39.895334Z digest=sha256:6992425698ecb700a943a8325556b685035715e1b4bb6d0bd37fb2cf94a3b906

Observation a641f918-90d0-4565-ac0b-88e965e5dc29 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.042083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.042083Z digest=sha256:1db08295d83acac2589e1bcabbc556dce9a71527c9469150d44b3e2bb1a18948

Observation 9d72c03d-8f2a-4a5f-a77a-4149ff0e054b · outbound

This paper cites Rlhf: A comprehensive survey for cultural, multimodal and low latency alignment methods.arXiv preprint arXiv:2511.03939, 2025.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Rlhf: A comprehensive survey for cultural, multimodal and low latency alignment methods.arXiv preprint arXiv:2511.03939, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.181922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.181922Z digest=sha256:081b0a81ebc14295d9dacdce89ba4f3ccc71d07badc09ed6851b078cc3db4a4f

Observation 78a6dedc-8cd2-4cb2-8b7c-7e324c731514 · outbound

This paper cites MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.294009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.294009Z digest=sha256:6f123fd0ea5c0c74ff8e6c6fd7f219214d180493f1a398eb56e0fe86fe9d9e6d

Observation 1d6d387b-8abf-4ccb-9373-7382e95a2bdf · outbound

This paper cites Confronting reward model overoptimization with constrained RLHF.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Confronting reward model overoptimization with constrained RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.437893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.437893Z digest=sha256:4f2bfa93d3c6e08f7b43b60a9b2a5739b54a8566ef1d3b06a13b72c8d4ccdd51

Observation 2a7161b2-cfe7-4a4f-aa95-d690a487ce79 · outbound

This paper cites Regularizing hidden states enables learning generalizable reward model for llms.Advances in Neural Information Processing Systems, 37:62279–62309, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Regularizing hidden states enables learning generalizable reward model for llms.Advances in Neural Information Processing Systems, 37:62279–62309, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.554543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.554543Z digest=sha256:226719509f8de7e8195547950d3516a7859eefae502110dfe4c33520a9224aad

Observation 811c9ace-3bff-4de1-b814-c5edaaea8256 · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.672240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.672240Z digest=sha256:6f626d0b40485e0b1cac24835f4003c9be7cda6b84da44a929557f4f33ef3880

Observation 9f1947b5-3dd1-4def-8064-62d01c363751 · outbound

This paper cites DPO meets PPO: Reinforced token optimization for RLHF.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF DPO meets PPO: Reinforced token optimization for RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.777408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.777408Z digest=sha256:1b96ebf9c75db20e928a09201b00db892df515309e0e162a634335d066d72e10

Observation c691c864-bdc1-4f50-a799-165888626e13 · outbound

This paper cites Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.883141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.883141Z digest=sha256:d1e3024318f53f318fb85a3676631c906dac6a14ed43b42fb4e54920fde3c059

Observation 466a9e3d-1a39-45d3-9d58-f200606e3a30 · outbound

This paper cites Capo: Towards enhancing llm reasoning through verifiable generative credit assignment.arXiv e-prints, pages arXiv–2508, 2025.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Capo: Towards enhancing llm reasoning through verifiable generative credit assignment.arXiv e-prints, pages arXiv–2508, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.971384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.971384Z digest=sha256:a0e8b98d0412982cfdcda91c0913219f59773d85eda98e6e69ff721fb02e5547

Observation 66bc4ba3-52a9-44f0-8c10-e91d630d0ece · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.076306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.076306Z digest=sha256:42fa9b4edfee8e71fc581dd78ab4d4e83ba95a16fcddff0b39f1f3ae247e41af

Observation 8b3e7637-bf81-47cb-905e-67f9f28ff4d7 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.205872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.205872Z digest=sha256:fcc976f7b0211bbf143e8a0b1ebe23dd67bb61312eeba3c39b1ed38d1579831f

Observation d68823f1-e86b-47a1-976e-7c63d2205a7d · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Finetuned Language Models Are Zero-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.317159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.317159Z digest=sha256:dcba2a2c8799ab7e3cc0cd0deba8facf5c8d0b7380c00d0a70201dc117dd067c

Observation af2e5afc-4658-4ede-8538-ee45c4e8f27b · outbound

This paper cites Learning to summarize with human feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Learning to summarize with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.419591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.419591Z digest=sha256:c467022d3e9b1cba89ec50b948b0887cc280c23c8fa51159a9545d843485f71a

Observation 5f21a7eb-3ce9-4467-86ea-e1bea4692410 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.532409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.532409Z digest=sha256:938df7bba1b7b96af84b0d7cacc1d72daaab403bfa4961ae407dcd7dfaef7bef

Observation f26ac867-d237-4dfe-8199-f00c18ea6625 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF The Alignment Problem from a Deep Learning Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.660778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.660778Z digest=sha256:453240ffbe27b547e39263e3f48d180664972313b1e73deebbc4c06039bfaf30

Observation 3930e48f-e5d7-44e7-930b-044532b77b35 · outbound

This paper cites Stepwise alignment for constrained language model policy optimization.Advances in Neural Information Processing Systems, 37:104471–104520, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Stepwise alignment for constrained language model policy optimization.Advances in Neural Information Processing Systems, 37:104471–104520, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.740345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.740345Z digest=sha256:f482dc6fec3e21f7f464cd30bc793ddd7d1cf8a3c67d2fa412ca2c8f0cab46aa

Observation db1f8bcc-148c-46c9-a36d-ed815553aac2 · outbound

This paper cites Panacea: Pareto alignment via preference adaptation for llms.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Panacea: Pareto alignment via preference adaptation for llms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.863440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.863440Z digest=sha256:12cc7d7ba198b8d64853648c53993ada95b51ca7bbba916874b6e6ac69fa5388

Observation 7bf01e07-aabb-4491-b581-22710502770a · outbound

This paper cites Preference learning for AI alignment: a causal perspective.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Preference learning for AI alignment: a causal perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.027109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.027109Z digest=sha256:8b965c607cc8b8c8d444448b9a4bc60ff384d2126dffb3f386c1defd73b4586f

Observation e0e8ae0d-f299-40f2-8fd8-e7a528581d5c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.150952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.150952Z digest=sha256:b461ccf30e0695b158acc215063c8e06dc9edda2d37ab1550915c63f6ae54c8f

Observation 382bfbfe-7516-4062-ae85-35d2c6448ef3 · outbound

This paper cites Contrastive Preference Learning: Learning from Human Feedback without RL.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.282065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.282065Z digest=sha256:6c0f2ef1c36392464629d162c50ddb67a54f479981228c2f864136ddf7ff9f12

Observation 79e82d91-7de6-41a1-b1bf-5bfc59681804 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ORPO: Monolithic Preference Optimization without Reference Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.386598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.386598Z digest=sha256:c532f82161fd1706622d563b4dd575459bead57d2d9db2c0f6a8d4af20589c7f

Observation 55d6b48d-d7e1-4024-be27-3c0143647f9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.521688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.521688Z digest=sha256:8101fdf6ce3f73bc896a47d8453421c73a780ddba78e31dc263ddbf875a7b39e

Observation 935ed82c-84b8-4118-b846-e028770ca025 · outbound

This paper cites Scaling laws for reward model overoptimization.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Scaling laws for reward model overoptimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.664558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.664558Z digest=sha256:18956e477e953e412834d2cf2719aa97a86b35d193af4f8c95f992156d67c9d2

Observation c9812646-da51-41eb-a4e6-fa5dfe955922 · outbound

This paper cites Learning to Understand Goal Specifications by Modelling Reward.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Learning to Understand Goal Specifications by Modelling Reward

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.782419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.782419Z digest=sha256:2f1bf94f74a3a2620e9430081e6dc04f85d31731b97bc1b4c1ce3138d8a3dcfb

Observation 2d5a709f-7554-4495-9096-3cfa72d7bed2 · outbound

This paper cites Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.937172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.937172Z digest=sha256:46600c71203a9055214b1da9afb7e0d64263b219b764568aebf952e7e6d173bd

Observation 2cc5c481-6322-44df-8ee8-1de5b8809906 · outbound

This paper cites Proximal Policy Optimization Algorithms.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.991216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.991216Z digest=sha256:2e2e762fbc9646566b5226eee0844264b730fa3e83982c94336c0f99e7d76f7c

Observation f3806584-e0f4-46d1-af92-90f3e81efe65 · outbound

This paper cites Privately Aligning Language Models with Reinforcement Learning.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Privately Aligning Language Models with Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.111907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.111907Z digest=sha256:194af68b3b85c446857009235d3f66e1e4d09167cbf0b899320f77942a59b7de

Observation 960fd674-aabb-4829-8aad-7e28efe7dab5 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Dense reward for free in reinforcement learning from human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.234422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.234422Z digest=sha256:64a21a9a5162d842a94c94ed1f5e9e4a4173a44c369688b857cae9007fd12067

Observation 91756c88-c717-4cf3-b916-2d149021361d · outbound

This paper cites Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.317305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.317305Z digest=sha256:679d1966158e4f947b5739a33a1879f6224a6670b4ba1dc046b9b09b3ebfd981

Observation 033c1848-ef05-4e5d-80ec-ad830e59491a · outbound

This paper cites SCAR: Shapley Credit Assignment for More Efficient RLHF.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF SCAR: Shapley Credit Assignment for More Efficient RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.421940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.421940Z digest=sha256:0132e0028bc9f23b419ad972f5721829d72235926db6c6719691b61d7f92d75a

Observation b6b31347-93a8-45a4-91e0-514061b8e5b6 · outbound

This paper cites RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.513582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.513582Z digest=sha256:b9ce4cccb02847706c9594402669e9c0c2d31595b07fdbdd25322ee8f45fc106

Observation 498542d1-fddb-41dd-8ee0-ee271e8119cc · outbound

This paper cites A critical look at tokenwise reward-guided text generation.arXiv preprint arXiv:2406.07780, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF A critical look at tokenwise reward-guided text generation.arXiv preprint arXiv:2406.07780, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.614063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.614063Z digest=sha256:bb509d2049f277a17d8a8724eee8ca971e5abe3bae4dc3b51b95f92adcb8cd56

Observation 396a5801-d84e-4043-b36b-9754894356f4 · outbound

This paper cites Judging llm-as-a-judgee with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Judging llm-as-a-judgee with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.712545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.712545Z digest=sha256:2bdb2ed8e1d7dffbc23a598a735657095c700d18bdb1eb508788f6e46cfb2bcb

Observation 49ecb033-fee4-43e5-8fe7-f7a3200cdc60 · outbound

This paper cites GPT-4o System Card.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF GPT-4o System Card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.813439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.813439Z digest=sha256:9270f65a8cb261028f2d62ee7f761040812e82617fe9db2eac7c5cb6a58be060

Observation 4c3221b3-a337-4f26-97bb-7572eddb06bc · outbound

This paper cites GPT-4 Technical Report.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF GPT-4 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.916499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.916499Z digest=sha256:32cc0987bcfe67a05ac88f74a0d364e3ec7ee45f031e84d4428b83fc5af2f6f6

Observation 8e76b477-3e0a-403b-91ae-c82d6ca1d08c · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Raft: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.960879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.960879Z digest=sha256:c5b7ee49d0a87bb8dc55d318ab077ebc6ac5c35fe55ba346a7d49abecdc323db

Observation a3cf37e1-420e-40a5-9ee9-a40c56a6fdbf · outbound

This paper cites Cambridge University Press, 1999.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Cambridge University Press, 1999

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.033578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.033578Z digest=sha256:9cf4178fa7b79a2c4c6aa232c5aee670cf8302bef879ce6f7f9e4b6d5e3d0dd9

Observation 21c71261-88d4-4fa8-b19b-c396ef0f6bd3 · outbound

This paper cites Multi-Task Learning as a Bargaining Game.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Multi-Task Learning as a Bargaining Game

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.111762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.111762Z digest=sha256:2d456c6389af0c0116b9bc15258af217e7518a78181e2965e60d29ebf3112919

Observation 8e27c840-2e87-4f65-a49d-bdf9865982a5 · outbound

This paper cites Fairness-aware meta-learning via nash bargaining.Advances in Neural Information Processing Systems, 37:83235–83267, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Fairness-aware meta-learning via nash bargaining.Advances in Neural Information Processing Systems, 37:83235–83267, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.224430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.224430Z digest=sha256:fd84dae318fee7b30f57ae80fc0e5b146a33a6ac8b04f004624598f74400643d

Observation c329ffb8-72cb-4586-a3e9-6ff234c50431 · outbound

This paper cites Hindsight-DICE: Stable Credit Assignment for Deep Reinforcement Learning.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Hindsight-DICE: Stable Credit Assignment for Deep Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.322878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.322878Z digest=sha256:f1756bbd6da14ed57c935b0c4965c54073d91d3518aea511c464d2353fb2c1a5

Observation 5a37414f-9070-4741-b658-b8279522a888 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.444288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.444288Z digest=sha256:e12db85c7b14aaa8dd3d0d60f1bf27db4f5f278a473183bdd92f8adc11c4100b

Observation 230a840b-9267-4b04-b53d-8fdb971f7559 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.543081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.543081Z digest=sha256:580def729995d2d332e2c51b245547ac1349f31dcd71048ae8885aa7a4732af4

Observation 832150d9-4421-4da7-935a-8d3d61499e3a · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.661562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.661562Z digest=sha256:b220285c9d4bd940815b0878812f7c232392258d38730dbab1f3147fba7a8076

Observation b1c598da-52c3-485c-99dc-b42c5ac3be2d · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.792769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.792769Z digest=sha256:c101db826d1282b4f4622ff488e76bafd26fcf61bca04d147bbbf9a4e086a429

Pith citing papers

No inbound Pith citation observations are available.