Pith. sign in

Paper Citation Record · LEDGER

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.18258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18258 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:58:44.792769Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1d01a30-cc13-41bc-87df-4428c04ff6a0 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:39.777727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:39.777727Z digest=sha256:51380db8062b73711739d81f1797ce86a3b7882a14565b98f9eba1365c01ed34

Observation e997443e-0234-45a4-966f-0222cca11a2f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:39.895334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:39.895334Z digest=sha256:888901a425009d9b1a223a1335e85b8f84d62d01cea64fe196555d1440608a2b

Observation a641f918-90d0-4565-ac0b-88e965e5dc29 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.042083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.042083Z digest=sha256:037a395f8a4fa42304532f29d17f1fd2e3316cfdd1d19bab7b71a71346296f7f

Observation 9d72c03d-8f2a-4a5f-a77a-4149ff0e054b · outbound

This paper cites Rlhf: A comprehensive survey for cultural, multimodal and low latency alignment methods.arXiv preprint arXiv:2511.03939, 2025.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Rlhf: A comprehensive survey for cultural, multimodal and low latency alignment methods.arXiv preprint arXiv:2511.03939, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.181922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.181922Z digest=sha256:b54be89e725f6295280c713335643f88b83bbf3c91d815292e6e6fb7229e442e

Observation 78a6dedc-8cd2-4cb2-8b7c-7e324c731514 · outbound

This paper cites MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.294009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.294009Z digest=sha256:ac7b711ffb59e67213cfd368551e8a09cca2344417e62242b7d6b0a338628493

Observation 1d6d387b-8abf-4ccb-9373-7382e95a2bdf · outbound

This paper cites Confronting reward model overoptimization with constrained RLHF.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Confronting reward model overoptimization with constrained RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.437893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.437893Z digest=sha256:b1f6d5e3aa6cf35350140b40ff48612974b18f1d7de69b7ea49ce254b5d2ec01

Observation 2a7161b2-cfe7-4a4f-aa95-d690a487ce79 · outbound

This paper cites Regularizing hidden states enables learning generalizable reward model for llms.Advances in Neural Information Processing Systems, 37:62279–62309, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Regularizing hidden states enables learning generalizable reward model for llms.Advances in Neural Information Processing Systems, 37:62279–62309, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.554543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.554543Z digest=sha256:94e284692a493990c75339487a020d1d16a00dd640f8350cb4151161209fa945

Observation 811c9ace-3bff-4de1-b814-c5edaaea8256 · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.672240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.672240Z digest=sha256:87279dec3c2c7668484b24246c9cb130ae523952455d1a9b7251676dae8a1827

Observation 9f1947b5-3dd1-4def-8064-62d01c363751 · outbound

This paper cites DPO meets PPO: Reinforced token optimization for RLHF.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF DPO meets PPO: Reinforced token optimization for RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.777408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.777408Z digest=sha256:76477794d65414e8529dff6cc29961d19209053bdd62b2b0a447516187874073

Observation c691c864-bdc1-4f50-a799-165888626e13 · outbound

This paper cites Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.883141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.883141Z digest=sha256:3bb99f5ba5248c450e84ba7f7a79ef4f0618123fe05acbdf7186cea1357bd250

Observation 466a9e3d-1a39-45d3-9d58-f200606e3a30 · outbound

This paper cites Capo: Towards enhancing llm reasoning through verifiable generative credit assignment.arXiv e-prints, pages arXiv–2508, 2025.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Capo: Towards enhancing llm reasoning through verifiable generative credit assignment.arXiv e-prints, pages arXiv–2508, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:40.971384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:40.971384Z digest=sha256:1ba393936a34d305447aa38a5be3237a459627166a7ccec0a8816341bd294f9c

Observation 66bc4ba3-52a9-44f0-8c10-e91d630d0ece · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.076306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.076306Z digest=sha256:a7ac5afc7c2c12e128ae8ae75ddf4e1e427315154d7d5e70a7bb36ba77ddaae8

Observation 8b3e7637-bf81-47cb-905e-67f9f28ff4d7 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.205872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.205872Z digest=sha256:80c2c50387c8b8206252b14cc5f51de6ca64a909aa89ea4a36bad7a1c63ca310

Observation d68823f1-e86b-47a1-976e-7c63d2205a7d · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Finetuned Language Models Are Zero-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.317159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.317159Z digest=sha256:e55970beeb7601c0d9b126a0eee50e73fb19154aed65047d7b0b20565de22acb

Observation af2e5afc-4658-4ede-8538-ee45c4e8f27b · outbound

This paper cites Learning to summarize with human feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Learning to summarize with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.419591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.419591Z digest=sha256:166aa325badc3c63de384e620d40804420b3f2176c2e50901fb5643b3de833f3

Observation 5f21a7eb-3ce9-4467-86ea-e1bea4692410 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.532409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.532409Z digest=sha256:bc9d035648a5adc808bc5517650bac57e1d517ccf8148a1080bfa1eddab02b9a

Observation f26ac867-d237-4dfe-8199-f00c18ea6625 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF The Alignment Problem from a Deep Learning Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.660778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.660778Z digest=sha256:fdd0c139e568dd847a7c1c3cbb4b6d855fd90ce1973ccb03e7f5b5f4b1fd41af

Observation 3930e48f-e5d7-44e7-930b-044532b77b35 · outbound

This paper cites Stepwise alignment for constrained language model policy optimization.Advances in Neural Information Processing Systems, 37:104471–104520, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Stepwise alignment for constrained language model policy optimization.Advances in Neural Information Processing Systems, 37:104471–104520, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.740345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.740345Z digest=sha256:e7a10d865c25235cd4088f06854718a5a0c7cca8292efb03d3839a7f883e8695

Observation db1f8bcc-148c-46c9-a36d-ed815553aac2 · outbound

This paper cites Panacea: Pareto alignment via preference adaptation for llms.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Panacea: Pareto alignment via preference adaptation for llms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.863440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.863440Z digest=sha256:996a0bfc4fa06e94ce954fd9144cb77d4f40bcdec271ec5cc463cd74b29282d1

Observation 7bf01e07-aabb-4491-b581-22710502770a · outbound

This paper cites Preference learning for AI alignment: a causal perspective.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Preference learning for AI alignment: a causal perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.027109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.027109Z digest=sha256:b207a89d615e64f5155b9ed0e445432b7c8061de8db0093188a9880ccad5b5ca

Observation e0e8ae0d-f299-40f2-8fd8-e7a528581d5c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.150952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.150952Z digest=sha256:b62e99253a277bf0e7c6824640e0fd590ccf512e45d3a0725009676fbede3a05

Observation 382bfbfe-7516-4062-ae85-35d2c6448ef3 · outbound

This paper cites Contrastive Preference Learning: Learning from Human Feedback without RL.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.282065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.282065Z digest=sha256:208a14c686e8d3bf8b6302e20397dbea44c785fe0b7d96277fd945bdf1c79d7a

Observation 79e82d91-7de6-41a1-b1bf-5bfc59681804 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ORPO: Monolithic Preference Optimization without Reference Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.386598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.386598Z digest=sha256:3eed50f59a3650e319a45f24b73ce6a825022754db31d53107e8eb2bad7b3f26

Observation 55d6b48d-d7e1-4024-be27-3c0143647f9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.521688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.521688Z digest=sha256:9fb09328c9f2df444488072e5e6a6cee1db318e71a886adbfaf7a39f5b79f82a

Observation 935ed82c-84b8-4118-b846-e028770ca025 · outbound

This paper cites Scaling laws for reward model overoptimization.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Scaling laws for reward model overoptimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.664558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.664558Z digest=sha256:32d2ab3e35d748415b4830b53762a45832b5c822f373a5b344d1d35d78a95bc5

Observation c9812646-da51-41eb-a4e6-fa5dfe955922 · outbound

This paper cites Learning to Understand Goal Specifications by Modelling Reward.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Learning to Understand Goal Specifications by Modelling Reward

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.782419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.782419Z digest=sha256:9afb5ba3e10229e211195a693744787faff16a2336753c894cf75404f462e2e1

Observation 2d5a709f-7554-4495-9096-3cfa72d7bed2 · outbound

This paper cites Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.937172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.937172Z digest=sha256:54240dc1f65731444f1a62d37600fa6b5588111591b839d4f25895da37bf8b8e

Observation 2cc5c481-6322-44df-8ee8-1de5b8809906 · outbound

This paper cites Proximal Policy Optimization Algorithms.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.991216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.991216Z digest=sha256:63642f1ed1a9b8113d965c6d7b830c400f7e4ed4da1304796a76c76bbe44fd2f

Observation f3806584-e0f4-46d1-af92-90f3e81efe65 · outbound

This paper cites Privately Aligning Language Models with Reinforcement Learning.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Privately Aligning Language Models with Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.111907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.111907Z digest=sha256:eee28984e34109e2b79c02cc6e984b4baefae20aa623f41648cc7c1d47797dfb

Observation 960fd674-aabb-4829-8aad-7e28efe7dab5 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Dense reward for free in reinforcement learning from human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.234422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.234422Z digest=sha256:954cb977298268dc4d137f885f990bfb6bc3ba66a9f04964267c74f4ae5ccc0b

Observation 91756c88-c717-4cf3-b916-2d149021361d · outbound

This paper cites Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.317305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.317305Z digest=sha256:34676ad617ff5e1add43894199074bcd9d93738ea3d6eb314a8de9119d0b79cf

Observation 033c1848-ef05-4e5d-80ec-ad830e59491a · outbound

This paper cites SCAR: Shapley Credit Assignment for More Efficient RLHF.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF SCAR: Shapley Credit Assignment for More Efficient RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.421940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.421940Z digest=sha256:aad413c596de86a03ca0503ffbc4c34e5d17343b22eef5e7af3af31305f862cc

Observation b6b31347-93a8-45a4-91e0-514061b8e5b6 · outbound

This paper cites RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.513582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.513582Z digest=sha256:a6086047759a7d6d4d0d471a6c00cc4f79cdafbf5f757877af982395c6536fbc

Observation 498542d1-fddb-41dd-8ee0-ee271e8119cc · outbound

This paper cites A critical look at tokenwise reward-guided text generation.arXiv preprint arXiv:2406.07780, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF A critical look at tokenwise reward-guided text generation.arXiv preprint arXiv:2406.07780, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.614063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.614063Z digest=sha256:05b0adde33ae3d5151ac2890a22960cd881be528ebb778738f578f44ec688a9c

Observation 396a5801-d84e-4043-b36b-9754894356f4 · outbound

This paper cites Judging llm-as-a-judgee with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Judging llm-as-a-judgee with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.712545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.712545Z digest=sha256:89f76175a42cfa9b8192eed13ced835fbb04fceaa31e83f7238897a91510ad22

Observation 49ecb033-fee4-43e5-8fe7-f7a3200cdc60 · outbound

This paper cites GPT-4o System Card.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF GPT-4o System Card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.813439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.813439Z digest=sha256:a81bc5f73009b8fa00453438fd28a6ded3203ba0466b8274ea568862a90b0f2c

Observation 4c3221b3-a337-4f26-97bb-7572eddb06bc · outbound

This paper cites GPT-4 Technical Report.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF GPT-4 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.916499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.916499Z digest=sha256:e6d0544850594fef29504e70018d086eebe7e684b044a7949dc45ec3e2fc336d

Observation 8e76b477-3e0a-403b-91ae-c82d6ca1d08c · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Raft: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.960879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.960879Z digest=sha256:0ffa87f413e9923a12fa2896c30d08b90f55b018c2be2bced090e7922f9b4888

Observation a3cf37e1-420e-40a5-9ee9-a40c56a6fdbf · outbound

This paper cites Cambridge University Press, 1999.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Cambridge University Press, 1999

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.033578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.033578Z digest=sha256:77eed7980a090ddf7c754975d1904392c6c1c2389e2dcda42a4c5d6d8eb62617

Observation 21c71261-88d4-4fa8-b19b-c396ef0f6bd3 · outbound

This paper cites Multi-Task Learning as a Bargaining Game.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Multi-Task Learning as a Bargaining Game

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.111762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.111762Z digest=sha256:3b922f4cfb36ecf4a0413ccc683d28e8682a301eb4bd884b80cb0b9c59dc024d

Observation 8e27c840-2e87-4f65-a49d-bdf9865982a5 · outbound

This paper cites Fairness-aware meta-learning via nash bargaining.Advances in Neural Information Processing Systems, 37:83235–83267, 2024.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Fairness-aware meta-learning via nash bargaining.Advances in Neural Information Processing Systems, 37:83235–83267, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.224430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.224430Z digest=sha256:c322176b6f8403c28fca43575347e8046c20c74df28e79478ca44d10c87e33f3

Observation c329ffb8-72cb-4586-a3e9-6ff234c50431 · outbound

This paper cites Hindsight-DICE: Stable Credit Assignment for Deep Reinforcement Learning.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Hindsight-DICE: Stable Credit Assignment for Deep Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.322878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.322878Z digest=sha256:99d8da966e7cd0d50b4d9606af025e776e56a8b2cc02c74466a86461bd17ac24

Observation 5a37414f-9070-4741-b658-b8279522a888 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.444288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.444288Z digest=sha256:f9a9b6bce7e5d11308ea0b594bbc53ac97a2ab11e64e8f60a1dd6373f9f31638

Observation 230a840b-9267-4b04-b53d-8fdb971f7559 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.543081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.543081Z digest=sha256:04b28e816d6fc0c4d79c666002a77c6f8b3038626cb75c0b9d6531788eb7b984

Observation 832150d9-4421-4da7-935a-8d3d61499e3a · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.661562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.661562Z digest=sha256:6c05654d347eb4c2f614750dd2fe99d60fb07ab19be8e35ffa84f7c957f58e31

Observation b1c598da-52c3-485c-99dc-b42c5ac3be2d · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:44.792769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:44.792769Z digest=sha256:bdfde627efa47218fe37a673110307385ec9ff8667bae30b9284a54070f04f37

Pith citing papers

No inbound Pith citation observations are available.