Pith. sign in

Paper Citation Record · LEDGER

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment

As of 13 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2411.09341.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09341 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:51:53.299361Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6544a8d-a5d6-45ec-9049-433565394c03 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.140481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.140481Z digest=sha256:717e3d18593a46e7cfed6aa1c3ab448a8bceab5369fddc5d089e8f729cec4704

Observation ec330355-af13-4987-84ff-9aec71cb526b · outbound

This paper cites write newline.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.146561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.146561Z digest=sha256:879d4350268d186ef8de59d53d219a41e5873f84d47813f6457940ce088ead7b

Observation 0e711594-8e85-4c93-9b0d-0a169ca5e309 · outbound

This paper cites GPT-4 Technical Report.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.152257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.152257Z digest=sha256:f9c5378a6b7e5d58ce204b530e032cde3d0a343c6db265f6baeecb64fb1cdf8f

Observation 07a35865-d9b9-46b2-a8eb-14a626b7c5be · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.158078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.158078Z digest=sha256:f17eaa518fa0ab958e8d8d5228943c113aefd76509afdf7f6f4332baad06e045

Observation af6357ba-414a-47e1-b329-5a5bf6b44c93 · outbound

This paper cites A.; and Terry, M.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment A.; and Terry, M

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.163436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.163436Z digest=sha256:17383cbee518cfa0a2d370c1e8396844285711fa5039b595d1dfdb480d383fa5

Observation 913daf76-fb90-4b3e-b938-856f51740ebc · outbound

This paper cites J.; and van der Schaar, M.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment J.; and van der Schaar, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:51:53.832025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T20:51:53.168178Z digest=sha256:b3e6661465195d984db44908da09a245be7d83a57882d80c18fb9bea97fa7173

Observation 79fa944b-b6a5-4a26-a5ec-b35439a72899 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.172925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.172925Z digest=sha256:7ccb5909d279705521c47aa20a1831f2466a7427a1600caf3b19a8ee8497f6f6

Observation e0a04680-6b45-42a7-9a5b-3c3d099f4132 · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.177784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.177784Z digest=sha256:f995a32e4b5f9e4ce38638cf6a33e4c2c1a76934b030d1b78dcda3034258d2ba

Observation 3dd6691b-88d4-4ef4-a09f-c38caea325b8 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.182857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.182857Z digest=sha256:5930dc04183759fc4d4ec6a9a7cc76098e4f8f3f3a052421ee1555e4d77e64d8

Observation 180e9686-c04d-49b6-9e74-8d0dc51b390b · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:51:53.808270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T20:51:53.187657Z digest=sha256:080b6f24f22f17d8793899fee0eb0a87bf61216427e78b997053c70b75301e68

Observation 6a1a0f39-1a00-4596-90d0-c3ecab994e9b · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment The Curious Case of Neural Text Degeneration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.191887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.191887Z digest=sha256:fb76ce0d46fdd5e24e200911e9bf8917488d0d6ff2a54c7b783ddb6004abe810

Observation d11b751f-1c02-46ca-9c8e-d1027f6df6b0 · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.196903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.196903Z digest=sha256:46179487a7e3ef757a2a203d8c6d2bbf9a7e49c836fe6a9fd0fea9314dc53bbb

Observation d4390c02-6ff5-4328-ba88-08fc5c607b18 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Playing Atari with Deep Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.201652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.201652Z digest=sha256:366906329b977eb0f1d4efdf0fd306f861b59147fbd3e4f82636ef0cf2c2b10b

Observation a3fd12f4-8589-4611-b78f-e21d994e7d5c · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment WebGPT: Browser-assisted question-answering with human feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.206287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.206287Z digest=sha256:4c3750e14e16af3f039367fc470f0ad2f70e23c4d265b344fc6915300ad0d5d3

Observation 0f3beb00-962d-4a3c-95e5-dc13be36a4b0 · outbound

This paper cites Y.; Russell, S.; et al.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Y.; Russell, S.; et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.210920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.210920Z digest=sha256:2d69c06dc04c661e4def123e4056f667cecd05f22799a08b926b2e50b5ae64aa

Observation a410275f-acb1-49cb-8dd7-7e14af80026c · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.215100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.215100Z digest=sha256:851e78a3d02d1978bc773a03cbf08f5f73c79ac235c659d8ad306ec254ad86fb

Observation ede472ff-1e61-4887-b132-1679e2df341b · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.219371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.219371Z digest=sha256:4dd8eb58b7a1b1e7b2f42117bbfe00a55c98c0941b3a47011d4d0a10019d70fd

Observation 92eb5f4e-8d06-4743-acbe-ff91a888c7b5 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment D.; Ermon, S.; and Finn, C

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.223593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.223593Z digest=sha256:1be34e0b615d7496f5fc572038a761ac95d9c8eebddbad5e15ab5c573b15470f

Observation 59597ecd-d41a-4535-a1cf-a4ef9e0c76e0 · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.227840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.227840Z digest=sha256:8942431d958aa91c5c6fe52cfaaafab0be6edf41f06dcc9d4067531dda6a038a

Observation 4d466063-c013-4035-a85b-dba08fc55272 · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:51:53.748521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T20:51:53.232500Z digest=sha256:3ad415faa3ffe57309c557f76edec9bdff1221f1dd4153eab76cd24e08622d8a

Observation a4512eda-ed6d-4102-8cf3-c51b8e3cd1fa · outbound

This paper cites Proximal Policy Optimization Algorithms.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.236719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.236719Z digest=sha256:2bcc5b52dd5d8548dff5bf4f2220821f4d5567e6c2f5831460ff2e68bbc710d4

Observation 88c54870-04ac-4b88-88ed-0e746332ae9e · outbound

This paper cites Large Language Model Alignment: A Survey.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Large Language Model Alignment: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.241261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.241261Z digest=sha256:868a66778a9a33e76f999d0cc7373ebbe111ec2eaa7de992a1abaa77f8440891

Observation 64667873-9296-4c99-8cca-8b6df2649084 · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.245842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.245842Z digest=sha256:fb8d3c041f9d7d1c25376ae38a1790bd58684661229e90083cc57a9ddc875041

Observation d98f7b19-1569-4d91-a1bd-3423859b9ecb · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.250055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.250055Z digest=sha256:045aa1cfc2312ed7c2277309474d6bd916f26e50c67583af3a6f399388b8206f

Observation 3e5b2aec-a4b9-4f0f-a803-cab55c4a8a75 · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.254243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.254243Z digest=sha256:5407e18f7c265d30f094fcc4bc48d6c965aa8c6f87a2237b78c9efb6063c4bea

Observation b34079ab-483f-49e5-9763-b1af369a2c45 · outbound

This paper cites Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.258371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.258371Z digest=sha256:26ffd4a26a67a2ad6a936f673f82e5b2bdf79c6efa0c4a97383eda0eb24caddd

Observation 17c4dcee-5634-4571-a3c4-435e1c433264 · outbound

This paper cites T.; and Pal, C.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment T.; and Pal, C

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:51:53.706679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T20:51:53.263108Z digest=sha256:17b6a176ce9156f7f0f61f0d4242806178f8972a8b0d8b842ae362a0d0e275b4

Observation 1f864896-899f-46a8-849d-b6da04f8975d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.267579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.267579Z digest=sha256:9006ad58dbfb49f2083fef828dc1042bafa563b49af918e97882fd811df1d4d0

Observation 6ba8518f-d633-4edd-ab5b-951a2990b1d9 · outbound

This paper cites N.; Kaiser, .; and Polosukhin, I.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment N.; Kaiser, .; and Polosukhin, I

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.272018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.272018Z digest=sha256:ab967df96a6414a0a4cbab6e8011c6bfc31030248f7664ddfaa25ccb41b48ddf

Observation 23ffa5b9-7eaa-41c1-83b8-fbf772db50e8 · outbound

This paper cites Ethical and social risks of harm from Language Models.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Ethical and social risks of harm from Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.276407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.276407Z digest=sha256:44c5e9b87c535d1cc9eed812165c9c22da5c0cd7ca4dc27a8c6376c42ba95911

Observation 853c1f88-dd45-4e82-9661-9e8fb6cff1a9 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Baichuan 2: Open Large-scale Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.281213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.281213Z digest=sha256:2a3cc0f86b210bce60ef639cb2237871fbf924d6957028874c3bad6c44bf8e87

Observation 2dbe5ef3-bce9-423e-8da1-daa8ebc825c6 · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.285782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.285782Z digest=sha256:a3fa200a513ca3880e7eba0315dd287d3d12c5c9b196d1e8b492b13220511781

Observation 64db0c9c-9e6c-4e0c-b0ab-cb8c096fa715 · outbound

This paper cites Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.290159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.290159Z digest=sha256:8203e7c902ceb9cc975fd16c1051ef8770c960da41f6c820ab7efde3a545a8ef

Observation 59bd2236-d0c8-4a6b-bd09-72a228229ad7 · outbound

This paper cites DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.294752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.294752Z digest=sha256:231ee92501566abb0566f2e3f15615d373d6c857df4a1dc1fd909338cd1e9d96

Observation f8e0967c-79b7-4f8f-bf22-c139f106c22d · outbound

This paper cites an unresolved cited work.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:51:53.682392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T20:51:53.299361Z digest=sha256:bf245215eeb595d0398f559210432d832ba7bfb5a11434e2d7ce5031f81841cd

Pith citing papers

No inbound Pith citation observations are available.