Pith. sign in

Paper Citation Record · LEDGER

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference

As of 17 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2602.15206.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.15206 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:59:38.022401Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c23f11f1-2fd9-4976-8cd6-1aa4647a3a10 · outbound

This paper cites an unresolved cited work.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Unresolved cited work

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-02T22:59:38.022401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:38.022401Z digest=sha256:c69260202b34a6ee977df56fd2471f9a81fcb7d524ea6b01cb0edd4593948f09

Observation dd1ee73f-9612-41ad-ac88-2d0b75328dbb · outbound

This paper cites Decoupled Weight Decay Regularization.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Decoupled Weight Decay Regularization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.082536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.082536Z digest=sha256:b4c5b7ee15f7ec65c0a05c881cf3cc0ce9cf941fe8fa2f0547e91716e15d3294

Observation f846c9cd-60c5-40d8-9c43-cb51d9fb91f2 · outbound

This paper cites URL https://doi.org/10.1145/3623384.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference URL https://doi.org/10.1145/3623384

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.379875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.379875Z digest=sha256:0943fb095c7c019e9f1a7ae385801210a49f3d4ede8d6b8de960e266770ddcc5

Observation 268e2b95-fc1e-44cd-a3ed-ac98b11670dc · outbound

This paper cites RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-02T23:04:09.648070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-02T22:59:37.478147Z digest=sha256:de59edc53e478b9833043c76c32977dacfe6862c0936ce78aeef2de7df0d7927

Observation ecc515b6-c2dd-428d-85e2-c4b2f8ccd695 · outbound

This paper cites Vivek Myers, Erdem Bıyık, Nima Anari, and Dorsa Sadigh.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Vivek Myers, Erdem Bıyık, Nima Anari, and Dorsa Sadigh

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.604747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.604747Z digest=sha256:2babcc9ea75ee069e3fb9da59d9b36e5853a36e5dd1371eeaab9afde21584b8b

Observation 35cae80d-7fd8-4431-8600-5ccbf5b974b7 · outbound

This paper cites Training language models to follow instructions with human feedback.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Training language models to follow instructions with human feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.730664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.730664Z digest=sha256:17ff6f9b462bbf8acfa186d908aed0d7fd634be0b3c7f05b52b9382affcb9f13

Observation 01800d95-c7cd-43c8-8e6c-b74caafb3306 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.894011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.894011Z digest=sha256:436e2bbb0ee042450b577c829d5eac87d7629018d339c024dfeca2beffa5cd00

Observation b2246660-9ab4-48c7-9e67-7fdeef9e3261 · outbound

This paper cites 13 A Implementation Details A.1 Model Architecture The reward encoder qθ and Q-value estimator Qϕ were implemented as two-layer MLPs with Leaky ReLU activations.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference 13 A Implementation Details A.1 Model Architecture The reward encoder qθ and Q-value estimator Qϕ were implemented as two-layer MLPs with Leaky ReLU activations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.960668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.960668Z digest=sha256:2230349198833e4ad89da5a81bc0281e9f7c6396c111604f7e404e161ab31043

Observation 9c228a7f-f9e2-4010-9580-9aea33351a72 · outbound

This paper cites Shaunak A.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Shaunak A

Reference 1980

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:37.216647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:37.216647Z digest=sha256:dfffe064fd6266478ba70656664db03dd8672472363e93bc90a6fe78c1eaac13

Observation d06f97f8-0610-4967-b1b9-6aab1ab2e25b · outbound

This paper cites ISBN 978-1-58113-838-2.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference ISBN 978-1-58113-838-2

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:35.977656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:35.977656Z digest=sha256:017fc04490e8ee8396dd608ae1349b377eaaedd1ba59fd6ff28e9026f51c3bec

Observation 8244124f-9067-429d-b58d-6bfdd43df72a · outbound

This paper cites ISBN 978-1-60558-658-8.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference ISBN 978-1-60558-658-8

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.914817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.914817Z digest=sha256:0d68df092f0a0b572e66b39f5602733545d34bc9f80be52afdf8a73dd9060f4b

Observation b57aefb9-b22b-45d8-a9d1-09c2763386ab · outbound

This paper cites Auto-Encoding Variational Bayes.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Auto-Encoding Variational Bayes

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.704749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.704749Z digest=sha256:f96bf364d1c3196b45b82d6030a9aa6a9436558c78e04ee23dbe98519f789326

Observation e40ae02c-c435-419f-a524-3fa31e9516e3 · outbound

This paper cites Unsolved Problems in ML Safety.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Unsolved Problems in ML Safety

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.474748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.474748Z digest=sha256:705233a930c7c8b67ea5ef398cc4c56ea743813c4da1330558539d282d96d493

Observation 61f29625-6a8f-4974-87b8-5875edba4c1e · outbound

This paper cites Reward-rational (implicit) choice: A unifying formalism for reward learning.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Reward-rational (implicit) choice: A unifying formalism for reward learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.591876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.591876Z digest=sha256:d2d872109f090e3573288b372ec93dd38e790b031051be9a3d03a7022262adb9

Observation 6b1e099b-a8b8-4ce1-870c-a28b782a31ad · outbound

This paper cites Learning Multimodal Rewards from Rankings.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Learning Multimodal Rewards from Rankings

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-02T23:04:09.540088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-02T22:59:37.663150Z digest=sha256:05fbc11f11f0364c77ee72f828f85fa73c9aa4249f9881b54491d1c50c7f81bf

Observation c9328504-8c06-47f4-9c62-3db756a303e8 · outbound

This paper cites doi: 10.1177/02783649211041652.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference doi: 10.1177/02783649211041652

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.202256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.202256Z digest=sha256:4746fc3843d2f880d49a49b75680110c295224b91b8ca651e9d7778684a77974

Observation 5eb5badb-d0db-496c-8459-160db2383dc8 · outbound

This paper cites doi: 10.1609/aaai.v37i5.25740.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference doi: 10.1609/aaai.v37i5.25740

Reference 2023

Resolution
verified exact
doi, observed 2026-08-02T23:04:09.842539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-02T22:59:36.363241Z digest=sha256:c55f41fddff75985641f7d8f586157ca15bfde76fdc596111338cce14dc1c1a0

Observation f93f15a9-5612-410e-a186-8e7972475d65 · outbound

This paper cites Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.074912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.074912Z digest=sha256:1e13b7dc1f29614d3cdc9352d6c1ae2369d28086ffe76701ed4c65b7c897f871

Observation 0b6b0a61-a1ed-4f96-ac42-0b337c3049d6 · outbound

This paper cites Scalable Bayesian inverse reinforcement learning.

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Scalable Bayesian inverse reinforcement learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:59:36.291389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:59:36.291389Z digest=sha256:16d5e225ef40bf57fc1830c03bed43d0b4134236ad1ca00a91e267ebb1e6b75a

Pith citing papers

No inbound Pith citation observations are available.