Pith. sign in

Paper Citation Record · LEDGER

Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.17466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.17466 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:53:09.027483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:44:00.901151Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b41119ee-d208-4a11-b610-82002cf7e145 · inbound

Multi-objective Large Language Model Alignment with Hierarchical Experts cites this paper.

Multi-objective Large Language Model Alignment with Hierarchical Experts Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:12.434324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:12.434324Z digest=sha256:38ddbbd712e9976c8b093cbab64af3c45ab9853ee256bfc9dccd6ddbf8cc5a8d

Observation 51ab15c2-aa1a-4d51-9507-df65c499ee88 · inbound

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach cites this paper.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.106357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.106357Z digest=sha256:4eab62aaa4bed156d16a2515ec54b5880e50d1bf6a7d3e6420ad8013c31b3dc7

Observation bf288d27-4fa1-4a5c-8e0b-ee6f1ca3725e · inbound

A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning cites this paper.

A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:29.689165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T04:04:19.891421Z digest=sha256:917caecc839f64c040ffbee0000d154c9dcfc164aa266b591e1d11dfe2c661b9

Observation 7d4ad250-cbd8-4c5f-9b50-c20acec4f67b · inbound

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems cites this paper.

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:18.942894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T16:26:04.152045Z digest=sha256:7fd69287e55dd6ca185c82c8d019eef27bc8568217d9a143b954c0567bfba671

Observation ce8807aa-1619-4c9c-816a-47c858360871 · inbound

Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization cites this paper.

Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:51.671927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T19:32:35.197431Z digest=sha256:4090a9c62ce5ba052636b033dc3ac2388133c69fc4b4446faae598349a51258f

Observation fdf2ba28-614f-43f0-b7f6-e57dda7c0738 · inbound

SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front cites this paper.

SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:44:00.902598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T06:42:15.135148Z digest=sha256:e3da64ca57f67aa7de0b16d9e6e38c8775e7330d43c332e54ce9fad25b58f154

Observation 573a5b80-221b-4ab5-b35d-a7b27625e9ea · inbound

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability cites this paper.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:6b7195212037a20b7475eb166809095a08726a684a933c374950c81911650265

Observation 1c10cf1e-909c-499b-9ac3-035e593912f3 · inbound

Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits cites this paper.

Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:53:09.027483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:53:09.027483Z digest=sha256:80f029797b75c9b4e3d7a28b967e204f4731481654a27e8e6e0a5ed2fa653fa9

Observation 0972ff51-7be1-487c-8a16-2764c24ce67a · inbound

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation cites this paper.

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:48:08.805698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:48:08.805698Z digest=sha256:294719dff212701458110db0fa586fd9385be6acfa40b391352323c2a8ee946f