Pith. sign in

Paper Citation Record · LEDGER

Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2405.07637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.07637 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:08:08.503214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:29:02.935467Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 699927c1-cb2f-4ee0-9007-a75377c529d0 · inbound

Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback cites this paper.

Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T00:08:08.503214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:08:08.503214Z digest=sha256:8e028ff5dd4ac766f13990687cc4ffd5d6069ec25c7ea3c390f322e6873a926e

Observation 8998aaef-1560-4dc3-926c-2f66599e2a49 · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:39.669228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:39.669228Z digest=sha256:b99cd5d6ee87c9d47f68a9d4961b3e77feece0ff4853aca6421d78df13412628

Observation 2235d6ee-d481-4f4f-9947-cf2a99a962ec · inbound

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? cites this paper.

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.936716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T22:06:24.412259Z digest=sha256:3ae873ae93808e6ffe319f2d0f10f13c423e06ae73a6227f78cd5ea4c956ab5c