Pith. sign in

Paper Citation Record · LEDGER

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.02016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02016 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:59:48.980316Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4847d24-0244-46fa-b7b9-5f18e4885b1c · outbound

This paper cites Game Theory: Analysis of Conflict.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Game Theory: Analysis of Conflict

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.175594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.547329Z digest=sha256:4c49d8050cfc9cc9d36cdf4497f88c1f8aedce462f225e604c5616a75cace553

Observation ca1e0a86-466d-4555-a3e9-4319b3998854 · outbound

This paper cites Three-player games are hard.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Three-player games are hard

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.130353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.577695Z digest=sha256:3b1e6e541a7e1485aed5d1372abf0163b2b7bd406086d52735dba087d6e70977

Observation 09042e0d-ea45-407c-8610-300d7272c35f · outbound

This paper cites MIT press, 1991.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning MIT press, 1991

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.585577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.585577Z digest=sha256:84f40aeec00a27bcf6c460dec1691a25cd83a31026eb029110847924f491e825

Observation 848b09e5-6c31-4789-be38-56417bcb46ce · outbound

This paper cites Recent developments of game theory and reinforcement learning approaches: A systematic review.IEEE Access, 12:9999–10011, 2024.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Recent developments of game theory and reinforcement learning approaches: A systematic review.IEEE Access, 12:9999–10011, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.078787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.597358Z digest=sha256:2b1d5189ea6e0ad1b00558582f4b21810ff15ad7b26e6c08fd33b91069ed0b78

Observation 4a0d0cf7-4197-439d-9d26-979a1da01723 · outbound

This paper cites Schapire.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Schapire

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.618480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.618480Z digest=sha256:4cb804ddc2e725257b98a3daffd7aa175e698b355a4269fa5fa3a93120267aa5

Observation bf46e130-0353-4eba-b79b-48a650236886 · outbound

This paper cites Explore no more: Improved high-probability regret bounds for non-stochastic bandits.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Explore no more: Improved high-probability regret bounds for non-stochastic bandits

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.017368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.654966Z digest=sha256:6bd4451d972426e2b754af1b4c85def321f02bd641a16703a74341d8783fd3ed

Observation 49094ab9-30a2-4735-be4b-54b906123b20 · outbound

This paper cites Eqilibrium approximation quality of current no-limit poker bots.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Eqilibrium approximation quality of current no-limit poker bots

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.958199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.662108Z digest=sha256:27f1b61bd0d34376302d23272794e7714a4aab0624cf53834d44bf8bef01c186

Observation 42ee4dc4-3d5e-4755-b10c-d4fd256a72c4 · outbound

This paper cites Cambridge University Press, 2007.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Cambridge University Press, 2007

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.855954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.697187Z digest=sha256:6179f6af38803dc2c0124ec4ef79057211097cf551e1253246261ca1b56d84a4

Observation 31f8c99c-8389-4bcc-ae1f-26d5f3691baf · outbound

This paper cites The complexity of approximate (coarse) correlated equilibrium for incomplete information games.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning The complexity of approximate (coarse) correlated equilibrium for incomplete information games

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.796293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.706982Z digest=sha256:a4848601f1480afecb0b825add8b15a410d965a0c680147cc2e725972991c9af

Observation fa5fe846-6aab-4c44-9f63-e93a9c76b7ce · outbound

This paper cites Coarse correlation in extensive- form games.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Coarse correlation in extensive- form games

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.718662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.712810Z digest=sha256:e3439a6ea782d715a04e22f1d545c72e26cd7c6e2b1d5b74741d830f236d2bd0

Observation 95ae39fb-bf7d-4462-980e-3b830e5e1525 · outbound

This paper cites Cambridge University Press, 2006.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Cambridge University Press, 2006

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.671647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.730237Z digest=sha256:dacf42fbde4f681b2f6fb110cf9bcc3a40b41082c6930c7498579c0ef5d37695

Observation ff2195f1-d24d-4549-bbb9-9d81a997c093 · outbound

This paper cites Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.630806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.737342Z digest=sha256:61296dcd79fb60d71700ee492466407e60627c4e4a7e0459db11118524dc29b5

Observation 2743a2ac-1fd2-4aec-9347-3624f55bf5f2 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.745863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.745863Z digest=sha256:9ee15172bfbe2d92fcd7295b96aea3c8b4dbb583b6e5a0d479caeaaaafe81d6c

Observation 0527151a-094a-4d79-891c-7b34810a5773 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.762187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.762187Z digest=sha256:9b2b2cd7fbc9e26054436fe692e272b628e8862d5a295a99b4a198d2721fd92a

Observation dece90e3-ea94-45d2-b68d-ac18fc6fa985 · outbound

This paper cites Trust Region Policy Optimization.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Trust Region Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.794755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.794755Z digest=sha256:ceb49adc62466f0fda673e54ef70fc8961334b0bfb2999bc4265352ef1c1af2a

Observation dc7ccef7-f6e1-4dce-a76b-e1ed706ce376 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.824823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.824823Z digest=sha256:50863b4c7d2f407336b067f9ff10ffd554cdee01dcfd44b717ec2166fbd14ac1

Observation 8685981c-d5ec-4895-bc7f-d122561f7578 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.Advances in neural information processing systems, 30, 2017.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Multi-agent actor-critic for mixed cooperative-competitive environments.Advances in neural information processing systems, 30, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.859418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.859418Z digest=sha256:ae257ad4f1f60c53c4e7be1bddf73431486bf5b47142616b469ee0ff82ca6972

Observation 708c3d88-6fc1-4112-85fd-80f9dd7296bc · outbound

This paper cites Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER, pages 105–130.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER, pages 105–130

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.519980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.865394Z digest=sha256:7fbb255c22926fd3c009abf0781ff84b4f0fdc23029e5bfaa4753394c37553c7

Observation 8bb01468-c5d0-4725-8992-a4b353f46fcc · outbound

This paper cites CybORG: A Gym for the Development of Autonomous Cyber Agents.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning CybORG: A Gym for the Development of Autonomous Cyber Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.889248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.889248Z digest=sha256:a56ad008759935c35b8bc41f5c42bab88d767df860e4009aa382196a8efc8608

Observation 2e132313-3a13-4f00-8e49-ecfac322a7d9 · outbound

This paper cites On Autonomous Agents in a Cyber Defence Environment.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning On Autonomous Agents in a Cyber Defence Environment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.907642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.907642Z digest=sha256:1c1cbef0852a72ac074ddd2d4e8b655608f298436a25e127d8c425640361cf54

Observation 1c6cfd61-7065-4898-8a75-cbb9ecb15e04 · outbound

This paper cites A Bradford Book, 2018.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning A Bradford Book, 2018

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.468476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.944754Z digest=sha256:a5c683bf878249d50704a3e4bb88b9c7383755220a71e49f9de350d93bfd4ce2

Observation 2d58370f-bfcb-4b07-b021-a76659ab6e08 · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.Journal of Machine Learning Research, 3(Nov):397–422, 2002.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Using confidence bounds for exploitation-exploration trade-offs.Journal of Machine Learning Research, 3(Nov):397–422, 2002

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.963482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.963482Z digest=sha256:10f2b5f43e547380cdec03bb38295925aad52643dd301be9792eab2bc11cc134

Observation 14cb441b-4c9d-414a-a3b9-3180250947f3 · outbound

This paper cites Flaxman, Adam Tauman Kalai, and H.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Flaxman, Adam Tauman Kalai, and H

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.410289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:59:48.980316Z digest=sha256:9405cfbe0b7e0c10d6729720daae999cf266bb24c37f17a99f52c96c5f013ea2

Pith citing papers

No inbound Pith citation observations are available.