Pith. sign in

Paper Citation Record · LEDGER

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.02016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02016 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:59:48.980316Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4847d24-0244-46fa-b7b9-5f18e4885b1c · outbound

This paper cites Game Theory: Analysis of Conflict.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Game Theory: Analysis of Conflict

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.175594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.547329Z digest=sha256:ffe357bfcb0490dc607203b77afad6604990df8cb44c79c97272a89c13776451

Observation ca1e0a86-466d-4555-a3e9-4319b3998854 · outbound

This paper cites Three-player games are hard.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Three-player games are hard

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.130353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.577695Z digest=sha256:47a8e7a5cf5b22cb5e506e933264b4fb975c4eeb4c32810a9e9fd595b06d3595

Observation 09042e0d-ea45-407c-8610-300d7272c35f · outbound

This paper cites MIT press, 1991.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning MIT press, 1991

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.585577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.585577Z digest=sha256:84f40aeec00a27bcf6c460dec1691a25cd83a31026eb029110847924f491e825

Observation 848b09e5-6c31-4789-be38-56417bcb46ce · outbound

This paper cites Recent developments of game theory and reinforcement learning approaches: A systematic review.IEEE Access, 12:9999–10011, 2024.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Recent developments of game theory and reinforcement learning approaches: A systematic review.IEEE Access, 12:9999–10011, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.078787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.597358Z digest=sha256:f1fb14a2561913e8c514623bc0b411304c83886cb1d4777f5396c9c32c394cd5

Observation 4a0d0cf7-4197-439d-9d26-979a1da01723 · outbound

This paper cites Schapire.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Schapire

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.618480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.618480Z digest=sha256:4cb804ddc2e725257b98a3daffd7aa175e698b355a4269fa5fa3a93120267aa5

Observation bf46e130-0353-4eba-b79b-48a650236886 · outbound

This paper cites Explore no more: Improved high-probability regret bounds for non-stochastic bandits.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Explore no more: Improved high-probability regret bounds for non-stochastic bandits

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:50.017368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.654966Z digest=sha256:1331e689fcccacbe83a84a9c487bf14c92b3385a4bfc546146a1034404fa5f59

Observation 49094ab9-30a2-4735-be4b-54b906123b20 · outbound

This paper cites Eqilibrium approximation quality of current no-limit poker bots.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Eqilibrium approximation quality of current no-limit poker bots

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.958199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.662108Z digest=sha256:9ff76672384267bbadb04e34c74b1b4d734d423544cba8ac8e896c19a777cc32

Observation 42ee4dc4-3d5e-4755-b10c-d4fd256a72c4 · outbound

This paper cites Cambridge University Press, 2007.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Cambridge University Press, 2007

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.855954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.697187Z digest=sha256:8fba94510bea0d3d08de63b50f910824df7fdc8f3f107d87359c6f877ae7789b

Observation 31f8c99c-8389-4bcc-ae1f-26d5f3691baf · outbound

This paper cites The complexity of approximate (coarse) correlated equilibrium for incomplete information games.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning The complexity of approximate (coarse) correlated equilibrium for incomplete information games

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.796293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.706982Z digest=sha256:1e6ac47be729c1e40787861032d0473e54dec8727556deae92de32d1c290d4d4

Observation fa5fe846-6aab-4c44-9f63-e93a9c76b7ce · outbound

This paper cites Coarse correlation in extensive- form games.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Coarse correlation in extensive- form games

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.718662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.712810Z digest=sha256:3d7d15bdd3ff80fb492f1f9305cbe43e51d15524cc9469d0228c930be637d0c4

Observation 95ae39fb-bf7d-4462-980e-3b830e5e1525 · outbound

This paper cites Cambridge University Press, 2006.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Cambridge University Press, 2006

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.671647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.730237Z digest=sha256:ce9c7c3b74f9b1be8eebdc1627f8dc585211c8cac55e0bd4ed0afaae13fff7da

Observation ff2195f1-d24d-4549-bbb9-9d81a997c093 · outbound

This paper cites Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.630806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.737342Z digest=sha256:c2335e0db92c59d359ac19ccbbd5095a87f3744d6233ce6a9127a567209e9436

Observation 2743a2ac-1fd2-4aec-9347-3624f55bf5f2 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.745863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.745863Z digest=sha256:9ee15172bfbe2d92fcd7295b96aea3c8b4dbb583b6e5a0d479caeaaaafe81d6c

Observation 0527151a-094a-4d79-891c-7b34810a5773 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.762187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.762187Z digest=sha256:9b2b2cd7fbc9e26054436fe692e272b628e8862d5a295a99b4a198d2721fd92a

Observation dece90e3-ea94-45d2-b68d-ac18fc6fa985 · outbound

This paper cites Trust Region Policy Optimization.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Trust Region Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.794755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.794755Z digest=sha256:ceb49adc62466f0fda673e54ef70fc8961334b0bfb2999bc4265352ef1c1af2a

Observation dc7ccef7-f6e1-4dce-a76b-e1ed706ce376 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.824823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.824823Z digest=sha256:50863b4c7d2f407336b067f9ff10ffd554cdee01dcfd44b717ec2166fbd14ac1

Observation 8685981c-d5ec-4895-bc7f-d122561f7578 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.Advances in neural information processing systems, 30, 2017.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Multi-agent actor-critic for mixed cooperative-competitive environments.Advances in neural information processing systems, 30, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.859418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.859418Z digest=sha256:ae257ad4f1f60c53c4e7be1bddf73431486bf5b47142616b469ee0ff82ca6972

Observation 708c3d88-6fc1-4112-85fd-80f9dd7296bc · outbound

This paper cites Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER, pages 105–130.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER, pages 105–130

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.519980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.865394Z digest=sha256:7908f7ed2c8e494412a73ef58dd4f4a3772ae55222d0305c6ccf863a21de623d

Observation 8bb01468-c5d0-4725-8992-a4b353f46fcc · outbound

This paper cites CybORG: A Gym for the Development of Autonomous Cyber Agents.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning CybORG: A Gym for the Development of Autonomous Cyber Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.889248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.889248Z digest=sha256:a56ad008759935c35b8bc41f5c42bab88d767df860e4009aa382196a8efc8608

Observation 2e132313-3a13-4f00-8e49-ecfac322a7d9 · outbound

This paper cites On Autonomous Agents in a Cyber Defence Environment.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning On Autonomous Agents in a Cyber Defence Environment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.907642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.907642Z digest=sha256:1c1cbef0852a72ac074ddd2d4e8b655608f298436a25e127d8c425640361cf54

Observation 1c6cfd61-7065-4898-8a75-cbb9ecb15e04 · outbound

This paper cites A Bradford Book, 2018.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning A Bradford Book, 2018

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.468476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.944754Z digest=sha256:d8aea431f61d6067c7eaadcf69c937a866be7f93d1d253bd45432bf255a56cdf

Observation 2d58370f-bfcb-4b07-b021-a76659ab6e08 · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.Journal of Machine Learning Research, 3(Nov):397–422, 2002.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Using confidence bounds for exploitation-exploration trade-offs.Journal of Machine Learning Research, 3(Nov):397–422, 2002

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:59:48.963482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:59:48.963482Z digest=sha256:10f2b5f43e547380cdec03bb38295925aad52643dd301be9792eab2bc11cc134

Observation 14cb441b-4c9d-414a-a3b9-3180250947f3 · outbound

This paper cites Flaxman, Adam Tauman Kalai, and H.

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Flaxman, Adam Tauman Kalai, and H

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:59:49.410289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:59:48.980316Z digest=sha256:a97516a41fbe6439c55e9aba600e6d2b8372a75712142f15d847d4667f79564b

Pith citing papers

No inbound Pith citation observations are available.