Pith. sign in

REVIEW 1 cited by

Scalable spectral representations for multi-agent reinforcement learning in network MDPs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.17221 v2 pith:TCXFQ2AS submitted 2024-10-22 cs.MA cs.LGcs.SYeess.SYmath.OC

classification cs.MAcs.LGcs.SYeess.SYmath.OC
keywords networklocalmdpsscalablerepresentationsspectralapproachexponential
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Network Markov Decision Processes (MDPs), a popular model for multi-agent control, pose a significant challenge to efficient learning due to the exponential growth of the global state-action space with the number of agents. In this work, utilizing the exponential decay property of network dynamics, we first derive scalable spectral local representations for network MDPs, which induces a network linear subspace for the local $Q$-function of each agent. Building on these local spectral representations, we design a scalable algorithmic framework for continuous state-action network MDPs, and provide end-to-end guarantees for the convergence of our algorithm. Empirically, we validate the effectiveness of our scalable representation-based approach on two benchmark problems, and demonstrate the advantages of our approach over generic function approximation approaches to representing the local $Q$-functions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

    cs.MA 2026-07 conditional novelty 6.0 of 10

    CDCPG claims epsilon^-2 shared-oracle complexity to a structural stationarity floor for continuous networked MARL under exponential decay and an assumed TD-excitation condition.

Pith tools