{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:3ZI67ZN37PGRQXARTVBKY643IZ","short_pith_number":"pith:3ZI67ZN3","schema_version":"1.0","canonical_sha256":"de51efe5bbfbcd185c119d42ac7b9b466550ee416328eda39c6e5a6e0eadd28a","source":{"kind":"arxiv","id":"2110.14555","version":1},"attestation_state":"computed","paper":{"title":"V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.GT","cs.MA","stat.ML"],"primary_cat":"cs.LG","authors_text":"Chi Jin, Qinghua Liu, Tiancheng Yu, Yuanhao Wang","submitted_at":"2021-10-27T16:25:55Z","abstract_excerpt":"A major challenge of multiagent reinforcement learning (MARL) is the curse of multiagents, where the size of the joint action space scales exponentially with the number of agents. This remains to be a bottleneck for designing efficient MARL algorithms even in a basic scenario with finitely many states and actions. This paper resolves this challenge for the model of episodic Markov games. We design a new class of fully decentralized algorithms -- V-learning, which provably learns Nash equilibria (in the two-player zero-sum setting), correlated equilibria and coarse correlated equilibria (in the"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2110.14555","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2021-10-27T16:25:55Z","cross_cats_sorted":["cs.AI","cs.GT","cs.MA","stat.ML"],"title_canon_sha256":"e368946e13b253a6743995839176f7271b20b914c80f369dde790f8d4c0df6ed","abstract_canon_sha256":"94b50860c54d8e15e75de442125b768d71376dcac58123d486b71f5034d66a32"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T03:26:29.135917Z","signature_b64":"f6p0G2yLjJKYdVRNa8T5hQL3ZTgKPWBaOrPyH7kdZaeqEpU2eQxcEpA3jw1SUHAo5mpgf6uj8nToCPlo15CVCA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"de51efe5bbfbcd185c119d42ac7b9b466550ee416328eda39c6e5a6e0eadd28a","last_reissued_at":"2026-07-05T03:26:29.135399Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T03:26:29.135399Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.GT","cs.MA","stat.ML"],"primary_cat":"cs.LG","authors_text":"Chi Jin, Qinghua Liu, Tiancheng Yu, Yuanhao Wang","submitted_at":"2021-10-27T16:25:55Z","abstract_excerpt":"A major challenge of multiagent reinforcement learning (MARL) is the curse of multiagents, where the size of the joint action space scales exponentially with the number of agents. This remains to be a bottleneck for designing efficient MARL algorithms even in a basic scenario with finitely many states and actions. This paper resolves this challenge for the model of episodic Markov games. We design a new class of fully decentralized algorithms -- V-learning, which provably learns Nash equilibria (in the two-player zero-sum setting), correlated equilibria and coarse correlated equilibria (in the"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2110.14555","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2110.14555/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2110.14555","created_at":"2026-07-05T03:26:29.135461+00:00"},{"alias_kind":"arxiv_version","alias_value":"2110.14555v1","created_at":"2026-07-05T03:26:29.135461+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2110.14555","created_at":"2026-07-05T03:26:29.135461+00:00"},{"alias_kind":"pith_short_12","alias_value":"3ZI67ZN37PGR","created_at":"2026-07-05T03:26:29.135461+00:00"},{"alias_kind":"pith_short_16","alias_value":"3ZI67ZN37PGRQXAR","created_at":"2026-07-05T03:26:29.135461+00:00"},{"alias_kind":"pith_short_8","alias_value":"3ZI67ZN3","created_at":"2026-07-05T03:26:29.135461+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.06486","citing_title":"Regret Minimization with Adaptive Opponents in Repeated Games","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2504.03353","citing_title":"Decentralized Collective World Model for Emergent Communication and Coordination","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17189","citing_title":"Sample-efficient inductive matrix completion with noise and inexact side-information","ref_index":158,"is_internal_anchor":false},{"citing_arxiv_id":"2603.28281","citing_title":"Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2605.03125","citing_title":"Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation","ref_index":6,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ","json":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ.json","graph_json":"https://pith.science/api/pith-number/3ZI67ZN37PGRQXARTVBKY643IZ/graph.json","events_json":"https://pith.science/api/pith-number/3ZI67ZN37PGRQXARTVBKY643IZ/events.json","paper":"https://pith.science/paper/3ZI67ZN3"},"agent_actions":{"view_html":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ","download_json":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ.json","view_paper":"https://pith.science/paper/3ZI67ZN3","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2110.14555&json=true","fetch_graph":"https://pith.science/api/pith-number/3ZI67ZN37PGRQXARTVBKY643IZ/graph.json","fetch_events":"https://pith.science/api/pith-number/3ZI67ZN37PGRQXARTVBKY643IZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ/action/storage_attestation","attest_author":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ/action/author_attestation","sign_citation":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ/action/citation_signature","submit_replication":"https://pith.science/pith/3ZI67ZN37PGRQXARTVBKY643IZ/action/replication_record"}},"created_at":"2026-07-05T03:26:29.135461+00:00","updated_at":"2026-07-05T03:26:29.135461+00:00"}