Pith. sign in

REVIEW 1 cited by

Building a 3-Player Mahjong AI using Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.12847 v3 pith:KBCJCJDG submitted 2022-02-25 cs.AI cs.LG

classification cs.AIcs.LG
keywords learningsanmamahjongreinforcementgamemeowjongplayerchallenging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mahjong is a popular multi-player imperfect-information game developed in China in the late 19th-century, with some very challenging features for AI research. Sanma, being a 3-player variant of the Japanese Riichi Mahjong, possesses unique characteristics including fewer tiles and, consequently, a more aggressive playing style. It is thus challenging and of great research interest in its own right, but has not yet been explored. In this paper, we present Meowjong, an AI for Sanma using deep reinforcement learning. We define an informative and compact 2-dimensional data structure for encoding the observable information in a Sanma game. We pre-train 5 convolutional neural networks (CNNs) for Sanma's 5 actions -- discard, Pon, Kan, Kita and Riichi, and enhance the major action's model, namely the discard model, via self-play reinforcement learning using the Monte Carlo policy gradient method. Meowjong's models achieve test accuracies comparable with AIs for 4-player Mahjong through supervised learning, and gain a significant further enhancement from reinforcement learning. Being the first ever AI in Sanma, we claim that Meowjong stands as a state-of-the-art in this game.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evolutionary Optimization of Deep Learning Agents for Sparrow Mahjong

    cs.NE 2025-08 unverdicted novelty 4.0 of 10

    A CMA-ES-optimized LSTM agent for Sparrow Mahjong is claimed to beat random and rule-based agents and match a PPO baseline, but the provided manuscript contains no verifiable experiments.

Pith tools