Pith. sign in

REVIEW 6 minor 55 references

RideGym is the first open-source Gym-style interface that decouples ride-sharing order dispatch from algorithms so methods can be fairly compared at city scale.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 13:44 UTC pith:7ADJZSOM

load-bearing objection Solid open Gym for city-scale ride-pooling plus a real finding that exploration noise reorders MARL rankings; worth engaging.

arxiv 2607.10173 v1 pith:7ADJZSOM submitted 2026-07-11 cs.MA

RideGym: A Standardized Interface for Real-World Large-Scale Ride-Sharing System

classification cs.MA
keywords Ride SharingSimulation GymStandardized InterfaceMulti-Agent Reinforcement LearningOrder DispatchTrip BundlingReal Road Networks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Progress on multi-agent order dispatch for ride-sharing has been slowed less by a lack of algorithms than by the lack of a shared, learning-friendly testbed: existing simulators are either operations-oriented and tightly coupled to one control pipeline, or algorithm-specific and often abstract the city into a grid. Most groups therefore rebuild custom environments, which blocks fair comparison and wastes engineering effort. RideGym supplies a standardized reset/step interface on real road networks, with heterogeneous fleets, multi-passenger orders, automatic shortest-path routing, and modular demand, planner, and reward components. One-hour Manhattan-scale episodes with thousands of vehicles and tens of thousands of orders finish in under a minute for every baseline tested. In the course of reproducing classical matchers and several MARL methods, the authors also show that the type of exploration noise used during training can change both absolute performance and which algorithm ranks best—an experimental factor prior work largely ignored.

Core claim

RideGym is the first open-source, standardized Gym-style interface tailored to multi-agent reinforcement learning order dispatch in real-world ride-sharing systems. By fully decoupling the environment from the dispatch algorithm, it enables model-based and learning-based methods to be developed and compared under identical, fully specified conditions on real road networks, with flexible vehicle and order configurations, city-scale efficiency (one-hour simulations completed within one minute), and the finding that exploration-noise choice can reorder MARL rankings.

What carries the argument

A Gym-style multi-agent reset/step interface that materializes a cooperative multi-agent MDP for vehicles, backed by a precomputed all-pairs shortest-path road network and a precedence-aware route planner that resequences pickups and drop-offs after each assignment while enforcing exclusive-order and capacity constraints.

Load-bearing premise

Vehicle travel is treated as deterministic free-flow motion along shortest paths with no congestion, accidents, or network dynamics, so every wait time, detour, and method ranking rests on that idealized mobility model.

What would settle it

If independent re-implementations of the same published MARL and matching baselines inside RideGym, under the paper’s stated reward, noise, and fleet settings, produce materially different relative rankings or service rates, the reproducibility claim fails; if load-dependent travel times from real congestion data reverse those rankings, the efficiency and ordering claims do not transfer to operations.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Model-based and MARL dispatch policies can be swapped behind one interface without rebuilding simulators for each study.
  • Head-to-head rankings of order-dispatch algorithms become possible under identical demand, fleet, network, and reward conditions.
  • City-scale one-hour episodes finish in under a minute, making long training runs and hyper-parameter search practical.
  • Exploration noise must be treated as a controlled experimental factor, because INF versus STD noise changes both service rates and relative method orderings.
  • Modular demand, road-network, planner, and reward APIs lower the cost of studying pooling, multi-passenger orders, and customized objectives.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same interface pattern could later absorb congestion models or joint relocation and pricing without forcing every group to reimplement the core loop.
  • Once exploration noise is standardized, published MARL rankings in ride-sharing may need re-evaluation under a common protocol.
  • Because wait and detour metrics rest on free-flow paths, any method that wins inside RideGym still needs validation under capacity-constrained traffic before operational use.
  • Precomputed distance matrices make the environment a natural substrate for testing whether centralized single-agent formulations can replace multi-agent ones at city scale.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 6 minor

Summary. The paper introduces RideGym, an open-source Gym-style multi-agent environment for order assignment and trip bundling in large-scale ride-sharing. It formulates the problem as a fully cooperative MAMDP (Section 2, Eqs. 1–11), implements modular components (precomputed OSM all-pairs shortest paths, precedence-aware nearest-feasible-stop route planner, configurable rewards and fleets), and exposes a standardized reset/step API that decouples the simulator from any particular dispatch policy. Validation consists of reproducing four classical model-based matchers and five MARL baselines on public Manhattan TLC data (1 000 vehicles, off-peak and on-peak hours), reporting service/completion rates, wait/detour times and wall-clock simulation times (Table 1). A secondary empirical finding is that the choice of exploration noise (INF versus the authors’ STD Gaussian scaling) materially changes both absolute performance and the relative ranking of the MARL methods.

Significance. If the claims hold, RideGym supplies the missing standardized, algorithm-agnostic testbed that the ride-sharing MARL community has lacked. The combination of real-road-network fidelity, heterogeneous fleets, multi-passenger orders, sub-minute city-scale episodes, and an open Gym-like interface is a genuine systems contribution that can reduce redundant engineering and improve reproducibility. The noise-ablation result is a useful cautionary finding for future benchmarking practice. Strengths that should be credited include the released code and pip package, the modular interface design, the explicit comparison table against prior simulators (Table 2), and the transparent reporting of means ± std over three seeds.

minor comments (6)
  1. Section 2, State Transition Function: the deterministic free-flow assumption is stated clearly, yet a short forward-looking sentence on planned SUMO coupling (already mentioned only in Appendix D) would help readers gauge transferability of the reported rankings.
  2. Table 1 caption and §4.3: clarify that “Simulation Time” is wall-clock per episode on the stated hardware; a one-line note that training-time overhead (especially for BMG-Q / MF-DDQN) is dominated by the policy network rather than the environment would prevent misreading the efficiency claim.
  3. Eq. (9) and §4.2: the four β weights are fixed without sensitivity analysis; a brief remark that the ranking conclusions are robust under modest re-weighting (or a pointer to the repository) would strengthen the reward-design discussion.
  4. Appendix C.1: the neighborhood size of 30 for BMG-Q and MF-DDQN is taken from the original papers; stating whether any retuning was performed under RideGym would improve reproducibility of the ranking result.
  5. Figure 2 code snippets: a few identifiers contain non-ASCII or oddly spaced characters (e.g., “d riv er _c ap aci ti es”); cleaning these would improve copy-paste usability.
  6. References: several arXiv preprints and conference papers appear with incomplete venue or year information; a final pass for bibliographic consistency is warranted.

Circularity Check

0 steps flagged

No significant circularity: RideGym is a systems/benchmark contribution whose claims rest on an open interface and empirical runs, not on a derivation that folds its inputs into its outputs.

full rationale

The paper's central claims are (i) the existence of an open, algorithm-agnostic Gym-style environment for MARL order dispatch on real maps and (ii) an empirical finding that exploration-noise type (INF vs STD) can reorder MARL rankings under that environment. Neither claim is a first-principles derivation that reduces to its own premises by construction. The MAMDP formulation (Section 2, Eqs. 1–11) defines state, constrained joint action, deterministic free-flow transitions, and a configurable reward; these are modeling choices, not predictions obtained by fitting and then re-labeling the same quantities. The reward betas and noise schedules are free hyperparameters, not parameters fitted to the reported metrics and then called predictions. Baselines that share authors (Assignment-Net and related food-delivery work) appear only as competitors evaluated under the shared interface; they are not invoked as uniqueness theorems or load-bearing premises that force the main result. The efficiency claim (one-hour city-scale episodes under one minute) and the noise-ranking result are measured outputs of the simulator, not tautologies of its definition. The idealized mobility model (no congestion) is an explicit limitation, not a circular step. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 2 invented entities

As a systems and benchmarking paper, load-bearing content is mostly modeling choices and hand-set hyperparameters rather than fitted physical constants or new particles. Free parameters are the reward weights, learning setup, and exploration schedules that shape Table 1. Axioms are standard MARL/MDP structure plus domain simplifications (deterministic routing, no congestion). No new physical entities are postulated; RideGym is a software artifact and STD noise is a simple Gaussian exploration rule.

free parameters (7)
  • reward weights β1, β2, β3, β4 = 1.0, 0.01, 0.04, 0.08
    Hand-set to 1.0, 0.01, 0.04, 0.08 in §4.2; they define the scalar reward that all MARL methods optimize and thus affect absolute and relative performance.
  • discount factor γ = 0.99
    Set to 0.99 for all MARL methods; standard but still a free choice that shapes long-horizon credit assignment.
  • Adam learning rate = 5e-4
    Fixed at 5e-4 for all MARL methods without reported sensitivity analysis.
  • exploration rate ε schedule and STD noise scale = ε: 1→0; scale = εσ
    ε decays from 1 to 0; STD noise uses ζ ~ N(0, (εσ)^2). The paper’s ranking claim depends on this design versus INF noise.
  • neighborhood size for BMG-Q / MF-DDQN = 30
    Fixed at 30 following original papers; authors note sensitivity in Appendix D but do not retune under both noise types.
  • order timeout θ and decision interval Δt = θ=3 min, Δt=1 min, H=60 min
    θ=3 min, Δt=1 min, horizon H=60 min; define when orders cancel and how often matching runs.
  • fleet size, capacity, speed = 1000 vehicles, cap 4, 35 km/h
    n=1000, capacity 4, speed 35 km/h; scenario parameters that set difficulty of off-peak vs on-peak tests.
axioms (6)
  • domain assumption Order dispatch is modeled as a fully cooperative multi-agent MDP with global state fully observed by every agent (O ≡ S).
    Section 2 MAMDP formulation; enables centralized training and joint bipartite matching on Q-values.
  • domain assumption Vehicle motion is deterministic along planned shortest paths; the only stochasticity is new order arrivals (and impatient cancellations).
    Section 2 State Transition; explicitly excludes congestion and accidents.
  • domain assumption Feasibility constraints (7a)–(7d): each order to at most one vehicle, capacity limits, and exclusive no-order action.
    Section 2 Action; environment enforces these after policy proposals.
  • ad hoc to paper Default route planner uses nearest-feasible-stop heuristic under pickup-before-drop-off precedence rather than exact constrained TSP.
    Section 3.1; justified by hot-path cost and bounded stop set, but affects ExtraTime and ExpectedTime in the reward.
  • domain assumption Road distances come from precomputed all-pairs Dijkstra on the largest strongly connected OSM component, with continuous points snapped to nearest nodes.
    Section 3.1 Precomputed Road Network; standard GIS practice but freezes topology and free-flow distances.
  • ad hoc to paper TLC zone-level OD pairs can be turned into point OD by placing requests at zone centers plus uniform 0.5 km noise.
    Section 4.2; approximation needed because public data lack exact coordinates.
invented entities (2)
  • RideGym environment (modular Gym-like ride-pooling simulator) independent evidence
    purpose: Provide a standardized, algorithm-agnostic reset/step interface and city-scale simulation stack for order dispatch research.
    Software artifact rather than a physical postulate; independent evidence is the public code and reproducible runs, not an external natural measurement.
  • STD exploration noise (Gaussian perturbation scaled by Q-value std and ε) no independent evidence
    purpose: Alternative to INF noise for bipartite-matching-based MARL exploration; used to show ranking sensitivity.
    Simple algorithmic choice introduced in §4.3; not independently validated outside this benchmark suite.

pith-pipeline@v1.1.0-grok45 · 25574 in / 4040 out tokens · 41468 ms · 2026-07-14T13:44:24.146363+00:00 · methodology

0 comments
read the original abstract

Ride-sharing has become an essential component of modern urban transportation and has attracted significant attention across computer science, transportation, and management science. While the field spans a broad range of problems, such as driver relocation, dynamic pricing, and vehicle charging or fueling dispatch, the core challenge remains order assignment and trip bundling, which directly affect urban traffic efficiency and carbon emissions. Despite its importance, existing simulation platforms are typically tailored to specific operational studies or tightly coupled to a particular dispatch algorithm, and rarely expose a standardized, learning-friendly interface. As a result, most researchers still build customized environments from scratch, raising serious concerns about reproducibility and fair comparison, and incurring substantial redundant effort. To address this gap, we present RideGym, the first open-source, standardized Gym-style interface tailored to MARL-based order dispatch in real-world ride-sharing systems. By fully decoupling the environment from the dispatch algorithm, RideGym enables diverse learning-based and model-based methods to be developed and compared under identical, fully specified conditions. It supports efficient, large-scale city-level simulations on real road networks, and offers flexible configurations for vehicle attributes, order specifications, and automatic shortest-path routing. We validate RideGym by reproducing several baselines, and demonstrate its high efficiency, with a one-hour simulation involving thousands of vehicles and tens of thousands of orders completed within one minute across all methods. Moreover, we reveal that the choice of exploration noise can significantly affect both the performance and the relative ranking of MARL solutions, an aspect often overlooked in prior work.

Figures

Figures reproduced from arXiv: 2607.10173 by Sen Li, Yulong Hu, Zijian Zhao.

Figure 1
Figure 1. Figure 1: Workflow of the proposed simulation framework. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Typical Python usage workflow [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: A single-frame visualization of three focused vehi [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references · 1 canonical work pages

  1. [1]

    Abubakr O Al-Abbasi, Arnob Ghosh, and Vaneet Aggarwal. 2019. Deeppool: Distributed model-free algorithm for ride-sharing using deep reinforcement learning.IEEE Transactions on Intelligent Transportation Systems20, 12 (2019), 4714–4727

  2. [2]

    Javier Alonso-Mora, Samitha Samaranayake, Alex Wallar, Emilio Frazzoli, and Daniela Rus. 2017. On-demand high-capacity ride-sharing via dynamic trip- vehicle assignment.Proceedings of the National Academy of Sciences114, 3 (2017), 462–467

  3. [3]

    Javier Alonso-Mora, Alex Wallar, and Daniela Rus. 2017. Predictive routing for autonomous mobility-on-demand systems with ride-sharing. InIEEE/RSJ International Conference on Intelligent Robots and Systems. 3583–3590

  4. [4]

    Michael Behrisch, Laura Bieker, Jakob Erdmann, and Daniel Krajzewicz. 2011. SUMO–simulation of urban mobility: an overview. InProceedings of SIMUL 2011, the third international conference on advances in system simulation. ThinkMind

  5. [5]

    Federico Berto, Chuanbo Hua, Junyoung Park, Laurin Luttmann, Yining Ma, Fanchen Bu, Jiarui Wang, Haoran Ye, Minsu Kim, Sanghyeok Choi, et al. 2025. Rl4co: an extensive reinforcement learning for combinatorial optimization bench- mark. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 5278–5289

  6. [6]

    Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schul- man, Jie Tang, and Wojciech Zaremba. 2016. Openai gym.arXiv preprint arXiv:1606.01540(2016)

  7. [7]

    Tingwei Chen, Yantao Wang, Hanzhi Chen, Zijian Zhao, Xinhao Li, Nicola Pi- ovesan, Guangxu Zhu, and Qingjiang Shi. 2025. Modelling the 5G Energy Con- sumption Using Real-world Data: Energy Fingerprint is All You Need. In2025 IEEE Globecom Workshops (GC Wkshps). 1675–1680. doi:10.1109/GCWkshps68340. 2025.11590940

  8. [8]

    Wang Chen, Hongzheng Shi, and Jintao Ke. 2025. HRSim: An agent-based simulation platform for high-capacity ride-sharing services.arXiv preprint arXiv:2505.17758(2025)

  9. [9]

    Zhong Chen. 2023. Understanding of Didi’s Trading Strategy: Driver Order Matching. https://mp.weixin.qq.com/s/i8IIaMYaFub0a9M9MtFgSw

  10. [10]

    Jean-François Cordeau and Gilbert Laporte. 2007. The dial-a-ride problem: models and algorithms.Annals of Operations Research153, 1 (2007), 29–46

  11. [11]

    Tobias Enders, James Harrison, Marco Pavone, and Maximilian Schiffer. 2023. Hybrid multi-agent deep reinforcement learning for autonomous mobility on demand systems. InLearning for Dynamics and Control Conference. PMLR, 1284– 1296

  12. [12]

    Roman Engelhardt, Florian Dandl, Arslan-Ali Syed, Yunfei Zhang, Fabian Fehn, Fynn Wolf, and Klaus Bogenberger. 2022. Fleetpy: A modular open-source simulation tool for mobility on-demand services.arXiv preprint arXiv:2207.14246 (2022)

  13. [13]

    Siyuan Feng, Taijie Chen, Yuhao Zhang, Jintao Ke, Zhengfei Zheng, and Hai Yang. 2024. A multi-functional simulation platform for on-demand ride service operations.Communications in Transportation Research4 (2024), 100141

  14. [14]

    David Gale and Lloyd S Shapley. 1962. College admissions and the stability of marriage.The American mathematical monthly69, 1 (1962), 9–15

  15. [15]

    Shuxin Ge, Xiaobo Zhou, and Tie Qiu. 2025. Marl-based pricing strategy via mutual attention for mod systems with ridesharing and repositioning. InIEEE INFOCOM 2025-IEEE Conference on Computer Communications. IEEE, 1–10

  16. [16]

    Mordechai Haklay and Patrick Weber. 2008. Openstreetmap: User-generated street maps.IEEE Pervasive computing7, 4 (2008), 12–18

  17. [17]

    Jiang Hao and Pradeep Varakantham. 2022. Hierarchical value decomposition for effective on-demand ride-pooling. InProceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems. 580–587

  18. [18]

    Joshua Holder, Natasha Jaques, and Mehran Mesbahi. 2025. Multi agent rein- forcement learning for sequential satellite assignment problems. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 26516–26524

  19. [19]

    Heiko Hoppe, Tobias Enders, Quentin Cappart, and Maximilian Schiffer. 2024. Global rewards in multi-agent deep reinforcement learning for autonomous mobility on demand systems. In6th Annual Learning for Dynamics & Control Conference. PMLR, 260–272

  20. [20]

    Yulong Hu, Tingting Dong, and Sen Li. 2025. Coordinating ride-pooling with public transit using Reward-Guided Conservative Q-Learning: An offline train- ing and online fine-tuning reinforcement learning framework.Transportation Research Part C: Emerging Technologies174 (2025), 105051

  21. [21]

    Yulong Hu, Siyuan Feng, and Sen Li. 2025. Bmg-q: Localized bipartite match graph attention q-learning for ride-pooling order dispatch.IEEE Transactions on Intelligent Transportation Systems(2025)

  22. [22]

    Cor AJ Hurkens and Gerhard J Woeginger. 2004. On the nearest neighbor rule for the traveling salesman problem.Operations Research Letters32, 1 (2004), 1–4

  23. [23]

    Scarlett T Jin, Hui Kong, Rachel Wu, and Daniel Z Sui. 2018. Ridesourcing, the sharing economy, and the future of cities.Cities76 (2018), 96–104

  24. [24]

    Weiqiang Jin, Hongyang Du, Biao Zhao, Xingwu Tian, Bohang Shi, and Guang Yang. 2025. A comprehensive survey on multi-agent cooperative decision- making: Scenarios, approaches, challenges and perspectives.arXiv preprint arXiv:2503.13415(2025)

  25. [25]

    Bala Kalyanasundaram and Kirk Pruhs. 1993. Online weighted matching.Journal of Algorithms14, 3 (1993), 478–488

  26. [26]

    Rafał Kucharski and Oded Cats. 2022. Simulating two-sided mobility platforms with MaaSSim.Plos one17, 6 (2022), e0269682

  27. [27]

    Harold W Kuhn. 1955. The Hungarian method for the assignment problem.Naval research logistics quarterly2, 1-2 (1955), 83–97

  28. [28]

    Moritz Laupichler, Robin Andre, Kim Kandler, Peter Sanders, and Peter Vor- tisch. 2026. Advancing Dynamic Ride-Pooling Simulation–A Highly Scalable Dispatcher.arXiv preprint arXiv:2605.11798(2026)

  29. [29]

    Minne Li, Zhiwei Qin, Yan Jiao, Yaodong Yang, Jun Wang, Chenxi Wang, Guobin Wu, and Jieping Ye. 2019. Efficient ridesharing order dispatching with mean field multi-agent reinforcement learning. InThe World Wide Web Conference. 983–994

  30. [30]

    Kaixiang Lin, Renyu Zhao, Zhe Xu, and Jiayu Zhou. 2018. Efficient large-scale fleet management via multi-agent deep reinforcement learning. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1774–1783

  31. [31]

    Michael L Littman. 1994. Markov games as a framework for multi-agent rein- forcement learning. InMachine learning proceedings 1994. Elsevier, 157–163

  32. [32]

    Yang Liu and Sen Li. 2025. Piggyback on idle ride-sourcing drivers for integrated on-demand and flexible intracity parcel delivery services.Transportation Science 59, 3 (2025), 494–517

  33. [33]

    Yang Liu, Yitong Shang, and Sen Li. 2025. Joint Infrastructure Planning and Order Assignment for On-Demand Food-Delivery Services with Coordinated Drones and Human Couriers.arXiv preprint arXiv:2501.14325(2025)

  34. [34]

    Pablo Alvarez Lopez, Michael Behrisch, Laura Bieker-Walz, Jakob Erdmann, Yun- Pang Flötteröd, Robert Hilbrich, Leonhard Lücken, Johannes Rummel, Peter Wagner, and Evamarie Wießner. 2018. Microscopic traffic simulation using sumo. In2018 21st international conference on intelligent transportation systems (ITSC). Ieee, 2575–2582

  35. [35]

    Dennis Luxen and Christian Vetter. 2011. Real-time routing with OpenStreetMap data. InProceedings of the 19th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems(Chicago, Illinois)(GIS ’11). ACM, New York, NY, USA, 513–516. doi:10.1145/2093973.2094062

  36. [36]

    Takuma Oda and Carlee Joe-Wong. 2018. MOVI: A model-free approach to dynamic fleet management. InIEEE INFOCOM 2018-IEEE Conference on Computer Communications. IEEE, 2708–2716

  37. [37]

    Zhiwei Qin, Xiaocheng Tang, Yan Jiao, Fan Zhang, Zhe Xu, Hongtu Zhu, and Jieping Ye. 2020. Ride-hailing order dispatching at didi via reinforcement learning. INFORMS Journal on Applied Analytics50, 5 (2020), 272–286

  38. [38]

    Connor Riley, Pascal Van Hentenryck, and Enpeng Yuan. 2021. Real-time dis- patching of large-scale ride-sharing systems: integrating optimization, machine learning, and model predictive control. InProceedings of the Twenty-Ninth Inter- national Conference on International Joint Conferences on Artificial Intelligence. 4417–4423

  39. [39]

    Susan Shaheen and Adam Cohen. 2019. Shared ride services in North America: definitions, impacts, and the future of pooling.Transport reviews39, 4 (2019), 427–442

  40. [40]

    Andrea Simonetto, Julien Monteil, and Claudio Gambella. 2019. Real-time city- scale ridesharing via linear assignment problems.Transportation Research Part C: Emerging Technologies101 (2019), 208–232

  41. [41]

    Jiahui Sun, Haiming Jin, Zhaoxing Yang, Lu Su, and Xinbing Wang. 2022. Optimiz- ing long-term efficiency and fairness in ride-hailing via joint order dispatching and driver repositioning. InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 3950–3960

  42. [42]

    Sien Yi Tan, Henokh Fibrianto, and Larry Lin. 2025. DispatchGym: Grab’s rein- forcement learning research framework. https://engineering.grab.com/techblog_- dispatchgym

  43. [43]

    Xiaocheng Tang, Zhiwei Qin, Fan Zhang, Zhaodong Wang, Zhe Xu, Yintai Ma, Hongtu Zhu, and Jieping Ye. 2019. A deep value-network based approach for multi-driver order dispatching. InProceedings of the 25th ACM SIGKDD interna- tional conference on knowledge discovery & data mining. 1780–1790

  44. [44]

    New York City Taxi and Limousine Commission. 2024. Nyc taxi and limousine commission-trip record data nyc. https://www.nyc.gov/site/tlc/about/tlc-trip- record-data.page

  45. [45]

    Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks.arXiv preprint arXiv:1710.10903(2017)

  46. [46]

    Cheng Wang, Yixuan Ding, and Jin Jiang. 2025. On-Demand Dynamic Intercity Shared-Taxi System With Re-Optimization.IEEE Transactions on Intelligent Transportation Systems(2025)

  47. [47]

    Zhe Xu, Zhixin Li, Qingwen Guan, Dingshui Zhang, Qiang Li, Junxiao Nan, Chunyang Liu, Wei Bian, and Jieping Ye. 2018. Large-scale order dispatch in on- demand ride-hailing platforms: A learning and planning approach. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 905–913

  48. [48]

    Shangyao Yan, Chun-Ying Chen, and Yu-Fang Lin. 2011. A model with a heuristic algorithm for solving the long-term many-to-many car pooling problem.IEEE Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Zhao et al. Transactions on Intelligent Transportation Systems12, 4 (2011), 1362–1373

  49. [49]

    Guixiang Yang, Hao Zhang, and Lin Qiu. 2026. Graph-based multi-agent rein- forcement learning with an enriched environment for joint ride-sharing and charging optimization.Applied Energy405 (2026), 127220

  50. [50]

    Xinlang Yue, Yiran Liu, Fangzhou Shi, Sihong Luo, Chen Zhong, Min Lu, and Zhe Xu. 2024. An end-to-end reinforcement learning based approach for micro-view order-dispatching in ride-hailing. InProceedings of the 33rd ACM international conference on information and knowledge management. 5054–5061

  51. [51]

    Zijian Zhao, Tingwei Chen, Zhijie Cai, Xiaoyang Li, Hang Li, Qimei Chen, and Guangxu Zhu. 2025. Crossfi: A cross domain wi-fi sensing framework based on siamese network.IEEE Internet of Things Journal(2025)

  52. [52]

    Zijian Zhao and Sen Li. 2025. The impacts of data privacy regulations on food- delivery platforms.Transportation Research Part C: Emerging Technologies181 (2025), 105364

  53. [53]

    Zijian Zhao and Sen Li. 2025. One step is enough: Multi-agent reinforcement learning based on one-step policy optimization for order dispatch on ride-sharing platforms.arXiv preprint arXiv:2507.15351(2025)

  54. [54]

    Zijian Zhao and Sen Li. 2026. Discriminatory order assignment and payment- setting of on-demand food-delivery platforms: A multi-action and multi-agent reinforcement learning framework.Transportation Research Part E: Logistics and Transportation Review208 (2026), 104653

  55. [55]

    learn-a-value- then-match

    Zijian Zhao and Sen Li. 2026. Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?. InThe Fourteenth International Conference on Learning Representations. A Literature Reviews A.1 Model-Based Order Dispatch Order dispatch has traditionally been cast as a combinatorial op- timization problem and solved with model-based methods....