Pith. sign in

REVIEW 2 major objections 5 minor 33 references

Cross-border data compliance is a sequential firm decision that can be learned as hard legal constraints inside a weekly MDP, and the resulting policy shows localization, front-loaded credentials, and an absorb-then-adjust cost pattern.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 10:24 UTC pith:YTUOIKF3

load-bearing objection Solid statute-to-MDP decision-support template with careful simulation; the absorb-then-adjust policy claim is real inside their friction form but is not a general institutional fact. the 2 major comments →

arxiv 2607.10620 v1 pith:YTUOIKF3 submitted 2026-07-12 cs.CY

Sequential compliance decisions of firms on cross-border data flows: An institutionally anchored decision support system

classification cs.CY
keywords Data governanceCross-border data flowsMarkov decision processDeep reinforcement learningAction maskingCompliance decision supportData localization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Firms that move personal or important data across borders do not face a one-shot legal checklist. Each week they must choose how much to transfer, whether to process locally instead, and whether to buy a costly compliance credential whose value depends on later history. This paper turns the Chinese data-export routing rules into a computable minimal-compliance map that defines the legal action set for every state, then models the firm’s year as a finite-horizon Markov decision process whose constraints are hard masks rather than soft penalties. Masked deep reinforcement learning, augmented with offline counterfactual path advantages, produces policies that beat simple rule baselines on simulated firms. The learned behavior concentrates local processing where small lawful transfers cannot cover compliance costs, front-loads credential purchases, and, under rising friction costs, shows expected rewards falling before localization or transfer volumes visibly change. If correct, the system supplies managers with readable signals and shallow decision trees while warning regulators that behavioral indicators alone can understate the burden firms already carry.

Core claim

Inside an institutionally anchored finite-horizon MDP whose legal actions are generated week-by-week by a statutory minimal-compliance mapping, masked D3QN with counterfactual path-advantage signals learns policies that outperform non-learning baselines and other masked learners; local processing concentrates where small lawful transfers fail to cover compliance costs, credential acquisition is front-loaded within the compliance year, and rising persistent-friction weight produces an absorb-then-adjust pattern in which expected reward declines before observable localization or transfer volume changes.

What carries the argument

The minimal compliance mapping Rmin that converts statutory priority layers into a state-dependent legal action set Alegal, enforced by action masking so that compliance is a hard boundary rather than a reward penalty; counterfactual path-advantage augmentation then supplies the four long-run path values that both guide learning and serve as manager-readable signals.

Load-bearing premise

All reported rewards and behavioral patterns rest on a hand-calibrated economic environment whose cost, value, and demand parameters are not estimated from real firm export and credential histories.

What would settle it

Estimate the same cost, demand, and credential-invalidation parameters from proprietary firm-level export logs, retrain the masked policy, and check whether localization still concentrates on low-demand non-exempt states, credentials remain front-loaded, and the absorb-then-adjust pattern under rising friction weight still appears.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper formulates firm-side cross-border data compliance under China’s 2024 Provisions as a finite-horizon MDP whose legal action set is generated by a priority-ordered minimal compliance mapping Rmin (Eqs. 1–5, 10). Compliance is enforced by action masking rather than reward penalties. Counterfactual path-advantage augmentation (CPAA) supplies offline long-run path values relative to a default continuation policy π0; masked D3QN+CPAA then learns weekly transfer, localization, and credential decisions. On held-out simulated firm libraries the learned policies beat non-learning baselines and other masked learners; localization concentrates where small lawful transfers do not cover compliance costs and shifts as assessment-tier exposure rises; credential acquisition is front-loaded; shallow trees recover the policy at high fidelity; and raising the persistent-friction weight κA produces an absorb-then-adjust pattern in which expected reward falls before localization or transfer volume change.

Significance. If the institutional mapping and sequential formulation hold, the paper supplies a reusable architecture—rule-to-legal-set conversion, hard-constraint masked RL, and dual interpretability via CPAA signals plus distilled trees—for firm-side compliance under history-dependent statutory floors. That framing is a clear advance over static localization/transfer characterizations and over RegTech that treats compliance as ex-post classification. The hard-constraint design, multi-learner comparison, five-seed evaluation, and explicit distillation fidelities (92.3% / 97.6%) are concrete strengths. The absorb-then-adjust welfare claim is policy-relevant for assessing data-governance costs, but its generality is limited by the additive friction specification and hand-calibrated economic parameters; the contribution is therefore strongest as a decision-support system and institutional modeling template rather than as a robust empirical law of regulatory burden.

major comments (2)
  1. [§5.6, Eqs. (14), (20), (27), Fig. 9] §5.6, Eqs. (14), (20), (27) and Fig. 9: the absorb-then-adjust claim is load-bearing for the abstract’s policy implication, yet ΔR is independent of κA by construction under the additive friction cost. The delayed behavioral response is therefore forced by the chosen functional form of Cfric rather than shown to be a general property of regulatory strictness. At minimum the paper should (i) state this dependence explicitly and (ii) report one alternative friction specification (e.g., multiplicative in transfer value, mechanism-specific, or entering the legal set) so readers can judge whether the welfare-before-behavior ordering survives. Without that, the claim should be scoped to the present cost structure.
  2. [Appendix B Table B-1; §5.5; §6] Appendix B Table B-1 and §6: all reported rewards, localization boundaries, front-loading, and the κA pattern are generated under a single hand-calibrated economic environment (β, κA, κσ, F(1), F(2), µ, α, pchg, qref). §5.5 varies only two firm-side parameters and leaves the friction and credential structure fixed. Because the authors themselves flag the lack of proprietary firm-level estimation as a main limitation, the central behavioral findings should be presented as simulator-conditional, and either a broader multi-parameter sensitivity or a clear external-validity caveat should be added before the policy language in the abstract and §5.6 is retained at full strength.
minor comments (5)
  1. [§4.1, Algorithm A-1] Clarify how the default continuation policy π0 is chosen for CPAA labels (Algorithm A-1 / §4.1) and whether results are sensitive to alternative deterministic references; a short ablation would strengthen the interpretability claims.
  2. [Table 3] Table 3: report confidence intervals or pairwise tests for reward differences among the top DQN-family policies so that ‘outperform the baselines’ is statistically transparent rather than mean-rank only.
  3. [Fig. 7] Fig. 7 trees use learned thresholds (e.g., friction At ≤ 0.22, demand ≤ 0.13 qref); state whether these thresholds are stable across seeds or scenario libraries.
  4. [Eq. (26); Table 2] Notation: the dual use of r_t for minimum compliance stringency and for single-period reward in the RL update (Eq. 26) is confusing; rename one of them.
  5. [Abstract; §6] Transferability claim (abstract, §6) would be more credible with a short sketch of how Rmin would be rewritten for one non-Chinese regime (e.g., GDPR SCCs / adequacy) rather than a generic assertion.

Circularity Check

1 steps flagged

No load-bearing circular derivation; mild model-structure tautology only on the κA single-period independence used to frame absorb-then-adjust.

specific steps
  1. self definitional [§5.6, Eqs. (20) and (27); Fig. 9 framing]
    "The current friction cost κA At depends only on the beginning-of-period friction stock At and is therefore common to both current choices. ... The formula does not contain κA, so the persistent-friction weight does not directly change the payoff difference between cross-border transfer and local processing within the week. Its effect is instead sequential: the current transfer decision changes the future friction stock, and a larger κA increases the cost associated with that stock in subsequent periods."

    Cfric_t = κA At + κσ xt/qref makes κA At a common additive term for transfer and LOCAL in the current week, so ΔR (Eq. 27) excludes κA by construction of the cost functional form. Presenting that independence as analytical support for the absorb-then-adjust welfare claim therefore restates a modeling definition rather than an independent empirical or first-principles result. The multi-period threshold itself is not purely definitional, but the static non-response to κA is.

full rationale

This is a simulation / decision-support paper, not a first-principles prediction paper. The minimal compliance mapping is a direct encoding of statutory routing rules into legal action sets; masked RL then optimizes within that fixed environment and is scored against non-learning baselines and ablations on held-out simulated firms. Those comparisons are not forced by definition: different learners and rules produce different rewards and path shares (Table 3). CPAA is an offline feature under a fixed default continuation policy, not a fit of the reported outcome. There is no self-citation uniqueness theorem, no ansatz smuggled from the authors’ prior work, and no renaming of an external empirical law. The only mild circularity is in §5.6: the single-period payoff gap ΔR is independent of κA by construction of the additive friction cost (Eqs. 20 and 27), so the claim that κA does not tilt the static transfer-vs-LOCAL comparison is definitional rather than an independent discovery. The multi-period absorb-then-adjust pattern in Fig. 9 still depends on sequential value and the learned policy, so it is not fully tautological; the residual issue is model-dependence (correctness risk), not a circular derivation chain. Score 2 reflects that one minor self-definitional step without collapsing the paper’s central RL and localization results.

Axiom & Free-Parameter Ledger

9 free parameters · 6 axioms · 3 invented entities

The central claims rest on a statute-derived hard-constraint MDP plus a fully simulated economic environment. Load-bearing free parameters are the hand-calibrated reward and friction constants in Table B-1; domain axioms include treating compliance as a hard mask and the concave transfer-value form; invented constructs are the minimal compliance mapping, the regulatory friction stock, and CPAA path-advantage signals. No independent firm-level estimation anchors the economic side.

free parameters (9)
  • value curvature β = 3.0 (baseline)
    Shapes business-value concavity g(z)=1-e^{-βz}; baseline 3.0, varied in robustness only after the main policy is trained.
  • persistent-friction weight κA = 0.50 (baseline)
    Continuous proxy for regulatory strictness; baseline 0.5, swept to 2.0 to produce the absorb-then-adjust claim.
  • per-period friction weight κσ = 0.30
    Immediate transfer friction cost; jointly with κA determines friction burden.
  • friction memory α = 0.85
    AR weight in friction stock update At+1; controls persistence of past transfers.
  • Level-1/Level-2 acquisition costs F(1), F(2) = 0.35 / 0.55
    One-time credential costs that drive investment timing and localization of small transfers.
  • shortfall rate µ = 0.30
    Opportunity-loss weight on unmet demand in net business value Vt.
  • credential invalidation probability pchg = 0.08
    Stochastic reset of credential holdings; affects reinvestment incentives.
  • demand normalization qref = 5e4
    Scale for transfer signal and value function; sets units of all cost/value terms.
  • default continuation policy π0 for CPAA labels = tier-dependent full-volume EXEMPT/SCC/SA-or-LOCAL rule
    Deterministic reference used to define offline path advantages Λk; choice of π0 shapes the interpretable signals and CPAA observations.
axioms (6)
  • domain assumption Statutory compliance floors are hard constraints on the action set (mask), not soft penalties in the reward.
    Stated in abstract and §3.3 (Eq. 10); central to distinguishing the formulation from penalty-based constrained RL.
  • domain assumption Business value of transfer is increasing and concave in normalized volume via g(z)=1-e^{-βz}.
    Eqs. 17–18, motivated by Goldfarb and Tucker (2019); drives when small transfers fail cost recovery.
  • ad hoc to paper Regulatory friction is a persistent clipped stock driven only by transfer volume, not mechanism type.
    Eq. 14; enables the κA strictness experiment and absorb-then-adjust narrative.
  • domain assumption Rmin priority layers (important data / exemptions / CIIO / volume thresholds) correctly and completely encode the 2024 Provisions.
    §3.2 Eqs. 2–5; legal fidelity of all legal action sets depends on this encoding.
  • ad hoc to paper Weekly tasks and firm attributes can be drawn from calibrated distributions so that 3000/300/300 firm libraries represent the decision problem.
    §5.1; no proprietary operational data; all performance and behavior metrics are conditional on this generative model.
  • standard math Finite-horizon discounted MDP with γ≈0.99 and T=52 is an adequate model of a compliance year.
    Standard sequential decision framework used in §3.3 and solution method.
invented entities (3)
  • Minimal compliance mapping Rmin no independent evidence
    purpose: Convert statutory export-routing logic into a computable state→{E,M,H} stringency and legal action set.
    Core institutional interface of the system; independent_evidence false because correctness is a legal-engineering claim not externally validated against court or firm outcomes.
  • Regulatory friction stock At no independent evidence
    purpose: Carry history-dependent operating cost of sustained export activity into future rewards.
    Model construct used as continuous strictness channel via κA; no external measurement protocol given.
  • Counterfactual path-advantage augmentation (CPAA) no independent evidence
    purpose: Offline long-run advantages of LOCAL/L0/L1/L2 paths as both learning features and manager-facing signals.
    Methodological add-on relative to plain masked D3QN; value depends on π0 and the same simulated rewards.

pith-pipeline@v1.1.0-grok45 · 29816 in / 4225 out tokens · 45675 ms · 2026-07-14T10:24:58.698076+00:00 · methodology

0 comments
read the original abstract

The economic value of data arises from its flow across organizations and national borders. Yet increasingly stringent data governance regimes are turning cross-border transfer into an institutionally constrained sequential decision, in which firms repeatedly weigh compliance costs against the value of data flows. From the perspective of a data-exporting firm, this paper develops an institutionally anchored decision support system. It converts regulatory rules into a computable minimal compliance mapping and models the firm's weekly decisions as a finite-horizon Markov decision process (MDP), with compliance represented as a hard constraint rather than a penalty term. The resulting problem is solved using masked deep reinforcement learning, while counterfactual path advantages provide interpretable signals to support the firm's cross-border data flow decisions. Experiments show that the policies learned within the system outperform the baselines considered and deliver interpretable, auditable decision support. Local processing concentrates in states where the business value of small lawful transfers does not cover their compliance costs, and the localization boundary shifts systematically as the regime tightens. Credential acquisition is front-loaded within the compliance year, and shallow decision trees reproduce the policy's decisions with high fidelity. Treating the persistent-friction weight as a continuous representation of regulatory strictness further reveals an absorb-then-adjust pattern, in which expected rewards decline before observable behavior changes, implying that assessments based only on behavioral indicators may understate the burden already borne by firms. Moreover, the system is not tied to any specific regulation and can be transferred to other jurisdictions and rule-based compliance problems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 19 canonical work pages

  1. [1]

    Another Digital Divide: The Rise of Data Realms and its Implications for the WTO

    “ Another Digital Divide: The Rise of Data Realms and its Implications for the WTO. ”Journal of International Economic Law21 (2): 245–272. https://doi.org/10.1093/jiel/jgy019. Abbasi, Ahmed, Conan Albrecht, Anthony Vance, and James Hansen

  2. [2]

    MetaFraud: A Meta- Learning Framework for Detecting Financial Fraud

    “MetaFraud: A Meta- Learning Framework for Detecting Financial Fraud. ”MIS Quarterly36 (4): 1293–1328. https: //doi.org/10.2307/41703508. Acemoglu, Daron, Ali Makhdoumi, Azarakhsh Malekian, and Asu Ozdaglar

  3. [3]

    Too Much Data: Prices and Inefficiencies in Data Markets

    “Too Much Data: Prices and Inefficiencies in Data Markets. ”American Economic Journal: Microeconomics14 (4): 218–256.https://doi.org/10.1257/mic.20200200. Baley, Isaac and Andrés Blanco

  4. [4]

    The Macroeconomics of Irreversibility

    “The Macroeconomics of Irreversibility. ”Review of Economic Studiesp. rdag001.https://doi.org/10.1093/restud/rdag001. Bao, Yang, Bin Ke, Bin Li, Y . Julia Yu, and Jie Zhang

  5. [5]

    Detecting Accounting Fraud in Publicly Traded U.S. Firms Using a Machine Learning Approach

    “Detecting Accounting Fraud in Publicly Traded U.S. Firms Using a Machine Learning Approach. ”Journal of Accounting Research58 (1): 199–235.https://doi.org/10.1111/1475-679X.12292. Bastani, Hamsa, Osbert Bastani, and Wichinpong Park Sinchaisri

  6. [6]

    Improving Human Sequential Decision Making with Reinforcement Learning

    “Improving Human Sequential Decision Making with Reinforcement Learning. ”Management Science72 (1): 733–755. https://doi.org/10.1287/mnsc.2022.02455. Bastani, Osbert, Yewen Pu, and Armando Solar-Lezama

  7. [7]

    Verifiable Reinforcement Learning via Policy Extraction

    “Verifiable Reinforcement Learning via Policy Extraction. ” InAdvances in Neural Information Processing Systems 31 (NeurIPS 2018), pp. 2499–2509 . Red Hook, NY: Curran Associates. Chen, Ji, Yifan Xu, Peiwen Yu, and Jun Zhang

  8. [8]

    A Reinforcement Learning Approach for Hotel Revenue Management with Evidence from Field Experiments

    “ A Reinforcement Learning Approach for Hotel Revenue Management with Evidence from Field Experiments. ”Journal of Operations Management69 (7): 1176–1201.https://doi.org/10.1002/joom.1246. Chisam, Natalie, Jordan W . Moffett, Frank Germann, and Robert W . Palmatier

  9. [9]

    Privacy Trade-Offs in International Markets

    “Privacy Trade-Offs in International Markets. ”Journal of International Business Studies. https://doi. org/10.1057/s41267-025-00837-4. Online first

  10. [10]

    Measures for the Security Assessment of Outbound Data Transfers

    “Measures for the Security Assessment of Outbound Data Transfers. ” , Cyberspace Administration of China, Beijing. https://www. cac.gov.cn/2022-07/07/c_1658811536396503.htm. Effective 1 September

  11. [11]

    Measures on the Standard Contract for the 36 Outbound Cross-Border Transfer of Personal Information

    “Measures on the Standard Contract for the 36 Outbound Cross-Border Transfer of Personal Information. ” , Cyberspace Administration of China, Beijing. https://www.cac.gov.cn/2023-02/24/c_1678884830036813.htm. Effective 1 June

  12. [12]

    Provisions on Promoting and Regulating the Cross-Border Flow of Data

    “Provisions on Promoting and Regulating the Cross-Border Flow of Data. ” , Cyberspace Administration of China, Beijing. https://www. cac.gov.cn/2024-03/22/c_1712776611775634.htm. Promulgated and effective 22 March

  13. [13]

    Data and Markets

    “Data and Markets. ”Annual Review of Economics 15 (1): 23–40.https://doi.org/10.1146/annurev-economics-082322-023244. Gijsbrechts, Joren, Robert N. Boute, Jan A. Van Mieghem, and Dennis J. Zhang

  14. [14]

    Can Deep Reinforcement Learning Improve Inventory Management? Performance on Lost Sales, Dual- Sourcing, and Multi-Echelon Problems

    “Can Deep Reinforcement Learning Improve Inventory Management? Performance on Lost Sales, Dual- Sourcing, and Multi-Echelon Problems. ”Manufacturing & Service Operations Management24 (3): 1349–1368.https://doi.org/10.1287/msom.2021.1064. Glanois, Claire, Paul Weng, Matthieu Zimmer, Dong Li, Tianpei Yang, Jianye Hao, and Wulong Liu

  15. [15]

    A Survey on Interpretable Reinforcement Learning

    “ A Survey on Interpretable Reinforcement Learning. ”Machine Learning113 (8): 5847–5890.https://doi.org/10.1007/s10994-024-06543-w. Goldfarb, Avi and Catherine Tucker. 2019 . “Digital Economics. ”Journal of Economic Literature57 (1): 3–43.https://doi.org/10.1257/jel.20171452. Harsha, Pavithra, Ashish Jagmohan, Jayant Kalagnanam, Brian Quanz, and Divya Singhvi

  16. [16]

    Deep Policy Iteration with Integer Programming for Inventory Management

    “Deep Policy Iteration with Integer Programming for Inventory Management. ”Manufacturing & Service Operations Management27 (2): 369–388. https://doi.org/10.1287/msom.2022

  17. [17]

    Are We Done with Business Process Compliance: State of the Art and Challenges Ahead

    “ Are We Done with Business Process Compliance: State of the Art and Challenges Ahead. ”Knowledge and Information Systems57 (1): 79–133.https://doi.org/10.1007/s10115-017-1142-1. Huang, Shengyi and Santiago Ontañón

  18. [18]

    A Closer Look at Invalid Action Masking in Policy Gra- dient Algorithms

    “ A Closer Look at Invalid Action Masking in Policy Gra- dient Algorithms. ” InProceedings of the 35th International Florida Artificial Intelligence Research Society Conference (FLAIRS-35).https://doi.org/10.32473/flairs.v35i.130584. Jia, Jian, Ginger Zhe Jin, and Liad Wagman

  19. [19]

    The Short-Run Effects of the General Data Protection Regulation on Technology Venture Investment

    “The Short-Run Effects of the General Data Protection Regulation on Technology Venture Investment. ”Marketing Science40 (4): 661–684. https://doi.org/10.1287/mksc.2020.1271. Johnson, Garrett A., Scott K. Shriver, and Samuel G. Goldberg

  20. [20]

    Privacy and Market Concen- tration: Intended and Unintended Consequences of the GDPR

    “Privacy and Market Concen- tration: Intended and Unintended Consequences of the GDPR. ”Management Science69 (10): 5695–5721.https://doi.org/10.1287/mnsc.2023.4709. Jones, Charles I. and Christopher Tonetti

  21. [21]

    Nonrivalry and the Economics of Data

    “Nonrivalry and the Economics of Data. ”American Economic Review110 (9): 2819–2858.https://doi.org/10.1257/aer.20191330. Krämer, Jan and Shiva Shekhar

  22. [22]

    Regulating Digital Platform Ecosystems Through Data Sharing and Data Siloing: Consequences for Innovation and Welfare

    “Regulating Digital Platform Ecosystems Through Data Sharing and Data Siloing: Consequences for Innovation and Welfare. ”MIS Quarterly49 (1): 123–154.https://doi.org/10.25300/MISQ/2024/18428. Ma, Guang and Hong Wu

  23. [23]

    Cross-Border Data Flow Supervision in China’s Free Trade Zones: Security and Compliance Rules

    “Cross-Border Data Flow Supervision in China’s Free Trade Zones: Security and Compliance Rules. ”Asia Pacific Law Review33 (2): 231–262. https: //doi.org/10.1080/10192557.2025.2471312. Mattoo, Aaditya and Joshua P . Meltzer

  24. [24]

    International Data Flows and Privacy: The Conflict 37 and Its Resolution

    “International Data Flows and Privacy: The Conflict 37 and Its Resolution. ”Journal of International Economic Law21 (4): 769–789 .https://doi.org/ 10.1093/jiel/jgy044. Ngai, E. W .T., Yong Hu, Y .H. Wong, Yijun Chen, and Xin Sun

  25. [25]

    The Application of Data Mining Techniques in Financial Fraud Detection: A Classification Framework and an Academic Review of Literature

    “The Application of Data Mining Techniques in Financial Fraud Detection: A Classification Framework and an Academic Review of Literature. ”Decision Support Systems50 (3): 559–569 .https://doi.org/10.1016/j.dss. 2010.08.006. OECD

  26. [26]

    The Nature, Evolution and Potential Implications of Data Localisation Measures

    “The Nature, Evolution and Potential Implications of Data Localisation Measures. ” , OECD Publishing, Paris.https://doi.org/10.1787/179f718a-en. OECD/WTO

  27. [27]

    Do People Around the World Care Where Their Data Are Stored?

    “Do People Around the World Care Where Their Data Are Stored?”Information Economics and Policy71: 101132. https://doi.org/10.1016/j. infoecopol.2025.101132. Rong, Ke, Yunshu Ling, Tianxi Yang, and Cheng Huang

  28. [28]

    Cross-Border Data Transfer: Patterns and Discrepancies

    “Cross-Border Data Transfer: Patterns and Discrepancies. ”Journal of International Business Policy8 (1): 10–32. https: //doi.org/10.1057/s42214-025-00209-7. Rudin, Cynthia. 2019 . “Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. ”Nature Machine Intelligence1 (5): 206–215. https://doi.org/10...

  29. [29]

    Explainability and Fairness of RegTech for Regulatory Enforcement: Automated Monitoring of Consumer Complaints

    “Explainability and Fairness of RegTech for Regulatory Enforcement: Automated Monitoring of Consumer Complaints. ”Decision Support Systems158: 113782. https: //doi.org/10.1016/j.dss.2022.113782. Standing Committee of the National People’s Congress (NPC)

  30. [30]

    Personal Information Protec- tion Law of the People’s Republic of China

    “Personal Information Protec- tion Law of the People’s Republic of China. ” , Standing Committee of the National People’s Congress, Beijing. http://www.npc.gov.cn/npc/c2/c30834/202108/t20210820_ 313088.html. Adopted 20 August 2021, effective 1 November

  31. [31]

    Excluding the Irrelevant: Focusing Reinforcement Learning Through Continu- ous Action Masking

    “Excluding the Irrelevant: Focusing Reinforcement Learning Through Continu- ous Action Masking. ” InAdvances in Neural Information Processing Systems 37 (NeurIPS 2024), pp. 95067–95094. Red Hook, NY: Curran Associates. https://doi.org/10.52202/079017-

  32. [32]

    A Survey of Constraint Formulations in Safe Reinforcement Learning

    “ A Survey of Constraint Formulations in Safe Reinforcement Learning. ” InProceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI-24), Survey Track, pp. 8262–8271.https://doi.org/10.24963/ijcai. 2024/913. Zhou, Weiwen, Hossein Fotouhi, and Elise Miller-Hooks

  33. [33]

    Decision Support Through Deep Reinforcement Learning for Maximizing a Courier’s Monetary Gain in a Meal Delivery Envi- ronment

    “Decision Support Through Deep Reinforcement Learning for Maximizing a Courier’s Monetary Gain in a Meal Delivery Envi- ronment. ”Decision Support Systems190: 114388. https://doi.org/10.1016/j.dss.2024. 114388. 38 Appendix A. Algorithms of the solution method The solution method comprises two modules (Fig. 2). The counterfactual path-prediction module com...