REVIEW 3 major objections 3 minor 72 references
Governing Agentic AI in FinTech
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper argues that the binding constraint on delegating consequential financial decisions to agentic AI is verifiability, not capability: an agent can stay accurate and stable while the institution cannot substantiate how it acted.
desk verdict Serious and useful paper on agentic AI governance, but its formal core rests on a Markov-chain assumption that the paper's own Study 2 architecture violates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Verifiability Gap, a shortfall in retained verification capacity indexed to a verifier, an evidentiary standard, and an audit lag, together with the reproducibility profile $R(t;\sigma) = (R_O, R_H, R_P, R_T; D_V)$ that separates current outcome, historical outcome, material process, and exact trace reproducibility and tracks verdict differentiation $D_V$ as a degeneracy diagnostic. The argument is carried by a serial-pipeline information model in which only the first layer observes the original case and the joint law factorizes as a Markov chain $C \to S_1 \to S_2 \to \cdots \to S_d \to Y$; on that chain the paper proves that discrimination contracts geometrically in depth (Theorem 1), that $\sigma$-reversibility is equivalent to equality in the data-processing inequality (Theorem 2), that deterministic summarization is incompatible with reversibility (Theorem 3), and that reversibility requires retaining at least $H([T]_\sigma \mid Y)$ bits of evidence (Theorem 4). These results convert the verbal claim that serial handoffs widen the gap into a bit-level accounting of what an institution must keep.
What would settle it
Run the Study 1 release-timeline and Study 2 orchestration protocols on a provider that exposes a dated pinned endpoint, a random seed, and temperature control: if baseline-modal historical fidelity stays at 1.000 across releases and exact trace reproducibility $R_T$ reaches 1 at any depth beyond a single agent, the provider-dominance and serial-contraction claims would be refuted.
Extended reading notes
Core claim
The paper's central claim is that an agentic AI system can remain highly capable, accurate, and behaviorally stable while the responsible institution becomes entirely unable to substantiate how that system exercised its delegated authority. It formalizes this as the Verifiability Gap, $G_{dq} = [\rho_\sigma(A_d) - V_{dq}]^+$, the positive shortfall between the verification required by delegated authority and the verification capacity retained in explainability and reproducibility, indexed to a named verifier, an evidentiary standard, and an audit lag. The formal core is a set of information-theoretic results on serial agent pipelines: discrimination $I(C;Y)$ contracts geometrically with pipeline depth via the data-processing inequality, a stage is $\sigma$-reversible exactly when it throws nothing away, a stage either summarizes or stays auditable but cannot do both, and the retained evidence must satisfy $H(E) \geq H([T]_\sigma \mid Y)$ — the institution must store at least as many bits as its pipeline destroyed. Empirically, three studies identify distinct mechanisms: provider releases silently shift risk posture while withdrawing temperature, top-p, top-k, and seed controls; orchestration acts as a latent policy layer in which no execution record repeats at any scale and apparent stability can be degenerate output collapse; and a preserved logistic credit model achieves perfect current reproducibility while failing to reproduce a historical decision that crossed a policy threshold. The conclusion is that reproducibility is a governance profile, not a scalar, and that delegated authority remains defensible only while retained evidence substantiates its exercise.
Load-bearing premise
The formal contraction and retention theorems assume that the pipeline factorizes as a Markov chain — only the first layer observes the original case and every downstream agent sees only the records transmitted along designated edges — so if production agentic systems let later agents see the original case, live tools, or shared memory, the serial information loss and the retention lower bound may not hold.
Editorial extensions
If this is right
- Provider releases, control-surface withdrawals, and orchestration changes are governance events: institutions should react with historical replay, revalidation, and reauthorization rather than treat them as routine maintenance.
- No threshold on current outcome reproducibility is a sufficient audit criterion: because $R_O \geq 1/L$ always while the trace record and case information can both be zero, outcome agreement alone can certify a system that has stopped distinguishing its cases.
- Institutions should retain an executable evidence bundle — decision-time inputs, component versions, prompts, tool responses, memory states, causally material handoffs, and policy thresholds — and secure contractual version pinning, escrow, or replay rights that follow audit horizons rather than release cycles.
- Delegation should be evidence-contingent: when retained verification capacity falls below the standard required by exercised authority, the institution should narrow the agent's mandate or route affected decisions to substantive human review.
- Deeper orchestration raises the evidentiary price of the same autonomy: because every compressing handoff destroys reconstructability, keeping the Verifiability Gap closed requires storing strictly more bits for every material compression stage.
Reading between the lines
- A cheap monitoring heuristic follows from Study 3's position-beats-magnitude finding: track each case's distance to the nearest policy cutoff, because cases near a boundary are the population at risk of historical-fidelity failure after any refresh, and audit those cases first.
- Theorem 6 implies a minimal supervisory screen: since a blanket-default system and a perfect system are statistically indistinguishable on within-case re-runs alone, any audit standard should include cross-case differentiation statistics comparing modal verdicts across cases with different correct actions.
- The Markov-chain restriction suggests a concrete research agenda: re-derive the retention bound for non-serial topologies with shared memory or original-case visibility, and measure empirically whether such topologies shrink the Verifiability Gap or merely relocate it.
- Transferred to clinical or public-sector delegation, the same profile would likely bind where the evidentiary standard is externally imposed by courts or regulators; measuring $R(t;\sigma)$ in those settings would test whether the gap's mechanisms are specific to provider-hosted financial models or general to any serial agent pipeline.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that the binding governance constraint on delegating consequential financial decisions to agentic AI is verifiability rather than capability. It defines the Verifiability Gap as the positive shortfall between verification demanded by delegated authority and retained explainability/reproducibility, and develops a multilevel governance theory with seven propositions at firm, regulatory, and network levels. The empirical component consists of three controlled mechanism tests: Study 1 documents provider release drift and withdrawal of replay controls; Study 2 shows orchestration changes decisions, decouples terminal verdicts from execution traces, and produces degenerate outcome reproducibility via output collapse; Study 3 shows a deterministic logistic credit model can perfectly reproduce its current action while failing to recover a historical action. The formal apparatus derives serial contraction and retention lower bounds from the data-processing inequality, and the paper defines reproducibility as a four-target governance profile with a degeneracy diagnostic.
Significance. If the central claim holds, the paper makes a useful contribution to the governance literature by shifting the unit of analysis from model transparency to retained evidence, and by distinguishing four reproducibility targets plus a degeneracy diagnostic. Strengths include the standard information-theoretic scaffold, the explicit bounding of the studies as counterexamples rather than prevalence estimates, the replication package with stored records, and the clean Study 3 demonstration that perfect current replay can coexist with historical-fidelity failure. The main concerns are formal: the Markov-chain assumption in Eq. (3) is violated by the architecture actually tested in Study 2, so the depth/retention theorems are not established for the tested configurations; and the bit-level gap in Corollary 2 does not fully implement the joint-failure construct of Eq. (1). These issues are load-bearing but fixable.
major comments (3)
- [§4.4, Eq. (3); §7.4, Fig. 6; Supplement S1.2] The formal core uses the chain C -> S1 -> ... -> Sd -> Y and justifies it by the rule that no layer beyond the first reads C. In the architecture actually tested in Study 2, the terminal Decider reads the layer-3 Quant/Customer reports and the layer-4 Critic report (Table S2, Fig. 5), so Y = g(S3, S4) and Y is not conditionally independent of S3 given S4. The joint law is therefore not the chain in Eq. (3). Consequently Theorem 1's geometric contraction, Theorem 5's I(C;Y) -> 0 in depth, and Corollary 3's retention-growth conclusion are not established for the tested configurations, and the claim in §7.4/Fig. 6 that the contraction exponent is 19 rather than 50 is not implied. The data-processing argument still gives I(C;Y) <= I(C;S3), so please either revise the architecture so the final action depends only on the last layer, or extend the theory to DAGs with skip edges and state which depth enters the bound.
- [§7.4.1 and Fig. 8a vs Table 7] The one-versus-ten-agent comparison is reported with inconsistent numbers. Section 7.4.1 and Figure 8a state that mean RO fell from 0.903 to 0.747, while Table 7 reports RO = 0.912 for one agent and 0.794 for ten agents, under what appears to be the same condition (released temperature, five repetitions, same model and cases). If the two analyses use different configurations, repetition counts, or execution conditions, state this explicitly; as written, the central quantitative result of Study 2 is ambiguous. Please reconcile the values and confirm that the family-level numbers in Fig. 8a are derivable from the stored records.
- [§4.4, Eq. (5), Corollary 2; Eq. (1)] Corollary 2 defines the Verifiability Gap in bits as [H([T]_sigma|Y) - H(E)]_+, which measures only the reproducibility route. Equation (1) defines the gap through V_dq = V_sigma(E_dq,R_dq), where the explainability route can compensate for reproducibility. If explainability alone is sufficient under sigma, Eq. (1) gives G_dq=0 even when H(E) < H([T]_sigma|Y). As stated, the bit-level corollary does not implement the joint-failure construct in Eq. (1) unless one assumes the explainability route is empty; please clarify that Eq. (5) formalizes a reproducibility-specific sub-gap, or extend the information-theoretic model to include the explanation evidence.
minor comments (3)
- [§3.3] Section 3.3 contains a duplicated phrase: 'They distinguish four targets: four distinct targets' should be 'They distinguish four targets.'
- [Fig. 6(b)] Figure 6(b)'s y-axis label 'count' is ambiguous; label the two series explicitly as 'number of agents' and 'sequential stages'.
- [§7.3.1] The statement that 29 of 32 modal decisions matched the baseline while five distinct cases changed at least once would benefit from a sentence clarifying that cases can change in one release and revert in another.
Circularity Check
No significant circularity: formal core rests on external DPI/entropy bounds; one definitional shortcut in the bits version of the gap.
-
self definitional
[Appendix S2.5, 'A floor on destroyed information']
"An institution retaining only the final action retains none of the destroyed information, so Gdq > 0 by Corollary 2 for every episode in the experiment."
Corollary 2 defines Gdq = [H([T]_sigma|Y) - H(E)]_+, i.e., the gap is defined as retained-evidence shortfall. If the institution retains only the final action, H(E)=0 and any trace variation makes H([T]_sigma|Y)>0, so Gdq>0 follows immediately from the definition. The empirical part independently establishes that traces never repeat, but the 'gap is positive' conclusion is a restatement of the operationalization rather than a separate empirical finding.
full rationale
The paper's formal derivation is largely self-contained and not circular. Theorem 1 and the retention lower bound (Theorem 4) are built on the classical data-processing inequality and entropy identities, with no fitted parameters; the proofs in Appendix S2 proceed from those external results. The empirical studies are presented as controlled mechanism tests and counterexamples (e.g., 'one deviating execution does not support a rate'), not as tuned predictions, so the fitted-input-called-prediction pattern does not apply. The construct V is partly defined in terms of reproducibility, and the paper's own bits-level operationalization makes the gap positive whenever retained evidence is incomplete; that is a definitional shortcut rather than a substantive derivation. A separate non-circularity concern is that the Markov-chain assumption in Eq. (3) is not satisfied by the paper's own Study 2 architecture, where the terminal Decider reads layer-3 reports directly, but this is a validity issue about whether the theorems apply, not a case of the paper deriving its conclusion from its own inputs.
Assumptions & free parameters
free parameters (3)
- Study 3 policy thresholds (approve p<0.45, refer 0.45<=p<0.55, deny p>=0.55)
- Study 3 five-variable logistic feature set
- Study 2 graph width (three specialists per layer) and role roster
assumptions (4)
- standard math Data-processing inequality and strong data-processing inequality for Markov chains.
- standard math Entropy and conditional entropy satisfy the chain rule, and conditioning reduces entropy.
- domain assumption The agentic pipeline factorizes as the Markov chain C -> S1 -> ... -> Sd -> Y, because only the first layer observes the original case and downstream agents observe only transmitted records.
- domain assumption A random seed constrains sampling only conditional on fixed model, instructions, tool responses, and upstream state.
invented entities (1)
-
Verifiability Gap
Cite this review
Pith. "Pith review of Governing Agentic AI in FinTech." pith.science (2026). https://pith.science/paper/XHX5OAPX
@misc{pith2026260811344,
author = {Pith},
title = {Pith review of: Governing Agentic AI in FinTech},
year = {2026},
howpublished = {\url{https://pith.science/paper/XHX5OAPX}},
note = {Machine review of arXiv:2608.11344}
}
read the original abstract
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity Checks for Saliency Maps. arXiv:1810.03292v3, 2018
arXiv 2018
-
[2]
Contestable AI by design: Towards a framework
Kars Alfrink, Ianus Keller, Gerd Kortuem, and Neelke Doorn. Contestable AI by design: Towards a framework. Minds and Machines, 33 0 (4): 0 613--639, 2023
work page 2023
-
[3]
Aleksandre Asatiani, Pekka Malo, Per R dberg Nagb l, Esko Penttinen, Tapani Rinta-Kahila, and Antti Salovaara. Sociotechnical envelopment of artificial intelligence: An approach to organizational deployment of inscrutable artificial intelligence systems. Journal of the Association for Information Systems, 22 0 (2): 0 325--352, 2021. doi:10.17705/1jais.00664
-
[4]
Berk At l, Sarp Aykent, Alexa Chittams, Lisheng Fu, Rebecca J. Passonneau, Evan Radcliffe, Guru Rajan Rajagopal, Adam Sloan, Tomasz Tudrej, Ferhan Ture, Zhe Wu, Lixinyu Xu, and Breck Baldwin. Non-determinism of ``deterministic'' LLM system settings in hosted environments. In Proceedings of the Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP...
arXiv 2025
-
[5]
Aaron Baird and Likoebe M. Maruping. The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45 0 (1): 0 315--341, 2021. doi:10.25300/MISQ/2021/15882
-
[6]
Consumer-lending discrimination in the FinTech Era
Robert Bartlett, Adair Morse, Richard Stanton, and Nancy Wallace. Consumer-lending discrimination in the FinTech Era. Journal of Financial Economics, 143 0 (1): 0 30--56, 2022. doi:10.1016/j.jfineco.2021.05.047
-
[7]
Managing artificial intelligence
Nicholas Berente, Bin Gu, Jan Recker, and Radhika Santhanam. Managing artificial intelligence. MIS Quarterly, 45 0 (3): 0 1433--1450, 2021
2021
-
[8]
On the Rise of FinTechs: Credit Scoring Using Digital Footprints
Tobias Berg, Valentin Burg, Ana Gombović, and Manju Puri. On the Rise of FinTechs: Credit Scoring Using Digital Footprints. The Review of Financial Studies, 33 0 (7): 0 2845--2897, 2020. doi:10.1093/rfs/hhz099
Show all 72 references
-
[9]
Supervisory guidance on model risk management ( SR 11-7)
Board of Governors of the Federal Reserve System and Office of the Comptroller of the Currency . Supervisory guidance on model risk management ( SR 11-7). Technical report, Board of Governors of the Federal Reserve System and Office of the Comptroller of the Currency, 2011
2011
-
[10]
Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jare...
2021 arXiv
-
[11]
Accounting for Variance in Machine Learning Benchmarks
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Naz Sepah, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Dmitriy Serdyuk, Tal Arbel, Chris Pal, Gaël Varoquaux, and Pascal Vincent. Accoun...
2021 arXiv
-
[12]
Cardinal, Sim B
Laura B. Cardinal, Sim B. Sitkin, and Chris P. Long. Balancing and Rebalancing in the Creation and Evolution of Organizational Control. Organization Science, 15 0 (4): 0 411--431, 2004. doi:10.1287/orsc.1040.0084
2004
-
[13]
Harms from Increasingly Agentic Algorithmic Systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstanti...
2023
-
[14]
Reviewable Automated Decision-Making: A Framework for Accountable Algorithmic Systems
Jennifer Cobbe, Michelle Seng Ah Lee, and Jatinder Singh. Reviewable Automated Decision-Making: A Framework for Accountable Algorithmic Systems. arXiv:2102.04201v2, 2021
2021 arXiv
-
[15]
Accountability in algorithmic decision making
Nicholas Diakopoulos. Accountability in algorithmic decision making. Communications of the ACM, 59 0 (2): 0 56--62, 2016. doi:10.1145/2844110
2016 doi
-
[16]
Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez and Been Kim. Towards A Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608v2, 2017
2017 arXiv
-
[17]
The rise of artificial intelligence: Benefits and risks for financial stability
European Central Bank . The rise of artificial intelligence: Benefits and risks for financial stability. Technical report, Financial Stability Review, Special Feature, May 2024
2024
-
[18]
Regulation ( EU ) 2024/1689 laying down harmonised rules on artificial intelligence ( Artificial Intelligence Act )
European Parliament and Council . Regulation ( EU ) 2024/1689 laying down harmonised rules on artificial intelligence ( Artificial Intelligence Act ). Official Journal of the European Union, L series, 12 July 2024
2024
-
[19]
Working and organizing in the age of the learning algorithm
Samer Faraj, Stella Pachidi, and Karla Sayegh. Working and organizing in the age of the learning algorithm. Information and Organization, 28 0 (1): 0 62--70, 2018. doi:10.1016/j.infoandorg.2018.02.005
2018 doi
-
[20]
Explainable machine learning challenge: Home equity line of credit ( HELOC ) dataset
FICO . Explainable machine learning challenge: Home equity line of credit ( HELOC ) dataset. Dataset and challenge documentation, 2018
2018
-
[21]
Emerging trend in GenAI : Observations on AI agents
Financial Industry Regulatory Authority . Emerging trend in GenAI : Observations on AI agents. FINRA, January 27, 2026
2026
-
[22]
Regulatory Notice 24-09: FINRA reminds members of regulatory obligations when using generative artificial intelligence and large language models
Financial Industry Regulatory Authority . Regulatory Notice 24-09: FINRA reminds members of regulatory obligations when using generative artificial intelligence and large language models. Technical report, FINRA, 2024
2024
-
[23]
The financial stability implications of artificial intelligence
Financial Stability Board . The financial stability implications of artificial intelligence. Technical report, Financial Stability Board, November 2024
2024
-
[24]
u gener, J \
Andreas F \"u gener, J \"o rn Grahl, Alok Gupta, and Wolfgang Ketter. Will humans-in-the-loop become borgs? M erits and pitfalls of working with AI . MIS Quarterly, 45 0 (3): 0 1527--1556, 2021
2021
-
[25]
Predictably unequal? the effects of machine learning on credit markets
Andreas Fuster, Paul Goldsmith-Pinkham, Tarun Ramadorai, and Ansgar Walther. Predictably unequal? the effects of machine learning on credit markets. Journal of Finance, 77 0 (1): 0 5--47, 2022
2022
-
[26]
Datasheets for Datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé, and Kate Crawford. Datasheets for Datasets. arXiv:1803.09010v8, 2018
2018 arXiv
-
[27]
Kate Goddard, Abdul Roudsari, and Jeremy C. Wyatt. Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19 0 (1): 0 121--127, 2012
2012
-
[28]
To FinTech and Beyond
Itay Goldstein, Wei Jiang, and G Andrew Karolyi. To FinTech and Beyond. The Review of Financial Studies, 32 0 (5): 0 1647--1661, 2019. doi:10.1093/rfs/hhz025
2019 doi
-
[29]
The Principles and Limits of Algorithm-in-the-Loop Decision Making
Ben Green and Yiling Chen. The Principles and Limits of Algorithm-in-the-Loop Decision Making. Proceedings of the ACM on Human-Computer Interaction, 3 0 (CSCW): 0 1--24, 2019. doi:10.1145/3359152
2019 doi
-
[30]
Explanations from intelligent systems: Theoretical foundations and implications for practice
Shirley Gregor and Izak Benbasat. Explanations from intelligent systems: Theoretical foundations and implications for practice. MIS Quarterly, 23 0 (4): 0 497--530, 1999
1999
-
[31]
Empirical asset pricing via machine learning
Shihao Gu, Bryan Kelly, and Dacheng Xiu. Empirical asset pricing via machine learning. Review of Financial Studies, 33 0 (5): 0 2223--2273, 2020
2020
-
[32]
State of the Art: Reproducibility in Artificial Intelligence
Odd Erik Gundersen and Sigbjørn Kjensmo. State of the Art: Reproducibility in Artificial Intelligence. Proceedings of the AAAI Conference on Artificial Intelligence, 32 0 (1), 2018. doi:10.1609/aaai.v32i1.11503
2018 doi
-
[33]
H. Han. Challenges of reproducible AI in biomedical data science. BMC Medical Genomics, 18 0 (Suppl 1): 0 8, 2025
2025
-
[34]
Alon Jacovi and Yoav Goldberg. Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4198--4205, 2020
2020
-
[35]
Letter to shareholders from Mary Callahan Erdoes: Asset and Wealth Management
JPMorganChase . Letter to shareholders from Mary Callahan Erdoes: Asset and Wealth Management. In 2025 Annual Report, JPMorgan Chase & Co., 2026
2025
-
[36]
Leakage and the reproducibility crisis in machine-learning-based science
Sayash Kapoor and Arvind Narayanan. Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4 0 (9): 0 100804, 2023
2023
-
[37]
Kellogg, Melissa A
Katherine C. Kellogg, Melissa A. Valentine, and Angéle Christin. Algorithms at Work: The New Contested Terrain of Control. Academy of Management Annals, 14 0 (1): 0 366--410, 2020. doi:10.5465/annals.2018.0174
2020
-
[38]
Khandani, Adlar J
Amir E. Khandani, Adlar J. Kim, and Andrew W. Lo. Consumer credit-risk models via machine-learning algorithms. Journal of Banking & Finance, 34 0 (11): 0 2767--2787, 2010. doi:10.1016/j.jbankfin.2010.06.001
2010 doi
-
[39]
Laurie S. Kirsch. Portfolios of Control Modes and IS Project Management. Information Systems Research, 8 0 (3): 0 215--239, 1997. doi:10.1287/isre.8.3.215
1997 doi
-
[40]
When justice is blind to algorithms: Multilayered blackboxing of algorithmic decision-making in the public sector
Charlotta Kronblad, Anna Ess \'e n, and Magnus M \"a hring. When justice is blind to algorithms: Multilayered blackboxing of algorithmic decision-making in the public sector. MIS Quarterly, 48 0 (4): 0 1637--1662, 2024
2024
-
[41]
Agentic artificial intelligence as a new frontier in information systems: Promise, peril, and research opportunities
Naveen Kumar, Xiahua Wei, and Han Zhang. Agentic artificial intelligence as a new frontier in information systems: Promise, peril, and research opportunities. Information & Management, 63 0 (3): 0 104317, 2026. doi:10.1016/j.im.2026.104317
2026
-
[42]
Bowman, and Ethan Perez
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamil \.e Luko s i \=u t \.e , Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, S...
2023 arXiv
-
[43]
Zachary C. Lipton. The Mythos of Model Interpretability. Queue, 16 0 (3): 0 31--57, 2018. doi:10.1145/3236386.3241340
2018
-
[44]
Logg, Julia A
Jennifer M. Logg, Julia A. Minson, and Don A. Moore. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151: 0 90--103, 2019. doi:10.1016/j.obhdp.2018.12.005
2019 doi
-
[45]
A Unified Approach to Interpreting Model Predictions
Scott Lundberg and Su-In Lee. A Unified Approach to Interpreting Model Predictions. arXiv:1705.07874v2, 2017
2017 arXiv
-
[46]
Mastercard unveils Agent Pay, pioneering agentic payments technology to power commerce in the age of AI
Mastercard . Mastercard unveils Agent Pay, pioneering agentic payments technology to power commerce in the age of AI . Company announcement, April 29, 2025
2025
-
[47]
The responsibility gap: Ascribing responsibility for the actions of learning automata
Andreas Matthias. The responsibility gap: Ascribing responsibility for the actions of learning automata. Ethics and Information Technology, 6 0 (3): 0 175--183, 2004
2004
-
[48]
Model Cards for Model Reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220--229...
2019
-
[49]
Alex Murray, Jen Rhymer, and David G. Sirmon. Humans and Technology: Forms of Conjoined Agency in Organizations. Academy of Management Review, 46 0 (3): 0 552--571, 2021. doi:10.5465/amr.2019.0186
2021
-
[50]
William G. Ouchi. Markets, Bureaucracies, and Clans. Administrative Science Quarterly, 25 0 (1): 0 129--141, 1980. doi:10.2307/2392231
1980 doi
-
[51]
Humans and Automation: Use, Misuse, Disuse, Abuse
Raja Parasuraman and Victor Riley. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors: The Journal of the Human Factors and Ergonomics Society, 39 0 (2): 0 230--253, 1997. doi:10.1518/001872097778543886
1997 doi
-
[52]
Roger D. Peng. Reproducible Research in Computational Science. Science, 334 0 (6060): 0 1226--1227, 2011. doi:10.1126/science.1213847
2011 doi
-
[53]
Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program)
Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivi \`e re, Alina Beygelzimer, Florence d'Alch \'e Buc, Emily Fox, and Hugo Larochelle. Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal o...
2019
-
[54]
A Step Toward Quantifying Independently Reproducible Machine Learning Research
Edward Raff. A Step Toward Quantifying Independently Reproducible Machine Learning Research. arXiv:1909.06674v1, 2019
1909 arXiv
-
[55]
Editor's comments: Next-generation digital platforms: Toward human-- AI hybrids
Arun Rai, Panos Constantinides, and Suprateek Sarker. Editor's comments: Next-generation digital platforms: Toward human-- AI hybrids. MIS Quarterly, 43 0 (1): 0 iii--ix, 2019
2019
-
[56]
White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes
Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. arXiv:2001.00973v1, 2020
2001 arXiv
-
[57]
``Why Should I Trust You?'': Explaining the Predictions of Any Classifier
Marco Ribeiro, Sameer Singh, and Carlos Guestrin. ``Why Should I Trust You?'': Explaining the Predictions of Any Classifier. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, pages 97--101, 201...
2016 doi
-
[58]
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1 0 (5): 0 206--215, 2019
2019
-
[59]
Decision provenance: Harnessing data flow for accountable systems
Jatinder Singh, Jennifer Cobbe, and Chris Norval. Decision provenance: Harnessing data flow for accountable systems. IEEE Access, 7: 0 6562--6574, 2019
2019
-
[60]
Skitka, Kathleen L
Linda J. Skitka, Kathleen L. Mosier, and Mark Burdick. Does automation bias decision-making? International Journal of Human-Computer Studies, 51 0 (5): 0 991--1006, 1999. doi:10.1006/ijhc.1999.0252
1999
-
[61]
Authenticated delegation and authorized AI agents
Tobin South, Samuele Marro, Thomas Hardjono, Robert Mahari, Cedric Deslandes Whitney, Dazza Greenwood, Alan Chan, and Alex Pentland. Authenticated delegation and authorized AI agents. arXiv preprint arXiv:2501.09674, 2025
2025 arXiv
-
[62]
Mike H. M. Teodorescu, Lily Morse, Yazeed Awwad, and Gerald C. Kane. Failures of Fairness in Automation Require a Deeper Understanding of Human–ML Augmentation. MIS Quarterly, 45 0 (3): 0 1483--1500, 2021. doi:10.25300/misq/2021/16535
2021 doi
-
[63]
The accountability horizon: An impossibility theorem for governing human-agent collectives
Haileleol Tibebu. The accountability horizon: An impossibility theorem for governing human-agent collectives. arXiv preprint arXiv:2604.07778, 2026
2026 arXiv
-
[64]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023
2023
-
[65]
Find and buy with AI : Visa unveils new era of commerce
Visa . Find and buy with AI : Visa unveils new era of commerce. Company announcement, April 30, 2025
2025
-
[66]
Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR
Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR. In SSRN Electronic Journal, 2017. doi:10.2139/ssrn.3063289
2017 doi
-
[67]
A Survey on Large Language Model based Autonomous Agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. A Survey on Large Language Model based Autonomous Agents. arXiv:2308.11432v7, 2023
2023 arXiv
-
[68]
Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35, pages 24824--24837, 2022. doi:10.52...
2022 doi
-
[69]
Control Configuration and Control Enactment in Information Systems Projects: Review and Expanded Theoretical Framework
Martin Wiener, Magnus Mähring, Ulrich Remus, and Carol Saunders. Control Configuration and Control Enactment in Information Systems Projects: Review and Expanded Theoretical Framework. MIS Quarterly, 40 0 (3): 0 741--774, 2016. doi:10.25300/misq/2016/40.3.11
2016 doi
-
[70]
What to account for when accounting for algorithms
Maranke Wieringa. What to account for when accounting for algorithms. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 1--18, 2020. doi:10.1145/3351095.3372833
2020
-
[71]
TradingAgents : Multi-agents LLM financial trading framework
Yijia Xiao, Edward Sun, Di Luo, and Wei Wang. TradingAgents : Multi-agents LLM financial trading framework. arXiv preprint arXiv:2412.20138, 2024
2024 arXiv
-
[72]
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629v3, 2022
2022 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.