Pith. sign in

REVIEW 3 major objections 3 minor 72 references

Governing Agentic AI in FinTech

T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper argues that the binding constraint on delegating consequential financial decisions to agentic AI is verifiability, not capability: an agent can stay accurate and stable while the institution cannot substantiate how it acted.

desk verdict Serious and useful paper on agentic AI governance, but its formal core rests on a Markov-chain assumption that the paper's own Study 2 architecture violates. read the letter →

arxiv 2608.11344 v2 pith:XHX5OAPX submitted 2026-08-11 cs.CY cs.AIq-fin.RM

classification cs.CYcs.AIq-fin.RM
keywords agenticAIgovernanceVerifiabilityGapdelegatedauthorityreproducibilityexplainabilityFinTechoutputcollapseevidence-contingentdelegation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Agentic AI in finance is new not because it is opaque — banks have validated opaque models for decades by re-running them — but because the empirical remedy for opacity may no longer work. The paper argues that the binding constraint on delegating consequential decisions is verifiability, not capability, and defines the Verifiability Gap as the shortfall between the verification that delegated authority demands and the explainability and reproducibility retained after a decision. Three controlled studies show the same gap opening from three directions: provider releases change historical financial actions while withdrawing the controls needed to replay them; orchestration changes the operative decision rule while no execution record ever repeats; and a deterministic credit model can reproduce its current action perfectly while failing to recover a historical one. The upshot is evidence-contingent delegation: an agent's authority is defensible only while retained evidence substantiates how it was exercised, and when that evidence falls short the institution should narrow authority or route decisions to substantive human review. A reader should care because banks, card networks, and regulators are already moving loan decisions, trades, and compliance dispositions into precisely these systems.

What carries the argument

The central object is the Verifiability Gap, a shortfall in retained verification capacity indexed to a verifier, an evidentiary standard, and an audit lag, together with the reproducibility profile $R(t;\sigma) = (R_O, R_H, R_P, R_T; D_V)$ that separates current outcome, historical outcome, material process, and exact trace reproducibility and tracks verdict differentiation $D_V$ as a degeneracy diagnostic. The argument is carried by a serial-pipeline information model in which only the first layer observes the original case and the joint law factorizes as a Markov chain $C \to S_1 \to S_2 \to \cdots \to S_d \to Y$; on that chain the paper proves that discrimination contracts geometrically in depth (Theorem 1), that $\sigma$-reversibility is equivalent to equality in the data-processing inequality (Theorem 2), that deterministic summarization is incompatible with reversibility (Theorem 3), and that reversibility requires retaining at least $H([T]_\sigma \mid Y)$ bits of evidence (Theorem 4). These results convert the verbal claim that serial handoffs widen the gap into a bit-level accounting of what an institution must keep.

What would settle it

Run the Study 1 release-timeline and Study 2 orchestration protocols on a provider that exposes a dated pinned endpoint, a random seed, and temperature control: if baseline-modal historical fidelity stays at 1.000 across releases and exact trace reproducibility $R_T$ reaches 1 at any depth beyond a single agent, the provider-dominance and serial-contraction claims would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that an agentic AI system can remain highly capable, accurate, and behaviorally stable while the responsible institution becomes entirely unable to substantiate how that system exercised its delegated authority. It formalizes this as the Verifiability Gap, $G_{dq} = [\rho_\sigma(A_d) - V_{dq}]^+$, the positive shortfall between the verification required by delegated authority and the verification capacity retained in explainability and reproducibility, indexed to a named verifier, an evidentiary standard, and an audit lag. The formal core is a set of information-theoretic results on serial agent pipelines: discrimination $I(C;Y)$ contracts geometrically with pipeline depth via the data-processing inequality, a stage is $\sigma$-reversible exactly when it throws nothing away, a stage either summarizes or stays auditable but cannot do both, and the retained evidence must satisfy $H(E) \geq H([T]_\sigma \mid Y)$ — the institution must store at least as many bits as its pipeline destroyed. Empirically, three studies identify distinct mechanisms: provider releases silently shift risk posture while withdrawing temperature, top-p, top-k, and seed controls; orchestration acts as a latent policy layer in which no execution record repeats at any scale and apparent stability can be degenerate output collapse; and a preserved logistic credit model achieves perfect current reproducibility while failing to reproduce a historical decision that crossed a policy threshold. The conclusion is that reproducibility is a governance profile, not a scalar, and that delegated authority remains defensible only while retained evidence substantiates its exercise.

Load-bearing premise

The formal contraction and retention theorems assume that the pipeline factorizes as a Markov chain — only the first layer observes the original case and every downstream agent sees only the records transmitted along designated edges — so if production agentic systems let later agents see the original case, live tools, or shared memory, the serial information loss and the retention lower bound may not hold.

Editorial extensions

If this is right

  • Provider releases, control-surface withdrawals, and orchestration changes are governance events: institutions should react with historical replay, revalidation, and reauthorization rather than treat them as routine maintenance.
  • No threshold on current outcome reproducibility is a sufficient audit criterion: because $R_O \geq 1/L$ always while the trace record and case information can both be zero, outcome agreement alone can certify a system that has stopped distinguishing its cases.
  • Institutions should retain an executable evidence bundle — decision-time inputs, component versions, prompts, tool responses, memory states, causally material handoffs, and policy thresholds — and secure contractual version pinning, escrow, or replay rights that follow audit horizons rather than release cycles.
  • Delegation should be evidence-contingent: when retained verification capacity falls below the standard required by exercised authority, the institution should narrow the agent's mandate or route affected decisions to substantive human review.
  • Deeper orchestration raises the evidentiary price of the same autonomy: because every compressing handoff destroys reconstructability, keeping the Verifiability Gap closed requires storing strictly more bits for every material compression stage.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A cheap monitoring heuristic follows from Study 3's position-beats-magnitude finding: track each case's distance to the nearest policy cutoff, because cases near a boundary are the population at risk of historical-fidelity failure after any refresh, and audit those cases first.
  • Theorem 6 implies a minimal supervisory screen: since a blanket-default system and a perfect system are statistically indistinguishable on within-case re-runs alone, any audit standard should include cross-case differentiation statistics comparing modal verdicts across cases with different correct actions.
  • The Markov-chain restriction suggests a concrete research agenda: re-derive the retention bound for non-serial topologies with shared memory or original-case visibility, and measure empirically whether such topologies shrink the Verifiability Gap or merely relocate it.
  • Transferred to clinical or public-sector delegation, the same profile would likely bind where the evidentiary standard is externally imposed by courts or regulators; measuring $R(t;\sigma)$ in those settings would test whether the gap's mechanisms are specific to provider-hosted financial models or general to any serial agent pipeline.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This paper argues that the binding governance constraint on delegating consequential financial decisions to agentic AI is verifiability rather than capability. It defines the Verifiability Gap as the positive shortfall between verification demanded by delegated authority and retained explainability/reproducibility, and develops a multilevel governance theory with seven propositions at firm, regulatory, and network levels. The empirical component consists of three controlled mechanism tests: Study 1 documents provider release drift and withdrawal of replay controls; Study 2 shows orchestration changes decisions, decouples terminal verdicts from execution traces, and produces degenerate outcome reproducibility via output collapse; Study 3 shows a deterministic logistic credit model can perfectly reproduce its current action while failing to recover a historical action. The formal apparatus derives serial contraction and retention lower bounds from the data-processing inequality, and the paper defines reproducibility as a four-target governance profile with a degeneracy diagnostic.

Significance. If the central claim holds, the paper makes a useful contribution to the governance literature by shifting the unit of analysis from model transparency to retained evidence, and by distinguishing four reproducibility targets plus a degeneracy diagnostic. Strengths include the standard information-theoretic scaffold, the explicit bounding of the studies as counterexamples rather than prevalence estimates, the replication package with stored records, and the clean Study 3 demonstration that perfect current replay can coexist with historical-fidelity failure. The main concerns are formal: the Markov-chain assumption in Eq. (3) is violated by the architecture actually tested in Study 2, so the depth/retention theorems are not established for the tested configurations; and the bit-level gap in Corollary 2 does not fully implement the joint-failure construct of Eq. (1). These issues are load-bearing but fixable.

major comments (3)
  1. [§4.4, Eq. (3); §7.4, Fig. 6; Supplement S1.2] The formal core uses the chain C -> S1 -> ... -> Sd -> Y and justifies it by the rule that no layer beyond the first reads C. In the architecture actually tested in Study 2, the terminal Decider reads the layer-3 Quant/Customer reports and the layer-4 Critic report (Table S2, Fig. 5), so Y = g(S3, S4) and Y is not conditionally independent of S3 given S4. The joint law is therefore not the chain in Eq. (3). Consequently Theorem 1's geometric contraction, Theorem 5's I(C;Y) -> 0 in depth, and Corollary 3's retention-growth conclusion are not established for the tested configurations, and the claim in §7.4/Fig. 6 that the contraction exponent is 19 rather than 50 is not implied. The data-processing argument still gives I(C;Y) <= I(C;S3), so please either revise the architecture so the final action depends only on the last layer, or extend the theory to DAGs with skip edges and state which depth enters the bound.
  2. [§7.4.1 and Fig. 8a vs Table 7] The one-versus-ten-agent comparison is reported with inconsistent numbers. Section 7.4.1 and Figure 8a state that mean RO fell from 0.903 to 0.747, while Table 7 reports RO = 0.912 for one agent and 0.794 for ten agents, under what appears to be the same condition (released temperature, five repetitions, same model and cases). If the two analyses use different configurations, repetition counts, or execution conditions, state this explicitly; as written, the central quantitative result of Study 2 is ambiguous. Please reconcile the values and confirm that the family-level numbers in Fig. 8a are derivable from the stored records.
  3. [§4.4, Eq. (5), Corollary 2; Eq. (1)] Corollary 2 defines the Verifiability Gap in bits as [H([T]_sigma|Y) - H(E)]_+, which measures only the reproducibility route. Equation (1) defines the gap through V_dq = V_sigma(E_dq,R_dq), where the explainability route can compensate for reproducibility. If explainability alone is sufficient under sigma, Eq. (1) gives G_dq=0 even when H(E) < H([T]_sigma|Y). As stated, the bit-level corollary does not implement the joint-failure construct in Eq. (1) unless one assumes the explainability route is empty; please clarify that Eq. (5) formalizes a reproducibility-specific sub-gap, or extend the information-theoretic model to include the explanation evidence.
minor comments (3)
  1. [§3.3] Section 3.3 contains a duplicated phrase: 'They distinguish four targets: four distinct targets' should be 'They distinguish four targets.'
  2. [Fig. 6(b)] Figure 6(b)'s y-axis label 'count' is ambiguous; label the two series explicitly as 'number of agents' and 'sequential stages'.
  3. [§7.3.1] The statement that 29 of 32 modal decisions matched the baseline while five distinct cases changed at least once would benefit from a sentence clarifying that cases can change in one release and revert in another.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity: formal core rests on external DPI/entropy bounds; one definitional shortcut in the bits version of the gap.

  1. self definitional [Appendix S2.5, 'A floor on destroyed information']
    "An institution retaining only the final action retains none of the destroyed information, so Gdq > 0 by Corollary 2 for every episode in the experiment."

    Corollary 2 defines Gdq = [H([T]_sigma|Y) - H(E)]_+, i.e., the gap is defined as retained-evidence shortfall. If the institution retains only the final action, H(E)=0 and any trace variation makes H([T]_sigma|Y)>0, so Gdq>0 follows immediately from the definition. The empirical part independently establishes that traces never repeat, but the 'gap is positive' conclusion is a restatement of the operationalization rather than a separate empirical finding.

full rationale

The paper's formal derivation is largely self-contained and not circular. Theorem 1 and the retention lower bound (Theorem 4) are built on the classical data-processing inequality and entropy identities, with no fitted parameters; the proofs in Appendix S2 proceed from those external results. The empirical studies are presented as controlled mechanism tests and counterexamples (e.g., 'one deviating execution does not support a rate'), not as tuned predictions, so the fitted-input-called-prediction pattern does not apply. The construct V is partly defined in terms of reproducibility, and the paper's own bits-level operationalization makes the gap positive whenever retained evidence is incomplete; that is a definitional shortcut rather than a substantive derivation. A separate non-circularity concern is that the Markov-chain assumption in Eq. (3) is not satisfied by the paper's own Study 2 architecture, where the terminal Decider reads layer-3 reports directly, but this is a validity issue about whether the theorems apply, not a case of the paper deriving its conclusion from its own inputs.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The formal core rests on standard information-theoretic inequalities and on the Markov-chain information-access rule of the controlled architecture. The empirical demonstrations use hand-chosen experimental parameters such as thresholds, feature sets, and graph width, but the central theoretical results do not depend on their exact values. The Verifiability Gap is a theoretical construct; it is not independently measurable in this paper.

free parameters (3)
  • Study 3 policy thresholds (approve p<0.45, refer 0.45<=p<0.55, deny p>=0.55)
    Hand-chosen to expose threshold-crossing. The authors report robustness across 39 admissible cutoff pairs, so the qualitative result does not depend on the exact values.
  • Study 3 five-variable logistic feature set
    Selected from the HELOC dataset without a formal selection procedure; different variable sets could change which cases move across thresholds, though the mechanism would remain.
  • Study 2 graph width (three specialists per layer) and role roster
    Researcher-designed architecture; the direction of conservative outcome shifts is partly conditioned by the roster, which the authors acknowledge.
assumptions (4)
  • standard math Data-processing inequality and strong data-processing inequality for Markov chains.
    Used in Theorem 1 to bound serial contraction of mutual information; standard results, stated without proof.
  • standard math Entropy and conditional entropy satisfy the chain rule, and conditioning reduces entropy.
    Used in Theorem 4 and the retention lower bound; standard information-theoretic facts.
  • domain assumption The agentic pipeline factorizes as the Markov chain C -> S1 -> ... -> Sd -> Y, because only the first layer observes the original case and downstream agents observe only transmitted records.
    This information-access rule is enforced in the authors' controlled architecture but is not shown to hold for deployed agentic systems; the contraction and reversibility theorems are conditional on it.
  • domain assumption A random seed constrains sampling only conditional on fixed model, instructions, tool responses, and upstream state.
    Used throughout to argue that seed control cannot repair changed execution environments; plausible and consistent with sampling theory, but treated as background.
invented entities (1)
  • Verifiability Gap
    purpose: Formal construct quantifying the shortfall between verification demanded by delegated authority and verification capacity retained after a decision (Equation 1).
    A theoretical construct defined via Equation 1; the paper measures reproducibility mechanisms rather than the construct itself, so no standalone falsifiable handle is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Governing Agentic AI in FinTech." pith.science (2026). https://pith.science/paper/XHX5OAPX

@misc{pith2026260811344,
  author       = {Pith},
  title        = {Pith review of: Governing Agentic AI in FinTech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XHX5OAPX}},
  note         = {Machine review of arXiv:2608.11344}
}
read the original abstract

Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.

Figures

Figures reproduced from arXiv: 2608.11344 by the authors.

Figure 1
Figure 1. Roadmap of the paper’s argument: from the phenomenon (agentic AI that acts in finance), [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Research design. Phenomenon specification defines the analytical boundary in FinTech; [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Conditions under which the Verifiability Gap becomes binding. The gap is most conse [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Serial dependence in a FinTech loan-denial workflow. Causally dependent handoffs can [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Controlled multi-agent governance test. Only Retrieval, Market, and Policy observe the [PITH_FULL_IMAGE:figures/full_fig_p026_5.png]
Figure 6
Figure 6. Figure 6: Agent count and chain depth are different quantities. [PITH_FULL_IMAGE:figures/full_fig_p027_6.png]
Figure 7
Figure 7. Figure 7: Provider dominance over the audit surface. [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 8
Figure 8. Figure 8: Orchestration changes both the decisions being reproduced and the records supporting [PITH_FULL_IMAGE:figures/full_fig_p031_8.png]
Figure 9
Figure 9. Figure 9: Why outcome reproducibility can recover while the underlying process remains nonre [PITH_FULL_IMAGE:figures/full_fig_p032_9.png]
Figure 10
Figure 10. Figure 10: The same sweep at two model scales. (a) Current outcome reproducibility. Both scales fall, then recover at forty agents; the larger model is lower throughout beyond ten agents. (b) Verdict differentiation. Both collapse at forty agents. The smaller model enters the ra…
Figure 11
Figure 11. Figure 11: What makes a financial action change under a model refresh. [PITH_FULL_IMAGE:figures/full_fig_p036_11.png]
Figure 12
Figure 12. Figure 12: A decision-level Verifiability Gap under current-only retention. Version 1 returns [PITH_FULL_IMAGE:figures/full_fig_p038_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 37 canonical work pages

  1. [1]

    Sanity Checks for Saliency Maps

    Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity Checks for Saliency Maps. arXiv:1810.03292v3, 2018

  2. [2]

    Contestable AI by design: Towards a framework

    Kars Alfrink, Ianus Keller, Gerd Kortuem, and Neelke Doorn. Contestable AI by design: Towards a framework. Minds and Machines, 33 0 (4): 0 613--639, 2023

  3. [3]

    Sociotechnical envelopment of artificial intelligence: An approach to organizational deployment of inscrutable artificial intelligence systems

    Aleksandre Asatiani, Pekka Malo, Per R dberg Nagb l, Esko Penttinen, Tapani Rinta-Kahila, and Antti Salovaara. Sociotechnical envelopment of artificial intelligence: An approach to organizational deployment of inscrutable artificial intelligence systems. Journal of the Association for Information Systems, 22 0 (2): 0 325--352, 2021. doi:10.17705/1jais.00664

  4. [4]

    Passonneau, Evan Radcliffe, Guru Rajan Rajagopal, Adam Sloan, Tomasz Tudrej, Ferhan Ture, Zhe Wu, Lixinyu Xu, and Breck Baldwin

    Berk At l, Sarp Aykent, Alexa Chittams, Lisheng Fu, Rebecca J. Passonneau, Evan Radcliffe, Guru Rajan Rajagopal, Adam Sloan, Tomasz Tudrej, Ferhan Ture, Zhe Wu, Lixinyu Xu, and Breck Baldwin. Non-determinism of ``deterministic'' LLM system settings in hosted environments. In Proceedings of the Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP...

  5. [5]

    Maruping

    Aaron Baird and Likoebe M. Maruping. The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45 0 (1): 0 315--341, 2021. doi:10.25300/MISQ/2021/15882

  6. [6]

    Consumer-lending discrimination in the FinTech Era

    Robert Bartlett, Adair Morse, Richard Stanton, and Nancy Wallace. Consumer-lending discrimination in the FinTech Era. Journal of Financial Economics, 143 0 (1): 0 30--56, 2022. doi:10.1016/j.jfineco.2021.05.047

  7. [7]

    Managing artificial intelligence

    Nicholas Berente, Bin Gu, Jan Recker, and Radhika Santhanam. Managing artificial intelligence. MIS Quarterly, 45 0 (3): 0 1433--1450, 2021

  8. [8]

    On the Rise of FinTechs: Credit Scoring Using Digital Footprints

    Tobias Berg, Valentin Burg, Ana Gombović, and Manju Puri. On the Rise of FinTechs: Credit Scoring Using Digital Footprints. The Review of Financial Studies, 33 0 (7): 0 2845--2897, 2020. doi:10.1093/rfs/hhz099

Show all 72 references
  1. [9]

    Supervisory guidance on model risk management ( SR 11-7)

    Board of Governors of the Federal Reserve System and Office of the Comptroller of the Currency . Supervisory guidance on model risk management ( SR 11-7). Technical report, Board of Governors of the Federal Reserve System and Office of the Comptroller of the Currency, 2011

  2. [10]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jare...

  3. [11]

    Accounting for Variance in Machine Learning Benchmarks

    Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Naz Sepah, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Dmitriy Serdyuk, Tal Arbel, Chris Pal, Gaël Varoquaux, and Pascal Vincent. Accoun...

  4. [12]

    Cardinal, Sim B

    Laura B. Cardinal, Sim B. Sitkin, and Chris P. Long. Balancing and Rebalancing in the Creation and Evolution of Organizational Control. Organization Science, 15 0 (4): 0 411--431, 2004. doi:10.1287/orsc.1040.0084

  5. [13]

    Harms from Increasingly Agentic Algorithmic Systems

    Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstanti...

  6. [14]

    Reviewable Automated Decision-Making: A Framework for Accountable Algorithmic Systems

    Jennifer Cobbe, Michelle Seng Ah Lee, and Jatinder Singh. Reviewable Automated Decision-Making: A Framework for Accountable Algorithmic Systems. arXiv:2102.04201v2, 2021

  7. [15]

    Accountability in algorithmic decision making

    Nicholas Diakopoulos. Accountability in algorithmic decision making. Communications of the ACM, 59 0 (2): 0 56--62, 2016. doi:10.1145/2844110

  8. [16]

    Towards A Rigorous Science of Interpretable Machine Learning

    Finale Doshi-Velez and Been Kim. Towards A Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608v2, 2017

  9. [17]

    The rise of artificial intelligence: Benefits and risks for financial stability

    European Central Bank . The rise of artificial intelligence: Benefits and risks for financial stability. Technical report, Financial Stability Review, Special Feature, May 2024

  10. [18]

    Regulation ( EU ) 2024/1689 laying down harmonised rules on artificial intelligence ( Artificial Intelligence Act )

    European Parliament and Council . Regulation ( EU ) 2024/1689 laying down harmonised rules on artificial intelligence ( Artificial Intelligence Act ). Official Journal of the European Union, L series, 12 July 2024

  11. [19]

    Working and organizing in the age of the learning algorithm

    Samer Faraj, Stella Pachidi, and Karla Sayegh. Working and organizing in the age of the learning algorithm. Information and Organization, 28 0 (1): 0 62--70, 2018. doi:10.1016/j.infoandorg.2018.02.005

  12. [20]

    Explainable machine learning challenge: Home equity line of credit ( HELOC ) dataset

    FICO . Explainable machine learning challenge: Home equity line of credit ( HELOC ) dataset. Dataset and challenge documentation, 2018

  13. [21]

    Emerging trend in GenAI : Observations on AI agents

    Financial Industry Regulatory Authority . Emerging trend in GenAI : Observations on AI agents. FINRA, January 27, 2026

  14. [22]

    Regulatory Notice 24-09: FINRA reminds members of regulatory obligations when using generative artificial intelligence and large language models

    Financial Industry Regulatory Authority . Regulatory Notice 24-09: FINRA reminds members of regulatory obligations when using generative artificial intelligence and large language models. Technical report, FINRA, 2024

  15. [23]

    The financial stability implications of artificial intelligence

    Financial Stability Board . The financial stability implications of artificial intelligence. Technical report, Financial Stability Board, November 2024

  16. [24]

    u gener, J \

    Andreas F \"u gener, J \"o rn Grahl, Alok Gupta, and Wolfgang Ketter. Will humans-in-the-loop become borgs? M erits and pitfalls of working with AI . MIS Quarterly, 45 0 (3): 0 1527--1556, 2021

  17. [25]

    Predictably unequal? the effects of machine learning on credit markets

    Andreas Fuster, Paul Goldsmith-Pinkham, Tarun Ramadorai, and Ansgar Walther. Predictably unequal? the effects of machine learning on credit markets. Journal of Finance, 77 0 (1): 0 5--47, 2022

  18. [26]

    Datasheets for Datasets

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé, and Kate Crawford. Datasheets for Datasets. arXiv:1803.09010v8, 2018

  19. [27]

    Kate Goddard, Abdul Roudsari, and Jeremy C. Wyatt. Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19 0 (1): 0 121--127, 2012

  20. [28]

    To FinTech and Beyond

    Itay Goldstein, Wei Jiang, and G Andrew Karolyi. To FinTech and Beyond. The Review of Financial Studies, 32 0 (5): 0 1647--1661, 2019. doi:10.1093/rfs/hhz025

  21. [29]

    The Principles and Limits of Algorithm-in-the-Loop Decision Making

    Ben Green and Yiling Chen. The Principles and Limits of Algorithm-in-the-Loop Decision Making. Proceedings of the ACM on Human-Computer Interaction, 3 0 (CSCW): 0 1--24, 2019. doi:10.1145/3359152

  22. [30]

    Explanations from intelligent systems: Theoretical foundations and implications for practice

    Shirley Gregor and Izak Benbasat. Explanations from intelligent systems: Theoretical foundations and implications for practice. MIS Quarterly, 23 0 (4): 0 497--530, 1999

  23. [31]

    Empirical asset pricing via machine learning

    Shihao Gu, Bryan Kelly, and Dacheng Xiu. Empirical asset pricing via machine learning. Review of Financial Studies, 33 0 (5): 0 2223--2273, 2020

  24. [32]

    State of the Art: Reproducibility in Artificial Intelligence

    Odd Erik Gundersen and Sigbjørn Kjensmo. State of the Art: Reproducibility in Artificial Intelligence. Proceedings of the AAAI Conference on Artificial Intelligence, 32 0 (1), 2018. doi:10.1609/aaai.v32i1.11503

  25. [33]

    H. Han. Challenges of reproducible AI in biomedical data science. BMC Medical Genomics, 18 0 (Suppl 1): 0 8, 2025

  26. [34]

    Alon Jacovi and Yoav Goldberg. Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4198--4205, 2020

  27. [35]

    Letter to shareholders from Mary Callahan Erdoes: Asset and Wealth Management

    JPMorganChase . Letter to shareholders from Mary Callahan Erdoes: Asset and Wealth Management. In 2025 Annual Report, JPMorgan Chase & Co., 2026

  28. [36]

    Leakage and the reproducibility crisis in machine-learning-based science

    Sayash Kapoor and Arvind Narayanan. Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4 0 (9): 0 100804, 2023

  29. [37]

    Kellogg, Melissa A

    Katherine C. Kellogg, Melissa A. Valentine, and Angéle Christin. Algorithms at Work: The New Contested Terrain of Control. Academy of Management Annals, 14 0 (1): 0 366--410, 2020. doi:10.5465/annals.2018.0174

  30. [38]

    Khandani, Adlar J

    Amir E. Khandani, Adlar J. Kim, and Andrew W. Lo. Consumer credit-risk models via machine-learning algorithms. Journal of Banking & Finance, 34 0 (11): 0 2767--2787, 2010. doi:10.1016/j.jbankfin.2010.06.001

  31. [39]

    Laurie S. Kirsch. Portfolios of Control Modes and IS Project Management. Information Systems Research, 8 0 (3): 0 215--239, 1997. doi:10.1287/isre.8.3.215

  32. [40]

    When justice is blind to algorithms: Multilayered blackboxing of algorithmic decision-making in the public sector

    Charlotta Kronblad, Anna Ess \'e n, and Magnus M \"a hring. When justice is blind to algorithms: Multilayered blackboxing of algorithmic decision-making in the public sector. MIS Quarterly, 48 0 (4): 0 1637--1662, 2024

  33. [41]

    Agentic artificial intelligence as a new frontier in information systems: Promise, peril, and research opportunities

    Naveen Kumar, Xiahua Wei, and Han Zhang. Agentic artificial intelligence as a new frontier in information systems: Promise, peril, and research opportunities. Information & Management, 63 0 (3): 0 104317, 2026. doi:10.1016/j.im.2026.104317

  34. [42]

    Bowman, and Ethan Perez

    Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamil \.e Luko s i \=u t \.e , Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, S...

  35. [43]

    Zachary C. Lipton. The Mythos of Model Interpretability. Queue, 16 0 (3): 0 31--57, 2018. doi:10.1145/3236386.3241340

  36. [44]

    Logg, Julia A

    Jennifer M. Logg, Julia A. Minson, and Don A. Moore. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151: 0 90--103, 2019. doi:10.1016/j.obhdp.2018.12.005

  37. [45]

    A Unified Approach to Interpreting Model Predictions

    Scott Lundberg and Su-In Lee. A Unified Approach to Interpreting Model Predictions. arXiv:1705.07874v2, 2017

  38. [46]

    Mastercard unveils Agent Pay, pioneering agentic payments technology to power commerce in the age of AI

    Mastercard . Mastercard unveils Agent Pay, pioneering agentic payments technology to power commerce in the age of AI . Company announcement, April 29, 2025

  39. [47]

    The responsibility gap: Ascribing responsibility for the actions of learning automata

    Andreas Matthias. The responsibility gap: Ascribing responsibility for the actions of learning automata. Ethics and Information Technology, 6 0 (3): 0 175--183, 2004

  40. [48]

    Model Cards for Model Reporting

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220--229...

  41. [49]

    Alex Murray, Jen Rhymer, and David G. Sirmon. Humans and Technology: Forms of Conjoined Agency in Organizations. Academy of Management Review, 46 0 (3): 0 552--571, 2021. doi:10.5465/amr.2019.0186

  42. [50]

    William G. Ouchi. Markets, Bureaucracies, and Clans. Administrative Science Quarterly, 25 0 (1): 0 129--141, 1980. doi:10.2307/2392231

  43. [51]

    Humans and Automation: Use, Misuse, Disuse, Abuse

    Raja Parasuraman and Victor Riley. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors: The Journal of the Human Factors and Ergonomics Society, 39 0 (2): 0 230--253, 1997. doi:10.1518/001872097778543886

  44. [52]

    Roger D. Peng. Reproducible Research in Computational Science. Science, 334 0 (6060): 0 1226--1227, 2011. doi:10.1126/science.1213847

  45. [53]

    Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program)

    Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivi \`e re, Alina Beygelzimer, Florence d'Alch \'e Buc, Emily Fox, and Hugo Larochelle. Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal o...

  46. [54]

    A Step Toward Quantifying Independently Reproducible Machine Learning Research

    Edward Raff. A Step Toward Quantifying Independently Reproducible Machine Learning Research. arXiv:1909.06674v1, 2019

  47. [55]

    Editor's comments: Next-generation digital platforms: Toward human-- AI hybrids

    Arun Rai, Panos Constantinides, and Suprateek Sarker. Editor's comments: Next-generation digital platforms: Toward human-- AI hybrids. MIS Quarterly, 43 0 (1): 0 iii--ix, 2019

  48. [56]

    White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes

    Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. arXiv:2001.00973v1, 2020

  49. [57]

    ``Why Should I Trust You?'': Explaining the Predictions of Any Classifier

    Marco Ribeiro, Sameer Singh, and Carlos Guestrin. ``Why Should I Trust You?'': Explaining the Predictions of Any Classifier. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, pages 97--101, 201...

  50. [58]

    Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead

    Cynthia Rudin. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1 0 (5): 0 206--215, 2019

  51. [59]

    Decision provenance: Harnessing data flow for accountable systems

    Jatinder Singh, Jennifer Cobbe, and Chris Norval. Decision provenance: Harnessing data flow for accountable systems. IEEE Access, 7: 0 6562--6574, 2019

  52. [60]

    Skitka, Kathleen L

    Linda J. Skitka, Kathleen L. Mosier, and Mark Burdick. Does automation bias decision-making? International Journal of Human-Computer Studies, 51 0 (5): 0 991--1006, 1999. doi:10.1006/ijhc.1999.0252

  53. [61]

    Authenticated delegation and authorized AI agents

    Tobin South, Samuele Marro, Thomas Hardjono, Robert Mahari, Cedric Deslandes Whitney, Dazza Greenwood, Alan Chan, and Alex Pentland. Authenticated delegation and authorized AI agents. arXiv preprint arXiv:2501.09674, 2025

  54. [62]

    Mike H. M. Teodorescu, Lily Morse, Yazeed Awwad, and Gerald C. Kane. Failures of Fairness in Automation Require a Deeper Understanding of Human–ML Augmentation. MIS Quarterly, 45 0 (3): 0 1483--1500, 2021. doi:10.25300/misq/2021/16535

  55. [63]

    The accountability horizon: An impossibility theorem for governing human-agent collectives

    Haileleol Tibebu. The accountability horizon: An impossibility theorem for governing human-agent collectives. arXiv preprint arXiv:2604.07778, 2026

  56. [64]

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023

  57. [65]

    Find and buy with AI : Visa unveils new era of commerce

    Visa . Find and buy with AI : Visa unveils new era of commerce. Company announcement, April 30, 2025

  58. [66]

    Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR

    Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR. In SSRN Electronic Journal, 2017. doi:10.2139/ssrn.3063289

  59. [67]

    A Survey on Large Language Model based Autonomous Agents

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. A Survey on Large Language Model based Autonomous Agents. arXiv:2308.11432v7, 2023

  60. [68]

    Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35, pages 24824--24837, 2022. doi:10.52...

  61. [69]

    Control Configuration and Control Enactment in Information Systems Projects: Review and Expanded Theoretical Framework

    Martin Wiener, Magnus Mähring, Ulrich Remus, and Carol Saunders. Control Configuration and Control Enactment in Information Systems Projects: Review and Expanded Theoretical Framework. MIS Quarterly, 40 0 (3): 0 741--774, 2016. doi:10.25300/misq/2016/40.3.11

  62. [70]

    What to account for when accounting for algorithms

    Maranke Wieringa. What to account for when accounting for algorithms. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 1--18, 2020. doi:10.1145/3351095.3372833

  63. [71]

    TradingAgents : Multi-agents LLM financial trading framework

    Yijia Xiao, Edward Sun, Di Luo, and Wei Wang. TradingAgents : Multi-agents LLM financial trading framework. arXiv preprint arXiv:2412.20138, 2024

  64. [72]

    ReAct: Synergizing Reasoning and Acting in Language Models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629v3, 2022

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.