Pith. sign in

REVIEW 2 major objections 5 minor 35 references

When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A certificate built from fresh randomized audits can guarantee that automated fraud decisions respect their declared current-risk limits with probability at least $1-\delta$.

desk verdict A careful, honest decision-support framework for fraud automation; the safety guarantee rests on an unfalsifiable temporal-stability assumption, but the paper says so and quantifies the cost. read the letter →

arxiv 2608.08577 v1 pith:QNY2AJHA submitted 2026-08-09 cs.LG stat.ML

classification cs.LGstat.ML
keywords financialfraudhuman–AIdecisionallocationdelayedfeedbackauditgovernanceriskcertificationreviewworkloadsupportfinite-samplecontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Fraud operations must decide which events a model may approve or block automatically and which must go to human analysts, but the labels that would reveal whether those actions are safe arrive late and only for audited cases. This paper argues that automation should be treated as an authorization decision, not just a score threshold: the question is whether the available evidence is current and representative enough to delegate an action. It proposes freshness-constrained audit capacity (FCAC), which certifies candidate approve and block regions only when mature randomized audits, adjusted for evidence age, fit inside a prespecified action-risk limit, and sends everything else to review. The central result is that under a governance-supplied bound on how fast risk may change, FCAC controls unsafe authorization simultaneously across time, thresholds, and actions: with probability at least $1-\delta$, no certified automated region exceeds its declared current-risk limit. This matters because it turns delayed, selectively observed labels into a defensible basis for delegating decisions while also accounting for the analyst workload that auditing itself consumes.

What carries the argument

The load-bearing mechanism is the fixed-limit KL certification index $U_{\mathrm{KL}}(k,n,\delta)$, the largest risk $p \ge k/n$ satisfying $n\,\mathrm{kl}(k/n \parallel p)=\log(1/\delta)$, backed by a Bernoulli test-martingale inequality that bounds the joint event of seeing few audit errors while the average predictable risk is high. Each candidate action region combines that statistical bound with a temporal allowance $L_a\cdot\mathrm{age}_{tja}$, and the region is authorized only when the sum stays at or below the action-risk limit $\alpha_a$. A prespecified $\alpha$-spending schedule distributes the total confidence $\delta$ across times, thresholds, and the two automated actions, while evidence windows chosen without error labels prevent uncounted multiplicity; the proof controls the entire rejection event by applying the martingale inequality at a fixed boundary.

What would settle it

Construct a simulated fraud stream in which a hidden regime change makes current action risk exceed the window-averaged predictable risk plus $L_a$ times the mean label age by a known margin, violating condition A4, and run FCAC repeatedly; if any certified region has risk above the limit in more than $\delta$ of trials, the simultaneous control of Proposition 2 fails. A field version waits until labels mature on deployed automated regions and tests whether more than $\delta$ of them exceeded their declared limits.

Watch

Extended reading notes

Core claim

The paper's central claim is that automation in fraud operations is an authorization decision that must be earned by evidence, not granted by a score. Mature randomized audits and current scores alone cannot certify current action risk: unless the evolution of unobserved labels is restricted, two data-generating laws that agree on everything observable can make the same automated action have risk zero or one, so no nontrivial certificate can be uniformly valid (Proposition 1). The framework therefore requires a prespecified temporal-transport allowance $L_a$ set before audit errors are seen, asserting that current action risk in a candidate region is at most the window-averaged predictable risk plus $L_a$ times the mean label age (condition A4). Under that condition, together with representative audits, label-independent evidence windows, and simultaneous confidence allocation (conditions A1-A6), the KL test-martingale index gives finite-sample control: the probability that FCAC certifies any time, threshold, or action whose current risk exceeds the limit $\alpha_a$ is at most $\delta$, so with probability at least $1-\delta$ every automated region in the implemented policy satisfies its declared risk limit (Proposition 2). If the declared allowance is understated relative to the true drift, the guaranteed limit is enlarged only by the shortfall times mean label age (Corollary 1).

Load-bearing premise

The load-bearing premise is that today's risk of an automated action is no greater than the average risk of its mature audited evidence plus a pre-agreed allowance per unit of label age; this carry-forward from old labels to current risk cannot be checked from any observed data and must be fixed by the organization before seeing audit errors.

Editorial extensions

If this is right

  • An operator can freeze a scorer, fix risk limits, confidence $\delta$, and a temporal allowance $L_a$, and then automate only the approve and block regions whose mature audits certify them; with probability at least $1-\delta$, every automated region obeys its declared current-risk limit.
  • Increasing the diagnostic audit rate does not monotonically reduce human workload: sparse auditing leaves larger manual regions because evidence is insufficient, while intensive auditing consumes analysts through diagnostics, so an interior audit rate can be the workload optimum.
  • Label delay weakens certification: longer delays reduce the mature audit sample and raise the mean label age, which consumes more of the risk budget and can remove otherwise feasible automation.
  • If the declared temporal allowance is set too low, any certified region may exceed its risk limit, but only by the understatement times the mean label age; overstating the allowance preserves the risk guarantee at the cost of less automation.
  • Count-risk control does not control value exposure: an approve region that satisfies its count-risk limit can still concentrate high-value fraud, so value-sensitive authorization requires separately specified value limits.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same authorization logic could apply to other delayed-feedback, human-in-the-loop settings such as medical triage, content moderation, or loan origination, wherever audits of automated decisions consume reviewer capacity and current risk cannot be read off from a score alone.
  • The candidate-specific feasibility frontier suggests a testable operational fallback rule: rather than a common drift fraction of the risk limit, an organization could revoke automation for a region whenever the realized allowance $L_a\bar{g}_{tja}$ exceeds the remaining budget $\alpha_a - U_{tja}$, making the BAF-style stress test pass candidate by candidate.
  • A natural untested extension is a value-weighted variant of the KL index that certifies limits on monetary exposure instead of count risk; the paper's own value-exposure results show count and value assessments can disagree by large factors.
  • The framework evaluates proposed audit rates rather than optimizing them; an optimization layer that moves audit capacity from already-certified regions to evidence-starved ones could plausibly shift the workload frontier, though the paper does not establish that.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes FCAC, a decision-support framework for fraud operations that treats automation of approve/block actions as an authorization decision constrained by evidence freshness and shared review capacity. The framework uses mature randomized diagnostic audits, a prespecified temporal-transport allowance L_a, and a single review-workload ledger to certify candidate score regions, with unsupported regions remaining in manual review. The paper proves an observational non-identifiability result (Prop. 1), a conditional finite-sample simultaneous control result (Prop. 2), a misspecification bound (Cor. 1), a zero-error sample-size formula (Cor. 2), and an exact candidate feasibility frontier (Prop. 3). It then reports chronological retrospective evaluations on IEEE-CIS, ULB-Worldline, Elliptic++, and a synthetic BAF stream, with simulated audits and label delays. The results show a non-monotone audit-capacity/workload trade-off, strong sensitivity to label delay, and a prespecified BAF stress test that fails and is reported transparently. The paper is explicit that all guarantees are conditional on A4 and that the temporal allowance is a governance input rather than an estimated quantity.

Significance. If the conditional guarantee is accepted, the paper makes a useful contribution to fraud decision support by separating predictive ranking from authority to automate and by making the evidence-freshness/workload coupling an explicit design object. The proofs are careful: Proposition 2 correctly applies a fixed-limit Bernoulli test-martingale inequality with simultaneous alpha spending, and Corollary 1 quantifies the effect of an understated temporal allowance. The paper is unusually transparent: it reports a prespecified stress test that failed, provides a post-hoc stress confirming the Corollary 1 asymmetry, and offers a verified reproduction package. The main limitation is that the safety guarantee rests on A4, an untestable temporal-transport assumption whose parameter L_a is a governance choice; however, this is acknowledged in the text and quantified in Corollary 1. The contribution is therefore a conditional decision-support framework rather than an unconditional safety certificate, and it should be presented as such.

major comments (2)
  1. [Section 5.2, A4 and Section 5.4, Proposition 2] Because Proposition 1 establishes that current action risk is unidentified from O_t, the bound in Proposition 2 is only as strong as A4, whose parameter L_a is a governance choice that no observable data can validate. The paper states this, but the abstract and the theorem statement still present the result as a certificate; I recommend adding an explicit sentence in both places that the guarantee is conditional on a maintained, untestable assumption and that the post-hoc stress in Section 7.4 does not estimate the true drift rate. I also recommend that Section 5.6 give at least one concrete protocol for fixing L_a (for example, a pre-registered stress grid or a regulator-specified envelope) rather than leaving the choice entirely open.
  2. [Section 7.7, BAF stress test] The conclusion that fallback thresholds must reflect candidate-specific evidence rather than a common fraction of the risk limit is only partially supported by the reported failure, because the prespecified 10% endpoint would not be expected to produce all-review on BAF: the reported post-hoc frontier Lcrit·age/α ranges from 71.1% to 91.3% of the risk budget, while L/α=10% per period with BAF's realized evidence ages is far below that range. Please report the realized staleness share L·age/α for BAF and compare it with the frontier, and rephrase the conclusion so that it follows from that comparison rather than from the failed endpoint alone.
minor comments (5)
  1. [Section 5.3, Eq. (5)] The symbol q is overloaded (it denotes the audit rate in Section 6.2 and the test threshold in Eq. (4)), and the definition of U_KL does not state that q is the realized empirical mean k/n; please use distinct notation such as \hat p for the empirical mean and clarify the definition.
  2. [Section 5.3, Eq. (7)] The expression '6δ/(π^2t^2J_a2)' is ambiguous; please write 3δ/(π^2t^2J_a) or 6δ/(π^2t^2(2J_a)) and verify the displayed sum equals δ.
  3. [Section 5.5, Eq. (10)-(11)] Please clarify whether the 'preselected evidence window' is a fixed set of periods or a rolling window whose composition changes with t; the proof conditions on G_tja, so the window-selection rule should be stated as fixed before the experiment.
  4. [Table 3, caption] The column 'Action-period exceed.' is defined only in the text; please add a footnote defining the event and reiterating that it is not the event controlled by Proposition 2.
  5. [Section 7.4, post-hoc stress] In the post-hoc stress, the declared-to-true allowance ratios 0, 0.5, 1, and 1.5 should be defined explicitly as \hat L/L* (noting that 1.5 is overstatement) to avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the finite-sample certificate is derived from an explicit maintained assumption and standard martingale bounds, not from a fitted input or self-citation.

full rationale

The paper's central formal claim, Proposition 2, is self-contained: it applies a standard fixed-limit martingale inequality (Eq. 4, attributed to Howard et al. [31]) to mature randomized audits, recentering the null risk by the predeclared temporal allowance L_a·age under condition A4. A4 is an explicit maintained assumption about the data-generating law, not a parameter fitted to the target outcome; Proposition 1 independently shows that some such restriction is necessary for current-risk authorization. The paper repeatedly states that L_a is a governance input chosen before inspecting audit errors, and it does not claim to estimate or identify it. Corollary 1 openly quantifies the degradation if A4 is misspecified, and the Limitations section explicitly says that neither Corollary 1 nor the post-hoc stress estimates the valid temporal rate. The experiments use prespecified grid values of L and a BAF protocol that failed its prespecified qualitative endpoint, which is inconsistent with hidden tuning of the framework to force desired results. Proposition 3 is an explicit feasibility characterization rather than a predictive claim, and Corollary 2 merely notes that the zero-error KL bound coincides with the familiar Clopper-Pearson expression. No equation reduces to its inputs by construction, and no load-bearing step relies on a self-citation. The unidentifiability of A4 and the sensitivity of the guarantee to the choice of L_a are honest limitations and governance concerns, not circularity.

Assumptions & free parameters 4 free parameters · 7 assumptions · 0 invented entities

The framework's guarantee rests on six explicit conditions. The most fragile is A4, the temporal-transport condition, which is a maintained assumption supplied by the organization, not identified from data. A2 (randomized audits) is also a design condition that production systems may not satisfy. The remaining conditions are standard statistical assumptions. No new physical or generative entities are introduced.

free parameters (4)
  • Temporal-transport allowance rate L_a = L/α ∈ {0, 0.5%, 1%, 2.5%, 5%, 10%} in experiments; no data-driven estimate
    The central guarantee is conditional on A4, which requires a predeclared L_a. The paper treats L_a as a governance risk-budget parameter, set via stress grid or historical envelope, not estimated from test data.
  • Approve action-risk limit α_A = 2% (IEEE-CIS, Elliptic++), 0.1% (ULB)
    Illustrative operating limits specified by the organization; the paper notes that regulatory choices require calibration.
  • Block legitimate-friction limit α_B = 5%
    Illustrative block limit in the experiments.
  • Diagnostic audit rate q = 0.05, 0.10, 0.20, 0.30 in grid; 0.10 (IEEE), 0.20 (ULB, Elliptic++) in operational study
    Operational parameter varied to study the audit-capacity trade-off.
assumptions (7)
  • domain assumption Randomized audits are independent of hidden action-error labels conditional on pre-audit information (A2).
    Requires representative constant-propensity audits; not typical of production selective audits.
  • domain assumption Audit errors form an adapted Bernoulli sequence with predictable conditional means (A3).
    Allows risk to adapt to past outcomes but requires the conditional-mean structure.
  • domain assumption Temporal-transport condition: current action risk R_t ≤ window average predictable risk + L_a * age, almost surely (A4).
    The load-bearing unverifiable link between past audited outcomes and current risk; the paper treats L_a as a governance input and Proposition 1 shows some such restriction is necessary.
  • domain assumption Evidence windows are selected without inspecting error labels (A5).
    Prevents selection bias; adaptive window search would need a multiplicity correction.
  • standard math Simultaneous confidence allocation via online alpha spending (A6).
    Union bound over times, thresholds, and actions using δ_{tja} = 6δ/(π^2 t^2 J_a^2).
  • standard math Fixed-limit Bernoulli test-martingale inequality (Equation 4).
    Standard result from Howard et al. [31], used in the proof of Proposition 2.
  • standard math Data-generating laws in Proposition 1 have regular conditional distributions for unmatured labels.
    Technical condition for the non-identifiability construction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload." pith.science (2026). https://pith.science/paper/QNY2AJHA

@misc{pith2026260808577,
  author       = {Pith},
  title        = {Pith review of: When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QNY2AJHA}},
  note         = {Machine review of arXiv:2608.08577}
}
read the original abstract

Fraud operations must allocate events among automatic approval, analyst review, and automatic blocking even though the labels needed to evaluate these actions are selective and delayed. Predictive scores order cases, but they do not show whether the evidence is current and representative enough to delegate an action to the model. We develop freshness-constrained audit capacity (FCAC), a decision-support framework that treats automation as an authorization decision constrained by action risk, evidence freshness, and shared review capacity. It evaluates candidate action regions from mature randomized audits and a prespecified temporal allowance. Supported regions are automated; unsupported regions remain in review. The resulting decision record reports evidence age, audit demand, total review workload, value exposure, and compatible temporal change. We show that current action risk is unidentified without restricting unobserved label evolution. Under representative randomized audits, label-independent evidence windows, and a prespecified condition linking historical and current action risk, we derive simultaneous finite-sample control of unsafe authorization. Chronological evaluations with simulated audits on IEEE-CIS, ULB-Worldline, and Elliptic++ yield zero-drift automation rates of 84.4%, 67.4%, and 81.3%, with total review workloads of 24.1%, 46.0%, and 43.1%. The experiments reveal an audit-capacity trade-off: sparse auditing delays authorization, whereas intensive auditing eventually increases workload. A separately specified BAF stress test further indicates that fallback thresholds must reflect candidate-specific evidence rather than a common fraction of the risk limit. These findings identify audit freshness and analyst capacity as joint design considerations for fraud decision support.

Figures

Figures reproduced from arXiv: 2608.08577 by the authors.

Figure 1
Figure 1. FCAC decision flow. Risk owners specify action limits, confidence, and a temporal-stability allowance; operations provide score bands, audit rate, label delay, and workload limits. The output records authorized regions, evidence age, review demand, and the candidate-specific feasibility frontier. FCAC uses finite-sample risk control within a wider decision procedure. Adjacent work develops risk-controlling sets, any… view at source ↗
Figure 2
Figure 2. Total human-review rate across audit-rate and delay scenarios. Values include the manual-decision [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Certified automation versus total human-review workload at label delay three. [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Operational FCAC response using realized audit counts, errors, and label ages. Lines are audit￾seed means; shading is the empirical 2.5th–97.5th percentile. The horizontal axis declares L/α per native period only; candidate decisions use Lg/α ¯ , so cross-domain slopes…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 20 canonical work pages

  1. [1]

    Dal Pozzolo, G

    A. Dal Pozzolo, G. Boracchi, O. Caelen, C. Alippi, G. Bontempi, Credit card fraud detection: A realistic modeling and a novel learning strategy, IEEE Transactions on Neural Networks and Learning Systems 29 (8) (2018) 3784–3797.doi:10.1109/TNNLS. 2017.2736643

  2. [2]

    Grzenda, H

    M. Grzenda, H. M. Gomes, A. Bifet, Delayed labelling evaluation for data streams, Data Mining and Knowledge Discovery 34 (2020) 1237–1266.doi:10.1007/ s10618-019-00654-y

  3. [3]

    Botacin, H

    M. Botacin, H. Gomes, Towards more realistic evaluations: The impact of label delays in malware detection pipelines, Computers & Security 148 (2025) 104122.doi:10. 1016/j.cose.2024.104122

  4. [4]

    Lakkaraju, J

    H. Lakkaraju, J. Kleinberg, J. Leskovec, J. Ludwig, S. Mullainathan, The selective labels problem: Evaluating algorithmic predictions in the presence of unobservables, in: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2017, pp. 275–284.doi:10.1145/3097983.3098066

  5. [5]

    Höppner, B

    S. Höppner, B. Baesens, W. Verbeke, T. Verdonck, Instance-dependent cost-sensitive learning for detecting transfer fraud, European Journal of Operational Research 297 (1) (2022) 291–300.doi:10.1016/j.ejor.2021.05.028

  6. [6]

    Bates, A

    S. Bates, A. Angelopoulos, L. Lei, J. Malik, M. I. Jordan, Distribution-free, risk- controlling prediction sets, Journal of the ACM 68 (6) (2021) 43:1–43:34.doi: 10.1145/3478535

  7. [7]

    Z. Xu, N. Karampatziakis, P. Mineiro, Active, anytime-valid risk-controlling pre- diction sets, in: Advances in Neural Information Processing Systems 37, 2024. doi:10.52202/079017-1920. URLhttps://papers.nips.cc/paper_files/paper/2024/hash/ 6eb05d8bc6bd7bb6868c64b5802125bd-Abstract-Conference.html 20

  8. [8]

    URLhttps://www.jmlr.org/papers/v26/24-0452.html

    Y.Bao, Y.Huo, H.Ren, C.Zou, CAP:Ageneralalgorithmforonlineselectiveconformal prediction with FCR control, Journal of Machine Learning Research 26 (287) (2025) 1– 74. URLhttps://www.jmlr.org/papers/v26/24-0452.html

Show all 35 references
  1. [9]

    Gibbs, E

    I. Gibbs, E. J. Candès, Conformal inference for online prediction with arbitrary distri- bution shifts, Journal of Machine Learning Research 25 (162) (2024) 1–36. URLhttps://www.jmlr.org/papers/v25/22-1218.html

  2. [10]

    Khosravi, X

    H. Khosravi, X. Huo, Conformal selective acting: Anytime-valid risk control for RLVR- trained LLMs (2026).arXiv:2605.20270,doi:10.48550/arXiv.2605.20270

  3. [11]

    Carcillo, Y.-A

    F. Carcillo, Y.-A. Le Borgne, O. Caelen, G. Bontempi, Streaming active learning strate- gies for real-life credit card fraud detection: Assessment and visualization, Interna- tional Journal of Data Science and Analytics 5 (4) (2018) 285–300.doi:10.1007/ s41060-018-0116-z

  4. [12]

    Carcillo, Y.-A

    F. Carcillo, Y.-A. Le Borgne, O. Caelen, Y. Kessaci, F. Oblé, G. Bontempi, Combin- ing unsupervised and supervised learning in credit card fraud detection, Information Sciences 557 (2021) 317–331.doi:10.1016/j.ins.2019.05.042

  5. [13]

    Baesens, S

    B. Baesens, S. Höppner, T. Verdonck, Data engineering for fraud detection, Decision Support Systems 150 (2021) 113492.doi:10.1016/j.dss.2021.113492

  6. [14]

    Hajek, J

    P. Hajek, J. Novotny, M. Munk, Financial statement fraud detection using topic-driven financial sentiment analysis, Decision Support Systems 203 (2026) 114615.doi:10. 1016/j.dss.2026.114615

  7. [15]

    Weber, G

    M. Weber, G. Domeniconi, J. Chen, D. K. I. Weidele, C. Bellei, T. Robinson, C. E. Leiserson, Anti-money laundering in bitcoin: Experimenting with graph convolutional networksforfinancialforensics, arXivpreprintarXiv:1908.02591(2019).doi:10.48550/ arXiv.1908.02591

  8. [16]

    Elmougy, L

    Y. Elmougy, L. Liu, Demystifying fraudulent transactions and illicit nodes in the bitcoin networkforfinancialforensics, in: Proceedingsofthe29thACMSIGKDDConferenceon Knowledge Discovery and Data Mining, 2023, pp. 3979–3990.doi:10.1145/3580305. 3599803

  9. [17]

    Dal Pozzolo, O

    A. Dal Pozzolo, O. Caelen, R. A. Johnson, G. Bontempi, Calibrating probability with undersampling for unbalanced classification, in: 2015 IEEE Symposium Series on Com- putational Intelligence, 2015, pp. 159–166.doi:10.1109/SSCI.2015.33

  10. [19]

    Nanduri, Y

    J. Nanduri, Y. Jia, A. Oka, J. Beaver, Y. Liu, Microsoft uses machine learning and optimization to reduce e-commerce fraud, INFORMS Journal on Applied Analytics 50 (1) (2020) 64–79.doi:10.1287/inte.2019.1017. 21

  11. [20]

    A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, T. Schuster, Conformal risk control, in: The Twelfth International Conference on Learning Representations, 2024. URLhttps://proceedings.iclr.cc/paper_files/paper/2024/hash/ f3549ef9b5ff520a7e41ff3cc306ab2b-Abstract-Conference.html

  12. [21]

    Prinster, S

    D. Prinster, S. D. Stanton, A. Liu, S. Saria, Conformal validity guarantees exist for any data distribution (and how to find them), in: Proceedings of the 41st International ConferenceonMachineLearning, Vol.235ofProceedingsofMachineLearningResearch, 2024, pp. 41086–41118. URLh...

  13. [22]

    Hultberg, D

    B. Hultberg, D. Zachariah, A. H. Ribeiro, Anytime-valid conformal risk control (2026). arXiv:2602.04364,doi:10.48550/arXiv.2602.04364

  14. [23]

    R. Zhu, X. Zhang, T. Wang, J. Liao, S. H. Chung, X. Zhang, DISCO: Decoupling representation learning and risk control for reliable credit card fraud detection, Decision Support Systems 208 (2026) 114717.doi:10.1016/j.dss.2026.114717

  15. [24]

    Xia, C.-S

    G. Xia, C.-S. Bouganis, Augmenting the softmax with additional confidence scores for improved selective classification with out-of-distribution data, International Journal of Computer Vision 132 (2024) 3714–3752.doi:10.1007/s11263-024-02029-3

  16. [25]

    G. J. Aguiar, A. Cano, Dynamic budget allocation for sparsely labeled drifting data streams, Information Sciences 654 (2024) 119821.doi:10.1016/j.ins.2023.119821

  17. [26]

    Wang, S.-H

    C.-A. Wang, S.-H. Huang, C.-T. Chen, Y.-T. Fang, Financial reinforcement learning under concept drift based on knowledge distillation and curriculum learning, Decision Support Systems 203 (2026) 114624.doi:10.1016/j.dss.2026.114624

  18. [27]

    D. R. Jones, D. Brown, The division of labor between human and computer in the presence of decision support system advice, Decision Support Systems 33 (4) (2002) 375–388.doi:10.1016/S0167-9236(02)00005-2

  19. [28]

    V. C. Storey, A. R. Hevner, V. Y. Yoon, The design of human-artificial intelligence systems in decision sciences: A look back and directions forward, Decision Support Systems 182 (2024) 114230.doi:10.1016/j.dss.2024.114230

  20. [29]

    H. M. Zolbanin, B. Davazdahemami, D. Delen, D. Wright, D. Crosby, Designing trans- parent, equitable, and efficient decision support systems for drug courts using machine learning, Decision Support Systems 204 (2026) 114635.doi:10.1016/j.dss.2026. 114635

  21. [30]

    Mozannar, D

    H. Mozannar, D. Sontag, Consistent estimators for learning to defer to an expert, in: Proceedings of the 37th International Conference on Machine Learning, Vol. 119 of Proceedings of Machine Learning Research, 2020, pp. 7076–7087. URLhttps://proceedings.mlr.press/v119/mozannar...

  22. [31]

    S. R. Howard, A. Ramdas, J. McAuliffe, J. Sekhon, Time-uniform, nonparametric, nonasymptotic confidence sequences, The Annals of Statistics 49 (2) (2021) 1055–1080. doi:10.1214/20-AOS1991

  23. [32]

    C. J. Clopper, E. S. Pearson, The use of confidence or fiducial limits illustrated in the case of the binomial, Biometrika 26 (4) (1934) 404–413.doi:10.1093/biomet/26.4. 404

  24. [33]

    URLhttps://www.kaggle.com/c/ieee-fraud-detection

    IEEE Computational Intelligence Society, Vesta Corporation, IEEE-CIS fraud detec- tion, Kaggle competition dataset, accessed 2026-07-13 (2019). URLhttps://www.kaggle.com/c/ieee-fraud-detection

  25. [34]

    URLhttps://www.kaggle.com/datasets/mlg-ulb/creditcardfraud

    Machine Learning Group, Université Libre de Bruxelles, Credit card fraud detection, Kaggle dataset, accessed 2026-07-13 (2013). URLhttps://www.kaggle.com/datasets/mlg-ulb/creditcardfraud

  26. [35]

    Jesus, J

    S. Jesus, J. Pombal, D. Alves, A. Cruz, P. Saleiro, R. P. Ribeiro, J. Gama, P. Bizarro, Turning the tables: Biased, imbalanced, dynamic tabular datasets for ml evaluation, in: Advances in Neural Information Processing Systems, Vol. 35, 2022. URLhttps://proceedings.neurips.cc/p...

  27. [36]

    T. Chen, C. Guestrin, XGBoost: A scalable tree boosting system, in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 785–794.doi:10.1145/2939672.2939785. 23

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.