Pith. sign in

REVIEW 4 major objections 5 minor 88 references

Model-Free Counterfactual Subset Selection at Scale

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper claims a single-pass streaming algorithm can select a diverse, label-balanced set of real counterfactual examples with logarithmic per-item updates and a constant-factor quality guarantee.

desk verdict The streaming problem is real and the preserve-set idea is plausible, but the utility functions break the submodularity assumption the whole guarantee rests on, so the paper's central claims do not hold. read the letter →

arxiv 2502.08326 v1 pith:YUO2RFUB submitted 2025-02-12 cs.LG cs.DBcs.DScs.IR

classification cs.LGcs.DBcs.DScs.IR MSC 68W2568W2790C2705B35
keywords counterfactualexplanationsstreamingalgorithmssubmodularmaximizationmatroidconstraintsdiversitymodel-freeexplainabilityfairnessonlineselection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to establish that counterfactual explanations — real, observed examples showing a different course of action that would change an outcome — can be selected on the fly from a live data stream rather than generated by a model or mined from a stored dataset. It formulates the task as maximizing a non-negative, monotone, submodular utility over subsets that obey a global cardinality bound and per-label lower and upper bounds, then gives a one-pass algorithm with an $O(\log k)$ update per arriving item and a constant-factor approximation guarantee. If correct, this is the first streaming, model-free way to return several diverse and actionable counterfactuals for a query, with fairness-style label constraints handled at the same time. The empirical section backs the claim with transport-cost, constraint-violation, runtime, sliding-window, and concept-drift experiments on census, bank, credit, and synthetic clinical data.

What carries the argument

The enabling object is the extensibility matroid $\hat{\mathcal{S}}$: the family of sets that can be extended to a feasible solution, characterized by per-label upper bounds $|S\cap D_l|\le \beta_l$ and the condition $\sum_l \max(|S\cap D_l|,\alpha_l)\le k$. By maintaining label counts $c_l$ and their capped sum $C$, the algorithm tests extensibility of a candidate extension in $O(1)$ time and finds the minimum-weight replaceable item using priority queues, which keeps each update at $O(\log k)$. The utility functions $f_1,f_2,f_3$ encode content-, sampling-, and clustering-based diversity; the guarantee applies to them only insofar as they are non-negative, monotone, and submodular.

What would settle it

Compute marginal gains for the three utilities on a small dataset and check whether $f(e\mid S)$ decreases as $S$ grows; for example, if adding an item to a larger set ever increases the marginal gain of another item, submodularity fails. For $f_2$, evaluate $\det(K_{S\cup\{e\}})-\det(K_S)$ on nested sets; non-increasing differences are required. If marginal-gain evaluation time scales with $k$, the $O(\log k)$ per-item update claim does not hold.

Watch

Extended reading notes

Core claim

The central claim is that selecting a diverse set of real counterfactuals can be solved in a single pass over the data, without storing the data or calling the classifier. The complete algorithm (Alg. 2) runs the streaming matroid-submodular maximization routine of Alg. 1 on the family of extensible sets, accepting an incoming item only when its marginal gain is at least $1+\lambda$ times the smallest gain of a removable candidate, while separately keeping up to $\alpha_l$ backup items per label so the final set satisfies the lower-bound constraints. The paper argues the returned set is feasible and inherits the warm-up algorithm's approximation ratio: the body's analysis gives $1/7.75$ at $\lambda=1$, and the contributions section announces $1/5.585$. It positions this as the first real-time, model-free counterfactual subset selection from streaming data.

Load-bearing premise

The approximation and complexity guarantees rely on the three proposed utility functions being non-negative, monotone, and submodular, and on the marginal gain of an item against a $k$-member set being computable in constant time; the paper assumes these properties rather than proving them.

Editorial extensions

If this is right

  • If the guarantee holds, counterfactual explanations can be served in real time from unbounded streams, with no stored dataset and no access to the underlying model.
  • The same matroid machinery applies to any streaming selection problem with per-group quotas, not only counterfactual explanations.
  • The experiments indicate users would need less effort: on the Customer dataset the method reduces transport cost by up to 33.18% compared with streaming baselines.
  • The algorithm keeps space $O(k)$, so a fixed-size summary of size $k$ is enough regardless of stream length.
  • Because the relaxed version (no lower bounds) is a matroid-constrained submodular maximization, its approximation bound transfers to other matroid feasibility constraints.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit consequence, not tested in the paper, is that the same procedure could be used for other quota-constrained streaming summarization tasks, such as diverse news digests or representative panel construction.
  • Since the guarantee depends on utility monotonicity and submodularity, a natural next step is to check which of the three proposed utilities satisfy these conditions; the determinant term in $f_2$ is the most likely place to fail.
  • The discrepancy between the advertised $1/5.585$ and the proved $1/7.75$ suggests the final ratio depends on problem-specific curvature or $\lambda$ tuning; a direct derivation of $5.585$ from the analysis would settle which number is the guaranteed one.
  • A user study measuring whether diverse real counterfactuals actually improve decision-making would be the natural behavioural test, since the paper's metrics are proxy costs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a streaming, model-free method for selecting a diverse and relevant subset of real counterfactual examples for a query item, subject to cardinality and label-balance constraints. The authors define three diversity-aware utility functions, present a greedy streaming algorithm with a swap threshold (Algorithm 1) and a complete algorithm that augments the result with lower-bound backup items (Algorithm 2), and claim a single-pass 1/5.585 approximation guarantee with O(log k) update complexity per item. The evaluation compares the method against offline, kNN, random, relaxed, and constraint-free baselines on three real datasets and a large synthetic dataset, reporting transport cost, constraint violations, runtime, and robustness to concept drift.

Significance. If the central algorithmic claims were correct, the paper would address a timely and practical gap: real-time, model-free counterfactual subset selection without storing the full dataset. The problem formulation, with label-diversity constraints and multiple diversity notions, is reasonable, and the experimental setup includes realistic datasets and a synthetic data generator at scale. However, the theoretical contribution is not currently supported: the utility functions used in the experiments are asserted to be monotone and submodular without proof, at least one of them is demonstrably not monotone, the claimed approximation constants are internally inconsistent, and the O(log k) per-item complexity is not established. The paper contains a useful heuristic and a substantial evaluation, but the advertised quality guarantee does not apply to the objective being optimized.

major comments (4)
  1. [§2.2, Eqs. (2)–(5); §3.2–3.3] The approximation guarantees for Algorithms 1 and 2 require the objective f to be non-negative, monotone, and submodular, but this property is never established for the proposed utilities. In fact, f1 in Eq. (2) is not monotone. For sim(e1,q)=0.9, sim(e2,q)=0.1, sim(e1,e2)=0.95, and λ1=0.5, Eq. (2) gives f1({e1})=0.9 and f1({e1,e2})=0.9+0.1−(0.5/4)·(2·0.95)=0.7625, so the marginal gain of e2 is −0.1375. Adding an item can therefore decrease the objective, contradicting the monotonicity assumption used in the greedy thresholding argument and in the augmentation step of §3.3. The determinant term in Eq. (3) and the |S|-scaled coverage term in Eq. (4) are not shown to be monotone submodular either; the assertion in §3.2 that the functions are 'well-designed' is not a substitute for proof. Because the central quality guarantee is stated for a monotone submodular objective, this is a load-bearing gap.
  2. [§3.2 and contribution bullet] The paper advertises a 1/5.585 approximation guarantee in the introduction, but §3.2 derives a minimum approximation factor ρ=7.75 and then states only 'check if 5.585/(1−cv(f))>7.75' without deriving this expression or relating it to the preceding analysis. If 5.585 is intended to follow from a curvature-aware bound, the relevant curvature parameter and proof are missing; if it is an empirical observation, it cannot be a worst-case guarantee. The text never reconciles the two numbers, so the stated quality guarantee is ambiguous.
  3. [§3.3, Runtime and memory] The O(log k) per-item claim is not supported by the algorithm description. Algorithm 1 makes two utility calls per item, and each marginal gain f(e|S) as defined in §2.2 requires interaction with all items currently in S; e.g., the pairwise sum in Eq. (2) is O(k) unless a sketch or index is described, and none is provided. The priority queues in §3.3 store item weights, but when S changes, the marginal gains—and hence the weights—of retained items generally change, and the paper does not explain how all affected weights are updated in O(log k) time. Without a mechanism to evaluate f(e|S) in O(1) or O(log k), the central complexity claim and the interpretation of Fig. 4d as O(log k) per item are unsubstantiated.
  4. [§4, Evaluation metrics] The empirical comparisons rely on metrics that largely coincide with components of the optimized objective. Transport cost (Eq. (6)) measures distance to the query and is essentially the similarity term appearing in Eqs. (2)–(5); constraint violations are enforced to zero by construction for the feasible algorithms; and Table 3 reports utility, which is the objective itself. The reported 'superior performance' is therefore partly self-referential and does not establish that the selected counterfactuals are more useful or actionable to humans than those of the baselines, a limitation the paper itself acknowledges in the case-study discussion. This weakens the empirical contribution, although it is not the primary reason for my recommendation.
minor comments (5)
  1. [§3.2] The sentence 'The complexity of involves runtime, utility calls, matroid queries, and space' is incomplete and should specify that it refers to Algorithm 1.
  2. [§3.3] The phrase 'Extending the idea in ?? to F being an extensibility matroid' contains a missing citation/reference marker '??' and needs to be completed.
  3. [§2.3] The solution space is defined with |S| < k, but all subsequent constraints and algorithms use |S| ≤ k; these should be made consistent.
  4. [Algorithm 2] The description of the augmentation step ('S = S 1 augmented with items in sets Pl') should specify how ties are broken when a label already satisfies its lower-bound constraint and multiple backup items are available.
  5. [Table 1 and Figure 4] The GeCo row in Table 1 contains an ambiguous combination of check/cross marks and a footnote; Figure 4 reports averages over 10 runs without error bars or variance information, despite the paper stating that variances are reported if appropriate.

Circularity Check

2 steps flagged · score 6.0 of 10

Algorithmic guarantees are adapted from independent streaming-submodular-maximization results, but two headline empirical claims—transport-cost superiority and zero constraint violations—are partly self-referential restatements of the optimized objective and the enforced solution space.

  1. self definitional [§2.2 Eq. (2) and §4 'Evaluation metrics' (Transport cost)]
    "Content-based utility: f1(S ) = X e∈S sim(e, q)− λ1 |S|2 X e X e′,e sim(e, e′) ... The first term is similar to the proximity in DiCE [9] ... Transport cost: ... cost = 1 |S| X e∈S dcon(e, q) + dcat(e, q) ... Users prefer counterfactual examples similar to the query example, minimizing the effort required."

    The dominant term of the optimized utility is the query-similarity term Σ sim(e,q), while the transport-cost metric is the average distance from S to q (and the paper defines dist = 1 − sim for the kernel in Eq. 3). Minimizing transport cost is therefore a monotone restatement of maximizing the similarity component of f1; reporting a transport-cost advantage over baselines is not an external check of explanation quality but a re-measurement of the very objective being optimized.

  2. self definitional [§2.3 Problem Statement, Alg. 2 (Line 9), and §4.1 'Constraint violations']
    "The solution spaceS is: S = {S ⊆ D : |S| < k,α l ≤ |S∩ Dl| ≤ βl,∀l = 1,..., L}. ... Alg. 2 returns a feasible solution belonging toS with the same approximation ratio as Alg. 1. ... Our approach, like the offline algorithm, had no violations (not shown for brevity)."

    The 'no constraint violations' result is enforced by the algorithm's definition: Alg. 2 adds backup items Pl exactly until αl ≤ |S∩Dl| and uses the extensibility matroid to respect βl and k, so every returned S lies in S by construction. Claiming zero violations as an empirical advantage is equivalent to asserting that the algorithm implements its own feasibility constraints.

full rationale

The formal derivation of the approximation ratio is not circular: Alg. 1 is explicitly inherited from Chakrabarti et al. and Huang et al. (independent streaming submodular-maximization results), and the quality argument in §3.2–3.3 is a standard reduction to MSIS/matroid submodular maximization. The utility functions are hand-designed rather than fitted, and the O(log k) data structure is an engineering contribution. The self-reference in the paper is concentrated in the evaluation: 'transport cost' largely inverts the query-similarity term of the optimized f1, and 'constraint violations' are impossible by the definition of the returned set. These two empirical claims are therefore partly self-definitional, which prevents a clean score of 0–2. However, the central algorithmic claim (one-pass, O(log k), 1/ρ guarantee for monotone submodular f) rests on independent external theory and is not itself a renamed input, so the paper is at score 6 rather than higher. Note also that the unsupported submodularity/monotonicity of f1–f3 is a correctness gap, not a circularity, and is not scored here.

Assumptions & free parameters 7 free parameters · 4 assumptions · 0 invented entities

The central algorithmic guarantee rests on unproven submodularity of the hand-designed utilities and on an unspecified O(1) marginal-gain oracle. The utility parameters and label bounds are free choices, and the similarity measure is under-specified.

free parameters (7)
  • lambda1 (content-based diversity weight)
    Appears in Eq. (2); value not reported in experiments.
  • lambda2 (sampling-based diversity weight)
    Appears in Eq. (3); value not reported.
  • lambda3 (clustering-based diversity weight)
    Appears in Eq. (4); value not reported.
  • Greedy threshold lambda = 0.717 or 1
    Chosen based on curvature cv(f) which is not measured; the paper says 'check if 5.585/(1-cv(f)) > 7.75' but does not provide a method to compute cv(f).
  • Label lower and upper bounds alpha_l, beta_l = e.g., 0.9*|D_l|/|D|*k and 1.1*|D_l|/|D|*k
    Set proportionally to label distribution in each dataset; hand-chosen constraints.
  • Similarity measure sim(.,.)
    Used throughout but not defined for continuous and categorical features; in evaluation l1 distances are used but the transformation to similarity is unspecified.
  • Diagonal perturbation for kernel determinant = unspecified
    Added to ensure well-defined determinant in sampling-based diversity; size not given.
assumptions (4)
  • ad hoc to paper The utility functions f1, f2, f3 are non-negative, monotone, and submodular.
    Section 2.2 defines f1-f3 and Section 3.1 assumes general submodularity, but no proof is given. In particular, the normalized pairwise penalty in f1 and the determinant term in f2 may break monotonicity and submodularity.
  • domain assumption A feasible solution exists: sum_l alpha_l <= k and alpha_l <= beta_l <= |D_l| for all l.
    Section 2.3 states 'We assume a feasible solution exists'.
  • standard math The family of extensible sets S_hat is a matroid.
    Upper bounds on label counts plus a global cardinality bound define a partition matroid; this is standard and used to invoke the streaming matroid algorithm.
  • ad hoc to paper Marginal gains f(e|S) can be computed in O(1) or O(log k) time.
    Required for the O(log k) complexity claim; no data structure is specified beyond priority queues, which only speed up the min-weight selection, not the marginal gain computation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Model-Free Counterfactual Subset Selection at Scale." pith.science (2026). https://pith.science/paper/YUO2RFUB

@misc{pith2026250208326,
  author       = {Pith},
  title        = {Pith review of: Model-Free Counterfactual Subset Selection at Scale},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YUO2RFUB}},
  note         = {Machine review of arXiv:2502.08326}
}
abstract

Ensuring transparency in AI decision-making requires interpretable explanations, particularly at the instance level. Counterfactual explanations are a powerful tool for this purpose, but existing techniques frequently depend on synthetic examples, introducing biases from unrealistic assumptions, flawed models, or skewed data. Many methods also assume full dataset availability, an impractical constraint in real-time environments where data flows continuously. In contrast, streaming explanations offer adaptive, real-time insights without requiring persistent storage of the entire dataset. This work introduces a scalable, model-free approach to selecting diverse and relevant counterfactual examples directly from observed data. Our algorithm operates efficiently in streaming settings, maintaining $O(\log k)$ update complexity per item while ensuring high-quality counterfactual selection. Empirical evaluations on both real-world and synthetic datasets demonstrate superior performance over baseline methods, with robust behavior even under adversarial conditions.

Figures

Figures reproduced from arXiv: 2502.08326 by the authors.

Figure 1
Figure 1. Counterfactual selection [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. presents the results, with the X-axis showing the explana￾tion size (k) and the Y-axis the transport cost as l1-distance. Our method outperformed streaming baselines (Random, Relaxed, NoConstraint) and the non-streaming baseline (kNN), matching the offline algorithm in some cases. For example, in the Cus￾tomer dataset (k = 10), our approach reduced transport costs by up to 33.18%, meaning users exerted only two-thir… view at source ↗
Figure 3
Figure 3. Constraint Violations in explanations. Runtime. We evaluated the total runtime of our approach against baselines. Non-streaming methods, such as offline and kNN, had significantly longer runtimes and are excluded for clarity. Fig. 4a, 4b, and 4c present results, with the X-axis showing the explanation size (k) and the Y-axis the runtime. We varied k up to 25 due to cognitive load limits. While our approach 5 [PITH_… view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: shows results, with the X-axis as the percentage of data processed and the Y-axis as utility change. Utility increases with data stream progress, confirming the algorithm’s sound￾ness. Larger windows (|b| = 5, 10) initially increase utility faster (up to 40-60%) as sub…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

88 extracted references · 77 canonical work pages

  1. [1]

    Dwivedi, D

    R. Dwivedi, D. Dave, H. Naik, S. Singhal, R. Omer, P. Patel, B. Qian, Z. Wen, T. Shah, G. Morgan, et al., Explainable ai (xai): Core ideas, techniques, and solutions, CSUR 55 (2023) 1–33

  2. [2]

    Wachter, et al., Counterfactual explanations without opening the black box: Automated decisions and the gdpr, Harv

    S. Wachter, et al., Counterfactual explanations without opening the black box: Automated decisions and the gdpr, Harv. JL & Tech. 31 (2017) 841

  3. [3]

    Barocas, A

    S. Barocas, A. D. Selbst, M. Raghavan, The hidden assumptions behind counterfactual explanations and principal reasons, in: FAccT, 2020, pp. 80–89

  4. [4]

    Schleich, Z

    M. Schleich, Z. Geng, Y . Zhang, D. Suciu, Geco: Quality counterfactual explanations in real time, PVLDB 14 (2021) 1681–1693

  5. [5]

    Kirkpatrick, Battling algorithmic bias: how do we ensure algorithms treat us fairly?, CACM 59 (2016) 16–17

    K. Kirkpatrick, Battling algorithmic bias: how do we ensure algorithms treat us fairly?, CACM 59 (2016) 16–17

  6. [6]

    Karimi, G

    A.-H. Karimi, G. Barthe, B. Schölkopf, I. Valera, A survey of algorithmic recourse: contrastive explanations and consequential recommendations, CSUR (2020)

  7. [7]

    Explainable Image Classification with Evidence Counterfactual

    T. Vermeire, D. Martens, Explainable image classification with evidence counterfactual, arXiv preprint arXiv:2004.07511 (2020)

  8. [8]

    Mehrabi, F

    N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, A. Galstyan, A survey on bias and fairness in machine learning, CSUR 54 (2021) 1–35

Show all 88 references
  1. [9]

    R. K. Mothilal, A. Sharma, C. Tan, Explaining machine learning classi- fiers through diverse counterfactual explanations, in: FAccT, 2020, pp. 607–617

  2. [10]

    Karimi, G

    A.-H. Karimi, G. Barthe, B. Balle, I. Valera, Model-agnostic counterfac- tual explanations for consequential decisions, in: AISTATS, 2020, pp. 895–905

  3. [11]

    N. Bui, D. Nguyen, V . A. Nguyen, Counterfactual plans under distribu- tional ambiguity, in: ICLR, 2022

  4. [12]

    Nguyen, N

    D. Nguyen, N. Bui, V . A. Nguyen, Feasible recourse plan via diverse interpolation, in: AISTATS, 2023, pp. 4679–4698

  5. [13]

    T. T. Nguyen, Q. V . H. Nguyen, M. Weidlich, K. Aberer, Result selection and summarization for web table search, in: ICDE, 2015, pp. 231–242

  6. [14]

    T. T. Nguyen, C. T. Duong, M. Weidlich, H. Yin, Q. V . H. Nguyen, Re- taining data from streams of social platforms with minimal regret, in: IJCAI, 2017, pp. 2850–2856

  7. [15]

    N. T. Tam, M. Weidlich, B. Zheng, H. Yin, N. Q. V . Hung, B. Stantic, From anomaly detection to rumour detection using data streams of social platforms, Proceedings of the VLDB Endowment 12 (2019) 1016–1029

  8. [16]

    T. T. Nguyen, T. D. Hoang, M. T. Pham, T. T. Vu, T. H. Nguyen, Q.-T. Huynh, J. Jo, Monitoring agriculture areas with satellite images and deep learning, Applied Soft Computing 95 (2020) 106565

  9. [17]

    T. T. Nguyen, M. Weidlich, H. Yin, B. Zheng, Q. V . H. Nguyen, B. Stan- tic, User guidance for e fficient fact checking, Proceedings of the VLDB Endowment 12 (2019) 850–863

  10. [18]

    N. T. Tam, H. T. Trung, H. Yin, T. Van Vinh, D. Sakong, B. Zheng, N. Q. V . Hung, Entity alignment for knowledge graphs with multi-order convolutional networks, TKDE 34 (2022) 4201–4214

  11. [19]

    T. T. Nguyen, M. T. Pham, T. T. Nguyen, T. T. Huynh, Q. V . H. Nguyen, T. T. Quan, et al., Structural representation learning for network align- ment with self-supervised anchor links, Expert Systems with Applica- tions 165 (2021) 113857

  12. [20]

    Stepin, J

    I. Stepin, J. M. Alonso, A. Catala, M. Pereira-Fariña, A survey of con- trastive and counterfactual explanation generation methods for explain- able artificial intelligence, IEEE Access 9 (2021) 11974–12001

  13. [21]

    C. T. Duong, T. T. Nguyen, H. Yin, M. Weidlich, T. S. Mai, K. Aberer, Q. V . H. Nguyen, E fficient and e ffective multi-modal queries through heterogeneous network embedding, IEEE Transactions on Knowledge and Data Engineering 34 (2022) 5307–5320

  14. [22]

    T. T. Nguyen, M. Weidlich, H. Yin, B. Zheng, Q. H. Nguyen, Q. V . H. Nguyen, Factcatch: Incremental pay-as-you-go fact checking with min- imal user e ffort, in: Proceedings of the 43rd International ACM SI- GIR Conference on Research and Development in Information Retrieval, 2...

  15. [23]

    N. Q. V . Hung, D. C. Thang, N. T. Tam, M. Weidlich, K. Aberer, H. Yin, X. Zhou, Answer validation for generic crowdsourcing tasks with mini- mal efforts, The VLDB Journal 26 (2017) 855–880

  16. [24]

    Q. V . H. Nguyen, C. T. Duong, T. T. Nguyen, M. Weidlich, K. Aberer, H. Yin, X. Zhou, Argument discovery via crowdsourcing, The VLDB Journal 26 (2017) 511–535

  17. [25]

    Z. Ren, T. T. Nguyen, W. Nejdl, Prototype learning for interpretable respiratory sound analysis, in: Proc. ICASSP, 2022, pp. 9087–9091

  18. [26]

    Q. V . H. Nguyen, K. Zheng, M. Weidlich, B. Zheng, H. Yin, T. T. Nguyen, B. Stantic, What-if analysis with conflicting goals: Recommending data ranges for exploration, in: ICDE, 2018, pp. 89–100

  19. [27]

    N. T. Toan, P. T. Cong, N. T. Tam, N. Q. V . Hung, B. Stantic, Diversifying group recommendation, IEEE Access 6 (2018) 17776–17786

  20. [28]

    Verma, V

    S. Verma, V . Boonsanong, M. Hoang, K. E. Hines, J. P. Dickerson, C. Shah, Counterfactual explanations and algorithmic recourses for ma- chine learning: A review, arXiv preprint arXiv:2010.10596 (2020)

  21. [29]

    R. M. Byrne, Counterfactuals in explainable artificial intelligence (xai): Evidence from human reasoning., in: IJCAI, 2019, pp. 6276–6282

  22. [30]

    Mertes, T

    S. Mertes, T. Huber, K. Weitz, A. Heimerl, E. André, Ganterfac- tual—counterfactual explanations for medical non-experts using gener- ative adversarial learning, Frontiers in artificial intelligence 5 (2022) 825565

  23. [31]

    J. Ma, R. Guo, S. Mishra, A. Zhang, J. Li, CLEAR: generative counter- factual explanations on graphs, in: NeurIPS, 2022

  24. [32]

    Dutta, J

    S. Dutta, J. Long, S. Mishra, C. Tilli, D. Magazzeni, Robust counterfac- tual explanations for tree-based ensembles, in: ICML, volume 162, 2022, pp. 5742–5756

  25. [33]

    Ustun, A

    B. Ustun, A. Spangher, Y . Liu, Actionable recourse in linear classifica- tion, in: FAT*, 2019, pp. 10–19

  26. [34]

    Van Looveren, J

    A. Van Looveren, J. Klaise, Interpretable counterfactual explanations guided by prototypes, in: ECML PKDD, 2021, pp. 650–665

  27. [35]

    Poyiadzi, K

    R. Poyiadzi, K. Sokol, R. Santos-Rodriguez, T. De Bie, P. Flach, Face: feasible and actionable counterfactual explanations, in: AIES, 2020, pp. 344–350

  28. [36]

    Kanamori, et al., Dace: Distribution-aware counterfactual explanation by mixed-integer linear optimization., in: IJCAI, 2020, pp

    K. Kanamori, et al., Dace: Distribution-aware counterfactual explanation by mixed-integer linear optimization., in: IJCAI, 2020, pp. 2855–2862

  29. [37]

    Höllig, A

    J. Höllig, A. F. Markus, et al., Semantic meaningfulness: Evaluating counterfactual approaches for real-world plausibility and feasibility, in: xAI, 2023, pp. 636–659

  30. [38]

    Moore, N

    J. Moore, N. Hammerla, C. Watkins, Explaining deep learning models with constrained adversarial examples, in: PRICAI, 2019, pp. 43–56

  31. [39]

    M. T. Lash, Q. Lin, N. Street, J. G. Robinson, J. Ohlmann, Generalized inverse classification, in: SDM, 2017, pp. 162–170

  32. [40]

    E. M. Kenny, M. T. Keane, On generating plausible counterfactual and semi-factual explanations for deep learning, in: AAAI, volume 35, 2021, pp. 11575–11585

  33. [41]

    M. T. Keane, B. Smyth, Good counterfactuals and where to find them: A case-based technique for generating counterfactuals for explainable ai (xai), in: ICCBR, 2020, pp. 163–178

  34. [42]

    B. Zhao, H. van der Aa, T. T. Nguyen, Q. V . H. Nguyen, M. Weidlich, Eires: E fficient integration of remote data in event stream processing, in: Proceedings of the 2021 International Conference on Management of Data, 2021, pp. 2128–2141

  35. [43]

    T. T. Huynh, C. T. Duong, T. T. Nguyen, V . T. Van, A. Sattar, H. Yin, Q. V . H. Nguyen, Network alignment with holistic embeddings, TKDE 35 (2021) 1881–1894

  36. [44]

    C. T. Duong, T. T. Nguyen, T.-D. Hoang, H. Yin, M. Weidlich, Q. V . H. Nguyen, Deep mincut: Learning node embeddings from detecting com- munities, Pattern Recognition (2022) 109126

  37. [45]

    T. T. Nguyen, T. C. Phan, M. H. Nguyen, M. Weidlich, H. Yin, J. Jo, Q. V . H. Nguyen, Model-agnostic and diverse explanations for streaming rumour graphs, Knowledge-Based Systems 253 (2022) 109438

  38. [46]

    T. T. Nguyen, T. T. Huynh, H. Yin, M. Weidlich, T. T. Nguyen, T. S. Mai, Q. V . H. Nguyen, Detecting rumours with latency guarantees using massive streaming data, The VLDB Journal (2022) 1–19

  39. [47]

    H. T. Trung, T. Van Vinh, N. T. Tam, J. Jo, H. Yin, N. Q. V . Hung, Learn- ing holistic interactions in lbsns with high-order, dynamic, and multi-role contexts, IEEE Transactions on Knowledge and Data Engineering 35 (2022) 5002–5016

  40. [48]

    T. T. Huynh, M. H. Nguyen, T. T. Nguyen, P. L. Nguyen, M. Weidlich, Q. V . H. Nguyen, K. Aberer, Efficient integration of multi-order dynamics and internal dynamics in stock movement prediction, in: Proceedings of the Sixteenth ACM International Conference on Web Search and Da...

  41. [49]

    D. C. Thang, H. T. Dat, N. T. Tam, J. Jo, N. Q. V . Hung, K. Aberer, Nature vs. nurture: Feature vs. structure for graph neural networks, PRL 159 (2022) 46–53

  42. [50]

    T. T. Nguyen, T. C. Phan, Q. V . H. Nguyen, K. Aberer, B. Stantic, Max- imal fusion of facts on the web with credibility guarantee, Information 7 Fusion 48 (2019) 55–66

  43. [51]

    T. T. Nguyen, T. T. Nguyen, T. T. Nguyen, B. V o, J. Jo, Q. V . H. Nguyen, Judo: Just-in-time rumour detection in streaming social platforms, Infor- mation Sciences 570 (2021) 70–93

  44. [52]

    T. T. Nguyen, T. T. Huynh, P. L. Nguyen, A. W.-C. Liew, H. Yin, Q. V . H. Nguyen, A survey of machine unlearning, arXiv preprint arXiv:2209.02299 (2022)

  45. [53]

    T. T. Nguyen, N. Quoc Viet Hung, T. T. Nguyen, T. T. Huynh, T. T. Nguyen, M. Weidlich, H. Yin, Manipulating recommender systems: A survey of poisoning attacks and countermeasures, ACM Computing Sur- veys 57 (2024) 1–39

  46. [54]

    T. T. Nguyen, T. T. Huynh, M. T. Pham, T. D. Hoang, T. T. Nguyen, Q. V . H. Nguyen, Validating functional redundancy with mixed generative adversarial networks, Knowledge-Based Systems 264 (2023) 110342

  47. [55]

    T. T. Nguyên, Debunking Misinformation on the Web: Detection, Valida- tion, and Visualisation, Technical Report, EPFL, 2019

  48. [56]

    Russell, E fficient search for diverse coherent explanations, in: FAccT, 2019, pp

    C. Russell, E fficient search for diverse coherent explanations, in: FAccT, 2019, pp. 20–28

  49. [57]

    M. E. Halabi, S. Mitrovic, A. Norouzi-Fard, J. Tardos, et al., Fairness in streaming submodular maximization: Algorithms and hardness, in: NIPS, 2020, pp. 1–14

  50. [58]

    Kulesza, B

    A. Kulesza, B. Taskar, et al., Determinantal point processes for machine learning, FTML 5 (2012) 123–286

  51. [59]

    Chakrabarti, S

    A. Chakrabarti, S. Kale, Submodular maximization meets streaming: Matchings, matroids, and more, Math. Program. 154 (2015) 225–247

  52. [60]

    Huang, N

    C.-C. Huang, N. Kakimura, S. Mauras, Y . Yoshida, Approximability of monotone submodular function maximization under cardinality and ma- troid constraints in the streaming model, arXiv preprint arXiv:2002.05477 (2020)

  53. [61]

    Celis, V

    E. Celis, V . Keswani, D. Straszak, A. Deshpande, T. Kathuria, N. Vishnoi, Fair and diverse dpp-based data summarization, in: ICML, 2018, pp. 716–725

  54. [62]

    S. Moro, P. Cortez, P. Rita, A data-driven approach to predict the success of bank telemarketing, Decision Support Systems 62 (2014) 22–31

  55. [63]

    Davenport, Lending club data analysis revisited with python, 2015

    K. Davenport, Lending club data analysis revisited with python, 2015

  56. [64]

    J. Chen, D. Chun, M. Patel, E. Chiang, J. James, The validity of syn- thetic clinical data: a validation study of a leading synthetic data generator (synthea) using clinical quality measures, MIDM 19 (2019) 1–9

  57. [65]

    Badanidiyuru, B

    A. Badanidiyuru, B. Mirzasoleiman, A. Karbasi, et al., Streaming sub- modular maximization: Massive data summarization on the fly, in: KDD, 2014, pp. 671–680

  58. [66]

    D. Ley, U. Bhatt, A. Weller, Diverse, global and amortised counterfactual explanations for uncertainty estimates, in: AAAI, 2022, pp. 7390–7398

  59. [67]

    Pawelczyk, T

    M. Pawelczyk, T. Datta, J. van den Heuvel, G. Kasneci, H. Lakkaraju, Probabilistically robust recourse: Navigating the trade-offs between costs and robustness in algorithmic recourse, in: ICLR, 2023

  60. [68]

    Guidotti, Counterfactual explanations and how to find them: literature review and benchmarking, DataMine (2022) 1–55

    R. Guidotti, Counterfactual explanations and how to find them: literature review and benchmarking, DataMine (2022) 1–55

  61. [69]

    Guidotti, S

    R. Guidotti, S. Ruggieri, Ensemble of counterfactual explainers, in: DS, 2021, pp. 358–368

  62. [70]

    T. T. Nguyen, Z. Ren, T. Pham, P. L. Nguyen, H. Yin, Q. V . H. Nguyen, Instruction-guided editing controls for images and multimedia: A survey in llm era, arXiv preprint arXiv:2411.09955 (2024)

  63. [71]

    T. T. Nguyen, T. T. Huynh, Z. Ren, T. T. Nguyen, P. L. Nguyen, H. Yin, Q. V . H. Nguyen, Privacy-preserving explainable ai: a survey, Science China Information Sciences 68 (2025) 111101

  64. [72]

    M. T. Pham, T. T. Huynh, T. T. Nguyen, T. T. Nguyen, T. T. Nguyen, J. Jo, H. Yin, Q. V . Hung Nguyen, A dual benchmarking study of facial forgery and facial forensics, CAAI Transactions on Intelligence Technology 9 (2024) 1377–1397

  65. [73]

    D. D. A. Nguyen, M. H. Nguyen, P. L. Nguyen, J. Jo, H. Yin, T. T. Nguyen, Multi-task learning of heterogeneous hypergraph representa- tions in lbsns, in: International Conference on Advanced Data Mining and Applications, Springer, 2024, pp. 161–177

  66. [74]

    T. T. Nguyen, T. T. Nguyen, M. Weidlich, J. Jo, Q. V . H. Nguyen, H. Yin, A. W.-C. Liew, Handling low homophily in recommender systems with partitioned graph transformer, IEEE Transactions on Knowledge and Data Engineering (2024)

  67. [75]

    Nguyen Thanh, N

    T. Nguyen Thanh, N. D. K. Quach, T. T. Nguyen, T. T. Huynh, V . H. Vu, P. L. Nguyen, J. Jo, Q. V . H. Nguyen, Poisoning gnn-based recommender systems with generative surrogate-based attacks, ACM Transactions on Information Systems 41 (2023) 1–24

  68. [76]

    T. T. Nguyen, T. C. Phan, H. T. Pham, T. T. Nguyen, J. Jo, Q. V . H. Nguyen, Example-based explanations for streaming fraud detection on graphs, Information Sciences 621 (2023) 319–340

  69. [77]

    Q. V . H. Nguyen, T. Nguyen Thanh, Z. Miklós, K. Aberer, Reconcil- ing schema matching networks through crowdsourcing, EAI Endorsed Transactions on Collaborative Computing 1 (2014) e2

  70. [78]

    Q. V . H. Nguyen, T. T. Nguyen, V . T. Chau, T. K. Wijaya, Z. Miklós, K. Aberer, A. Gal, M. Weidlich, Smart: A tool for analyzing and recon- ciling schema matching networks, in: ICDE, 2015, pp. 1488–1491

  71. [79]

    D. C. Thang, N. T. Tam, N. Q. V . Hung, K. Aberer, An evaluation of diversification techniques, in: DEXA, 2015, pp. 215–231

  72. [80]

    Q. V . H. Nguyen, S. T. Do, T. T. Nguyen, K. Aberer, Tag-based paper retrieval: minimizing user effort with diversity awareness, in: DASFAA, 2015, pp. 510–528

  73. [81]

    N. Q. V . Hung, M. Weidlich, N. T. Tam, Z. Miklós, K. Aberer, A. Gal, B. Stantic, Handling probabilistic integrity constraints in pay-as-you-go reconciliation of data models, Information Systems 83 (2019) 166–180

  74. [82]

    C. Yang, W. Yuan, L. Qu, T. T. Nguyen, Pdc-frs: Privacy-preserving data contribution for federated recommender system, in: International Conference on Advanced Data Mining and Applications, Springer, 2024, pp. 65–79

  75. [83]

    Sakong, V

    D. Sakong, V . H. Vu, T. T. Huynh, P. Le Nguyen, H. Yin, Q. V . H. Nguyen, T. T. Nguyen, Higher-order knowledge-enhanced recommen- dation with heterogeneous hypergraph multi-attention, Information Sci- ences 680 (2024) 121165

  76. [84]

    T. T. Huynh, T. B. Nguyen, P. L. Nguyen, T. T. Nguyen, M. Weidlich, Q. V . H. Nguyen, K. Aberer, Fast-fedul: A training-free federated un- learning with provable skew resilience, in: Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer, ...

  77. [85]

    T. T. Huynh, T. B. Nguyen, T. T. Nguyen, P. L. Nguyen, H. Yin, Q. V . H. Nguyen, T. T. Nguyen, Certified unlearning for federated recommenda- tion, ACM Transactions on Information Systems (2025)

  78. [86]

    T. T. Nguyen, T. T. Nguyen, T. H. Nguyen, H. Yin, T. T. Nguyen, J. Jo, Q. V . H. Nguyen, Isomorphic graph embedding for progressive maximal frequent subgraph mining, ACM Transactions on Intelligent Systems and Technology 15 (2023) 1–26

  79. [87]

    T. T. Nguyen, Z. Ren, T. T. Nguyen, J. Jo, Q. V . H. Nguyen, H. Yin, Portable graph-based rumour detection against multi-modal heterophily, Knowledge-Based Systems 284 (2024) 111310

  80. [88]

    Z. Ren, Y . Chang, T. T. Nguyen, Y . Tan, K. Qian, B. W. Schuller, A comprehensive survey on heart sound analysis in the deep learning era, IEEE Computational Intelligence Magazine 19 (2024) 42–57. 8

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.