REVIEW 4 major objections 5 minor 88 references
Model-Free Counterfactual Subset Selection at Scale
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper claims a single-pass streaming algorithm can select a diverse, label-balanced set of real counterfactual examples with logarithmic per-item updates and a constant-factor quality guarantee.
desk verdict The streaming problem is real and the preserve-set idea is plausible, but the utility functions break the submodularity assumption the whole guarantee rests on, so the paper's central claims do not hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The enabling object is the extensibility matroid $\hat{\mathcal{S}}$: the family of sets that can be extended to a feasible solution, characterized by per-label upper bounds $|S\cap D_l|\le \beta_l$ and the condition $\sum_l \max(|S\cap D_l|,\alpha_l)\le k$. By maintaining label counts $c_l$ and their capped sum $C$, the algorithm tests extensibility of a candidate extension in $O(1)$ time and finds the minimum-weight replaceable item using priority queues, which keeps each update at $O(\log k)$. The utility functions $f_1,f_2,f_3$ encode content-, sampling-, and clustering-based diversity; the guarantee applies to them only insofar as they are non-negative, monotone, and submodular.
What would settle it
Compute marginal gains for the three utilities on a small dataset and check whether $f(e\mid S)$ decreases as $S$ grows; for example, if adding an item to a larger set ever increases the marginal gain of another item, submodularity fails. For $f_2$, evaluate $\det(K_{S\cup\{e\}})-\det(K_S)$ on nested sets; non-increasing differences are required. If marginal-gain evaluation time scales with $k$, the $O(\log k)$ per-item update claim does not hold.
Extended reading notes
Core claim
The central claim is that selecting a diverse set of real counterfactuals can be solved in a single pass over the data, without storing the data or calling the classifier. The complete algorithm (Alg. 2) runs the streaming matroid-submodular maximization routine of Alg. 1 on the family of extensible sets, accepting an incoming item only when its marginal gain is at least $1+\lambda$ times the smallest gain of a removable candidate, while separately keeping up to $\alpha_l$ backup items per label so the final set satisfies the lower-bound constraints. The paper argues the returned set is feasible and inherits the warm-up algorithm's approximation ratio: the body's analysis gives $1/7.75$ at $\lambda=1$, and the contributions section announces $1/5.585$. It positions this as the first real-time, model-free counterfactual subset selection from streaming data.
Load-bearing premise
The approximation and complexity guarantees rely on the three proposed utility functions being non-negative, monotone, and submodular, and on the marginal gain of an item against a $k$-member set being computable in constant time; the paper assumes these properties rather than proving them.
Editorial extensions
If this is right
- If the guarantee holds, counterfactual explanations can be served in real time from unbounded streams, with no stored dataset and no access to the underlying model.
- The same matroid machinery applies to any streaming selection problem with per-group quotas, not only counterfactual explanations.
- The experiments indicate users would need less effort: on the Customer dataset the method reduces transport cost by up to 33.18% compared with streaming baselines.
- The algorithm keeps space $O(k)$, so a fixed-size summary of size $k$ is enough regardless of stream length.
- Because the relaxed version (no lower bounds) is a matroid-constrained submodular maximization, its approximation bound transfers to other matroid feasibility constraints.
Reading between the lines
- An implicit consequence, not tested in the paper, is that the same procedure could be used for other quota-constrained streaming summarization tasks, such as diverse news digests or representative panel construction.
- Since the guarantee depends on utility monotonicity and submodularity, a natural next step is to check which of the three proposed utilities satisfy these conditions; the determinant term in $f_2$ is the most likely place to fail.
- The discrepancy between the advertised $1/5.585$ and the proved $1/7.75$ suggests the final ratio depends on problem-specific curvature or $\lambda$ tuning; a direct derivation of $5.585$ from the analysis would settle which number is the guaranteed one.
- A user study measuring whether diverse real counterfactuals actually improve decision-making would be the natural behavioural test, since the paper's metrics are proxy costs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a streaming, model-free method for selecting a diverse and relevant subset of real counterfactual examples for a query item, subject to cardinality and label-balance constraints. The authors define three diversity-aware utility functions, present a greedy streaming algorithm with a swap threshold (Algorithm 1) and a complete algorithm that augments the result with lower-bound backup items (Algorithm 2), and claim a single-pass 1/5.585 approximation guarantee with O(log k) update complexity per item. The evaluation compares the method against offline, kNN, random, relaxed, and constraint-free baselines on three real datasets and a large synthetic dataset, reporting transport cost, constraint violations, runtime, and robustness to concept drift.
Significance. If the central algorithmic claims were correct, the paper would address a timely and practical gap: real-time, model-free counterfactual subset selection without storing the full dataset. The problem formulation, with label-diversity constraints and multiple diversity notions, is reasonable, and the experimental setup includes realistic datasets and a synthetic data generator at scale. However, the theoretical contribution is not currently supported: the utility functions used in the experiments are asserted to be monotone and submodular without proof, at least one of them is demonstrably not monotone, the claimed approximation constants are internally inconsistent, and the O(log k) per-item complexity is not established. The paper contains a useful heuristic and a substantial evaluation, but the advertised quality guarantee does not apply to the objective being optimized.
major comments (4)
- [§2.2, Eqs. (2)–(5); §3.2–3.3] The approximation guarantees for Algorithms 1 and 2 require the objective f to be non-negative, monotone, and submodular, but this property is never established for the proposed utilities. In fact, f1 in Eq. (2) is not monotone. For sim(e1,q)=0.9, sim(e2,q)=0.1, sim(e1,e2)=0.95, and λ1=0.5, Eq. (2) gives f1({e1})=0.9 and f1({e1,e2})=0.9+0.1−(0.5/4)·(2·0.95)=0.7625, so the marginal gain of e2 is −0.1375. Adding an item can therefore decrease the objective, contradicting the monotonicity assumption used in the greedy thresholding argument and in the augmentation step of §3.3. The determinant term in Eq. (3) and the |S|-scaled coverage term in Eq. (4) are not shown to be monotone submodular either; the assertion in §3.2 that the functions are 'well-designed' is not a substitute for proof. Because the central quality guarantee is stated for a monotone submodular objective, this is a load-bearing gap.
- [§3.2 and contribution bullet] The paper advertises a 1/5.585 approximation guarantee in the introduction, but §3.2 derives a minimum approximation factor ρ=7.75 and then states only 'check if 5.585/(1−cv(f))>7.75' without deriving this expression or relating it to the preceding analysis. If 5.585 is intended to follow from a curvature-aware bound, the relevant curvature parameter and proof are missing; if it is an empirical observation, it cannot be a worst-case guarantee. The text never reconciles the two numbers, so the stated quality guarantee is ambiguous.
- [§3.3, Runtime and memory] The O(log k) per-item claim is not supported by the algorithm description. Algorithm 1 makes two utility calls per item, and each marginal gain f(e|S) as defined in §2.2 requires interaction with all items currently in S; e.g., the pairwise sum in Eq. (2) is O(k) unless a sketch or index is described, and none is provided. The priority queues in §3.3 store item weights, but when S changes, the marginal gains—and hence the weights—of retained items generally change, and the paper does not explain how all affected weights are updated in O(log k) time. Without a mechanism to evaluate f(e|S) in O(1) or O(log k), the central complexity claim and the interpretation of Fig. 4d as O(log k) per item are unsubstantiated.
- [§4, Evaluation metrics] The empirical comparisons rely on metrics that largely coincide with components of the optimized objective. Transport cost (Eq. (6)) measures distance to the query and is essentially the similarity term appearing in Eqs. (2)–(5); constraint violations are enforced to zero by construction for the feasible algorithms; and Table 3 reports utility, which is the objective itself. The reported 'superior performance' is therefore partly self-referential and does not establish that the selected counterfactuals are more useful or actionable to humans than those of the baselines, a limitation the paper itself acknowledges in the case-study discussion. This weakens the empirical contribution, although it is not the primary reason for my recommendation.
minor comments (5)
- [§3.2] The sentence 'The complexity of involves runtime, utility calls, matroid queries, and space' is incomplete and should specify that it refers to Algorithm 1.
- [§3.3] The phrase 'Extending the idea in ?? to F being an extensibility matroid' contains a missing citation/reference marker '??' and needs to be completed.
- [§2.3] The solution space is defined with |S| < k, but all subsequent constraints and algorithms use |S| ≤ k; these should be made consistent.
- [Algorithm 2] The description of the augmentation step ('S = S 1 augmented with items in sets Pl') should specify how ties are broken when a label already satisfies its lower-bound constraint and multiple backup items are available.
- [Table 1 and Figure 4] The GeCo row in Table 1 contains an ambiguous combination of check/cross marks and a footnote; Figure 4 reports averages over 10 runs without error bars or variance information, despite the paper stating that variances are reported if appropriate.
Circularity Check
Algorithmic guarantees are adapted from independent streaming-submodular-maximization results, but two headline empirical claims—transport-cost superiority and zero constraint violations—are partly self-referential restatements of the optimized objective and the enforced solution space.
-
self definitional
[§2.2 Eq. (2) and §4 'Evaluation metrics' (Transport cost)]
"Content-based utility: f1(S ) = X e∈S sim(e, q)− λ1 |S|2 X e X e′,e sim(e, e′) ... The first term is similar to the proximity in DiCE [9] ... Transport cost: ... cost = 1 |S| X e∈S dcon(e, q) + dcat(e, q) ... Users prefer counterfactual examples similar to the query example, minimizing the effort required."
The dominant term of the optimized utility is the query-similarity term Σ sim(e,q), while the transport-cost metric is the average distance from S to q (and the paper defines dist = 1 − sim for the kernel in Eq. 3). Minimizing transport cost is therefore a monotone restatement of maximizing the similarity component of f1; reporting a transport-cost advantage over baselines is not an external check of explanation quality but a re-measurement of the very objective being optimized.
-
self definitional
[§2.3 Problem Statement, Alg. 2 (Line 9), and §4.1 'Constraint violations']
"The solution spaceS is: S = {S ⊆ D : |S| < k,α l ≤ |S∩ Dl| ≤ βl,∀l = 1,..., L}. ... Alg. 2 returns a feasible solution belonging toS with the same approximation ratio as Alg. 1. ... Our approach, like the offline algorithm, had no violations (not shown for brevity)."
The 'no constraint violations' result is enforced by the algorithm's definition: Alg. 2 adds backup items Pl exactly until αl ≤ |S∩Dl| and uses the extensibility matroid to respect βl and k, so every returned S lies in S by construction. Claiming zero violations as an empirical advantage is equivalent to asserting that the algorithm implements its own feasibility constraints.
full rationale
The formal derivation of the approximation ratio is not circular: Alg. 1 is explicitly inherited from Chakrabarti et al. and Huang et al. (independent streaming submodular-maximization results), and the quality argument in §3.2–3.3 is a standard reduction to MSIS/matroid submodular maximization. The utility functions are hand-designed rather than fitted, and the O(log k) data structure is an engineering contribution. The self-reference in the paper is concentrated in the evaluation: 'transport cost' largely inverts the query-similarity term of the optimized f1, and 'constraint violations' are impossible by the definition of the returned set. These two empirical claims are therefore partly self-definitional, which prevents a clean score of 0–2. However, the central algorithmic claim (one-pass, O(log k), 1/ρ guarantee for monotone submodular f) rests on independent external theory and is not itself a renamed input, so the paper is at score 6 rather than higher. Note also that the unsupported submodularity/monotonicity of f1–f3 is a correctness gap, not a circularity, and is not scored here.
Assumptions & free parameters
free parameters (7)
- lambda1 (content-based diversity weight)
- lambda2 (sampling-based diversity weight)
- lambda3 (clustering-based diversity weight)
- Greedy threshold lambda =
0.717 or 1
- Label lower and upper bounds alpha_l, beta_l =
e.g., 0.9*|D_l|/|D|*k and 1.1*|D_l|/|D|*k
- Similarity measure sim(.,.)
- Diagonal perturbation for kernel determinant =
unspecified
assumptions (4)
- ad hoc to paper The utility functions f1, f2, f3 are non-negative, monotone, and submodular.
- domain assumption A feasible solution exists: sum_l alpha_l <= k and alpha_l <= beta_l <= |D_l| for all l.
- standard math The family of extensible sets S_hat is a matroid.
- ad hoc to paper Marginal gains f(e|S) can be computed in O(1) or O(log k) time.
Cite this review
Pith. "Pith review of Model-Free Counterfactual Subset Selection at Scale." pith.science (2026). https://pith.science/paper/YUO2RFUB
@misc{pith2026250208326,
author = {Pith},
title = {Pith review of: Model-Free Counterfactual Subset Selection at Scale},
year = {2026},
howpublished = {\url{https://pith.science/paper/YUO2RFUB}},
note = {Machine review of arXiv:2502.08326}
}
abstract
Ensuring transparency in AI decision-making requires interpretable explanations, particularly at the instance level. Counterfactual explanations are a powerful tool for this purpose, but existing techniques frequently depend on synthetic examples, introducing biases from unrealistic assumptions, flawed models, or skewed data. Many methods also assume full dataset availability, an impractical constraint in real-time environments where data flows continuously. In contrast, streaming explanations offer adaptive, real-time insights without requiring persistent storage of the entire dataset. This work introduces a scalable, model-free approach to selecting diverse and relevant counterfactual examples directly from observed data. Our algorithm operates efficiently in streaming settings, maintaining $O(\log k)$ update complexity per item while ensuring high-quality counterfactual selection. Empirical evaluations on both real-world and synthetic datasets demonstrate superior performance over baseline methods, with robust behavior even under adversarial conditions.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Dwivedi, D
R. Dwivedi, D. Dave, H. Naik, S. Singhal, R. Omer, P. Patel, B. Qian, Z. Wen, T. Shah, G. Morgan, et al., Explainable ai (xai): Core ideas, techniques, and solutions, CSUR 55 (2023) 1–33
2023
-
[2]
Wachter, et al., Counterfactual explanations without opening the black box: Automated decisions and the gdpr, Harv
S. Wachter, et al., Counterfactual explanations without opening the black box: Automated decisions and the gdpr, Harv. JL & Tech. 31 (2017) 841
2017
-
[3]
Barocas, A
S. Barocas, A. D. Selbst, M. Raghavan, The hidden assumptions behind counterfactual explanations and principal reasons, in: FAccT, 2020, pp. 80–89
2020
-
[4]
Schleich, Z
M. Schleich, Z. Geng, Y . Zhang, D. Suciu, Geco: Quality counterfactual explanations in real time, PVLDB 14 (2021) 1681–1693
2021
-
[5]
Kirkpatrick, Battling algorithmic bias: how do we ensure algorithms treat us fairly?, CACM 59 (2016) 16–17
K. Kirkpatrick, Battling algorithmic bias: how do we ensure algorithms treat us fairly?, CACM 59 (2016) 16–17
2016
-
[6]
Karimi, G
A.-H. Karimi, G. Barthe, B. Schölkopf, I. Valera, A survey of algorithmic recourse: contrastive explanations and consequential recommendations, CSUR (2020)
2020
-
[7]
Explainable Image Classification with Evidence Counterfactual
T. Vermeire, D. Martens, Explainable image classification with evidence counterfactual, arXiv preprint arXiv:2004.07511 (2020)
work page Pith review arXiv 2020
-
[8]
Mehrabi, F
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, A. Galstyan, A survey on bias and fairness in machine learning, CSUR 54 (2021) 1–35
2021
Show all 88 references
-
[9]
R. K. Mothilal, A. Sharma, C. Tan, Explaining machine learning classi- fiers through diverse counterfactual explanations, in: FAccT, 2020, pp. 607–617
2020
-
[10]
Karimi, G
A.-H. Karimi, G. Barthe, B. Balle, I. Valera, Model-agnostic counterfac- tual explanations for consequential decisions, in: AISTATS, 2020, pp. 895–905
2020
-
[11]
N. Bui, D. Nguyen, V . A. Nguyen, Counterfactual plans under distribu- tional ambiguity, in: ICLR, 2022
2022
-
[12]
Nguyen, N
D. Nguyen, N. Bui, V . A. Nguyen, Feasible recourse plan via diverse interpolation, in: AISTATS, 2023, pp. 4679–4698
2023
-
[13]
T. T. Nguyen, Q. V . H. Nguyen, M. Weidlich, K. Aberer, Result selection and summarization for web table search, in: ICDE, 2015, pp. 231–242
2015
-
[14]
T. T. Nguyen, C. T. Duong, M. Weidlich, H. Yin, Q. V . H. Nguyen, Re- taining data from streams of social platforms with minimal regret, in: IJCAI, 2017, pp. 2850–2856
2017
-
[15]
N. T. Tam, M. Weidlich, B. Zheng, H. Yin, N. Q. V . Hung, B. Stantic, From anomaly detection to rumour detection using data streams of social platforms, Proceedings of the VLDB Endowment 12 (2019) 1016–1029
2019
-
[16]
T. T. Nguyen, T. D. Hoang, M. T. Pham, T. T. Vu, T. H. Nguyen, Q.-T. Huynh, J. Jo, Monitoring agriculture areas with satellite images and deep learning, Applied Soft Computing 95 (2020) 106565
2020
-
[17]
T. T. Nguyen, M. Weidlich, H. Yin, B. Zheng, Q. V . H. Nguyen, B. Stan- tic, User guidance for e fficient fact checking, Proceedings of the VLDB Endowment 12 (2019) 850–863
2019
-
[18]
N. T. Tam, H. T. Trung, H. Yin, T. Van Vinh, D. Sakong, B. Zheng, N. Q. V . Hung, Entity alignment for knowledge graphs with multi-order convolutional networks, TKDE 34 (2022) 4201–4214
2022
-
[19]
T. T. Nguyen, M. T. Pham, T. T. Nguyen, T. T. Huynh, Q. V . H. Nguyen, T. T. Quan, et al., Structural representation learning for network align- ment with self-supervised anchor links, Expert Systems with Applica- tions 165 (2021) 113857
2021
-
[20]
Stepin, J
I. Stepin, J. M. Alonso, A. Catala, M. Pereira-Fariña, A survey of con- trastive and counterfactual explanation generation methods for explain- able artificial intelligence, IEEE Access 9 (2021) 11974–12001
2021
-
[21]
C. T. Duong, T. T. Nguyen, H. Yin, M. Weidlich, T. S. Mai, K. Aberer, Q. V . H. Nguyen, E fficient and e ffective multi-modal queries through heterogeneous network embedding, IEEE Transactions on Knowledge and Data Engineering 34 (2022) 5307–5320
2022
-
[22]
T. T. Nguyen, M. Weidlich, H. Yin, B. Zheng, Q. H. Nguyen, Q. V . H. Nguyen, Factcatch: Incremental pay-as-you-go fact checking with min- imal user e ffort, in: Proceedings of the 43rd International ACM SI- GIR Conference on Research and Development in Information Retrieval, 2...
2020
-
[23]
N. Q. V . Hung, D. C. Thang, N. T. Tam, M. Weidlich, K. Aberer, H. Yin, X. Zhou, Answer validation for generic crowdsourcing tasks with mini- mal efforts, The VLDB Journal 26 (2017) 855–880
2017
-
[24]
Q. V . H. Nguyen, C. T. Duong, T. T. Nguyen, M. Weidlich, K. Aberer, H. Yin, X. Zhou, Argument discovery via crowdsourcing, The VLDB Journal 26 (2017) 511–535
2017
-
[25]
Z. Ren, T. T. Nguyen, W. Nejdl, Prototype learning for interpretable respiratory sound analysis, in: Proc. ICASSP, 2022, pp. 9087–9091
2022
-
[26]
Q. V . H. Nguyen, K. Zheng, M. Weidlich, B. Zheng, H. Yin, T. T. Nguyen, B. Stantic, What-if analysis with conflicting goals: Recommending data ranges for exploration, in: ICDE, 2018, pp. 89–100
2018
-
[27]
N. T. Toan, P. T. Cong, N. T. Tam, N. Q. V . Hung, B. Stantic, Diversifying group recommendation, IEEE Access 6 (2018) 17776–17786
2018
-
[28]
Verma, V
S. Verma, V . Boonsanong, M. Hoang, K. E. Hines, J. P. Dickerson, C. Shah, Counterfactual explanations and algorithmic recourses for ma- chine learning: A review, arXiv preprint arXiv:2010.10596 (2020)
2020 arXiv
-
[29]
R. M. Byrne, Counterfactuals in explainable artificial intelligence (xai): Evidence from human reasoning., in: IJCAI, 2019, pp. 6276–6282
2019
-
[30]
Mertes, T
S. Mertes, T. Huber, K. Weitz, A. Heimerl, E. André, Ganterfac- tual—counterfactual explanations for medical non-experts using gener- ative adversarial learning, Frontiers in artificial intelligence 5 (2022) 825565
2022
-
[31]
J. Ma, R. Guo, S. Mishra, A. Zhang, J. Li, CLEAR: generative counter- factual explanations on graphs, in: NeurIPS, 2022
2022
-
[32]
Dutta, J
S. Dutta, J. Long, S. Mishra, C. Tilli, D. Magazzeni, Robust counterfac- tual explanations for tree-based ensembles, in: ICML, volume 162, 2022, pp. 5742–5756
2022
-
[33]
Ustun, A
B. Ustun, A. Spangher, Y . Liu, Actionable recourse in linear classifica- tion, in: FAT*, 2019, pp. 10–19
2019
-
[34]
Van Looveren, J
A. Van Looveren, J. Klaise, Interpretable counterfactual explanations guided by prototypes, in: ECML PKDD, 2021, pp. 650–665
2021
-
[35]
Poyiadzi, K
R. Poyiadzi, K. Sokol, R. Santos-Rodriguez, T. De Bie, P. Flach, Face: feasible and actionable counterfactual explanations, in: AIES, 2020, pp. 344–350
2020
-
[36]
Kanamori, et al., Dace: Distribution-aware counterfactual explanation by mixed-integer linear optimization., in: IJCAI, 2020, pp
K. Kanamori, et al., Dace: Distribution-aware counterfactual explanation by mixed-integer linear optimization., in: IJCAI, 2020, pp. 2855–2862
2020
-
[37]
Höllig, A
J. Höllig, A. F. Markus, et al., Semantic meaningfulness: Evaluating counterfactual approaches for real-world plausibility and feasibility, in: xAI, 2023, pp. 636–659
2023
-
[38]
Moore, N
J. Moore, N. Hammerla, C. Watkins, Explaining deep learning models with constrained adversarial examples, in: PRICAI, 2019, pp. 43–56
2019
-
[39]
M. T. Lash, Q. Lin, N. Street, J. G. Robinson, J. Ohlmann, Generalized inverse classification, in: SDM, 2017, pp. 162–170
2017
-
[40]
E. M. Kenny, M. T. Keane, On generating plausible counterfactual and semi-factual explanations for deep learning, in: AAAI, volume 35, 2021, pp. 11575–11585
2021
-
[41]
M. T. Keane, B. Smyth, Good counterfactuals and where to find them: A case-based technique for generating counterfactuals for explainable ai (xai), in: ICCBR, 2020, pp. 163–178
2020
-
[42]
B. Zhao, H. van der Aa, T. T. Nguyen, Q. V . H. Nguyen, M. Weidlich, Eires: E fficient integration of remote data in event stream processing, in: Proceedings of the 2021 International Conference on Management of Data, 2021, pp. 2128–2141
2021
-
[43]
T. T. Huynh, C. T. Duong, T. T. Nguyen, V . T. Van, A. Sattar, H. Yin, Q. V . H. Nguyen, Network alignment with holistic embeddings, TKDE 35 (2021) 1881–1894
2021
-
[44]
C. T. Duong, T. T. Nguyen, T.-D. Hoang, H. Yin, M. Weidlich, Q. V . H. Nguyen, Deep mincut: Learning node embeddings from detecting com- munities, Pattern Recognition (2022) 109126
2022
-
[45]
T. T. Nguyen, T. C. Phan, M. H. Nguyen, M. Weidlich, H. Yin, J. Jo, Q. V . H. Nguyen, Model-agnostic and diverse explanations for streaming rumour graphs, Knowledge-Based Systems 253 (2022) 109438
2022
-
[46]
T. T. Nguyen, T. T. Huynh, H. Yin, M. Weidlich, T. T. Nguyen, T. S. Mai, Q. V . H. Nguyen, Detecting rumours with latency guarantees using massive streaming data, The VLDB Journal (2022) 1–19
2022
-
[47]
H. T. Trung, T. Van Vinh, N. T. Tam, J. Jo, H. Yin, N. Q. V . Hung, Learn- ing holistic interactions in lbsns with high-order, dynamic, and multi-role contexts, IEEE Transactions on Knowledge and Data Engineering 35 (2022) 5002–5016
2022
-
[48]
T. T. Huynh, M. H. Nguyen, T. T. Nguyen, P. L. Nguyen, M. Weidlich, Q. V . H. Nguyen, K. Aberer, Efficient integration of multi-order dynamics and internal dynamics in stock movement prediction, in: Proceedings of the Sixteenth ACM International Conference on Web Search and Da...
2023
-
[49]
D. C. Thang, H. T. Dat, N. T. Tam, J. Jo, N. Q. V . Hung, K. Aberer, Nature vs. nurture: Feature vs. structure for graph neural networks, PRL 159 (2022) 46–53
2022
-
[50]
T. T. Nguyen, T. C. Phan, Q. V . H. Nguyen, K. Aberer, B. Stantic, Max- imal fusion of facts on the web with credibility guarantee, Information 7 Fusion 48 (2019) 55–66
2019
-
[51]
T. T. Nguyen, T. T. Nguyen, T. T. Nguyen, B. V o, J. Jo, Q. V . H. Nguyen, Judo: Just-in-time rumour detection in streaming social platforms, Infor- mation Sciences 570 (2021) 70–93
2021
-
[52]
T. T. Nguyen, T. T. Huynh, P. L. Nguyen, A. W.-C. Liew, H. Yin, Q. V . H. Nguyen, A survey of machine unlearning, arXiv preprint arXiv:2209.02299 (2022)
2022 arXiv
-
[53]
T. T. Nguyen, N. Quoc Viet Hung, T. T. Nguyen, T. T. Huynh, T. T. Nguyen, M. Weidlich, H. Yin, Manipulating recommender systems: A survey of poisoning attacks and countermeasures, ACM Computing Sur- veys 57 (2024) 1–39
2024
-
[54]
T. T. Nguyen, T. T. Huynh, M. T. Pham, T. D. Hoang, T. T. Nguyen, Q. V . H. Nguyen, Validating functional redundancy with mixed generative adversarial networks, Knowledge-Based Systems 264 (2023) 110342
2023
-
[55]
T. T. Nguyên, Debunking Misinformation on the Web: Detection, Valida- tion, and Visualisation, Technical Report, EPFL, 2019
2019
-
[56]
Russell, E fficient search for diverse coherent explanations, in: FAccT, 2019, pp
C. Russell, E fficient search for diverse coherent explanations, in: FAccT, 2019, pp. 20–28
2019
-
[57]
M. E. Halabi, S. Mitrovic, A. Norouzi-Fard, J. Tardos, et al., Fairness in streaming submodular maximization: Algorithms and hardness, in: NIPS, 2020, pp. 1–14
2020
-
[58]
Kulesza, B
A. Kulesza, B. Taskar, et al., Determinantal point processes for machine learning, FTML 5 (2012) 123–286
2012
-
[59]
Chakrabarti, S
A. Chakrabarti, S. Kale, Submodular maximization meets streaming: Matchings, matroids, and more, Math. Program. 154 (2015) 225–247
2015
-
[60]
Huang, N
C.-C. Huang, N. Kakimura, S. Mauras, Y . Yoshida, Approximability of monotone submodular function maximization under cardinality and ma- troid constraints in the streaming model, arXiv preprint arXiv:2002.05477 (2020)
2020 arXiv
-
[61]
Celis, V
E. Celis, V . Keswani, D. Straszak, A. Deshpande, T. Kathuria, N. Vishnoi, Fair and diverse dpp-based data summarization, in: ICML, 2018, pp. 716–725
2018
-
[62]
S. Moro, P. Cortez, P. Rita, A data-driven approach to predict the success of bank telemarketing, Decision Support Systems 62 (2014) 22–31
2014
-
[63]
Davenport, Lending club data analysis revisited with python, 2015
K. Davenport, Lending club data analysis revisited with python, 2015
2015
-
[64]
J. Chen, D. Chun, M. Patel, E. Chiang, J. James, The validity of syn- thetic clinical data: a validation study of a leading synthetic data generator (synthea) using clinical quality measures, MIDM 19 (2019) 1–9
2019
-
[65]
Badanidiyuru, B
A. Badanidiyuru, B. Mirzasoleiman, A. Karbasi, et al., Streaming sub- modular maximization: Massive data summarization on the fly, in: KDD, 2014, pp. 671–680
2014
-
[66]
D. Ley, U. Bhatt, A. Weller, Diverse, global and amortised counterfactual explanations for uncertainty estimates, in: AAAI, 2022, pp. 7390–7398
2022
-
[67]
Pawelczyk, T
M. Pawelczyk, T. Datta, J. van den Heuvel, G. Kasneci, H. Lakkaraju, Probabilistically robust recourse: Navigating the trade-offs between costs and robustness in algorithmic recourse, in: ICLR, 2023
2023
-
[68]
Guidotti, Counterfactual explanations and how to find them: literature review and benchmarking, DataMine (2022) 1–55
R. Guidotti, Counterfactual explanations and how to find them: literature review and benchmarking, DataMine (2022) 1–55
2022
-
[69]
Guidotti, S
R. Guidotti, S. Ruggieri, Ensemble of counterfactual explainers, in: DS, 2021, pp. 358–368
2021
-
[70]
T. T. Nguyen, Z. Ren, T. Pham, P. L. Nguyen, H. Yin, Q. V . H. Nguyen, Instruction-guided editing controls for images and multimedia: A survey in llm era, arXiv preprint arXiv:2411.09955 (2024)
2024 arXiv
-
[71]
T. T. Nguyen, T. T. Huynh, Z. Ren, T. T. Nguyen, P. L. Nguyen, H. Yin, Q. V . H. Nguyen, Privacy-preserving explainable ai: a survey, Science China Information Sciences 68 (2025) 111101
2025
-
[72]
M. T. Pham, T. T. Huynh, T. T. Nguyen, T. T. Nguyen, T. T. Nguyen, J. Jo, H. Yin, Q. V . Hung Nguyen, A dual benchmarking study of facial forgery and facial forensics, CAAI Transactions on Intelligence Technology 9 (2024) 1377–1397
2024
-
[73]
D. D. A. Nguyen, M. H. Nguyen, P. L. Nguyen, J. Jo, H. Yin, T. T. Nguyen, Multi-task learning of heterogeneous hypergraph representa- tions in lbsns, in: International Conference on Advanced Data Mining and Applications, Springer, 2024, pp. 161–177
2024
-
[74]
T. T. Nguyen, T. T. Nguyen, M. Weidlich, J. Jo, Q. V . H. Nguyen, H. Yin, A. W.-C. Liew, Handling low homophily in recommender systems with partitioned graph transformer, IEEE Transactions on Knowledge and Data Engineering (2024)
2024
-
[75]
Nguyen Thanh, N
T. Nguyen Thanh, N. D. K. Quach, T. T. Nguyen, T. T. Huynh, V . H. Vu, P. L. Nguyen, J. Jo, Q. V . H. Nguyen, Poisoning gnn-based recommender systems with generative surrogate-based attacks, ACM Transactions on Information Systems 41 (2023) 1–24
2023
-
[76]
T. T. Nguyen, T. C. Phan, H. T. Pham, T. T. Nguyen, J. Jo, Q. V . H. Nguyen, Example-based explanations for streaming fraud detection on graphs, Information Sciences 621 (2023) 319–340
2023
-
[77]
Q. V . H. Nguyen, T. Nguyen Thanh, Z. Miklós, K. Aberer, Reconcil- ing schema matching networks through crowdsourcing, EAI Endorsed Transactions on Collaborative Computing 1 (2014) e2
2014
-
[78]
Q. V . H. Nguyen, T. T. Nguyen, V . T. Chau, T. K. Wijaya, Z. Miklós, K. Aberer, A. Gal, M. Weidlich, Smart: A tool for analyzing and recon- ciling schema matching networks, in: ICDE, 2015, pp. 1488–1491
2015
-
[79]
D. C. Thang, N. T. Tam, N. Q. V . Hung, K. Aberer, An evaluation of diversification techniques, in: DEXA, 2015, pp. 215–231
2015
-
[80]
Q. V . H. Nguyen, S. T. Do, T. T. Nguyen, K. Aberer, Tag-based paper retrieval: minimizing user effort with diversity awareness, in: DASFAA, 2015, pp. 510–528
2015
-
[81]
N. Q. V . Hung, M. Weidlich, N. T. Tam, Z. Miklós, K. Aberer, A. Gal, B. Stantic, Handling probabilistic integrity constraints in pay-as-you-go reconciliation of data models, Information Systems 83 (2019) 166–180
2019
-
[82]
C. Yang, W. Yuan, L. Qu, T. T. Nguyen, Pdc-frs: Privacy-preserving data contribution for federated recommender system, in: International Conference on Advanced Data Mining and Applications, Springer, 2024, pp. 65–79
2024
-
[83]
Sakong, V
D. Sakong, V . H. Vu, T. T. Huynh, P. Le Nguyen, H. Yin, Q. V . H. Nguyen, T. T. Nguyen, Higher-order knowledge-enhanced recommen- dation with heterogeneous hypergraph multi-attention, Information Sci- ences 680 (2024) 121165
2024
-
[84]
T. T. Huynh, T. B. Nguyen, P. L. Nguyen, T. T. Nguyen, M. Weidlich, Q. V . H. Nguyen, K. Aberer, Fast-fedul: A training-free federated un- learning with provable skew resilience, in: Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer, ...
2024
-
[85]
T. T. Huynh, T. B. Nguyen, T. T. Nguyen, P. L. Nguyen, H. Yin, Q. V . H. Nguyen, T. T. Nguyen, Certified unlearning for federated recommenda- tion, ACM Transactions on Information Systems (2025)
2025
-
[86]
T. T. Nguyen, T. T. Nguyen, T. H. Nguyen, H. Yin, T. T. Nguyen, J. Jo, Q. V . H. Nguyen, Isomorphic graph embedding for progressive maximal frequent subgraph mining, ACM Transactions on Intelligent Systems and Technology 15 (2023) 1–26
2023
-
[87]
T. T. Nguyen, Z. Ren, T. T. Nguyen, J. Jo, Q. V . H. Nguyen, H. Yin, Portable graph-based rumour detection against multi-modal heterophily, Knowledge-Based Systems 284 (2024) 111310
2024
-
[88]
Z. Ren, Y . Chang, T. T. Nguyen, Y . Tan, K. Qian, B. W. Schuller, A comprehensive survey on heart sound analysis in the deep learning era, IEEE Computational Intelligence Magazine 19 (2024) 42–57. 8
2024
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.