Pith. sign in

REVIEW 4 major objections 5 minor 111 references

MEDebiaser: A Human-AI Feedback System for Mitigating Bias in Multi-label Medical Image Classification

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Adding an attention loss that steers a multi-label chest X-ray classifier's Grad-CAM heatmaps toward physician-drawn polygons raises rare-label AUC from 0.613 to 0.675, and the surrounding system lets physicians do this without ML training.

desk verdict A genuinely useful HCI systems paper with a solid user study and an honest limitations section, but the central bias-mitigation claim rests on one un-replicated AUC gain that the paper over-asserts. read the letter →

arxiv 2507.10044 v3 pith:NRI576BW submitted 2025-07-14 cs.HC

classification cs.HC
keywords multi-labelclassificationmedicalimagehuman-AIfeedbackinteractivemachinelearningattentionlossGrad-CAMbiasmitigationexplainableAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MEDebiaser is an interactive system that lets physicians correct a multi-label medical image classifier by drawing polygons on the model's Grad-CAM heatmaps, with no machine-learning expertise required. The paper's central claim is that fine-tuning with a combination of prediction loss and attention loss—where the attention loss is the mean-squared error between the model's heatmap and the physician's mask—meaningfully reduces bias for underrepresented labels. The mechanism study supports this with a rare-label AUC for Pleural_Thickening that rises from 0.613 with prediction loss alone to 0.675 when attention loss is added. A user study on an ear-endoscopy dataset adds evidence that both physicians and engineers find the system usable, that it lowers the annotation workload by ranking the most problematic images first, and that it shortens the physician-engineer feedback loop.

What carries the argument

The load-bearing object is the dynamically weighted joint loss: prediction loss (taken from the explanation-guided learning framework, as in MAGI) plus an attention loss defined as the mean squared error between the model's Grad-CAM heatmap and the physician-drawn polygon mask, with attention weight scaled by label frequency to protect rare labels. Grad-CAM, a heatmap overlay of the image regions most influential for a label's prediction, serves not only as the explanation shown to the physician but also as the optimization target. Complementing the loss is a customized three-tier ranking strategy—label prediction accuracy, heatmap concentration, and inverse-frequency co-occurrence dependency—that prioritizes which validation images the physician should annotate, reducing the annotation budget and making the iterative fine-tuning loop feasible.

What would settle it

Fine-tune the same ChestX-ray14 model with deliberately wrong physician masks (random or shifted polygons) under the joint prediction-plus-attention loss, and compare AUC on Pleural_Thickening: if wrong masks still raise AUC to 0.675, the improvement is not caused by clinical attention correction; if they drop it below the 0.613 prediction-only baseline, the attention mechanism is doing real work.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes a new division of labor: instead of physicians reporting model errors to engineers who then adjust losses or data, physicians directly annotate correct regions on validation images using polygon masks, and those masks become the optimization target. The model is fine-tuned with a dynamically weighted joint loss composed of prediction loss plus attention loss, where the attention loss is the mean squared error between the Grad-CAM heatmap and the physician-provided mask. On a one-tenth subset of ChestX-ray14, this raises AUC for the rare label Pleural_Thickening from 0.613 to 0.675, with precision rising from 0.140 to 0.160, while a comparable prediction-loss-only fine-tuning stays at 0.613. The user study with otolaryngologists then shows six physician-engineer teams completing bias-correction rounds on the EOF label, with all six physicians able to operate the system without ML background and with engineers reporting reduced workload. The conclusion the authors draw is that local explanation can carry expert feedback directly into the training loop, reducing both bias and reliance on engineers.

Load-bearing premise

The whole approach depends on the heatmap overlay honestly showing which image regions the model used for that label; if the heatmap points at the wrong places, physicians will correct the wrong areas and the fine-tuning will lock the bias in.

Editorial extensions

If this is right

  • Rare-label bias in a multi-label chest X-ray classifier can be reduced by physician attention correction alone, without expanding the dataset or reweighting samples.
  • A physician with no ML background can run the whole observe-annotate-fine-tune loop, so model iteration no longer has to wait for an engineer.
  • Because the attention loss is model-agnostic for CNN backbones with Grad-CAM, the same fine-tuning recipe transfers to other multi-label medical imaging tasks, such as dermatology, digital pathology, and brain MRI.
  • Ranking images by prediction accuracy, heatmap concentration, and co-occurrence dependency lets a fixed annotation budget buy more bias reduction than random annotation, as the breadth-and-depth experiment shows.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If attention loss genuinely steers the model's heatmaps toward physician masks on held-out images, then the same mechanism should also improve localization quality, such as the overlap between final heatmaps and physician segmentations, a metric the paper does not report; that is a testable extension.
  • The co-occurrence-dependency ranking is model-driven, and the paper's own limitation discussion notes physicians often prefer clinically similar cases; an untested alternative is a medical-feature-similarity ranking, which could change which images get annotated and how much bias drops per annotation.
  • The mechanism study isolates one rare label on a one-tenth data subset; generalizing to several rare labels on full ChestX-ray14 and on the EarEndo dataset would establish whether the AUC gain is a systematic property of attention-guided fine-tuning or specific to Pleural_Thickening.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents MEDebiaser, an interactive human-in-the-loop system for multi-label medical image classification. The system allows physicians to inspect Grad-CAM explanations, annotate biased images with polygon masks, and fine-tune the model with a joint prediction-plus-attention loss. The authors report a formative study with five physicians and four engineers, a mechanism study on ChestX-ray14 showing an AUC improvement for the rare label Pleural_Thickening (0.613 to 0.675) when attention loss is added, and a user study with six physician-engineer teams on an ear endoscopy dataset, reporting subjective usability and collaboration benefits. The paper claims that MEDebiaser enables physicians to mitigate label bias without technical expertise and reduces the need for engineers in the revision loop.

Significance. If the central claims were robustly established, MEDebiaser would be a meaningful HCI contribution: it offers a concrete, intuitive way for clinicians to inject spatial expertise into multi-label medical classifiers, with a plausible attention-alignment mechanism. The formative study and interface design are thoughtful, and the user study provides encouraging usability signals. However, the quantitative evidence is currently insufficient. The headline AUC gain rests on a single un-replicated run for one label, the Grad-CAM faithfulness assumption is not validated, and the user study lacks a control condition and objective model metrics. The manuscript does not ship code or data, further limiting reproducibility. The core idea is defensible, but the empirical support is too thin to justify the abstract's claim of demonstrated bias reduction.

major comments (4)
  1. [5.1.2, Table 3] The central quantitative result, AUC for Pleural_Thickening improving from 0.613 to 0.675, is based on a single fine-tuning run on a single label with no reported variance, confidence intervals, or significance tests. Given typical run-to-run variance in fine-tuning deep networks, this difference may not be reproducible. The abstract's claim that the system 'effectively reduces biases' requires multiple random seeds and a statistical test (e.g., a paired test over seeds). Without this, the observed gain could be seed noise or an artifact of the particular 100 annotated images.
  2. [4.2.1-4.2.2] The mechanism assumes that Grad-CAM heatmaps from the last convolutional layer provide faithful per-label localization that physicians can correct and that the attention loss can then steer. The paper provides no localization evaluation (e.g., IoU between heatmaps and anatomical regions or physician masks) and no evidence that Grad-CAM is reliable in this multi-label chest X-ray setting, where co-occurring labels such as Effusion and Pleural_Thickening may produce overlapping or diffuse heatmaps. If Grad-CAM highlights the wrong or overly broad regions, physicians may correct non-biased attention, and the attention loss could reinforce rather than mitigate the intended bias. Please add a sanity check of Grad-CAM faithfulness, or cite existing validation studies for this architecture and dataset.
  3. [5.2-5.3] The user study has no control condition and reports no objective model metrics after fine-tuning. All claims about bias mitigation, collaboration efficiency, and workload reduction are based on self-reported Likert scores without significance tests or comparisons to a baseline workflow. Consequently, the study cannot support the statement that MEDebiaser 'effectively reduces biases' in real use; it only shows that participants found the system usable. At minimum, report per-team model performance (e.g., AUC or F1 for the target label) before and after each round of physician annotation, and compare against a control team using the model without the interactive feedback loop.
  4. [4.2.2-4.2.3] The joint loss function is described qualitatively, but the attention loss weight and the schedule for 'dynamically weighted' adjustment are not specified. Similarly, the positive-label co-occurrence threshold in Eq. (2) and the stability constant 0.01 are defined, but the threshold value itself is not reported. These degrees of freedom make the results in Tables 3 and 4 difficult to reproduce or interpret, and they are load-bearing because the attention-loss contribution is exactly what the mechanism study claims to validate.
minor comments (5)
  1. [Figure 3] The label 'Co-occurance' should be spelled 'Co-occurrence'.
  2. [Table 5] The footnotes '1 Experience' and '2 Experience' are ambiguous; clarify that they refer to experience with engineers/AI and with physicians, respectively.
  3. [Eq. (2)] The summation index in the denominator of Eq. (2) is not explicitly restricted to positive labels; please specify that the sum runs over the same positive-label set as the numerator.
  4. [5.3.1] Report standard deviations (and ideally confidence intervals) for the Likert-scale results; the current text gives only means, which obscures the spread across the six participants.
  5. [5.2.2] Abbreviations such as EOF, OME, and EOM are introduced in Table 13 and the results section but not defined at first use; define them near the dataset description.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found.

full rationale

The paper's central claim—that physicians can reduce multi-label bias by correcting Grad-CAM attention with polygon masks—is not circular. The quantitative support is Experiment I (§5.1.2), where 100 physician-annotated Pleural_Thickening images are used to fine-tune with attention loss (MSE between Grad-CAM heatmap and physician mask) versus prediction loss alone, and AUC on a held-out test set improves from 0.613 to 0.675. The attention-loss target (physician polygon) is not the evaluation target (test-set AUC), so the improvement is not forced by construction. The prediction-loss baseline controls for the extra supervision, further reducing any concern that the gain merely reflects additional labeled data. The paper's other claims (usability, workload, collaboration) rest on user-study questionnaires and interviews, which are subjective but not definitionally tied to the system's outputs. Self-citations (e.g., [63], [64], [85]) appear only in related work and discussion and are not load-bearing for the mechanism. The acknowledged limitation in §6.3.1—that the recommendation strategy is model-driven rather than clinically validated—is a validity concern, not a circularity. The unverified assumption that Grad-CAM faithfully localizes per-label decisions is an empirical risk, not a reduction of the derivation to its own inputs.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central result depends on assumed validity of Grad-CAM, physician annotations, and report-derived labels. The attention loss weighting and co-occurrence thresholds are unspecified free choices. No new physical entities are introduced.

free parameters (3)
  • Attention loss weight lambda_att = not reported; dynamically scaled by label frequency
    Joint loss in Section 4.2.2 balances prediction and attention terms, but the exact weight schedule is not given and is tuned by the authors.
  • Positive-label co-occurrence threshold = not reported
    Equation (2) in Section 4.2.3 includes only labels whose co-occurrence with the target exceeds a defined threshold; the threshold value is unspecified.
  • Inverse frequency stability constant = 0.01
    Added to Equation (1) to prevent division by zero; chosen manually and affects dependency scores.
assumptions (4)
  • domain assumption Grad-CAM faithfully localizes model decisions in multi-label medical images.
    Invoked in Section 4.2.1 as the basis for physician review and as the target for attention loss in Section 4.2.2.
  • domain assumption Physician polygon annotations identify correct regions for model attention.
    The fine-tuning loss in Section 4.2.2 assumes that aligning the Grad-CAM heatmap with physician masks improves generalization.
  • domain assumption ChestX-ray14 and EarEndo labels extracted from reports are reliable ground truth.
    Used without independent verification in Sections 4.1.1 and 5.2.2.
  • domain assumption ImageNet pretrained DenseNet provides a suitable starting point for chest X-ray and ear endoscopy classification.
    Adopted in Sections 4.1.2 and 5.1.1 as standard transfer learning practice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MEDebiaser: A Human-AI Feedback System for Mitigating Bias in Multi-label Medical Image Classification." pith.science (2026). https://pith.science/paper/NRI576BW

@misc{pith2026250710044,
  author       = {Pith},
  title        = {Pith review of: MEDebiaser: A Human-AI Feedback System for Mitigating Bias in Multi-label Medical Image Classification},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NRI576BW}},
  note         = {Machine review of arXiv:2507.10044}
}
read the original abstract

Medical images often contain multiple labels with imbalanced distributions and co-occurrence, leading to bias in multi-label medical image classification. Close collaboration between medical professionals and machine learning practitioners has significantly advanced medical image analysis. However, traditional collaboration modes struggle to facilitate effective feedback between physicians and AI models, as integrating medical expertise into the training process via engineers can be time-consuming and labor-intensive. To bridge this gap, we introduce MEDebiaser, an interactive system enabling physicians to directly refine AI models using local explanations. By combining prediction with attention loss functions and employing a customized ranking strategy to alleviate scalability, MEDebiaser allows physicians to mitigate biases without technical expertise, reducing reliance on engineers, and thus enhancing more direct human-AI feedback. Our mechanism and user studies demonstrate that it effectively reduces biases, improves usability, and enhances collaboration efficiency, providing a practical solution for integrating medical expertise into AI-driven healthcare.

Figures

Figures reproduced from arXiv: 2507.10044 by the authors.

Figure 1
Figure 1. Traditional modeling tasks face challenges such as the high workload of annotating entire datasets and the inefficiencies [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. A The Panel View provides components for uploading datasets, selecting models, and setting training parameters. B The Label View includes a table displaying label distribution and a chord diagram showing co-occurrence. C The Attention View displays local explanations for the selected labels. D The Modification View includes D1 an Editing Area for fine-tuning attention, D2 a Recommendation Area for sorting based on m… view at source ↗
Figure 3
Figure 3. The MEDebiaser workflow includes three main stages: Loading Dataset and Model, Observing and Modifying Attention, and Evaluating Model Performance. maps, offering clearer and more intuitive explanations for physi￾cians. Specifically, Grad-CAM calculates the gradients of a target label with respect to the feature maps of the selected convolutional layer. These gradients are then globally averaged to determine the imp… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: The procedure of the User Study [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Likert Results on [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 7
Figure 7. Figure 7: The results of two rounds of fine-tuning on EOF [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 6
Figure 6. Figure 6: 1 Original Image. 2 Polygon annotated by UD5. 3 Polygon annotated by UD4. 5.3.2 RQ5: What’s the impact of the system on workload, physi￾cian diagnosis, collaboration, and medical knowledge integration? MEDebiaser’s interface effectively facilitates iterative feed￾back …
Figure 8
Figure 8. Figure 8: Likert Results on Feedback & Communication and Workload. lead, with engineers offering support. Both engineers and physi￾cians generally agree that MEDebiaser facilitates their collaboration (physicians: mean = 5.33; engineers: mean = 5.5). UP5 mentioned, “This essenti…
Figure 9
Figure 9. Figure 9: The label distribution of ChestX-ray14 [PITH_FULL_IMAGE:figures/full_fig_p025_9.png]
Figure 10
Figure 10. Figure 10: The comparison of different visualization and interpretability methods is as follows: [PITH_FULL_IMAGE:figures/full_fig_p026_10.png]
Figure 11
Figure 11. Figure 11: The label distribution of EarEndo [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

111 extracted references · 39 canonical work pages

  1. [1]

    and Subhashini R

    Saranya A. and Subhashini R. 2023. A systematic review of Explainable Artificial Intelligence models and applications: Recent developments and future trends. Decision Analytics Journal 7 (2023), 100230. https://doi.org/10.1016/j.dajour. 2023.100230

  2. [2]

    Albahri, Ali M

    A.S. Albahri, Ali M. Duhaim, Mohammed A. Fadhel, Alhamzah Alnoor, Noor S. Baqer, Laith Alzubaidi, O.S. Albahri, A.H. Alamoodi, Jinshuai Bai, Asma Salhi, Jose Santamaría, Chun Ouyang, Ashish Gupta, Yuantong Gu, and Muhammet Deveci. 2023. A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias ri...

  3. [3]

    Pinar Barlas, Kyriakos Kyriakou, Olivia Guest, Styliani Kleanthous, and Jahna Otterbacher. 2021. To "See" is to Stereotype: Image Tagging Algorithms, Gender Recognition, and the Accuracy-Fairness Trade-off. Proc. ACM Hum.-Comput. Interact. 4, CSCW3, Article 232 (jan 2021), 31 pages. https://doi.org/10.1145/ 3432931

  4. [4]

    Markus Bertl, Toomas Klementi, Gunnar Piho, Peeter Ross, and Dirk Draheim

  5. [5]

    Friedrich, and Felix Nensa

    Katarzyna Borys, Yasmin Alyssa Schmitt, Meike Nauta, Christin Seifert, Nicole Krämer, Christoph M. Friedrich, and Felix Nensa. 2023. Explainable AI in medical imaging: An overview for clinical practitioners – Beyond saliency- based XAI approaches. European Journal of Radiology 162 (2023), 110786. https: //doi.org/10.1016/j.ejrad.2023.110786

  6. [6]

    Serdar Bozyel, Evrim Şimşek, Duygu Koçyiğit Burunkaya, Arda Güler, Yetkin Korkmaz, Mehmet Şeker, Mehmet Ertürk, and Nurgül Keser. 2024. Artifi- cial Intelligence-Based Clinical Decision Support Systems in Cardiovascu- lar Diseases. Anatolian Journal of Cardiology 28, 2 (January 7 2024), 74–86. https://doi.org/10.14744/AnatolJCardiol.2023.3685 PMID: 38168009

  7. [7]

    John Brooke. 2013. SUS: a retrospective. J. Usability Studies 8, 2 (feb 2013), 29–40

  8. [8]

    Nunes, and Jacinto C

    Francisco Maria Calisto, João Maria Abrantes, Carlos Santiago, Nuno J. Nunes, and Jacinto C. Nascimento. 2025. Personalized explanations for clinician-AI interaction in breast imaging diagnosis by adapting communication to expertise levels. International Journal of Human-Computer Studies 197 (2025), 103444. https://doi.org/10.1016/j.ijhcs.2025.103444

Show all 111 references
  1. [9]

    Nascimento

    Francisco Maria Calisto, João Fernandes, Margarida Morais, Carlos Santi- ago, João Maria Abrantes, Nuno Nunes, and Jacinto C. Nascimento. 2023. Assertiveness-based Agent Communication for a Personalized Medicine on Medical Imaging Diagnosis. In Proceedings of the 2023 CHI Conf...

  2. [10]

    Nascimento

    Francisco Maria Calisto, Nuno Nunes, and Jacinto C. Nascimento. 2022. Modeling adoption of intelligent agents in medical imaging. International Journal of Human-Computer Studies 168 (2022), 102922. https://doi.org/10.1016/j.ijhcs. 2022.102922

  3. [11]

    Yidong Chai, Hongyan Liu, Jie Xu, Sagar Samtani, Yuanchun Jiang, and Haoxin Liu. 2023. A Multi-Label Classification with an Adversarial-Based Denoising Autoencoder for Medical Image Annotation. ACM Trans. Manage. Inf. Syst. 14, 2, Article 19 (jan 2023), 21 pages. https://doi.o...

  4. [12]

    Bingzhi Chen, Jinxing Li, Guangming Lu, Hongbing Yu, and David Zhang. 2020. Label Co-Occurrence Learning With Graph Convolutional Networks for Multi- Label Chest X-Ray Image Classification. IEEE Journal of Biomedical and Health Informatics 24, 8 (2020), 2292–2302. https://doi....

  5. [13]

    Changjian Chen, Jun Yuan, Yafeng Lu, Yang Liu, Hang Su, Songtao Yuan, and Shixia Liu. 2021. OoDAnalyzer: Interactive Analysis of Out-of-Distribution Samples. IEEE Transactions on Visualization and Computer Graphics 27, 7 (July 2021), 3335–3349. https://doi.org/10.1109/TVCG.202...

  6. [14]

    Haomin Chen, Catalina Gomez, Chien-Ming Huang, and Mathias Unberath

  7. [15]

    Tianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu, and Liang Lin. 2019. Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . 522–531. https://doi.org/10.1109/ICCV.2019.00061

  8. [16]

    Zhao-Min Chen, Xiu-Shen Wei, Peng Wang, and Yanwen Guo. 2019. Multi- Label Image Recognition With Graph Convolutional Networks. In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5172–5181. https: //doi.org/10.1109/CVPR.2019.00532

  9. [17]

    Zhao-Min Chen, Xiu-Shen Wei, Peng Wang, and Yanwen Guo. 2023. Learning Graph Convolutional Networks for Multi-Label Recognition and Applications. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 6 (2023), 6969–6983. https://doi.org/10.1109/TPAMI.2021.3063496

  10. [18]

    Alexandra Chouldechova and Aaron Roth. 2018. The Frontiers of Fair- ness in Machine Learning. https://doi.org/10.48550/arXiv.1810.08810 arXiv:1810.08810 [cs.LG]

  11. [19]

    Haluk Demirkan and Dursun Delen. 2013. Leveraging the capabilities of service- oriented decision support systems: Putting analytics and big data in cloud.Decis. Support Syst. 55, 1 (apr 2013), 412–421. https://doi.org/10.1016/j.dss.2012.05.048

  12. [20]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference UIST ’25, September 28-October 1, 2025, Busan, Republic of Korea Shaohan Shi, Yuheng Shao, Haoran Jiang, Yunjie Yao, Zhijun...

  13. [21]

    Bernstein, Alex Berg, and Li Fei-Fei

    Jia Deng, Olga Russakovsky, Jonathan Krause, Michael S. Bernstein, Alex Berg, and Li Fei-Fei. 2014. Scalable multi-label annotation. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14). Association for Computing Mac...

  14. [22]

    Joseph Donia and James A. Shaw. 2021. Co-design and ethical artificial intelli- gence for health: An agenda for critical research and practice.Big Data & Society 8, 2 (2021), 20539517211065248. https://doi.org/10.1177/20539517211065248

  15. [23]

    Dudley and Per Ola Kristensson

    John J. Dudley and Per Ola Kristensson. 2018. A Review of User Interface Design for Interactive Machine Learning. ACM Trans. Interact. Intell. Syst. 8, 2, Article 8 (jun 2018), 37 pages. https://doi.org/10.1145/3185517

  16. [24]

    Leivon, Trupti Kolur, Vivek Shetty, Vidya Bushan, Rohan M

    Kevin Figueroa, Bofan Song, Sumsum Sunny, Shaobai Li, Keerthi Gurushanth, Pramila Mendonca, Nirza Mukhia, Sanjana Patrick, Shubha Gurudath, Sub- hashini Raghavan, Imchen Tsusennaro, Shirley T. Leivon, Trupti Kolur, Vivek Shetty, Vidya Bushan, Rohan M. Ramesh, Vijay Pillai, Pet...

  17. [25]

    Hiroshi Fukui, Tsubasa Hirakawa, Takayoshi Yamashita, and Hironobu Fu- jiyoshi. 2019. Attention Branch Network: Learning of Attention Mechanism for Visual Explanation. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10697–10706. https://doi.org/1...

  18. [26]

    Yuyang Gao, Siyi Gu, Junji Jiang, Sungsoo Ray Hong, Dazhou Yu, and Liang Zhao. 2024. Going Beyond XAI: A Systematic Survey for Explanation-Guided Learning. ACM Comput. Surv. 56, 7, Article 188 (apr 2024), 39 pages. https: //doi.org/10.1145/3644073

  19. [27]

    Yuyang Gao, Tong Steven Sun, Liang Zhao, and Sungsoo Ray Hong. 2022. Align- ing Eyes between Humans and Deep Neural Network through Interactive At- tention Alignment. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 489 (nov 2022), 28 pages. https://doi.org/10.1145/3555590

  20. [28]

    Shizhan Gong, Cheng Chen, Yuqi Gong, Nga Yan Chan, Wenao Ma, Calvin Hoi-Kwan Mak, Jill Abrigo, and Qi Dou. 2023. Diffusion model based semi- supervised learning on brain hemorrhage images for efficient midline shift quantification. In International Conference on Information Pr...

  21. [29]

    Liang Gou, Lincan Zou, Nanxiang Li, Michael Hofmann, Arvind Kumar Shekar, Axel Wendt, and Liu Ren. 2021. VATLD: A Visual Analytics System to Assess, Understand and Improve Traffic Light Detection. IEEE Transactions on Visual- ization and Computer Graphics 27, 2 (2021), 261–271...

  22. [30]

    Xuan Guo, Qi Yu, Rui Li, Cecilia Ovesdotter Alm, Cara Calvelli, Pengcheng Shi, and Anne Haake. 2016. An Expert-in-the-loop Paradigm for Learning Medical Image Grouping. In Advances in Knowledge Discovery and Data Mining , James Bailey, Latifur Khan, Takashi Washio, Gill Dobbie...

  23. [31]

    Shivam Gupta, Sachin Modgil, Samadrita Bhattacharyya, and Indranil Bose. 2021. Artificial intelligence for decision support systems in the field of operations research: review and future scope of research. Annals of Operations Research 308 (2021), 215 – 274. https://doi.org/10...

  24. [32]

    Meng Han, Hongxin Wu, Zhiqiang Chen, Muhan Li, and Xilong Zhang. 2022. A survey of multi-label classification based on supervised and semi-supervised learning. International Journal of Machine Learning and Cybernetics 14 (2022), 697–724. https://doi.org/10.1007/s13042-022-01658-9

  25. [33]

    Allan Hanbury, Henning Mller, and Georg Langs. 2017. Cloud-Based Bench- marking of Medical Image Analysis (1st ed.). Springer Publishing Company, Incorporated. https://doi.org/10.1007/978-3-319-49644-3

  26. [34]

    Hart and Lowell E

    Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Human Mental Workload, Peter A. Hancock and Najmedin Meshkati (Eds.). Advances in Psychology, Vol. 52. North-Holland, 139–183. https://doi...

  27. [35]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 770–778. https://doi.org/10.1109/CVPR.2016.90

  28. [36]

    Sauter, Fides Regina Schwartz, Andreas Termer, Felix Wagner, Hannes Götz Kenngott, and Lena Maier-Hein

    Eric Heim, Tobias Roß, Alexander Seitel, Keno März, Bram Stieltjes, Matthias Eisenmann, Johannes Lebert, Jasmin Metzger, Gregor Sommer, Alexander W. Sauter, Fides Regina Schwartz, Andreas Termer, Felix Wagner, Hannes Götz Kenngott, and Lena Maier-Hein. 2018. Large-scale medica...

  29. [37]

    Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. 2018. Women Also Snowboard: Overcoming Bias in Captioning Mod- els. In Computer Vision – ECCV 2018: 15th European Conference, Munich, Germany, September 8–14, 2018, Proceedings, Part III (Munich...

  30. [38]

    Fred Hohman, Andrew Head, Rich Caruana, Robert DeLine, and Steven M. Drucker. 2019. Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk)...

  31. [39]

    Summers, George Shih, Zhangyang Wang, and Yifan Peng

    Gregory Holste, Yiliang Zhou, Song Wang, Ajay Jaiswal, Mingquan Lin, Sherry Zhuge, Yuzhe Yang, Dongkyun Kim, Trong-Hieu Nguyen-Mau, Minh-Triet Tran, Jaehyup Jeong, Wongi Park, Jongbin Ryu, Feng Hong, Arsh Verma, Yosuke Yamagishi, Changhyun Kim, Hyeryeong Seo, Myungjoo Kang, Le...

  32. [40]

    Feng Hong, Tianjie Dai, Jiangchao Yao, Ya Zhang, and Yanfeng Wang. 2023. Bag of Tricks for Long-Tailed Multi-Label Classification on Chest X-Rays. https: //doi.org/10.48550/arXiv.2308.08853 arXiv:2308.08853 [cs.CV]

  33. [41]

    Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. https://doi.org/10.1145...

  34. [42]

    Weinberger

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger

  35. [44]

    Liu Jiang, Shixia Liu, and Changjian Chen. 2019. Recent research advances on interactive machine learning. J. Vis. 22, 2 (apr 2019), 401–417. https: //doi.org/10.1007/s12650-018-0531-1

  36. [45]

    Mohamed Khalifa and Mona Albadawy. 2024. AI in diagnostic imaging: Rev- olutionising accuracy and efficiency. Computer Methods and Programs in Biomedicine Update 5 (2024), 100146. https://doi.org/10.1016/j.cmpbup.2024. 100146

  37. [46]

    Shahzeb Khan and Jawwad Ahmed Shamsi. 2021. Health Quest: A generalized clinical decision support system with multi-label classification. Journal of King Saud University - Computer and Information Sciences 33, 1 (2021), 45–53. https://doi.org/10.1016/j.jksuci.2018.11.003

  38. [50]

    Philip Chen

    Qi Lai, Jianhang Zhou, Yanfen Gan, Chi-Man Vong, and C.L. Philip Chen. 2024. Single-Stage Broad Multi-Instance Multi-Label Learning (BMIML) With Diverse Inter-Correlations and Its Application to Medical Image Classification. IEEE Transactions on Emerging Topics in Computationa...

  39. [51]

    Xiang Li, Menglin Cui, Jingpeng Li, Ruibin Bai, Zheng Lu, and Uwe Aickelin

  40. [52]

    Weiwei Liu, Haobo Wang, Xiaobo Shen, and Ivor W. Tsang. 2022. The Emerging Trends of Multi-Label Learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 11 (2022), 7955–7974. https://doi.org/10.1109/TPAMI. 2021.3119334

  41. [53]

    Martínez-Trinidad, Jesús Ariel Carrasco- Ochoa, and Milton García-Borroto

    Octavio Loyola-González, José Fco. Martínez-Trinidad, Jesús Ariel Carrasco- Ochoa, and Milton García-Borroto. 2016. Study of the impact of resampling meth- ods for contrast pattern based classifiers in imbalanced databases. Neurocomput. 175, PB (jan 2016), 935–947. https://doi...

  42. [54]

    Lundberg and Su-In Lee

    Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 4768–47...

  43. [55]

    Maximilian Mackeprang, Claudia Müller-Birn, and Maximilian Timo Stauss

  44. [56]

    Martín-Noguerol, F

    T. Martín-Noguerol, F. Paulano-Godino, R. López-Ortega, J.M. Górriz, R.F. Ri- ascos, and A. Luna. 2021. Artificial intelligence in radiology: relevance of collaborative work between radiologists and engineers for building a multidisci- plinary team. Clinical Radiology 76, 5 (2...

  45. [57]

    Dud- ley

    Riccardo Miotto, Fei Wang, Shuang Wang, Xiaoqian Jiang, and Joel T. Dud- ley. 2018. Deep learning for healthcare: review, opportunities and challenges. Briefings in Bioinformatics 19, 6 (2018), 1236–1246. https://doi.org/10.1093/bib/ bbx044 arXiv:PMC6455466 PMID: 28481991

  46. [58]

    Irwansyah

    Eka Miranda, Mediana Aryuni, and E. Irwansyah. 2016. A survey of medical image classification techniques. In 2016 International Conference on Informa- tion Management and Technology (ICIMTech). 56–61. https://doi.org/10.1109/ ICIMTech.2016.7930302

  47. [59]

    Elham Nasarian, Roohallah Alizadehsani, U.Rajendra Acharya, and Kwok-Leung Tsui. 2024. Designing interpretable ML system to enhance trust in healthcare: A systematic review to proposed responsible clinician-AI-collaboration framework. Information Fusion 108 (2024), 102412. htt...

  48. [61]

    Julianne S. Oktay. 2012. Grounded Theory. Oxford University Press. https: //doi.org/10.1093/acprof:oso/9780199753697.001.0001

  49. [62]

    Pedro Osorio, Guillermo Jimenez-Perez, Javier Montalt-Tordera, Jens Hooge, Guillem Duran-Ballester, Shivam Singh, Moritz Radbruch, Ute Bach, Sabrina Schroeder, Krystyna Siudak, Julia Vienenkoetter, Bettina Lawrenz, and Sadegh Mohammadi. 2024. Latent Diffusion Models with Image...

  50. [63]

    Yang Ouyang, Yuchen Wu, He Wang, Chenyang Zhang, Furui Cheng, Chang Jiang, Lixia Jin, Yuanwu Cao, and Quan Li. 2024. Leveraging Historical Medical Records as a Proxy via Multimodal Modeling and Visualization to Enrich Medical Diagnostic Learning. IEEE Transactions on Visualiza...

  51. [64]

    Yang Ouyang, Chenyang Zhang, He Wang, Tianle Ma, Chang Jiang, Yuheng Yan, Zuoqin Yan, Xiaojuan Ma, Chuhan Shi, and Quan Li. 2024. A Two- Phase Visualization System for Continuous Human-AI Collaboration in Se- quelae Analysis and Modeling. https://doi.org/10.48550/arXiv.2407.14...

  52. [65]

    Meghana Padmanabhan, Pengyu Yuan, Govind Chada, and Hien Van Nguyen

  53. [66]

    Wongi Park, Inhyuk Park, Sungeun Kim, and Jong Bin Ryu. 2023. Robust Asymmetric Loss for Multi-Label Long-Tailed Learning. 2023 IEEE/CVF Inter- national Conference on Computer Vision Workshops (ICCVW) (2023), 2703–2712. https://doi.org/10.48550/arXiv.2308.05542

  54. [67]

    John Pavlopoulos, Vasiliki Kougia, and Ion Androutsopoulos. 2019. A Survey on Biomedical Image Captioning. In Proceedings of the Second Workshop on Shortcomings in Vision and Language , Raffaella Bernardi, Raquel Fernandez, Spandana Gella, Kushal Kafle, Christopher Kanan, Stef...

  55. [68]

    Gupta, Xiaojiang Chen, and Xin Wang

    Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B. Gupta, Xiaojiang Chen, and Xin Wang. 2021. A Survey of Deep Active Learning. ACM Comput. Surv. 54, 9, Article 180 (Oct. 2021), 40 pages. https://doi.org/10.1145/ 3472291

  56. [69]

    Why Should I Trust You?

    Marco Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, John DeNero, Ma...

  57. [70]

    Tal Ridnik, Emanuel Ben-Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor. 2021. Asymmetric Loss For Multi-Label Classification. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV). 82–91. https://doi.org/10.1109/ICCV48922.2021.00015

  58. [71]

    Russell, Antonio Torralba, Kevin P

    Bryan C. Russell, Antonio Torralba, Kevin P. Murphy, and William T. Free- man. 2008. LabelMe: A Database and Web-Based Tool for Image Annota- tion. International Journal of Computer Vision 77 (2008), 157–173. https: //doi.org/10.1007/s11263-007-0090-8

  59. [72]

    Journal of Clinical Medicine 8, 7 (2019)

    Physician-Friendly Machine Learning: A Case Study with Cardiovascular Disease Risk Prediction. Journal of Clinical Medicine 8, 7 (2019). https://doi. org/10.3390/jcm8071050

  60. [74]

    Chuhan Shi, Yicheng Hu, Shenan Wang, Shuai Ma, Chengbo Zheng, Xiaojuan Ma, and Qiong Luo. 2023. RetroLens: A Human-AI Collaborative System for Multi-step Retrosynthetic Route Planning. In Proceedings of the 2023 CHI Con- ference on Human Factors in Computing Systems (Hamburg, ...

  61. [75]

    https://doi.org/10.18653/v1/W19-1803

  62. [76]

    Benjamin Shickel, Patrick James Tighe, Azra Bihorac, and Parisa Rashidi. 2018. Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis. IEEE Journal of Biomedical and Health Informatics 22, 5 (2018), 1589–1604. https://doi....

  63. [77]

    Aram Siamak, Roozbeh Sadeghian, Iheb Abdellatif, and Stanley Nwoji. 2019. Diagnosing Heart Disease Types from Chest X-Rays Using a Deep Learning Approach. In 2019 International Conference on Computational Science and Com- putational Intelligence (CSCI). 910–913. https://doi.or...

  64. [78]

    Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Mar- tin Riedmiller. 2015. Striving for Simplicity: The All Convolutional Net. arXiv:1412.6806 [cs.LG] https://arxiv.org/abs/1412.6806

  65. [79]

    George Sun and Yi-Hui Zhou. 2023. AI in healthcare: navigating opportunities and challenges in digital communication. Frontiers in Digital Health 5 (2023), 1291132. https://doi.org/10.3389/fdgth.2023.1291132

  66. [80]

    Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra

    Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual Expla- nations from Deep Networks via Gradient-Based Localization. In 2017 IEEE International Conference on Computer Vision (ICCV) . 618–626. ht...

  67. [81]

    Tatiana Tommasi, Novi Patricia, Barbara Caputo, and Tinne Tuytelaars. 2015. A Deeper Look at Dataset Bias. arXiv e-prints , Article arXiv:1505.01257 (May 2015), arXiv:1505.01257 pages. https://doi.org/10.48550/arXiv.1505.01257 arXiv:1505.01257 [cs.CV]

  68. [82]

    Torralba and A

    A. Torralba and A. A. Efros. 2011. Unbiased look at dataset bias. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR ’11). IEEE Computer Society, USA, 1521–1528. https://doi.org/10.1109/CVPR. 2011.5995347

  69. [83]

    Wenqi Shi, Li Tong, Yuanda Zhu, and May D. Wang. 2021. COVID-19 Automatic Diagnosis With Radiographic Imaging: Explainable Attention Transfer Deep Neural Networks. IEEE Journal of Biomedical and Health Informatics 25, 7 (2021), 2376–2387. https://doi.org/10.1109/JBHI.2021.3074893

  70. [84]

    Guoli Wang, Pingping Wang, and Benzheng Wei. 2024. Multi-label local aware- ness and global co-occurrence priori learning improve chest X-ray classification. Multim. Syst. 30 (2024), 132. https://doi.org/10.1007/s00530-024-01321-z

  71. [85]

    He Wang, Yang Ouyang, Yuchen Wu, Chang Jiang, Lixia Jin, Yuanwu Cao, and Quan Li. 2024. KMTLabeler: An Interactive Knowledge-Assisted Labeling Tool for Medical Text Classification. IEEE Transactions on Visualization and Computer Graphics (2024), 1–18. https://doi.org/10.1109/T...

  72. [86]

    Kai Wang, Shuqi He, Wenlu Wang, Jinbei Yu, Yu Liu, and Lingyun Yu. 2024. CHORDination: Evaluating Visual Design Choices in Chord Diagrams for Net- work Data. In Proceedings of the 17th International Symposium on Visual Infor- mation Communication and Interaction (VINCI ’24) . ...

  73. [87]

    Minku, and Xin Yao

    Shuo Wang, Leandro L. Minku, and Xin Yao. 2015. Resampling-Based Ensemble Methods for Online Class Imbalance Learning. IEEE Transactions on Knowledge and Data Engineering 27, 5 (2015), 1356–1368. https://doi.org/10.1109/TKDE. 2014.2345380

  74. [88]

    Sutton, David Pincock, Daniel C

    Reed T. Sutton, David Pincock, Daniel C. Baumgart, Daniel C. Sadowski, Richard N. Fedorak, and Karen I. Kroeker. 2020. An overview of clinical decision support systems: benefits, risks, and strategies for success. npj Digital Medicine 3, 1 (February 2020), 17. https://doi.org/...

  75. [89]

    Tong Wu, Qingqiu Huang, Ziwei Liu, Yu Wang, and Dahua Lin. 2020. Distribution-Balanced Loss for Multi-label Classification in Long-Tailed Datasets. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV (Glasgow, United ...

  76. [90]

    Yao Xie, Melody Chen, David Kao, Ge Gao, and Xiang ’Anthony’ Chen. 2020. CheXplain: Enabling Physicians to Explore and Understand Data-Driven, AI- Enabled Medical Imaging Analysis. In Proceedings of the 2020 CHI Conference on UIST ’25, September 28-October 1, 2025, Busan, Repu...

  77. [91]

    Arsh Verma. 2023. How Can We Tame the Long-Tail of Chest X-ray Datasets? https://doi.org/10.48550/arXiv.2309.04293 arXiv:2309.04293 [eess.IV]

  78. [92]

    Weikai Yang, Yukai Guo, Jing Wu, Zheng Wang, Lan-Zhe Guo, Yu-Feng Li, and Shixia Liu. 2024. Interactive Reweighting for Mitigating Label Quality Issues. IEEE Transactions on Visualization and Computer Graphics 30, 3 (2024), 1837–1852. https://doi.org/10.1109/TVCG.2023.3345340

  79. [93]

    Weikai Yang, Mengchen Liu, Zheng Wang, and Shixia Liu. 2024. Foundation models meet visualizations: Challenges and opportunities.Computational Visual Media 10, 3 (2024), 399–424. https://doi.org/10.1007/s41095-023-0393-x

  80. [94]

    Jin Ye, Junjun He, Xiaojiang Peng, Wenhao Wu, and Yu Qiao. 2020. Attention- Driven Dynamic Graph Convolutional Network for Multi-label Image Recogni- tion. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI (Glasgow...

  81. [95]

    Seid Muhie Yimam, Chris Biemann, Ljiljana Majnaric, Šefket Šabanović, and Andreas Holzinger. 2015. Interactive and Iterative Annotation for Biomedical Entity Recognition. InBrain Informatics and Health, Yike Guo, Karl Friston, Faisal Aldo, Sean Hill, and Hanchuan Peng (Eds.). ...

  82. [96]

    Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M. Summers. 2017. ChestX-Ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Com- mon Thorax Diseases. In 2017 IEEE Conference on Compute...

  83. [97]

    Jun Yuan, Changjian Chen, Weikai Yang, Mengchen Liu, Jiazhi Xia, and Shixia Liu. 2021. A survey of visual analytics techniques for machine learning. Com- putational Visual Media 7, 1 (Mar 2021), 3–36. https://doi.org/10.1007/s41095- 020-0191-7

  84. [98]

    Alwin Yaoxian Zhang, Sean Shao Wei Lam, Nan Liu, Yan Pang, Ling Ling Chan, and Phua Hwee Tang. 2018. Development of a Radiology Decision Support System for the Classification of MRI Brain Scans. In 2018 IEEE/ACM 5th International Conference on Big Data Computing Applications a...

  85. [99]

    Yosuke Yamagishi and Shohei Hanaoka. 2023. Effect of Stage Training for Long-Tailed Multi-Label Image Classification. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). 2713–2720. https://doi.org/ 10.1109/ICCVW60793.2023.00287

  86. [100]

    Min-Ling Zhang and Zhi-Hua Zhou. 2014. A Review on Multi-Label Learning Algorithms. IEEE Transactions on Knowledge and Data Engineering 26, 8 (2014), 1819–1837. https://doi.org/10.1109/TKDE.2013.39

  87. [101]

    Xiaoyu Zhang, Xiwei Xuan, Alden Dima, Thurston Sexton, and Kwan-Liu Ma

  88. [102]

    Yu Zhang, Jing Chen, Xiangxun Ma, Gang Wang, Uzair Aslam Bhatti, and Mengxing Huang. 2024. Interactive medical image annotation using improved Attention U-net with compound geodesic distance. Expert Systems with Appli- cations 237 (2024), 121282. https://doi.org/10.1016/j.eswa...

  89. [103]

    Yifei Zhang, Siyi Gu, Yuyang Gao, Bo Pan, Xiaofeng Yang, and Liang Zhao

  90. [104]

    Chien Wen (Tina) Yuan, Nanyi Bi, Ya-Fang Lin, and Yuen-Hsien Tseng. 2023. Contextualizing User Perceptions about Biases for Human-Centered Explainable Artificial Intelligence. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (C...

  91. [105]

    Xuehan Zhao, Jiaqi Liu, Zhiwen Yu, and Bin Guo. 2024. HADT: Human-AI Diagnostic Team via Hierarchical Reinforcement Learning. In Proceedings of the 2024 SIAM International Conference on Data Mining (SDM) . 860–868. https: //doi.org/10.1137/1.9781611978032.98

  92. [106]

    Jiayi Zhou, Renzhong Li, Junxiu Tang, Tan Tang, Haotian Li, Weiwei Cui, and Yingcai Wu. 2024. Understanding Nonlinear Collaboration between Human and AI Agents: A Co-design Framework for Creative Design. In Proceedings of the CHI Conference on Human Factors in Computing System...

  93. [107]

    Jie Zhang and Zong-ming Zhang. 2023. Ethics and governance of trustworthy medical artificial intelligence. BMC medical informatics and decision making 23, 1 (2023), 7. https://doi.org/10.1186/s12911-023-02103-9

  94. [110]

    In 2023 IEEE 16th Pacific Visualization Symposium (PacificVis)

    LabelVizier: Interactive Validation and Relabeling for Technical Text Annotations. In 2023 IEEE 16th Pacific Visualization Symposium (PacificVis) . 167–176. https://doi.org/10.1109/PacificVis56936.2023.00026

  95. [113]

    In Proceedings of the IEEE/CVF International Conference on Computer Vision

    Magi: Multi-annotated explanation-guided learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1977–1987. https: //doi.org/10.1109/ICCV51070.2023.00189

  96. [114]

    Yuhan Zhang, Luyang Luo, Qi Dou, and Pheng-Ann Heng. 2023. Triplet attention and dual-pool contrastive learning for clinic-driven multi-label medical image classification. Medical Image Analysis 86 (2023), 102772. https://doi.org/10. 1016/j.media.2023.102772

  97. [117]

    Oren Zuckerman, Viva Sarah Press, Ehud Barda, Benny Megidish, and Hadas Erel. 2022. Tangible Collaboration: A Human-Centered Approach for Sharing Control With an Actuated-Interface. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, L...

  98. [839]

    https://doi.org/10.1109/TETCI.2023.3287978

  99. [2017]

    In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

    Densely Connected Convolutional Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2261–2269. https://doi.org/10. 1109/CVPR.2017.243

  100. [2019]

    Discovering the Sweet Spot of Human-Computer Configurations: A Case Study in Information Extraction. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 195 (nov 2019), 30 pages. https://doi.org/10.1145/3359297

  101. [2021]

    Neurocomputing 443 (2021), 345–355

    A hybrid medical text classification framework: Integrating attentive rule construction and neural network. Neurocomputing 443 (2021), 345–355. https://doi.org/10.1016/j.neucom.2021.02.069

  102. [2022]

    npj Digital Medicine 5, 1 (2022), 156

    Explainable medical imaging AI needs human-centered design: guidelines and evidence from a systematic review. npj Digital Medicine 5, 1 (2022), 156. https://doi.org/10.1038/s41746-022-00699-2

  103. [2023]

    How Domain Engineering Can Help to Raise Adoption Rates of Artifi- cial Intelligence in Healthcare. In Information Integration and Web Intelligence , Pari Delir Haghighi, Eric Pardede, Gillian Dobbie, Vithya Yogarajan, Ngurah Agus Sanjaya ER, Gabriele Kotsis, and Ismail Khalil...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.