REVIEW 4 major objections 5 minor 53 references
QuMAB: Query-based Multi-Annotator Behavior Modeling with Reliability under Sparse Labels
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper's central claim is that multi-annotator learning should shift from sample-wise aggregation to annotator-wise behavior modeling, with per-annotator queries reconstructing missing labels, improving consensus under sparse data, and…
desk verdict Genuinely new annotator-wise query architecture and two dense datasets, but the headline accuracy claims are undercut by uncontrolled backbone differences and missing uncertainty reporting. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a set of learnable query tokens, one per annotator, inside a Q-Former-style module. Queries first attend to each other in a shared self-attention layer, which makes inter-annotator similarities act as a regularizer, and then attend to input features in multi-head cross-attention, producing annotator-specific representations that feed separate classifiers. Because the queries are vectors rather than full networks, the per-annotator cost stays low. The cross-attention weight maps are the same mechanism used for both prediction and explanation, since they indicate which input regions each annotator focuses on. The training objective is simply the sum of per-annotator cross-entropy losses.
What would settle it
Ask the same annotators to re-label a held-out set weeks after their original labels, then compare QuMAB's predicted labels to the second-round labels: if agreement with second-round labels is no higher than agreement with a different annotator's original labels, the stable-behavior premise fails.
Extended reading notes
Core claim
QuMAB is a query-based architecture for modeling individual annotator behavior rather than aggregating labels sample by sample. For each annotator, a small learnable query token is passed through a shared self-attention layer, then through cross-attention with features from a frozen image or video encoder, and finally into that annotator's own classifier. The paper's central hypothesis is that annotator judgment differences arise from varying focus on different regions of the input; the cross-attention weights therefore encode each annotator's behavior pattern and double as a visualization of which image patches or video frames they rely on. Shared self-attention lets annotator queries influence each other, capturing inter-annotator correlations as implicit structural regularization that prevents per-annotator models from overfitting to small label sets while preserving individual differences. The paper claims that this design outperforms aggregation-oriented baselines on individual annotator prediction, that majority voting over per-annotator predictions beats direct aggregation when test labels are missing, and that per-annotator queries degrade less under 40% label removal.
Load-bearing premise
The load-bearing premise is that each annotator's judgment is stable over time and expressible as a consistent focus pattern over input regions, so that a model trained on their past labels can predict their labels on new samples.
Editorial extensions
If this is right
- With enough dense per-annotator labels, each annotator's model can predict their label on unannotated samples, so the annotation matrix can be completed rather than averaged over disjoint subsets.
- Consensus prediction by majority vote over reconstructed per-annotator predictions stays accurate when 20% to 40% of test labels are missing, where direct majority vote on remaining labels loses accuracy.
- Training under sparse annotations (40% of labels removed) costs QuMAB a 20.4% average accuracy drop versus 27.4% for the strongest baseline, indicating inter-annotator regularization helps under data scarcity.
- Ablation results show that disabling shared self-attention between annotator queries lowers performance, supporting the role of inter-annotator correlations as implicit regularization.
- Attention visualizations link focus differences to label differences: annotators who attend to a dog in a street scene rate happiness higher, and annotators who attend to early versus late video frames label different emotions.
Reading between the lines
- If the stable-focus hypothesis holds, the same query architecture could be used to detect annotator drift: a query whose attention pattern shifts markedly as new labels arrive would signal that an annotator's judgment criteria have changed and that their past labels should be down-weighted.
- The reconstruction step suggests an active-labeling loop: use per-annotator uncertainties to choose which samples each annotator should label next, converting saved annotation cost into even denser coverage of the most informative cells of the annotation matrix.
- Since AMER is multimodal (audio, video, text), the focus-region story could be extended across modalities to ask whether disagreement is explained by which modality an annotator weights, not just which frame or patch they attend to.
- If per-annotator models are accurate, the consensus label could be replaced by a distribution over annotator types, which might be more honest than majority vote for subjective tasks without ground truth.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces QuMAB, a query-based architecture for multi-annotator learning in which each annotator is represented by a learnable query token processed by a shared Q-Former with cross-attention to input features and self-attention among annotator queries. The stated goal is a paradigm shift from sample-wise aggregation to annotator-wise behavior modeling, with claimed benefits in reconstructing unlabeled labels, improving majority-vote consensus, and providing explainable visualizations of annotator focus regions. The authors also introduce two datasets, STREET (urban image impressions) and AMER (multimodal video emotion), and report experiments against D-LEMA, PADL, and MaDL on individual annotator accuracy/F1, consensus prediction, sparse-label robustness, ablations, and qualitative attention visualizations.
Significance. The annotator-wise perspective and the proposed query-based mechanism are conceptually interesting and could be valuable to the multi-annotator learning community if the empirical claims are supported. The datasets, if released, would also be a useful resource because per-annotator longitudinal labels are rare. The ablation study (Table 5) gives some internal evidence that the annotator queries, self-attention, and per-annotator classifiers contribute to the method's performance, and the consensus-reconstruction experiment in Table 6 is a legitimate prediction task. However, the paper does not provide code or data at submission time, the accuracy comparisons are not backbone-matched, and the AMER density claim is contradicted by the paper's own supplementary statistics. These issues currently prevent the central claims from being established as stated.
major comments (4)
- [Section 5.1 and Tables 2–3] The main accuracy comparisons do not control for the backbone network. QuMAB uses a frozen EVA-CLIP ViT-G/14 encoder and an InstructBLIP-initialized Q-Former, while the baselines are not described as using the same backbone. The supplementary efficiency analysis (Section 7.4, Table 8) explicitly states that 'different methods use different backbone networks' and standardizes to ResNet-34 only for that efficiency comparison, not for the accuracy experiments. Because ViT-G/14 is a very large, heavily pretrained model, the reported gains in Tables 2 and 3 may be due to the encoder rather than the proposed annotator-wise query modeling. Please rerun the baselines with the same EVA-CLIP/Q-Former backbone, or otherwise provide a controlled comparison where the only difference is the annotator modeling mechanism.
- [Tables 2–4 and Table 6] No error bars, confidence intervals, or significance tests are reported for any of the quantitative results. Several reported differences are small (e.g., AMER average accuracy 0.84 vs. 0.80 for MaDL; consensus CoPr 0.60 vs. 0.57), and the sparse-annotation results in Table 4 and the missing-label simulations in Table 6 are based on random removal without multiple runs or seeds. Please provide repeated runs with standard deviations and appropriate statistical tests, or at least mean±std over several random splits and masking seeds, so the reader can assess whether the differences are reliable.
- [Abstract, Section 1, Table 1, and Supplementary Table 7] The paper repeatedly states that AMER has 'average 3,118 labels per annotator' and calls the datasets 'dense per-annotator labels'. However, Supplementary Table 7 reports the average as 1,999.1 labels per annotator, with annotators A1–A10 each having roughly 1,000 labels (about 79–81% missing) and only A11–A13 having near-complete labels. The abstract and introduction should be corrected to match these numbers, and the term 'dense' should be qualified so that it does not overstate the coverage of most annotators.
- [Section 7.4, Table 8] The efficiency comparison is not consistent with the model evaluated in the main experiments. Table 8 standardizes all methods to ResNet-34, whereas Section 5.1 states that QuMAB uses EVA-CLIP ViT-G/14 and an InstructBLIP-initialized Q-Former. As a result, the reported 106.02M parameters for 'Ours' does not reflect the actual model in Tables 2 and 3, whose frozen encoder alone is far larger. The claim that the approach is 'lightweight' should either be limited to the query mechanism and trainable parameters, or the efficiency table should include the backbone parameters actually used in the main comparisons.
minor comments (5)
- [Table 3 caption] The caption says 'k = 1, ..., 13' for the STREET dataset, but STREET has 10 annotators; this should be corrected to 10.
- [Section 5.2] The text says 'as in AMER (1,040 vs. 5,195 labels for annotators 1–10 vs. 11–13)', but Supplementary Table 7 gives different per-annotator counts (e.g., A1 has 1,096, A11 has 5,187). Please align the numbers.
- [References] References [21] and [22] appear to be duplicate entries for the same CIFAR-10H paper, and references [26] and [27] also appear to be the same paper on learning from multiple annotators. Please deduplicate and cite the original venues correctly.
- [Section 5.3, Table 4] The statement that under 40% removal 'our model’s average performance drops by 20.4%, whereas the best baseline PADL experiences a larger drop of 27.4%' should specify the exact per-method values it is computed from and how the average is taken across the STREET perspectives and AMER.
- [Section 7.6 and Figure 7] In the caption of Figure 7, the frame ranges and annotator numbers are described inconsistently with the main text (e.g., which annotators focus on early versus late frames). Please unify the descriptions.
Circularity Check
No circular derivation: central claims are held-out empirical predictions; the few self-citations are not load-bearing.
full rationale
The paper is an empirical systems paper; there is no mathematical derivation whose output is equivalent to its input. The central claim—per-annotator query models trained on dense longitudinal labels can predict that annotator's labels on unseen samples and improve consensus—is tested by held-out evaluation. Table 6 masks test-set labels after training and fills them with model predictions; the masked labels are not used to fit the models, so the 'reconstruction' is a genuine out-of-sample prediction rather than a reintroduction of a fitted quantity. The loss in Eq. (1) is the standard sum of per-annotator cross-entropies; it defines the training objective, not a prediction derived from itself. Ablations (Table 5) compare architectural variants and are consistent with the method's assumptions. The cross-attention visualizations are post hoc interpretations of learned weights, not predictions claimed to be derived from first principles. Several references are self-citations (SimLabel [44], MicroEmo [43,46], and others), but none is load-bearing: they appear in related-work and efficiency discussions, while the accuracy and consensus results are benchmarked against external methods (D-LEMA, PADL, MaDL) on newly collected datasets. Therefore no circular step reduces the central claim to its inputs. The uncontrolled-backbone concern is a potential confound in empirical comparison, not a circularity.
Assumptions & free parameters
free parameters (3)
- Annotator query tokens (learnable embeddings) =
learned during training, one per annotator
- Compression queries for the video pipeline =
32
- Annotator retention criterion =
13 of 15 annotators retained
assumptions (5)
- domain assumption Annotator judgment differences arise from varying degrees of focus on different regions of the input content
- domain assumption Each annotator's behavior is longitudinally consistent, so models trained on dense per-annotator labels generalize to unlabeled samples
- domain assumption Majority vote over raw annotations is a valid proxy for consensus ground truth
- ad hoc to paper Shared self-attention across annotator queries captures inter-annotator correlations without explicit annotation-similarity supervision
- domain assumption AMER is a valid new dataset despite re-annotating videos from MER2024
Cite this review
Pith. "Pith review of QuMAB: Query-based Multi-Annotator Behavior Modeling with Reliability under Sparse Labels." pith.science (2026). https://pith.science/paper/VO2C32L5
@misc{pith2026250717653,
author = {Pith},
title = {Pith review of: QuMAB: Query-based Multi-Annotator Behavior Modeling with Reliability under Sparse Labels},
year = {2026},
howpublished = {\url{https://pith.science/paper/VO2C32L5}},
note = {Machine review of arXiv:2507.17653}
}
read the original abstract
Multi-annotator learning traditionally aggregates diverse annotations to approximate a single ground truth, treating disagreements as noise. However, this paradigm faces fundamental challenges: subjective tasks often lack absolute ground truth, and sparse annotation coverage makes aggregation statistically unreliable. We introduce a paradigm shift from sample-wise aggregation to annotator-wise behavior modeling. By treating annotator disagreements as valuable information rather than noise, modeling annotator-specific behavior patterns can reconstruct unlabeled data to reduce annotation cost, enhance aggregation reliability, and explain annotator decision behavior. To this end, we propose QuMAB (Query-based Multi-Annotator Behavior Pattern Learning), which uses light-weight queries to model individual annotators while capturing inter-annotator correlations as implicit regularization, preventing overfitting to sparse individual data while maintaining individualization and improving generalization, with a visualization of annotator focus regions offering an explainable analysis of behavior understanding. We contribute two large-scale datasets with dense per-annotator labels: STREET (4,300 labels/annotator) and AMER (average 3,118 labels/annotator), the first multimodal multi-annotator dataset. Extensive experiments demonstrate the superiority of our QuMAB in modeling individual annotators' behavior patterns, their utility for consensus prediction, and applicability under sparse annotations.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Ahmed Almazroa, Sami Alodhayb, Essameldin Osman, Eslam Ramadan, Mo- hammed Hummadi, Mohammed Dlaim, Muhannad Alkatee, Kaamran Raahemifar, and Vasudevan Lakshminarayanan. 2017. Agreement among ophthalmologists in marking the optic disc and optic cup in fundus images. International ophthal- mology 37 (2017), 701–717
work page 2017
-
[2]
Samuel G Armato III, Geoffrey McLennan, Luc Bidaut, Michael F McNitt-Gray, Charles R Meyer, Anthony P Reeves, Binsheng Zhao, Denise R Aberle, Claudia I Henschke, Eric A Hoffman, et al . 2011. The lung image database consortium (LIDC) and image database resource initiative (IDRI): a completed reference database of lung nodules on CT scans. Medical physics ...
work page 2011
-
[3]
Peng Cao, Yilun Xu, Yuqing Kong, and Yizhou Wang. 2019. Max-mig: an in- formation theoretic approach for joint learning from crowds. arXiv preprint arXiv:1905.13436 (2019)
work page Pith review arXiv 2019
-
[4]
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su. 2019. This looks like that: deep learning for interpretable image recognition. Advances in neural information processing systems 32 (2019)
work page 2019
-
[5]
Junfan Chen, Richong Zhang, Jie Xu, Chunming Hu, and Yongyi Mao. 2023. A Neural Expectation-Maximization Framework for Noisy Multi-Label Text Classification. IEEE Transactions on Knowledge and Data Engineering 35, 11 (2023), 10992–11003. https://doi.org/10.1109/TKDE.2022.3223067
-
[6]
Yuan-Chia Cheng, Zu-Yun Shiau, Fu-En Yang, and Yu-Chiang Frank Wang. 2023. TAX: Tendency-and-Assignment Explainer for Semantic Segmentation with Multi-Annotators. arXiv preprint arXiv:2302.09561 (2023)
arXiv 2023
-
[7]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. In Thirty- seventh Conference on Neural Information Processing Systems . https://openreview. net/forum?id=vvoWPYqZJA
work page 2023
-
[8]
Alexander Philip Dawid and Allan M Skene. 1979. Maximum likelihood esti- mation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics) 28, 1 (1979), 20–28
work page 1979
Show all 53 references
-
[9]
Marek Herde, Denis Huseljic, and Bernhard Sick. 2023. Multi-annotator Deep Learning: A Probabilistic Framework for Classification. arXiv preprint arXiv:2304.02539 (2023)
2023 arXiv
-
[10]
Wei Ji, Shuang Yu, Junde Wu, Kai Ma, Cheng Bian, Qi Bi, Jingjing Li, Hanruo Liu, Li Cheng, and Yefeng Zheng. 2021. Learning Calibrated Medical Image Segmentation via Multi-rater Agreement Modeling. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) ....
2021
-
[11]
Wei Ji, Shuang Yu, Junde Wu, Kai Ma, Cheng Bian, Qi Bi, Jingjing Li, Hanruo Liu, Li Cheng, and Yefeng Zheng. 2021. Learning calibrated medical image segmentation via multi-rater agreement modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...
2021
-
[12]
Uthman Jinadu, Jesse Annan, Shanshan Wen, and Yi Ding. 2023. Loss Modeling for Multi-Annotator Datasets. arXiv preprint arXiv:2311.00619 (2023)
2023 arXiv
-
[13]
Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. (Jan 2009)
2009
-
[14]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742
2023
-
[15]
Zheng Lian, Haiyang Sun, Licai Sun, Kang Chen, Mngyu Xu, Kexin Wang, Ke Xu, Yu He, Ying Li, Jinming Zhao, et al . 2023. Mer 2023: Multi-label learning, modality robustness, and semi-supervised learning. In Proceedings of the 31st ACM International Conference on Multimedia . 9610–9614
2023
-
[16]
Zheng Lian, Haiyang Sun, Licai Sun, Zhuofan Wen, Siyuan Zhang, Shun Chen, Hao Gu, Jinming Zhao, Ziyang Ma, Xie Chen, et al . 2024. MER 2024: Semi- Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emo- tion Recognition. In MRAC’24: Proceedings of the 2nd In...
2024
-
[17]
Zehui Liao, Shishuai Hu, Yutong Xie, and Yong Xia. 2024. Modeling annotator preference and stochastic annotation error for medical image segmentation. Medical Image Analysis 92 (2024), 103028
2024
-
[18]
Bjoern Menze, Leo Joskowicz, Spyridon Bakas, Andras Jakab, Ender Konukoglu, Anton Becker, and et al. 2020. Quantification of uncertainties in biomedical image quantification challenge. https://qubiq.grand-challenge.org/
2020
-
[19]
Zahra Mirikharaji, Kumar Abhishek, Saeed Izadi, and Ghassan Hamarneh. 2021. D-lema: Deep learning ensembles from multiple annotations-application to skin lesion segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1837–1846
2021
-
[20]
Jeppe Nørregaard and Leon Derczynski. 2022. Sparse Probability of Agreement. arXiv preprint arXiv:2208.06161 (2022)
2022 arXiv
-
[21]
Joshua Peterson, Ruairidh Battleday, Thomas Griffiths, and Olga Russakovsky
-
[22]
Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Rus- sakovsky. 2019. Human uncertainty makes classification more robust. In Proceed- ings of the IEEE/CVF international conference on computer vision . 9617–9626
2019
-
[23]
Pranav Rajpurkar, Jeremy Irvin, Aarti Bagul, Daisy Ding, Tony Duan, Hershel Mehta, Brandon Yang, Kaylie Zhu, Dillon Laird, Robyn L Ball, et al. 2017. Mura: Large dataset for abnormality detection in musculoskeletal radiographs. arXiv preprint arXiv:1712.06957 (2017)
2017 arXiv
-
[24]
Raykar, Shipeng Yu, Linda H
Vikas C. Raykar, Shipeng Yu, Linda H. Zhao, Gerardo Hermosillo Valadez, Charles Florin, Luca Bogoni, and Linda Moy. 2010. Learning From Crowds. Journal of Machine Learning Research 11, 43 (2010), 1297–1322. http://jmlr.org/papers/v11/ raykar10a.html
2010
-
[25]
Filipe Rodrigues and Francisco Pereira. 2018. Deep learning from crowds. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
2018
-
[27]
Filipe Rodrigues, Francisco Pereira, and Bernardete Ribeiro. 2013. Learning from multiple annotators: distinguishing good from random labelers. Pattern Recognition Letters 34, 12 (2013), 1428–1436
2013
-
[28]
Filipe Rodrigues, Francisco Pereira, and Bernardete Ribeiro. 2014. Gaussian Pro- cess Classification and Active Learning with Multiple Annotators. In Proceedings of the 31st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 32), Eric ...
2014
-
[29]
Mike Schaekermann, Graeme Beaton, Minahz Habib, Andrew Lim, Kate Larson, and Edith Law. 2019. Understanding Expert Disagreement in Medical Data Analysis through Structured Adjudication. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 76 (Nov. 2019), 23 pages. https://doi.org...
2019 doi
-
[30]
Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita, and Yuki Uranishi. 2024. 3DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion Policy. arXiv preprint arXiv:2409.10848 (2024)
2024 arXiv
-
[31]
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. 2008. Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, Mirella Lapata and Hwee ...
2008
-
[32]
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389 (2023)
2023 arXiv
-
[33]
Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C Alexander, and Nathan Silberman. 2019. Learning from noisy labels by regularized estimation of annotator confusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 11244–11253
2019
-
[34]
Alexander, and Nathan Silberman
Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C. Alexander, and Nathan Silberman. 2019. Learning From Noisy Labels by Regularized Esti- mation of Annotator Confusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[35]
2020-2025
Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Li- ubimov. 2020-2025. Label Studio: Data labeling software. https:// github.com/HumanSignal/label-studio Open source software available from https://github.com/HumanSignal/label-studio. , , Liyun Zhang, Zheng Lian...
2020
-
[36]
Chongyang Wang, Yuan Gao, Chenyou Fan, Junjie Hu, Tin Lum Lam, Nicholas D Lane, and Nadia Bianchi-Berthouze. 2023. Learn2agree: Fitting with multiple annotators without objective ground truth. In International Workshop on Trust- worthy Machine Learning for Healthcare . Springe...
2023
-
[37]
Peter Welinder, Steve Branson, Pietro Perona, and Serge Belongie. 2010. The multidimensional wisdom of crowds. Advances in neural information processing systems 23 (2010)
2010
-
[38]
Jacob Whitehill, Ting-fan Wu, Jacob Bergsma, Javier Movellan, and Paul Ruvolo
-
[39]
Yan Yan, Rómer Rosales, Glenn Fung, Ramanathan Subramanian, and Jennifer Dy. 2014. Learning from multiple annotators with varying expertise. Machine learning 95 (2014), 291–327. QuMAB: Query-based Multi-Annotator Behavior Modeling with Reliability under Sparse Labels , , Annot...
2014
-
[40]
Xueying Zhan, Yaowei Wang, Yanghui Rao, and Qing Li. 2019. Learning from multi-annotator data: A noise-aware classification framework. ACM Transactions on Information Systems (TOIS) 37, 2 (2019), 1–28
2019
-
[41]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 543–553
2023
-
[42]
Liyun Zhang. 2024. Integrating Panoptic-Level to Image Translation . Ph. D. Dis- sertation. PhD Dissertation
2024
-
[43]
Liyun Zhang. 2024. MicroEmo: Time-Sensitive Multimodal Emotion Recog- nition with Micro-Expression Dynamics in Video Dialogues. arXiv preprint arXiv:2407.16552 (2024)
2024 arXiv
-
[44]
Liyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe, and Yuta Nakashima. 2025. SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing Labels. arXiv preprint arXiv:2504.09525 (2025)
2025 arXiv
-
[45]
Liyun Zhang, Nanyan Liu, Yuanbin Hou, and Xiaojian Liu. 2014. Uneven il- lumination image segmentation based on multi-threshold S-F. Opto-Electronic Engineering 41, 7 (2014), 81–87
2014
-
[46]
Liyun Zhang, Zhaojie Luo, Shuqiong Wu, and Yuta Nakashima. 2024. MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Subtle Clue Dynamics in Video Dialogues. In Proceedings of the 2nd International Workshop on Multimodal and Responsible Affective Computing . 110–115
2024
-
[47]
Liyun Zhang, Photchara Ratsamee, Zhaojie Luo, Yuki Uranishi, Manabu Hi- gashida, and Haruo Takemura. 2023. Panoptic-level image-to-image translation for object recognition and visual odometry enhancement. IEEE Transactions on Circuits and Systems for Video Technology 34, 2 (20...
2023
-
[48]
Liyun Zhang, Photchara Ratsamee, Yuki Uranishi, Manabu Higashida, and Haruo Takemura. 2022. Thermal-to-Color Image Translation for Enhancing Visual , , Liyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe, and Yuta Nakashima Odometry of Thermal Vision. In 2022 IEEE International...
2022
-
[49]
Liyun Zhang, Photchara Ratsamee, Bowen Wang, Zhaojie Luo, Yuki Uranishi, Manabu Higashida, and Haruo Takemura. 2023. Panoptic-aware image-to-image translation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. 259–268
2023
-
[50]
Le Zhang, Ryutaro Tanno, Moucheng Xu, Yawen Huang, Kevin Bronik, Chen Jin, Joseph Jacob, Yefeng Zheng, Ling Shao, Olga Ciccarelli, et al. 2023. Learning from multiple annotators for medical image segmentation. Pattern Recognition 138 (2023), 109400
2023
-
[51]
Yifei Zhang, Siyi Gu, Yuyang Gao, Bo Pan, Xiaofeng Yang, and Liang Zhao
-
[2009]
Advances in neural information processing systems 22 (2009)
Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. Advances in neural information processing systems 22 (2009)
2009
-
[2019]
In 2019 IEEE/CVF International Conference on Computer Vision (ICCV)
Human uncertainty makes classification more robust. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . https://doi.org/10.1109/iccv. 2019.00971
2019
-
[2023]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Magi: Multi-annotated explanation-guided learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1977–1987
1977
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.