REVIEW 4 major objections 6 minor 41 references
Scientific figure quality is decided by how the figure aligns with the manuscript’s claims, and that alignment has to be learned—not prompted.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 12:12 UTC pith:XZATH76T
load-bearing objection Solid systems paper with a real corpus and clean ablations; the 59% MAE cut is internally credible, but unreported rater agreement caps how hard you can lean on the headline claim. the 4 major comments →
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Scientific figure assessment for peer review requires learned multimodal alignment between the published figure crop and manuscript evidence. Under paper-level splits on a human-rated test set of 396 figures, SciFigAlign reaches macro MAE 0.3524 and within-paper pairwise accuracy 81.64%, a 59% relative error reduction versus the best LLM-as-judge baseline (MAE 0.864). Ablations show that caption and citing context, citing-context denoising, and a within-paper ranking objective are all necessary; prompting state-of-the-art VLMs does not close the gap.
What carries the argument
SciFigAlign: a fine-tuned multi-stream scorer that encodes the figure with CLIP and caption/citing/abstract/metadata with SciBERT, aligns each text stream to image patches via per-modality cross-attention, fuses the five streams with CubeMLP, and jointly trains SmoothL1 score regression with a within-paper ranking hinge so absolute scores and same-paper order are optimized together.
Load-bearing premise
The four-dimension human and GPT-assisted 1–5 labels on this machine-learning conference corpus are treated as a stable enough gold standard of peer-review figure utility that a model fit to them measures real review quality.
What would settle it
On a fresh, fully human-rated test set drawn from different venues or disciplines, with paper-level splits and the same CL/RE/IN/ST rubric, check whether SciFigAlign still cuts macro MAE by roughly half versus a full-context LLM/VLM judge and still exceeds ~80% within-paper pairwise accuracy; a collapse to judge-level error or chance ordering would falsify the central claim.
If this is right
- Review and authoring tools can triage which figure in a submission is weaker by ranking same-paper figures rather than trusting a single absolute score.
- Generic IQA, CLIPScore-style similarity, and zero-shot VLM judges are insufficient baselines for published-figure quality once caption and citing paragraphs are available.
- Citing-context cleaning (dropping bare “see Fig. k” boilerplate) is a first-class design choice, not a minor preprocessing detail.
- Exportable cross-attention maps and modality weights give a coarse audit trail of which evidence streams drive each dimension score.
- The 3,857-figure, paper-split benchmark becomes a shared testbed for manuscript-grounded figure scoring beyond generation or chart-QA tasks.
Where Pith is reading between the lines
- The same binding of crop to index-resolved citing text could support automated “figure–claim mismatch” flags during submission checks, not only post-hoc quality scores.
- Because Structure and OCR-dense plots remain the hardest cases, layout- and OCR-aware visual backbones are the most direct path to further gains without changing the rubric.
- If mid–high score concentration is typical of accepted papers, ranking objectives may matter more than absolute regression for any review-triage metric, not only figures.
- Extending the protocol beyond CS conference PDFs would test whether “peer-review figure utility” is domain-general or venue-specific.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that scientific figure quality for peer review is a manuscript-grounded multimodal problem, not a pure perceptual IQA or generic CLIP-matching task. It releases a corpus of 3,857 figures from ICLR/NeurIPS/ICML PDFs bound to captions and index-resolved citing paragraphs, each scored on Clarity, Relevance, Informativeness, and Structure (1–5). SciFigAlign fine-tunes CLIP ViT-B/32 and SciBERT with per-modality cross-attention and CubeMLP fusion, jointly optimizing SmoothL1 regression and a within-paper ranking hinge. On a paper-level held-out human test set (n=396) it reports macro MAE 0.3524 and within-paper pairwise accuracy 81.64%, a large gain over constant, similarity-ridge, and zero-shot LLM/VLM judge baselines (best judge MAE 0.864). Ablations attribute gains to caption/context binding, citing-context denoising, and ranking supervision.
Significance. If the result holds under a stable annotation target, the work is a solid systems and benchmark contribution for review assistants and authoring tools: it correctly identifies that published-figure utility depends on claim–visual alignment, supplies a manuscript-bound corpus with paper-level splits, and shows that supervised multimodal fusion plus within-paper ranking substantially outperforms prompt-only VLM judges on the same fields. Strengths include clear multi-family baselines, bootstrap CIs, an interpretable ablation ladder (image→caption→context→denoising→ranking), exportable attention/modality weights, and an explicit pairwise metric (PA) matched to triage use. The main interpretive claim—that learned visual–manuscript alignment is required rather than prompting alone—is consequential for how the community builds figure-assessment tools, provided the gold labels are shown to be reliable peer-review signal rather than protocol-specific fit.
major comments (4)
- [Appendix B, G; §4.2; Abstract] Appendix B and G state that the authors do not report human–human agreement ceilings on this corpus, while training mixes 1,875 GPT-4o-assisted labels accepted after pilot screens (Krippendorff’s α>0.6) with 1,982 human labels; the canonical test set is human-only (n=396). The headline 59% relative MAE reduction (0.864→0.3524) and the claim that assessment “requires learned alignment … rather than prompting alone” treat the CL/RE/IN/ST labels as a stable peer-review gold standard. Without double-annotation agreement (overall and per head, especially RE/IN/ST) on the held-out human test set, one cannot tell whether MAE 0.3524 is near rater noise, whether the judge’s higher error partly reflects disagreement with an idiosyncratic protocol, or how much PA 81.64% is transferable triage signal. Please report multi-rater agreement (or a second independent pass) on a substantial test subset and
- [Table 4; §4.2 Ranking vs. Absolute Error; §5] Table 4 reports SciFigAlign SRCC 0.3088 alongside PA 81.64% (gap≥0.5, |P|=365). The manuscript correctly notes mid–high score concentration (mean 3.54) and that PA and SRCC answer different questions, but the abstract and conclusion still lead with a strong necessity claim about learned alignment. Given compressed labels, moderate global rank correlation weakens the interpretation that the model recovers a well-ordered peer-review utility scale rather than mainly separating clearer within-paper winners under τ=0.5. Please either strengthen ordinal evidence (e.g., SRCC/PA stratified by dimension and score gap, calibration plots) or narrow the claim to within-paper triage under this rubric rather than general superiority of learned alignment for scientific figure assessment.
- [Table 5; §4.3; Appendix C.10, F] Table 5’s test ladder shows large MAE drops from adding caption/context and ranking, but the matched validation protocol (Appendix F / Table 4 in supplement) finds Image+Caption best on absolute MAE (0.2733) while Full Denoised+λ=0.10 peaks Spearman; the deployed checkpoint is chosen by validation loss with λ=0.2 and lands at test MAE 0.3524. This val/test and objective trade-off is acknowledged but under-discussed relative to the claim that full manuscript grounding is critical. Please make checkpoint selection and the λ trade-off primary (not appendix-only), report the selected model’s validation metrics next to test, and avoid equating the validation Image+Caption MAE with the canonical full-model test result in any summary comparison.
- [Abstract; §5; Appendix A, G] Scope and external validity: the corpus is restricted to ICLR/NeurIPS/ICML PDF figures (Appendix G), with architecture/qualitative/other mixes typical of ML papers. Concurrent SIQA/SIU2A/AIBench lines are reasonably distinguished in Appendix A, but the necessity claim in the abstract is phrased as about scientific figure assessment generally. Either add a small out-of-venue or out-of-domain stress set, or explicitly limit claims to ML conference figures under this extraction and rubric pipeline so the 59% reduction is not over-read as field-wide.
minor comments (6)
- [Figure 3] Figure 3 architecture diagram labels Informativeness as “(N)” rather than “(IN)”; fix for consistency with the CL/RE/IN/ST notation used everywhere else.
- [Table 4] Table 4 footnote: some baselines use n=185 while main rows use n=396. State clearly in the table body or caption which methods share the exact same test instances to avoid apples-to-oranges reading of MAE gaps among non-learned baselines.
- [§3.2 Eq. (1)] Eq. (1) / PA definition uses overall mean scores; briefly restate in §3.2 that yi are scalar means over four dimensions so readers do not confuse vector vs. scalar pairwise comparisons.
- [§2; Appendix A] Related work and Appendix A cite several 2026 preprints as concurrent; ensure final versions, venues, and reproducibility status are updated at camera-ready and that non-replicable closed benchmarks are not implied to have been run.
- [§1, §3] Minor prose issues: spacing anomalies (“afigure”, “thecaption”, “within-paperranking”) appear in the compiled text; a full copy-edit pass would help.
- [Appendix D] Dashboard screenshots in Appendix D (Figures 1–4) are useful for reproducibility narrative but are secondary; if space-constrained, keep only metrics-export description and move UI screenshots to supplementary material only.
Circularity Check
Standard supervised multimodal regression on held-out human labels; no derivation-by-construction circularity.
full rationale
SciFigAlign does not present a first-principles derivation whose outputs reduce to its inputs by definition. The load-bearing chain is ordinary supervised learning: annotate figures on a CL/RE/IN/ST rubric, fine-tune CLIP+SciBERT with cross-attention and CubeMLP under SmoothL1 plus a within-paper ranking hinge, and evaluate MAE/SRCC/PA on a paper-level held-out human-rated test set (n=396) against frozen similarity ridges and zero-shot LLM judges that share no weights and are not distilled into the model. The ranking term (Eq. 5–6 / L_pair) uses gold same-paper orderings as supervision—the standard ranking-loss pattern, not a fit renamed as an independent prediction. Ablations vary inputs and λ under matched splits; reported gains are empirical comparisons, not tautologies. Mild residual concern is only that 1,875 GPT-4o-assisted labels expand training after pilot screens, but the canonical test is human-only and assisted labels are not used as a teacher—so this is annotation-protocol risk, not circular reduction of the claimed result to its inputs. No self-citation uniqueness theorem, ansatz smuggling, or self-definitional step anchors the 59% MAE claim. Score 1 only for that weak training-label mix, not for a circular derivation.
Axiom & Free-Parameter Ledger
free parameters (4)
- ranking loss weight λ =
0.2 (deployed)
- ranking margin m and pair gap τ =
m=1.0, τ=0.5
- CubeMLP/cross-attn width and depths =
d=256, 3 CubeMLP blocks, 4 heads
- GPT-4o-assisted label acceptance threshold =
Krippendorff α>0.6 pilot screen
axioms (5)
- domain assumption Peer-review figure utility is adequately captured by four continuous 1–5 dimensions CL/RE/IN/ST and their mean overall score.
- domain assumption Index-resolved citing-paragraph binding (plus denoising of bare 'see Fig. k' and OCR headers) yields the correct manuscript evidence for Relevance/Informativeness.
- domain assumption Paper-level held-out human scores are an unbiased estimate of generalization for the deployed multimodal regressor.
- ad hoc to paper SmoothL1 on rubric scores plus margin hinge on mean-score pairs is a valid surrogate for reviewer comparative judgments.
- domain assumption Standard encoder inductive biases (CLIP ViT-B/32 vision, SciBERT text) remain appropriate after end-to-end fine-tuning for print-scale scientific figures.
invented entities (2)
-
SciFigAlign corpus (3,857 manuscript-bound figures with CL/RE/IN/ST labels)
no independent evidence
-
SciFigAlign model (per-modality cross-attention + CubeMLP fusion + four score heads + ranking objective)
no independent evidence
read the original abstract
Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper's scientific argument; CLIP-based methods assess generic image-text correspondence, yet lack understanding of manuscript context; and zero-shot LLM/VLM judges, when repurposed for figure scoring, often yield overly concentrated scores with limited fusion of visual and textual evidence. We introduce an annotated dataset of 3,857 scientific figures from peer-reviewed conference papers, each rated along four peer-review-oriented dimensions: Clarity, Relevance, Informativeness, and Structure. We propose SciFigAlign, a fine-tuned multimodal scorer that grounds figure quality assessment in manuscript evidence. Given a figure crop, caption, citing paragraphs, and light paper context, SciFigAlign fine-tunes CLIP and SciBERT end-to-end with per-modality cross-attention and CubeMLP fusion, jointly optimizing SmoothL1 regression with a within-paper ranking hinge loss. Under paper-level splits, SciFigAlign achieves a macro MAE of 0.3524 and a within-paper pairwise accuracy of 81.64% on the test set with n = 396, a 59% relative error reduction over the best LLM-as-judge baseline with MAE 0.864. Ablations confirm that manuscript-grounded inputs, citing-context denoising, and ranking supervision are all critical, showing that scientific figure assessment requires learned alignment between visual content and manuscript evidence rather than prompting alone, even with state-of-the-art VLMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the 38th International Conference on Machine Learning , series =
Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and Krueger, Gretchen and Sutskever, Ilya , title =. Proceedings of the 38th International Conference on Machine Learning , series =. 2021 , url =
2021
-
[2]
Beltagy, Iz and Lo, Kyle and Cohan, Arman , title =. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , pages =. 2019 , publisher =. doi:10.18653/v1/D19-1371 , url =
-
[3]
Lee and Huang, Ting-Hao , title =
Hsu, Ting-Yao and Giles, C. Lee and Huang, Ting-Hao , title =. Findings of the Association for Computational Linguistics: EMNLP 2021 , pages =. 2021 , publisher =. doi:10.18653/v1/2021.findings-emnlp.277 , url =
-
[4]
Journal of Natural Language Processing , volume =
Yang, Zhishen and Dabre, Raj and Tanaka, Hideki and Okazaki, Naoaki , title =. Journal of Natural Language Processing , volume =. 2024 , doi =
2024
-
[5]
Advances in Neural Information Processing Systems , volume =
Roberts, Jonathan and Han, Kai and Houlsby, Neil and Albanie, Samuel , title =. Advances in Neural Information Processing Systems , volume =. 2024 , note =. doi:10.52202/079017-0593 , url =
-
[6]
Findings of the Association for Computational Linguistics: ACL 2022 , pages =
Masry, Ahmed and Long, Do Xuan and Tan, Jia Qing and Joty, Shafiq and Hoque, Enamul , title =. Findings of the Association for Computational Linguistics: ACL 2022 , pages =. 2022 , publisher =. doi:10.18653/v1/2022.findings-acl.177 , url =
-
[7]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =. 2023 , publisher =. doi:10.18653/v1/2023.emnlp-main.153 , url =
-
[8]
Findings of the Association for Computational Linguistics: ACL 2024 , pages =
Lee, Seongyun and Kim, Seungone and Park, Sue Hyun and Kim, Geewook and Seo, Minjoon , title =. Findings of the Association for Computational Linguistics: ACL 2024 , pages =. 2024 , publisher =. doi:10.18653/v1/2024.findings-acl.672 , url =
-
[9]
, title =
Mittal, Anish and Moorthy, Anush Krishna and Bovik, Alan C. , title =. IEEE Transactions on Image Processing , volume =. 2012 , doi =
2012
-
[10]
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages =
Hessel, Jack and Holtzman, Ari and Forbes, Maxwell and Bras, Ronan Le and Choi, Yejin , title =. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages =. 2021 , publisher =. doi:10.18653/v1/2021.emnlp-main.595 , url =
-
[11]
and Artzi, Yoav , title =
Zhang, Tianyi and Kishore, Varsha and Wu, Felix and Weinberger, Kilian Q. and Artzi, Yoav , title =. International Conference on Learning Representations , year =
-
[12]
2023 IEEE/CVF International Conference on Computer Vision (ICCV) , pages =
Zhai, Xiaohua and Mustafa, Basil and Kolesnikov, Alexander and Beyer, Lucas , title =. 2023 IEEE/CVF International Conference on Computer Vision (ICCV) , pages =. 2023 , doi =
2023
-
[13]
Proceedings of the 30th ACM International Conference on Multimedia , pages =
Sun, Hao and Wang, Hongyi and Liu, Jiaqing and Chen, Yen-Wei and Lin, Lanfen , title =. Proceedings of the 30th ACM International Conference on Multimedia , pages =. 2022 , publisher =. doi:10.1145/3503161.3548025 , url =
arXiv 2022
-
[14]
Advances in Neural Information Processing Systems , volume =
Lu, Jiasen and Batra, Dhruv and Parikh, Devi and Lee, Stefan , title =. Advances in Neural Information Processing Systems , volume =. 2019 , url =
2019
-
[15]
Li, Wenzhe and Chen, Liang and Wang, Junying and Guo, Yijing and Shen, Ye and Wen, Farong and Li, Chunyi and Zhang, Zicheng and Zhai, Guangtao , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.06700 , url =
-
[16]
Liao, Zhaohe and Jiang, Kaixun and Liu, Zhihang and Wei, Yujie and Yu, Junqiu and Li, Quanhao and Yu, Hong-Tao and Li, Pandeng and Wang, Yuzheng and Xing, Zhen and Zhang, Shiwei and Xie, Chen-Wei and Zheng, Yun and Liu, Xihui , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.28068 , url =
-
[17]
Towards Characterizing Scientific Image Utility and Upgradability
Li, WenZhe and Yan, Qihang and Chen, Liang and Wang, Junying and Wen, Farong and Guo, Yijin and Li, Chunyi and Zhang, Zicheng and Zhai, Guangtao , title =. 2026 , eprint =. doi:10.48550/arXiv.2606.03401 , url =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2606.03401 2026
-
[18]
Zhu, Dawei and Meng, Rui and Song, Yale and Wei, Xiyu and Li, Sujian and Pfister, Tomas and Yoon, Jinsung , title =. 2026 , eprint =. doi:10.48550/arXiv.2601.23265 , url =
-
[19]
The Fourteenth International Conference on Learning Representations , year =
Zhu, Minjun and Lin, Zhen and Weng, Yixuan and Lu, Panzhong and Xie, Qiujie and Wei, Yifan and Liu, Sifan and Sun, QiYao and Zhang, Yue , title =. The Fourteenth International Conference on Learning Representations , year =
-
[20]
Guan, Yaohan and Wang, Pristina and Dehak, Najim and Yuille, Alan and Chen, Jieneng and Khashabi, Daniel , title =. 2026 , eprint =. doi:10.48550/arXiv.2604.04172 , url =
-
[21]
Ding, Junpeng and Tang, Zichen and E, Haihong and Ji, Mengyuan and Liu, Yang and Tian, Haolin and Sun, Haiyang and Sun, Pengqi and Xu, Yang and Liu, Yichen and Gao, Haocheng and Xi, Zijie and Jiang, Ruomeng and Zhao, Peizhi and Li, Rongjin and Li, Yuanze and Liu, Jiacheng and Yang, Zhongjun and Chen, Jintong and Lin, Siying , title =. Proceedings of the 6...
doi:10.18653/v1/2026 2026
-
[22]
The Fourteenth International Conference on Learning Representations , year =
Xie, Yupeng and Zhang, Zhiyang and Wu, Yifan and Lu, Sirong and Zhang, Jiayi and Yu, Zhaoyang and Wang, Jinlin and Hong, Sirui and Liu, Bang and Wu, Chenglin and Luo, Yuyu , title =. The Fourteenth International Conference on Learning Representations , year =
-
[23]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages =
Ye, Guanghui and Zhao, Huan and Zhao, Zhixue and Ma, Tengfei and Wang, Kehan and Eger, Steffen and Jiang, Zhihua , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages =. 2026 , url =
2026
-
[24]
Advances in Neural Information Processing Systems , volume =
Xu, Jiazheng and Liu, Xiao and Wu, Yuchen and Tong, Yuxuan and Li, Qinkai and Ding, Ming and Tang, Jie and Dong, Yuxiao , title =. Advances in Neural Information Processing Systems , volume =. 2023 , doi =
2023
-
[25]
and Vo, Azalea A
Borkin, Michelle A. and Vo, Azalea A. and Bylinskii, Zoya and Isola, Phillip and Sunkavalli, Shashank and Oliva, Aude and Pfister, Hanspeter , title =. IEEE Transactions on Visualization and Computer Graphics , volume =. 2013 , doi =
2013
-
[26]
and Sheikh, Hamid R
Wang, Zhou and Bovik, Alan C. and Sheikh, Hamid R. and Simoncelli, Eero P. , title =. IEEE Transactions on Image Processing , volume =. 2004 , doi =
2004
-
[27]
Learning Visual Importance for Graphic Designs and Data Visualizations , booktitle =
Bylinskii, Zoya and Kim, Nam Wook and O'Donovan, Peter and Alsheikh, Sami and Madan, Spandan and Pfister, Hanspeter and Durand, Fr. Learning Visual Importance for Graphic Designs and Data Visualizations , booktitle =. 2017 , publisher =. doi:10.1145/3126594.3126653 , url =
arXiv 2017
-
[28]
Journal of Vision , volume =
Rosenholtz, Ruth and Li, Yuanzhen and Nakano, Lisa , title =. Journal of Vision , volume =. 2007 , doi =
2007
-
[29]
2011 , url =
Krippendorff, Klaus , title =. 2011 , url =
2011
-
[30]
2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages =
Prabhu, Viraj and Purushwalkam, Senthil and Yan, An and Xiong, Caiming and Xu, Ran , title =. 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages =. 2025 , doi =
2025
-
[31]
Wang, He and Guo, Longteng and Huo, Pengkang and Lin, Xuanxu and Yuan, Yichen and Jiang, Jie and Liu, Jing , title =. 2026 , eprint =. doi:10.48550/arXiv.2601.00264 , url =
-
[32]
Sun, Hao and Shen, Yunyi and van der Schaar, Mihaela , title =. 2025 , eprint =. doi:10.48550/arXiv.2505.21537 , url =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2505.21537 2025
-
[33]
2025 3rd International Conference on Inventive Computing and Informatics (ICICI) , pages =
Raval, Parth and Bhaidasna, Hetal , title =. 2025 3rd International Conference on Inventive Computing and Informatics (ICICI) , pages =. 2025 , publisher =. doi:10.1109/ICICI65870.2025.11069873 , url =
arXiv 2025
-
[34]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Advances in Neural Information Processing Systems , volume =. 2023 , note =. doi:10.52202/075280-2020 , url =
-
[35]
Proceedings of the 41st International Conference on Machine Learning , series =
Wu, Haoning and Zhang, Zicheng and Zhang, Weixia and Chen, Chaofeng and Liao, Liang and Li, Chunyi and Gao, Yixuan and Wang, Annan and Zhang, Erli and Sun, Wenxiu and Yan, Qiong and Min, Xiongkuo and Zhai, Guangtao and Lin, Weisi , title =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , url =
2024
-
[36]
Advances in Neural Information Processing Systems , volume =
Wang, Zirui and Xia, Mengzhou and He, Luxi and Chen, Howard and Liu, Yitao and Zhu, Richard and Liang, Kaiqu and Wu, Xindi and Liu, Haotian and Malladi, Sadhika and Chevalier, Alexis and Arora, Sanjeev and Chen, Danqi , title =. Advances in Neural Information Processing Systems , volume =. 2024 , note =. doi:10.52202/079017-3609 , url =
-
[37]
and Kumar, Pratyush , title =
Methani, Nitesh and Ganguly, Pritha and Khapra, Mitesh M. and Kumar, Pratyush , title =. 2020 IEEE Winter Conference on Applications of Computer Vision (WACV) , pages =. 2020 , doi =
2020
-
[38]
Workshop Track of the Sixth International Conference on Learning Representations , year =
Kahou, Samira Ebrahimi and Michalski, Vincent and Atkinson, Adam and K. Workshop Track of the Sixth International Conference on Learning Representations , year =
-
[39]
Advances in Neural Information Processing Systems , volume =
Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae , title =. Advances in Neural Information Processing Systems , volume =. 2023 , doi =
2023
-
[40]
Liu, Fangyu and Piccinno, Francesco and Krichene, Syrine and Pang, Chenxi and Lee, Kenton and Joshi, Mandar and Altun, Yasemin and Collier, Nigel and Eisenschlos, Julian , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2023 , publisher =. doi:10.18653/v1/2023.acl-long.714 , url =
-
[41]
Li, Shengzhi and Tajbakhsh, Nima , title =. 2023 , eprint =. doi:10.48550/arXiv.2308.03349 , url =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.