Pith. sign in

REVIEW 4 major objections 5 minor 32 references

Image-blind LLM judging of 200k disaster captions yields high-precision, low-noise anchors for data-free distillation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 12:41 UTC pith:WBHMEKX2

load-bearing objection Solid 100k multimodal disaster resource and a clean image-blind judge design; the DFKD-suitability claim is motivated but not demonstrated. the 4 major comments →

arxiv 2607.28269 v1 pith:WBHMEKX2 submitted 2026-07-30 cs.CV cs.AIcs.MM

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

classification cs.CV cs.AIcs.MM
keywords Vision-Language ModelsData-Free Knowledge DistillationLLM-as-a-JudgeDisaster ManagementDataset GenerationMixture-of-ExpertsIncidents1MMultimodal Captioning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Disaster-response vision models need clean image-text pairs so knowledge can be transferred without original training images, yet the best visual incident dataset has no text and social-media alternatives are noisy. This paper recovers 100,000 images from Incidents1M, captions them with both a dense 4B and a 35B Mixture-of-Experts vision-language model, and then validates the captions with a third model that never sees the image. The blind judge recovers original multi-labels from text alone, scoring 78.65/100 agreement between the two captioners and revealing a conservative profile: roughly 77% precision and 46% recall. High precision means few invented disasters leak into the text, which the authors argue is exactly the property a student model needs when text is its only semantic anchor. The same pipeline also flags inconsistent human labels in the original ground truth, turning the validation step into a dataset audit.

Core claim

An image-blind LLM-as-a-Judge applied to 200,000 captions over 100,000 recovered Incidents1M images shows high semantic agreement (78.65/100) between dense and MoE captioners and a conservative high-precision (~77.6%), low-recall (~46%) profile that supplies reliable, low-false-positive semantic anchors for multimodal data-free knowledge distillation while also exposing human annotation errors in the original ground truth.

What carries the argument

Image-blind LLM-as-a-Judge: a 9B language model that never sees the source image and must recover Incidents1M multi-labels from caption text alone, thereby measuring whether essential visual information survived the image-to-text transition under the same modality gap a data-free student faces.

Load-bearing premise

That a text-only judge recovering original multi-labels is a faithful stand-in for the modality gap a real data-free student will face, so high judge precision can be taken as proof the captions are good distillation anchors.

What would settle it

Run the promised end-to-end data-free distillation with these captions as the sole anchors and measure whether the student’s disaster-classification accuracy on held-out images rises or collapses relative to a supervised or noisy-text baseline.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The released 100k-image, 200k-caption set becomes a usable multimodal resource for training compact disaster models without original images.
  • Conservative high-precision captions can be preferred over exhaustive ones when the downstream goal is data-free transfer.
  • Mixture-of-Experts captioners improve rare long-tail disaster classes without harming common ones.
  • Modern vision-language models can serve as auditors that surface and correct crowdsourced label errors before distillation.
  • The same image-blind validation recipe can be reused on other vision-only benchmarks that need text for cross-modal transfer.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the same-family judge and generators share systematic blind spots, measured precision may overstate true caption fidelity for an unrelated student architecture.
  • Spatial-coordinate prompting may be doing as much work as model scale; ablating it would isolate how much of the DFKD readiness comes from the prompt rather than the VLM.
  • Extending the audit to the full million-image Incidents1M corpus could produce a cleaned multi-label benchmark whose value exceeds the caption set itself.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript reconstructs 100,000 images from the vision-only Incidents1M benchmark via a fault-tolerant download pipeline, generates 200,000 spatially constrained captions with Qwen3.5-4B (dense) and Qwen3.5-35B-A3B (MoE), and validates them with an image-blind Qwen3.5-9B LLM-as-a-Judge. Qualitative comparison yields mean semantic agreement of 78.65/100 between the two captioners. Quantitative label recovery against 173,179 Incidents1M multi-label pairs shows conservative behaviour (Precision ≈77.6%, Recall ≈46%). The authors argue that image-blind label recovery simulates the DFKD student modality gap and that high-precision captions therefore supply reliable semantic anchors for multimodal data-free distillation; they also surface human annotation inconsistencies in the original ground truth and note MoE gains on the long tail of rare disaster classes.

Significance. A clean, large-scale multimodal extension of Incidents1M would be a genuine community resource for disaster-response VLMs, where existing multimodal sets (e.g., CrisisMMD) suffer text–image misalignment. The atomic reconstruction pipeline, dense-vs-MoE captioning comparison under continuous batching, per-class long-tail analysis (Tables 1–2, Figs. 6–7), and the demonstration that modern VLMs can audit crowdsourced labels (§6) are concrete, reusable contributions. The image-blind judge protocol is a thoughtful design choice for measuring text-only information retention. These strengths stand even if the DFKD-suitability headline is not yet experimentally closed; the work is therefore significant as dataset and methodology, provided claims about distillation anchors are either validated or appropriately scoped.

major comments (4)
  1. [§4.2, §5.2, §7] Abstract, §4.2, §5.2 and §7 assert that the captions are “reliable semantic anchors for DFKD” because the image-blind judge recovers Incidents1M labels at ~77.6% precision. No end-to-end DFKD run (surrogate image synthesis or student training/evaluation) is reported. Without that link, suitability for distillation remains an untested design hypothesis rather than a demonstrated result. Either a minimal DFKD experiment on a subset or a clear reframing of the contribution as dataset construction plus text-retention validation is required for the central claim to hold.
  2. [§4.2–4.3, §5.1] Generators (4B, 35B-A3B) and judge (9B) all belong to the Qwen3.5 family. Inter-model agreement (78.65/100) and label-recovery precision can therefore partly reflect shared lexical and taxonomic priors rather than independent visual grounding. An ablation with a non-Qwen or cross-family judge (or a small human correlation study on a stratified subset) is needed to show that the reported scores are not family-internal.
  3. [§4.2] §4.2 equates “judge never sees the image” with “faithfully simulates the modality gap a DFKD student faces.” A DFKD student typically consumes text to drive a generative visual model and is then evaluated on real images; label recovery from text alone does not measure whether the retained semantics suffice to synthesize useful surrogate visuals or to train a competent student. The proxy should be justified with evidence or demoted from operational equivalence to a necessary-but-not-sufficient retention check.
  4. [Table 1, §5.2, §6] Table 1 reports high Precision / low Recall (~46%). §5.2 claims this conservative profile is “highly advantageous” for DFKD because false-positive noise is worse than omissions. That preference is plausible but unsupported by any distillation or synthesis ablation (e.g., high-P vs. higher-R caption sets). Given the authors’ own finding of GT omissions and inconsistencies (§6), the precision figure is also a lower bound against a noisy reference; it does not establish sufficiency of the retained text. Soften or evidence this claim.
minor comments (5)
  1. [§3.3, Fig. 2] Figure 2 captions state generations were translated from Italian; the system prompt language and any effect of translation on the English judge evaluation should be stated explicitly in §3.3.
  2. [§5.1] The Final Score weighting that produces 78.65/100 from the four 0–10 criteria is not specified; a short formula or weight table would aid reproducibility.
  3. [Fig. 4, Fig. 5] Fig. 4 and Fig. 5 are informative but axis labels and the exact definition of “errors per image” could be clearer in the captions for standalone reading.
  4. [§1, §4.1] Minor typography/spacing issues appear in the compiled text (e.g., missing spaces after periods in “Thecontemporarylandscape”, “Toovercomethesebottlenecks”); a pass over PDF line-break artifacts would improve readability.
  5. [§2.4] Related work on multimodal DFKD and hallucination mitigation is adequate; a brief pointer to recent caption-faithfulness or VLM-as-judge calibration papers would strengthen §2.4.

Circularity Check

1 steps flagged

No derivation-by-construction circularity; DFKD suitability is an unvalidated proxy framing, not a result forced by the paper’s own inputs.

specific steps
  1. self definitional [Section 4.2; Abstract; Section 5.2]
    "By making sure the judge cannot see the image, we measure exactly what the DFKD process requires, which is whether the essential information survived the transition from the visual domain to the textual representation. ... To ensure the generated captions provide reliable semantic anchoring for DFKD, we introduce an image-blind LLM-as-a-Judge validation pipeline"

    Minor only: the paper operationally defines ‘reliable DFKD anchor’ as image-blind recovery of Incidents1M labels from text, then reports that metric as establishing DFKD suitability. The numerical Precision/Recall themselves are not forced by this definition (they use external GT), and no DFKD student performance is predicted from a fitted quantity. This is proxy framing without independent verification, not Eq.X = Eq.Y by construction; scored as a 1, not a central circular derivation.

full rationale

This is an empirical dataset-construction and automated-evaluation paper, not a first-principles derivation. The quantitative scores (Precision/Recall/F1 against 173,179 Incidents1M label pairs; 78.65/100 inter-captioner agreement) are measurements against an external multi-label ground truth and between two independently generated caption sets; they are not algebraically or statistically forced by a fitted parameter that is then re-reported as a prediction. There is no self-citation load-bearing uniqueness theorem, no ansatz imported from overlapping-author prior work, and no renaming of a known closed-form result. The only mild definitional move is equating image-blind label recovery with ‘what DFKD requires’ (Sec. 4.2) and then treating high judge precision as evidence of ‘reliable semantic anchoring for DFKD’ without an end-to-end distillation run—this is an unvalidated operational proxy (a validity/correctness gap), not circularity by construction. Same-family Qwen generators and judge raise a shared-prior confound, again a validity concern rather than a reduction of output to input. Honest finding: no significant circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 2 invented entities

The load-bearing story rests on methodological assumptions about what an image-blind same-family LLM judge measures, on the untested claim that high-precision/low-recall text is the right objective for DFKD, and on engineering choices (temperature, prompt spatial mandates, dynamic URL reservation) rather than on free physical parameters. No new particles or forces are invented; the invented entities are protocol constructs.

free parameters (3)
  • generative_temperature = 0.2
    Fixed at 0.2 to reduce hallucination/entropy; choice is hand-set and directly shapes caption conservatism and reported precision/recall.
  • final_semantic_agreement_weighting
    Final Score is a weighted average of four 0-10 judge criteria (Objects, Position/Action, Style/Detail, Global); exact weights are not fully specified yet the 78.65/100 headline depends on them.
  • concurrent_worker_count_and_batching = 48 workers
    Up to 48 async workers and continuous batching on H100s affect which generations succeed at scale; operational, not scientific, but conditions the realized 100k set.
axioms (5)
  • ad hoc to paper Recovering multi-label disaster categories from caption text alone, without the image, faithfully simulates the information available to a DFKD student and therefore measures anchor quality for distillation.
    Core design claim of Sections 4.2–4.3; never validated by training a student.
  • domain assumption Qwen3.5-9B as LLM-as-a-Judge produces reliable binary label decisions and 0-10 semantic scores that correlate with human judgment for disaster captions.
    Invokes the broader LLM-as-a-Judge literature (Zheng et al., G-Eval, Prometheus) without a paper-specific human correlation study.
  • ad hoc to paper High precision / low recall (conservative captioning) is preferable to higher-recall captions for multimodal DFKD because false-positive semantic noise is more harmful than omissions.
    Stated in Section 5.2 as advantageous; no ablation on student accuracy.
  • domain assumption Incidents1M multi-labels, despite crowdsourcing noise the paper itself documents, remain a usable external reference for precision/recall.
    Quantitative Table 1 treats GT as reference while Section 6 shows systematic human inconsistencies.
  • domain assumption Strict spatial language in captions (foreground/background/corners) is necessary to transfer spatial structure into surrogate images under DFKD.
    Prompt design claim in Section 3.3; plausible but untested in a generation loop.
invented entities (2)
  • image-blind LLM-as-a-Judge validation pipeline for DFKD anchors no independent evidence
    purpose: Score caption fidelity and label retention without giving the judge pixels, to mimic student modality gap.
    Central methodological construct; independent evidence would be human correlation or successful downstream DFKD, neither shipped.
  • Theia multimodal Incidents1M extension (100k images, 200k captions) no independent evidence
    purpose: Provide clean text anchors for disaster DFKD and cross-modal training.
    Primary artifact; existence is demonstrated inside the paper’s pipeline but public release/DOI is not evidenced in the text.

pith-pipeline@v1.2.0-daily-grok45 · 18249 in / 3803 out tokens · 92892 ms · 2026-07-31T12:41:49.232972+00:00 · methodology

0 comments
read the original abstract

The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for transferring knowledge via Data-Free Knowledge Distillation (DFKD). However, existing datasets in this domain either entirely lack descriptive text, such as Incidents1M, or suffer from severe text-image semantic misalignment, such as CrisisMMD. In this work, we present a novel methodology to construct and automatically validate a large-scale multimodal dataset for disaster response. Starting from the vision-only Incidents1M, we successfully recovered 100,000 images and generated high-fidelity textual descriptions using two distinct Qwen3.5 architectures: a 4B dense model and a 35B Mixture-of-Experts (MoE) model. To ensure the generated captions provide reliable semantic anchoring for DFKD, we introduce an image-blind LLM-as-a-Judge validation pipeline leveraging Qwen3.5-9B. By intentionally obscuring the original image from the judge, this evaluator accurately simulates the modality gap of the student model during data-free distillation. Our evaluation across 173,179 label pairs demonstrates a high semantic agreement (78.65/100) between the two architectures. Furthermore, the automated evaluation reveals a conservative captioning behaviour, characterized by a high Precision (77.6%) and low Recall (46.0%). This minimizes the false positive noise, while simultaneously exposing underlying human annotation inconsistencies in the original ground truth. This work provides a scalable, LLM-validated multimodal dataset and a reproducible framework to advance cross-modal knowledge distillation.

Figures

Figures reproduced from arXiv: 2607.28269 by Adriano Mancini, Alessandro Galdelli, Lorenzo Severini, Simone Giano.

Figure 1
Figure 1. Figure 1: Image examples extracted from Incidents1M [16]. [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Qualitative comparison between dense and MoE generations regarding an image [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of the Final Semantic Agreement Score (0-100) between the 4B and [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Distribution of positive and negative labels across the 43 incident categories in [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Log-scale distribution of correctly identified labels per image (left) and total [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: F1-score comparison between the Qwen3.5-4B and 35B-A3B models across the [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Variation of the F1-score per label (35B-A3B minus 4B). The MoE architecture [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Inconsistent ground truth labeling in Incidents1M [16]. Both images depict [PITH_FULL_IMAGE:figures/full_fig_p018_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Ground truth omissions. In both images, the VLMs correctly described severe [PITH_FULL_IMAGE:figures/full_fig_p019_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 2 canonical work pages

  1. [1]

    C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, T. Duerig, Scaling up visual and vision-language rep- resentation learning with noisy text supervision, CoRR abs/2102.05918 (2021). arXiv:2102.05918. URLhttps://arxiv.org/abs/2102.05918

  2. [2]

    Mensch, K

    J.-B.Alayrac, J.Donahue, P.Luc, A.Miech, I.Barr, Y.Hasson, K.Lenc, A. Mensch, K. Millicah, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, K. Simonyan, Flamingo: a visual language model for few-sh...

  3. [3]

    J. Li, D. Li, S. Savarese, S. Hoi, Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models, in: Proceedings of the 40th International Conference on Machine Learning, ICML’23, JMLR.org, 2023

  4. [4]

    H. Liu, C. Li, Q. Wu, Y. J. Lee, Visual instruction tuning, in: Thirty- seventh Conference on Neural Information Processing Systems, 2023. URLhttps://openreview.net/forum?id=w0H2xGHlkw

  5. [5]

    M. Du, A. Ramisa, A. K. K C, S. Chanda, M. Wang, N. Rajesh, S. Li, Y. Hu, T. Zhou, N. Lakshminarayana, S. Tran, D. Gray, Amazon shop the look: A visual search system for fashion and home, KDD ’22, As- sociation for Computing Machinery, New York, NY, USA, 2022, p. 2822–2830. doi:10.1145/3534678.3539071. URLhttps://doi.org/10.1145/3534678.3539071

  6. [6]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning transferable visual models from natural language supervision, CoRR abs/2103.00020 (2021). arXiv:2103.00020. URLhttps://arxiv.org/abs/2103.00020 21

  7. [7]

    Zitkovich, T

    B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T.-W. E. Lee, I. Leal, Y. Kuang, D. Kalash- nikov, R. Julian...

  8. [8]

    Driess, F

    D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, P. Florence, Palm-e: an embodiedmultimodallanguagemodel, in: Proceedingsofthe40thInter- national Conference on Ma...

  9. [9]

    T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P.-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena, A. Palepu, B. Mustafa, A. Chowdhery, Y. Liu, S. Kornblith, D. Fleet, P. Mansfield, S. Prakash, R. Wong, S. Virmani, C. Semturs, S. S. Mahdavi, B. Green, E. Domi- nowska, B. A. y Arcas, J. Barral, D. Webster, G. S. Corrado, Y. Matias, K. Singhal, P. F...

  10. [10]

    M. Y. Lu, B. Chen, D. F. K. Williamson, R. J. Chen, I. Liang, T. Ding, G. Jaume, I. Odintsov, L. P. Le, G. Gerber, A. V. Parwani, A. Zhang, F. Mahmood, A visual-language foundation model for computational pathology, Nature Medicine 30 (3) (2024) 863–874. doi:10.1038/s41591- 024-02856-4. 22

  11. [11]

    Gurari, Q

    D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, J. P. Bigham, Vizwiz grand challenge: Answering visual questions from blind people, CoRR abs/1802.08218 (2018). arXiv:1802.08218. URLhttp://arxiv.org/abs/1802.08218

  12. [12]

    Gurari, Y

    D. Gurari, Y. Zhao, M. Zhang, N. Bhattacharya, Captioning images taken by people who are blind, in: A. Vedaldi, H. Bischof, T. Brox, J.-M. Frahm (Eds.), Computer Vision – ECCV 2020, Springer International Publishing, Cham, 2020, pp. 417–434

  13. [13]

    F. Ofli, F. Alam, M. Imran, Analysis of social media data using multi- modal deep learning for disaster response, in: 17th International Con- ference on Information Systems for Crisis Response and Management, ISCRAM, ISCRAM, 2020

  14. [14]

    Weber, N

    E. Weber, N. Marzo, D. P. Papadopoulos, A. Biswas, A. Lapedriza, F. Ofli, M. Imran, A. Torralba, Detecting natural disasters, damage, and incidents in the wild, in: A. Vedaldi, H. Bischof, T. Brox, J.-M. Frahm (Eds.), Computer Vision – ECCV 2020, Springer International Publishing, Cham, 2020, pp. 331–350

  15. [15]

    F. Alam, F. Ofli, M. Imran, Crisismmd: Multimodal twitter datasets from natural disasters, Proceedings of the International AAAI Conference on Web and Social Media 12 (1) (Jun. 2018). doi:10.1609/icwsm.v12i1.14983. URLhttps://ojs.aaai.org/index.php/ICWSM/article/view/14983

  16. [16]

    Weber, D

    E. Weber, D. P. Papadopoulos, A. Lapedriza, F. Ofli, M. Imran, A. Torralba, Incidents1m: A large-scale dataset of images with nat- ural disasters, damage, and incidents, IEEE Transactions on Pat- tern Analysis and Machine Intelligence 45 (4) (2023) 4768–4781. doi:10.1109/TPAMI.2022.3191996

  17. [17]

    J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, J. Zhou, Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond (2023). arXiv:2308.12966. URLhttps://arxiv.org/abs/2308.12966

  18. [18]

    Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, 23 J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, Z....

  19. [19]

    Hinton, O

    G. Hinton, O. Vinyals, J. Dean, Distilling the knowledge in a neural network (2015). arXiv:1503.02531. URLhttps://arxiv.org/abs/1503.02531

  20. [20]

    J. Gou, B. Yu, S. J. Maybank, D. Tao, Knowledge distillation: A survey, CoRR abs/2006.05525 (2020). arXiv:2006.05525. URLhttps://arxiv.org/abs/2006.05525

  21. [21]

    H. Chen, Y. Wang, C. Xu, Z. Yang, C. Liu, B. Shi, C. Xu, C. Xu, Q. Tian, Data-free learning of student networks, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 3513–

  22. [22]

    H. Yin, P. Molchanov, J. M. Alvarez, Z. Li, A. Mallya, D. Hoiem, N. K. Jha, J. Kautz, Dreaming to distill: Data-free knowledge trans- fer via deepinversion, in: 2020 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2020, pp. 8712–8721. doi:10.1109/CVPR42600.2020.00874

  23. [23]

    G. Fang, K. Mo, X. Wang, J. Song, S. Bei, H. Zhang, M. Song, Up to 100x faster data-free knowledge distillation, Proceedings of the AAAI Conference on Artificial Intelligence 36 (6) (2022) 6597–6604. doi:10.1609/aaai.v36i6.20613. URLhttps://ojs.aaai.org/index.php/AAAI/article/view/20613

  24. [25]

    Shazeer, A

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. V. Le, G. E. Hinton, J. Dean, Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, CoRR abs/1701.06538 (2017). arXiv:1701.06538. URLhttp://arxiv.org/abs/1701.06538

  25. [26]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bam- ford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, W. E. Sayed, Mixtral of experts (2024). arXiv:2...

  26. [27]

    Zheng, W.-L

    L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, I. Stoica, Judging llm-as-a-judge with mt-bench and chatbot arena, in: Proceedings of the 37th International Conference on Neural Information Processing Sys- tems, NIPS ’23, Curran Associates Inc., Red Hook, NY, USA, 2023

  27. [28]

    Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, C. Zhu, G-eval: NLG evalua- tion using gpt-4 with better human alignment, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Meth- odsinNaturalLanguageProcessing, AssociationforComputationalLin- guistics, Singapore, 2023, pp. 2511–2522. doi:10.18653/v1/2023.emnlp- main.153. URLh...

  28. [29]

    S. Kim, J. Shin, Y. Cho, J. Jang, S. Longpre, H. Lee, S. Yun, S. Shin, S. Kim, J. Thorne, M. Seo, Prometheus: Inducing fine-grained evalua- tion capability in language models, in: The Twelfth International Con- ference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=8euJaTveKw

  29. [30]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: a method for auto- matic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL ’02, Association for Computational Linguistics, USA, 2002, p. 311–318. doi:10.3115/1073083.1073135. URLhttps://doi.org/10.3115/1073083.1073135 25

  30. [31]

    Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004, pp

    C.-Y. Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004, pp. 74–81. URLhttps://aclanthology.org/W04-1013/

  31. [32]

    Lavie, A

    A. Lavie, A. Agarwal, Meteor: an automatic metric for mt evaluation with high levels of correlation with human judgments, in: Proceedings of the Second Workshop on Statistical Machine Translation, StatMT ’07, Association for Computational Linguistics, USA, 2007, p. 228–231. 26

  32. [3521]

    doi:10.1109/ICCV.2019.00361