{"id":"b32c38ac-f087-4cdb-a4fc-13e72668adfd","arxiv_id":"2505.22387","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A training-time module with learnable spatial masks and frequency-based pseudo-domain labels improves dataset condensation on multi-domain data without increasing images per class.","lead":"Dataset condensation methods distill large training sets into tiny synthetic ones, but they assume the data comes from one visual style. This paper adds a training-time module that teaches each synthetic image to carry several domain styles at once, improving accuracy on mixed-domain benchmarks without enlarging the condensed dataset.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MTT comparisons are confounded by a separate hyperparameter search and Gaussian-noise initialization: the large MTT+DAM gains (e.g., PACS 39.7→46.6) may reflect tuning, not DAM.","rationale":"The reader's weakest assumption was the FFT-based pseudo-domain labeling. That is a reasonable concern, but Table 6 already provides an ablation showing FFT labels outperform random labels, and even random labels improve over the baseline, so the empirical claim has some support independent of whether the FFT statistic captures true domain structure. A more load-bearing weakness is the fairness of the MTT experiments. The paper's strongest quantitative evidence includes MTT improvements (e.g., PACS 39.7→46.6, CIFAR-10 10 IPC 50.7→57.9), but Appendix F explicitly states that MTT required a separate hyperparameter search when combined with DAM and Gaussian initialization, while no such re-tuning is reported for the vanilla MTT baseline. Section 4.2 claims prior-method hyperparameters were set identically, which is contradicted by Appendix F. Furthermore, the main tables use Gaussian-noise initialization, which is not MTT's standard initialization and sharply lowers vanilla MTT performance, as shown by the supplementary real-initialization results. Under real initialization, MTT+DAM gains are often marginal (+0.1 to +0.2 on several settings), so the consistent-improvement claim is much weaker than the main tables suggest. This is an internal consistency and experimental-control issue, not a disagreement with field consensus, and it can be settled by matched-hyperparameter re-runs. The DC and DM evidence is more solid, which is why I recommend CONDITIONAL rather than REJECT: the paper should either present controlled MTT comparisons or qualify the claim to DC/DM. I partially agree with the reader because they flagged missing hyperparameters, but not the sharper confound that MTT's baseline and DAM variant used different hyperparameters under a nonstandard initialization.","tokens_in":16453,"tokens_out":10519,"duration_ms":130041,"concrete_test":"Run the PACS 1 IPC MTT comparison in both directions: (1) train vanilla MTT with Gaussian-noise initialization using the exact Table H hyperparameters used for MTT+DAM, but without DAM; (2) train MTT+DAM using the original MTT hyperparameters that produced the 39.7 baseline. If (1) closes most of the gap toward 46.6, or (2) triggers NaNs or severe degradation, the Table 2 gain is a tuning/initialization artifact. Repeat for CIFAR-10 10 IPC and also report real-image-initialization results for all multi-domain Table 2 settings to verify the gains persist outside supplementary Table A.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that DAM consistently improves MTT is not supported by a controlled comparison. Section 4.2 states that all hyperparameters from prior methods are set identically, but Appendix F reveals that MTT required a separate hyperparameter search when combined with DAM and Gaussian-noise initialization, and Table H lists hyperparameters for DAM-with-MTT. No evidence is given that the vanilla MTT baseline was re-tuned under the same Gaussian-noise initialization, so the baseline may be operating with default hyperparameters that are poorly suited to that setting. This is especially serious because all main experiments use Gaussian-noise synthetic initialization, which depresses vanilla MTT: in the supplementary Table A, real-image initialization gives MTT 65.3 on CIFAR-10 at 10 IPC, versus 50.7 in the main Table 1. Under real-image initialization, DAM's MTT gains shrink to +0.2 on CIFAR-10 10 IPC and +0.1 on PACS 1 IPC. Thus headline results such as MTT PACS 1 IPC 39.7→46.6 and CIFAR-10 10 IPC 50.7→57.9 may be artifacts of comparing a default-hyperparameter Gaussian-init baseline against a tuned DAM variant. The DC and DM comparisons are less affected, but the paper's claim spans MTT, so the MTT evidence is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Domain-Aware Module (DAM), a training-time module for dataset condensation that equips each synthetic image with D learnable spatial masks, supervised by pseudo-domain labels derived from low-frequency FFT amplitude sorting. The module is plugged into DC, DM, and MTT; the final condensed images themselves remain the same size and IPC. Experiments on CIFAR-10/100, Tiny ImageNet, PACS, VLCS, Office-Home, and DomainNet report in-domain, cross-architecture, and leave-one-domain-out gains, along with ablations of pseudo-label strategies, domain count, and hyperparameters.","tokens_in":16766,"tokens_out":5909,"duration_ms":59797,"significance":"If the results hold, DAM is a simple and useful plug-in for dataset condensation in heterogeneous and multi-domain data, with a plausible mechanism (spatial domain masks) and a practical pseudo-labeling scheme that needs no explicit domain annotations. The paper is strengthened by systematic ablations, error bars over 10 runs, cross-architecture evaluation, and an honest limitation discussion of condensation-time overhead. The main risk is the MTT comparison, which is not controlled; this must be fixed before the general claim of consistent improvement is accepted.","major_comments":[{"comment":"The MTT comparisons are not controlled, and they are load-bearing for the paper's central claim. Section 4.2 states that 'all of the hyperparameters introduced in each prior method are set identically,' but Appendix F reports that MTT required a separate hyperparameter search when combined with DAM and Gaussian-noise initialization, with chosen values listed in Table H. No evidence is provided that the vanilla MTT baseline was re-tuned under the same Gaussian-noise initialization or the same search budget. This matters because Table A shows that real-image initialization gives vanilla MTT 65.3 on CIFAR-10 at 10 IPC, versus 50.7 in Table 1 under Gaussian initialization; under real initialization, MTT+DAM improves by only +0.2 on CIFAR-10 10 IPC and +0.1 on PACS 1 IPC. The headline MTT+DAM gains (e.g., PACS 1 IPC 39.7 to 46.6, CIFAR-10 10 IPC 50.7 to 57.9) may therefore reflect hyperparameter tuning rather than DAM. Please provide a controlled comparison in which vanilla MTT is re-tuned with Gaussian-noise initialization under the same search budget, or report both initialization settings with matched hyperparameters for all compared methods.","section":"Section 4.2, Appendix F, Table H"},{"comment":"The domain loss is under-specified. D_dom_syn is defined as {(x̃_d_m, ỹ_m)} with class labels ỹ_m, but the text says that for the domain loss the real batch is grouped by pseudo-domain label. If L_base is the same loss as the class loss (e.g., gradient matching or distribution matching), it needs a consistent label set for both the real and synthetic batches. As written, the reader cannot tell whether the synthetic masked images are labeled with pseudo-domain labels, class labels, or both; the notation suggests class labels, which would make the domain loss inconsistent with the claimed pseudo-domain grouping. Please specify exactly how L_base is adapted to the pseudo-domain task, including the label sets for D_real and D_dom_syn, and how the per-image masks α_d_m are associated with pseudo-domain indices.","section":"Section 3.3, Eq. (7)"},{"comment":"The choice of Gaussian-noise initialization deserves more careful treatment in the main text. The paper justifies the choice by privacy-preserving goals, which is reasonable, but the supplementary results show that the MTT+DAM advantage is much smaller under real-image initialization. Since the main claim is 'consistently improves,' the presentation should make clear that the MTT gains in Tables 1 and 2 depend on the interaction of DAM with Gaussian initialization and the separately tuned hyperparameters, rather than presenting them as unconditional improvements. At minimum, the main text should reference the magnitude of the real-initialization results when the MTT rows are discussed.","section":"Section 4.2 and Table A"}],"minor_comments":[{"comment":"The checkmark table for pseudo-domain labeling strategies is difficult to parse because the column headers do not clearly separate the feature extractor (FFT, log-Var) from the clustering/ordering strategy (Mean-Sort, K-Means). Please label the columns explicitly and define what each checkmark combination means.","section":"Table 6"},{"comment":"The sentence 'Due to observed instability at 1 IPC, we omit MTT + DAM results for IPC 10 in this ablation' appears in several places, but the omitted rows are for IPC 10 while the stated instability is at 1 IPC. Please clarify the wording so the reason for omitting IPC 10 is unambiguous.","section":"Supplementary C, Tables B-D"},{"comment":"The caption refers to a dashed red line and a solid blue curve, but the rendered figure appears to be in grayscale; please ensure the legend remains readable in print and that the visual distinction between the baseline and DAM curves is clear.","section":"Figure 4"},{"comment":"There is a typo in the sentence about the effect of the number of pseudo-domains: 'even with random number of domains D' should be 'even with a random number of domains D' or 'for any number of domains D.' Also, 'basline' appears once in the same section and should be corrected.","section":"Section 5"},{"comment":"The pseudo-domain assignment formula uses N/D and integer brackets, but it is not stated how ties are broken or how N is handled when it is not divisible by D; adding a sentence to define the rounding convention would improve reproducibility.","section":"Section 3.2, Eq. (6)"}],"recommendation":"major_revision","confidential_remarks":"The MTT comparison is the decisive issue. If the authors can supply a matched-hyperparameter Gaussian-init MTT baseline and the gains shrink to roughly 0.1-0.2 accuracy points, the claim of consistent MTT improvement should be withdrawn or substantially softened. The DC and DM results are more robust and could support a revised, more focused claim. I do not see circularity or fabrication; the supplementary material is unusually transparent, which is to the authors' credit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this paper proposes Multi-Domain Dataset Condensation (MDDC), asking whether you can condense a heterogeneous dataset into a small set of synthetic images that work across domains without explicit domain labels. The proposed DAM module trains learnable spatial masks that decompose each synthetic image into several pseudo-domains, with pseudo-labels derived from low-frequency FFT amplitude. That's a clean extension of dataset condensation, and the reconstruction identity (sum of masked images equals original) is a nice touch. The experiments are broad: five datasets, three base methods, multiple architectures, plus leave-one-domain-out. For DC and DM, the gains are consistent, around 1-6 points, and the ablation shows FFT pseudo-labels beat random labels and even match actual domain labels. That's real evidence the mechanism does something.\n\nThe problem is with MTT. Section 4.2 says all prior hyperparameters were set identically, but Appendix F says MTT needed a separate hyperparameter search when combined with DAM under Gaussian-noise initialization, and Table H lists different hyperparameters. The paper doesn't show that the vanilla MTT baseline was re-tuned under the same Gaussian-init. The supplementary Table A reveals why this matters: with real-image initialization, vanilla MTT jumps from 50.7 to 65.3 on CIFAR-10 10 IPC, and DAM's gains shrink to +0.2 there and +0.1 on PACS 1 IPC. So the headline MTT numbers (e.g., PACS 39.7→46.6, CIFAR-10 50.7→57.9) likely reflect a default-hyperparameter baseline against a tuned model. That's a load-bearing flaw for the claim that DAM consistently improves MTT. The DC and DM results are far less affected, but the MTT story as told is not supported.\n\nThere are smaller issues too. The domain loss in Eq. 7 is under-specified: the synthetic side carries class labels while the real side is grouped by pseudo-domain labels, and it's unclear what the loss is actually matching. No code is released, and a few hyperparameters (sort order, beta) are not reported. These are fixable in revision.\n\nOverall, the paper is worth engaging. The task is well-motivated, the idea is simple, and the DC/DM evidence suggests it works. But the MTT claims need either a properly re-tuned baseline or a scaled-back claim. I'd recommend sending it to a competent referee, with the request to pin down the MTT setup and the domain loss before any acceptance.","headline":"Introduces a useful new task and a plausible module, but the MTT evidence is confounded by an unfair baseline; the core idea is still worth a serious look.","tokens_in":17280,"tokens_out":2926,"would_cite":false,"duration_ms":28839,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A plug-in domain-aware module improves dataset condensation by embedding spatial domain masks into synthetic images, boosting in-domain, out-of-domain, and cross-architecture accuracy without adding images per class.","keywords":["dataset condensation","multi-domain dataset condensation","domain-aware module","pseudo-domain labeling","frequency-based labeling","spatial masks","domain generalization","cross-architecture generalization"],"falsifier":"Train DAM on a purpose-built multi-domain dataset where the domain shift is deliberately placed in high-frequency content (e.g., texture swaps) so that mean low-frequency amplitude is nearly constant across domains; if DAM does not beat the base condensation method there, the FFT pseudo-labeling assumption is the failing component.","tokens_in":16272,"feed_emoji":"🎯","tokens_out":7855,"duration_ms":64541,"temperature":0.7,"pith_summary":"Dataset condensation squeezes a large training set into a tiny synthetic one, but existing methods assume the source imagery shares a single visual style. This paper argues that real-world datasets are multi-domain, with photos, cartoons, sketches, and paintings mixed together, and that ignoring this causes condensed data to collapse toward dominant styles and lose accuracy. It introduces Multi-Domain Dataset Condensation (MDDC) and the Domain-Aware Module (DAM), a plug-in training-time component that superimposes learnable spatial masks onto each synthetic image so different domains occupy different regions. Because explicit domain labels are usually unavailable, DAM derives pseudo-domain labels from the mean amplitude of low-frequency Fourier components of the real images. The claim is that this consistently improves in-domain, out-of-domain, and cross-architecture accuracy across DC, DM, and MTT baselines without increasing images per class.","feed_headline":"Dataset condensation gains up to 6.9 points with domain masks","feed_subtitle":"Plug-in module keeps images per class while boosting in-domain, out-of-domain, and cross-architecture accuracy.","key_machinery":"The load-bearing object is the spatial domain mask $\\alpha^{d,i}_m \\in \\mathbb{R}^{H \\times W \\times 3}$, computed as a per-pixel temperature softmax over D learnable mask tensors $z^{d,i}_m$. Multiplying the synthetic image $\\tilde{x}^i_m$ element-wise by each mask gives domain-specific views $\\tilde{x}^{d,i}_m$, and the identity $\\tilde{x}^i_m = \\sum_{d=0}^{D-1} \\tilde{x}^{d,i}_m$ holds because the softmax outputs sum to 1, so no information is lost when the image is split into domain views. Domain supervision comes from frequency-based pseudo-domain labels: each real image is assigned one of D bins by ranking its mean low-frequency FFT amplitude, a heuristic borrowed from Fourier-based domain adaptation and generalization. The domain branch of the loss uses the same base condensation loss (gradient matching, distribution matching, or trajectory matching) but grouped by pseudo-domain instead of by class. This machinery lets domain structure be embedded during training while leaving the final synthetic dataset and its IPC unchanged.","core_discovery":"The paper's central claim is that domain diversity can be encoded inside a synthetic dataset without extra images or labels. DAM attaches D learnable spatial masks to each synthetic image, normalizes them with a per-pixel softmax so they sum to one, and multiplies each mask against the image to produce D domain-specific views that reconstruct the original image exactly. A second branch of the condensation loss trains these masks against pseudo-domain labels, obtained by sorting real images by mean low-frequency FFT amplitude and slicing the sorted list into D bins. During condensation the synthetic image is updated by both class and domain losses, but after condensation the masks are discarded, so downstream models see only the original unmodified synthetic images. The reported result is consistent improvement over three prior methods on five datasets, with the largest gains on multi-domain benchmarks such as PACS where 1-IPC MTT accuracy rises from 39.7 to 46.6.","pith_inferences":["If low-frequency amplitude sorting really captures the domain structure that matters, the same pseudo-labeling could be reused beyond condensation, for example to guide data selection or augmentation schedules in other data-efficient learning pipelines.","The exact reconstruction identity means DAM is a form of input-space regularization; one testable extension is whether the learned masks can be reused or transferred to new condensation runs rather than discarded.","Because DAM improves out-of-domain generalization, it could serve as a cheap proxy for domain-generalization benchmarking before collecting more data, though the paper only tests this indirectly with leave-one-domain-out evaluation.","The optimal number of pseudo-domains D appears dataset-dependent, as the CIFAR-10 sweep shows, suggesting that pseudo-domain granularity itself carries task-relevant information that could be tuned automatically."],"forward_implications":["DAM can be plugged into gradient-matching, distribution-matching, and trajectory-matching condensation methods without changing the number of synthetic images or their class budget.","Condensed data becomes more transferable: accuracy on unseen domains improves, as in leave-one-domain-out tests on PACS, VLCS, and Office-Home.","Cross-architecture generalization improves, including to ViT-Tiny and ViT-Small, which the original condensation papers did not evaluate.","Multi-domain condensation no longer requires explicit domain labels or per-domain condensation, which would inflate the synthetic dataset size linearly with the number of domains.","The performance gap between single-domain and multi-domain condensation narrows, with some target domains matching single-domain performance."],"supporting_citations":[{"why":"Defines the dataset distillation task that this work extends.","marker":"[1]"},{"why":"Supplies the gradient-matching baseline DC that DAM plugs into.","marker":"[2]"},{"why":"Supplies the distribution-matching baseline DM that DAM plugs into.","marker":"[3]"},{"why":"Supplies the trajectory-matching baseline MTT that DAM plugs into.","marker":"[4]"},{"why":"Provides the Fourier amplitude heuristic that pseudo-domain labeling relies on.","marker":"[7]"},{"why":"Supports using low-frequency amplitude statistics as domain-specific features.","marker":"[8]"},{"why":"Provides the PACS multi-domain benchmark used in main and ablation experiments.","marker":"[6]"},{"why":"Provides the VLCS multi-domain benchmark used in evaluation.","marker":"[26]"},{"why":"Provides the Office-Home multi-domain benchmark used in evaluation.","marker":"[27]"}],"fun_headline_variants":["Domain-aware module lifts multi-domain dataset condensation","Condensed data gets multi-domain gains via learnable masks","Dataset condensation improves across domains with DAM","Domain-aware module adds up to 6.9 points to condensed data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a single low-frequency image statistic can split a mixed dataset into pseudo-domains that capture the visual variation that matters for classification.","fun_headline_variants_meta":{"raw":{"variants":["Domain-aware module lifts multi-domain dataset condensation","Condensed data gets multi-domain gains via learnable masks","Dataset condensation improves across domains with DAM","Domain-aware module adds up to 6.9 points to condensed data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001055,"raw_usage":{"total_tokens":4406,"prompt_tokens":898,"completion_tokens":3508,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":3445}},"tokens_in":514,"tokens_out":3508,"duration_ms":28066,"temperature":1.0,"reasoning_tokens":3445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:08:22.526815+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train DAM on a purpose-built multi-domain dataset where the domain shift is deliberately placed in high-frequency content (e.g., texture swaps) so that mean low-frequency amplitude is nearly constant across domains; if DAM does not beat the base condensation method there, the FFT pseudo-labeling assumption is the failing component.","supporting_citations":[{"cited_title":"Fda: Fourier domain adaptation for semantic segmentation,","cited_arxiv_id":null,"evidence_quote":"Provides the Fourier amplitude heuristic that pseudo-domain labeling relies on."},{"cited_title":"A fourier-based framework for domain generalization,","cited_arxiv_id":null,"evidence_quote":"Supports using low-frequency amplitude statistics as domain-specific features."},{"cited_title":"Deeper, broader and artier domain generalization,","cited_arxiv_id":null,"evidence_quote":"Provides the PACS multi-domain benchmark used in main and ablation experiments."},{"cited_title":"Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias,","cited_arxiv_id":null,"evidence_quote":"Provides the VLCS multi-domain benchmark used in evaluation."},{"cited_title":"Deep hashing network for unsu- pervised domain adaptation,","cited_arxiv_id":null,"evidence_quote":"Provides the Office-Home multi-domain benchmark used in evaluation."},{"cited_title":"Dataset condensation with distribution matching,","cited_arxiv_id":null,"evidence_quote":"Supplies the distribution-matching baseline DM that DAM plugs into."},{"cited_title":"Dataset distillation by matching training trajectories,","cited_arxiv_id":null,"evidence_quote":"Supplies the trajectory-matching baseline MTT that DAM plugs into."},{"cited_title":"Dataset condensation with gradient matching,","cited_arxiv_id":null,"evidence_quote":"Supplies the gradient-matching baseline DC that DAM plugs into."}],"review_version":1}