{"id":"0b371d41-d5b5-4771-93e2-eb627d76cfc4","arxiv_id":"2608.07713","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"On 64x64 ChestMNIST, tokenizer quality for generation is not separable from the generator and sampler, and a new token-predictability statistic predicts which tokenizers will generate well.","lead":"This paper tests, on low-resolution chest X-rays, whether image tokenizers and generators can be chosen independently, and finds they cannot: the best quantizer changes with the generator, and sampling settings change the apparent ranking. It also introduces a token-only statistic that predicts which tokenizers will generate well, where reconstruction quality does not.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The seeded 'FSQ beats LFQ under MaskGIT' flip reverses under standard FID-2048, so the interaction's specific direction is metric-dependent.","rationale":"The paper is carefully designed, honestly scoped, and supported by released code and a three-seed replication, so outright rejection is not warranted. However, the strongest claim as stated by the reader includes a specific flip—FSQ best under MaskGIT—that reverses under standard FID-2048 in the paper's own Table 17. This is more pointed than the reader's general concern about FID-192 fidelity: the exact contrast carrying the interaction is the one on which the headline metric and the standard metric disagree. The broad interaction claim (no single quantizer is best across generators) is metric-robust, so the paper's main thesis survives, but the concrete direction of the interaction is not established. The paper should therefore report the seeded block under a standard or domain metric, or explicitly demote the specific flip to an observation under FID-192. This supports a conditional verdict: accept with revisions that align the strongest claim with the metric-robust evidence.","tokens_in":31500,"tokens_out":7354,"duration_ms":69798,"concrete_test":"Compute domain-FID (ChestMNIST ResNet-18 penultimate features) and standard FID-2048 for all ten cells of the three-seed vocabulary-1024 block, reporting the per-generator best quantizer under each metric across seeds. If the MaskGIT winner under FID-192 (FSQ) is not the winner under FID-2048 or domain-FID, restrict the central claim to the metric-robust statement that no quantizer is best across all generators, and remove the specific FSQ-over-LFQ-under-MaskGIT flip from the abstract and Section 4.1.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claimed finding—'best quantizer flips from LFQ under AR and D3PM to FSQ under MaskGIT' (Section 4.1, Table 9)—is not stable across the two FID variants the paper itself reports. Table 17 shows the MaskGIT row on identical generated samples: FID-192 ranks FSQ (1.15) better than LFQ (1.91), while standard FID-2048 ranks LFQ (78) far better than FSQ (195). The authors acknowledge that this MaskGIT ordering is metric-dependent and say they do not claim it, but the abstract and the seeded-block summary assert exactly that flip. A reader cannot simultaneously use FID-192 as the headline metric and disclaim the one ordering that carries the headline interaction. The broader non-separability conclusion survives under FID-2048, but the precise interaction pattern—which quantizer wins under which generator—is load-bearing and is not supported once the metric is changed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a controlled factorial study on ChestMNIST at 64x64, crossing three discrete quantizer families (VQ, LFQ, FSQ) at three vocabulary sizes, six discrete generator families (AR, MaskGIT, DFM, D3PM, SE-D3PM, BFN), and continuous VAE/AE reference cells with LDM/RF, all under a shared latent grid. The central claims are that tokenizer, generator, and sampler rankings interact, that the best quantizer changes with the generator, that reconstruction PSNR is a poor predictor of generation quality, and that a new generator-free statistic (neighbour-conditional predictive gain) separates quantizer families by downstream generation quality. The interaction is verified with a three-seed block at vocabulary 1024, and the wider grid is scoped as single-seed. The paper also reports validation-selected sampling hyperparameter sweeps that substantially improve D3PM and SE-D3PM on LFQ-1024, and it releases two software libraries for reproducing the comparisons.","tokens_in":31747,"tokens_out":9380,"duration_ms":87482,"significance":"If the main conclusions hold, the paper makes a useful methodological point: tokenizer selection for latent medical-image generation should not be based on reconstruction alone, and the tokenizer-generator-sampler triple is the appropriate experimental unit even in a low-resolution controlled setting. The strengths are the matched-vocabulary factorial design, the three-seed variance block, the extensive metric cross-checks, the explicit scoping of single-seed results, and the release of reproducible libraries and sweep scripts. The significance is tempered, however, by the metric dependence of the paper's flagship interaction example, which the authors partly acknowledge but do not fully reconcile with the main-text claims.","major_comments":[{"comment":"The seed-verified illustration of the interaction, stated as 'the best quantizer changes with the generator (LFQ lowest under AR and D3PM; FSQ lowest under MaskGIT, where LFQ is worst)', is contradicted by Table 17 in the same paper. On identical generated samples, the MaskGIT row ranks FSQ best under FID-192 (1.15 vs. LFQ 1.91) but LFQ best under FID-2048 (78 vs. FSQ 195). Section 6.3(iii) explicitly lists this as a metric-dependent ordering that is 'not claimed under either metric', yet Section 4.1 and the abstract's generic 'best quantizer changes' claim rely on it as the seeded interaction evidence. The non-separability conclusion itself does survive under FID-2048 (best quantizer is FSQ for AR and LFQ for MaskGIT and D3PM), so the fix is achievable: state the interaction in a metric-robust form, remove the FSQ-under-MaskGIT assertion from the seeded-block summary, or provide a pre-specified aggregation across metrics.","section":"Section 4.1 and Table 9 vs. Table 17 and Section 6.3(iii)"},{"comment":"The paper uses Spearman rho = 0.80 between FID-192 and FID-2048 as the license for treating FID-192 as a reliable internal ranking metric. A rank correlation of 0.80 over 12 cells leaves room for pairwise reversals, and Table 17 shows a reversal in the very MaskGIT cell that carries the seeded interaction example. The main text should either report FID-2048 counterparts for the seeded-block and HP-sweep headline cells, or restrict all pairwise ordering claims to those that are robust across both metrics. At minimum, Section 4.1 should not present pairwise quantizer orderings from FID-192 without explicitly flagging the metric dependence shown in Table 17.","section":"Section 3.3 and Appendix D"},{"comment":"The rank-AUC 1.00 for neighbour-conditional predictive gain is based on only n = 9 tokenizers, and the statistic was selected after inspecting several candidate predictors with a validation-selected smoothing constant. The main text appropriately calls the result 'directional' with 'no p-values', but the abstract and Section 1 state it unqualified as 'separates the quantizer families'. Please add a permutation-based confidence interval or otherwise quantify the uncertainty, and move the directional caveat to the abstract. As written, the abstract overstates the evidence for this contribution.","section":"Section 4.3, Table 10"}],"minor_comments":[{"comment":"The text says 'seed noise measured above (~0.024)', but Table 9 reports a median per-cell standard deviation of 0.047. Please correct the numerical inconsistency.","section":"Section 4.3"},{"comment":"The symbol N is used for inference step count in Table 4, while later text and Table 15 use k for top-k truncation. A short notation list near Table 4 would prevent confusion between step counts and top-k values.","section":"Section 3.2 and Table 4"},{"comment":"The PneumoniaMNIST column compares 10,000 generated samples against 624 real test images. The real-vs-real floor is reported, but a sentence noting that FID gaps below roughly 0.05 are not resolvable in that column would help readers interpret the 2.73 vs. 1.00 difference.","section":"Table 12"}],"recommendation":"major_revision","confidential_remarks":"This is a serious empirical study with unusually honest scoping and reusable code. The main risk is the internal contradiction between Section 4.1 and Table 17: the flagship seeded interaction example is metric-dependent in the paper's own data. If the authors revise the interaction claim to a metric-robust form and temper the predictive-gain abstract claim, this could be suitable for publication. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious, well-scoped controlled study of tokenizer-generator-sampler interactions for low-resolution medical image generation. The core empirical claim holds: the best quantizer depends on the generator, so reconstruction-only selection is unreliable. The code release is strong and the limitations section is honestly written.\n\nWhat's new: a matched VQ/LFQ/FSQ vocabulary grid crossed with six discrete generator families plus continuous references, all on ChestMNIST-64. The seed-verified vocabulary-1024 interaction block is the right way to put the interaction on inferential footing, and the domain-FID and classifier-two-sample cross-checks give the FID-192 rankings some independent support.\n\nSoft spots, in proportion. The stress-test note lands: the specific 'FSQ beats LFQ under MaskGIT' flip reverses under FID-2048 in Table 17. The authors acknowledge this in the limitations and say they do not claim that ordering, but Section 4.1 and Table 9 assert it. That tension should be fixed, either by reporting both FID variants in the main interaction table or by explicitly scoping the directional claim to FID-192. The broad non-separability conclusion survives under both metrics, so this is a presentation flaw rather than a refutation, but it is load-bearing for the headline.\n\nThe neighbour-conditional predictive gain (rank-AUC 1.00) is computed on n=9 with no p-values; the paper later scopes it as directional, but the abstract overstates it by omitting those caveats. The tuned discrete-versus-continuous comparison is design-unequal because the continuous references got no sampler sweep; the authors disclose this. The single-seed wider grid and the unexplained VAE-c8 anomaly are also disclosed.\n\nWho this is for: anyone working on latent generation for low-res medical images, and anyone doing tokenizer benchmarking in constrained domains. It deserves a serious referee. I'd send it out, asking for a fix to the metric-dependence contradiction and a more careful abstract.","headline":"A careful, well-scoped empirical study of tokenizer-generator-sampler interactions in low-res medical image generation; the core non-separability claim holds, but one headline flip is metric-dependent and the abstract oversells a small-n predictor.","tokens_in":32306,"tokens_out":3791,"would_cite":true,"duration_ms":31328,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The best image quantizer depends on the generator, not reconstruction quality alone.","keywords":["tokenizer-generator interaction","medical image generation","vector quantization","lookup-free quantization","finite scalar quantization","discrete diffusion","FID-192","rate-distortion-modelability"],"falsifier":"Retrain the vocabulary-1024 interaction block at additional seeds and evaluate with a radiomics-based metric or at higher resolution: if the best quantizer no longer flips with generator (LFQ lowest under AR and D3PM, FSQ lowest under MaskGIT) or all pairwise gaps fall within seed noise, the central interaction claim collapses.","tokens_in":31246,"feed_emoji":"🩻","tokens_out":3273,"duration_ms":30804,"temperature":0.7,"pith_summary":"This paper tries to establish that tokenizer selection for medical image generation cannot be done by reconstruction quality alone, because the best quantizer changes with the generator and the sampling configuration. In a controlled ChestMNIST-64 study crossing three discrete quantizers, three vocabulary sizes, six discrete generators, and two continuous generators, the authors find that rankings depend jointly on the tokenizer, generator, and sampler. The interaction is placed on inferential footing by retraining the vocabulary-1024 block at three seeds: six of nine pairwise quantizer comparisons exceed three pooled seed standard deviations, with the best quantizer flipping from LFQ under autoregressive and D3PM generators to FSQ under MaskGIT. The paper also introduces a generator-free statistic, neighbour-conditional predictive gain, that separates quantizer families by downstream generation quality (rank-AUC 1.00), where reconstruction PSNR and marginal token entropy do not. If correct, the result invalidates reconstruction-only tokenizer selection and makes the tokenizer-generator-sampler triple the proper experimental unit.","feed_headline":"Best quantizer flips with generator in medical image generation","feed_subtitle":"Controlled ChestMNIST study shows reconstruction alone cannot choose a tokenizer; the sampling step matters too.","key_machinery":"The load-bearing object is the controlled factorial design: a shared 8x8 latent grid on which VQ, LFQ, and FSQ tokenizers at matched vocabularies are crossed with six discrete generators (AR, MaskGIT, DFM, D3PM, SE-D3PM, BFN) and two continuous generators (LDM, RF), evaluated with FID-192. The inferential core is the three-seed retraining of the vocabulary-1024 interaction block, whose per-cell standard deviations bound which rank orderings are real. The paper's new positive statistic, neighbour-conditional predictive gain, measures how much a token's left or upper neighbour reduces its entropy using bigram counts fit on training data and scored as held-out cross-entropy, thereby rewarding generalisable spatial structure rather than overfitted tables.","core_discovery":"The central claim is that reconstruction quality is not a reliable criterion for choosing a tokenizer in low-resolution medical image generation, and that the best quantizer and the best sampler are conditional on the generator. In the seed-verified vocabulary-1024 block, LFQ is lowest-FID under AR and D3PM while FSQ is lowest under MaskGIT, and no single quantizer can be ranked independently of the generator; six of nine pairwise comparisons exceed three pooled seed standard deviations. Retuning D3PM and SE-D3PM on LFQ-1024 through a held-out validation sweep moves them from default FID-192 of 0.44 and 0.41 to 0.09 and 0.10 at fewer sampling steps, replicated across seeds. The paper further introduces neighbour-conditional predictive gain, a generator-free statistic computed directly from tokenized data, which separates the three VQ tokenizers from the six LFQ/FSQ tokenizers with rank-AUC 1.00, whereas reconstruction PSNR (0.61) and marginal token entropy (0.50, chance) do not.","pith_inferences":["The quantizer-generator interaction likely generalises beyond ChestMNIST, but its magnitude may shrink or grow with resolution and dataset scale; the paper's own cross-dataset check already shows the winning tokenizer is dataset-dependent.","Neighbour-conditional predictive gain could be validated as a pre-training screening tool: compute it on tokenized data before training any generator, and use it to shortlist quantizers for the full factorial sweep.","The rate-distortion-modelability framing suggests that tokenizer benchmarks for generative modelling should report joint sweeps over at least one generator and sampler, since a single-generator FID or reconstruction metric cannot order quantizers consistently.","A natural next test is whether the same interaction holds for class-conditional generation; the paper is unconditional only, and conditioning may change which token distributions generators can exploit."],"forward_implications":["Tokenizers for medical image generation should be selected jointly with the generator and sampler, not by reconstruction PSNR or codebase defaults.","Default 1,000-step sampling budgets for D3PM/SE-D3PM can mis-rank iterative discrete generators; tuning to 100 to 500 steps on codebook-free tokens yields 4 to 10 times better FID in this setting.","Neighbour-conditional predictive gain offers a cheap generator-free screening statistic: more spatially predictable token fields (VQ) tend to generate worse, while near-independent high-entropy fields (LFQ) are easier for generators to model.","The best tokenizer is dataset-dependent, as shown by VQ-1024 overtaking LFQ-1024 on PneumoniaMNIST but not on ChestMNIST or OrganAMNIST, so per-dataset matched sweeps are needed.","Tuned iterative generators close much of the gap to the autoregressive transformer: on LFQ-1024, tuned D3PM reaches FID 0.09 versus AR's 0.33, at 100 function evaluations instead of 64 autoregressive steps."],"supporting_citations":[{"why":"Supplies the VQ tokenizer method and straight-through/EMA codebook training that the paper uses as one of its three discrete quantizer families.","marker":"van den Oord et al., 2017"},{"why":"Supplies finite scalar quantization (FSQ), the codebook-free fixed-grid quantizer contrasted with VQ and LFQ.","marker":"Mentzer et al., 2024"},{"why":"Supplies lookup-free quantization (LFQ) and the prior result that upgrading only the tokenizer can improve generation with the same generator, which the paper tests and extends.","marker":"Yu et al., 2024"},{"why":"Supplies the D3PM discrete diffusion framework whose 1,000-step default the paper retunes and whose training the paper implements matrix-free.","marker":"Austin et al., 2021"},{"why":"Supplies MaskGIT, the masked generative transformer under which the best quantizer flips to FSQ in the seed-verified block.","marker":"Chang et al., 2022"},{"why":"Supplies the score-entropy objective used to define SE-D3PM, which the paper compares against D3PM under matched backbone and sampler.","marker":"Lou et al., 2024"},{"why":"Supplies the latent diffusion model (LDM) reference cell and the VAE-style continuous tokenizer setup used in the channel and KL sweeps.","marker":"Rombach et al., 2022"},{"why":"Supplies rectified flow (RF), the continuous generator that pairs with LDM on identical latents and achieves the lowest continuous FID in the study.","marker":"Liu et al., 2023"},{"why":"Defines the FID metric family; the paper's FID-192 variant and its bootstrap noise floor build on this baseline.","marker":"Heusel et al., 2017"},{"why":"Provides the rate-distortion-modelability framing the paper uses to interpret why reconstruction quality does not predict generation quality.","marker":"Dieleman, 2025"}],"fun_headline_variants":["Reconstruction PSNR can't pick tokenizers for medical image gen","Best quantizer depends on generator in medical image gen","Sampler tuning flips tokenizer ranking in medical image gen","For medical image gen, tokenizer choice hinges on generator","Tokenizers: reconstruction quality isn't enough for medical image gen"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that FID-192, computed from InceptionV3's second max-pooling layer, ranks generation quality correctly at 64x64 resolution; it cross-checks against FID-2048, a classifier two-sample test, and a domain-trained ResNet-18 feature space, but those checks cover only a handful of cells and do not validate every fine-grained ordering that the interaction claims rest on.","fun_headline_variants_meta":{"raw":{"variants":["Reconstruction PSNR can't pick tokenizers for medical image gen","Best quantizer depends on generator in medical image gen","Sampler tuning flips tokenizer ranking in medical image gen","For medical image gen, tokenizer choice hinges on generator","Tokenizers: reconstruction quality isn't enough for medical image gen"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000461,"raw_usage":{"total_tokens":2403,"prompt_tokens":1140,"completion_tokens":1263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":756,"completion_tokens_details":{"reasoning_tokens":1178}},"tokens_in":756,"tokens_out":1263,"duration_ms":8022,"temperature":1.0,"reasoning_tokens":1178,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:22:59.224121+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the vocabulary-1024 interaction block at additional seeds and evaluate with a radiomics-based metric or at higher resolution: if the best quantizer no longer flips with generator (LFQ lowest under AR and D3PM, FSQ lowest under MaskGIT) or all pairwise gaps fall within seed noise, the central interaction claim collapses.","supporting_citations":[],"review_version":1}