{"id":"87d2cbaf-f776-4d1d-a49a-8e7a2c1dc017","arxiv_id":"2506.09275","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.","lead":"A survey of simulators for training large AI models across many computers, grouped into three categories and compared with cost and carbon models. It helps engineers choose tools and highlights open gaps in validation, energy modeling, and integration.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table VI's Error column compares non-comparable validation statistics, undermining the survey's tool-selection support.","rationale":"The reader's weakest assumption correctly pinpoints the Error column. My read of the paper confirms that this column is presented without a header definition, and the footnotes explicitly mix maximum error, average error, and a correlation-derived score. Because the paper's stated purpose is to support tool selection, and Table VI is the centerpiece of the comparison, this is the most load-bearing concern. It is independently corroborated by internal inconsistencies elsewhere in the paper, such as Section III.C describing Calculon as IR-based while Table VI lists it as config-based, and Takeaway III.8 crediting vTrain with indirect TCO modeling while Table VI marks no TCO aspects; these further reduce confidence in the table's reliability. These issues are fixable, so a conditional verdict is appropriate. I do not see a reason to move to accept or reject; the survey's synthesis and taxonomy remain useful once the comparison columns are made commensurable and the row-level misstatements are corrected.","tokens_in":27147,"tokens_out":4778,"duration_ms":53317,"concrete_test":"Re-annotate every row of Table VI with the exact error definition and validation setup taken from the cited source paper, including metric type, models, hardware, and scale. If any two rows use different metric types or different validation workloads, add a per-row 'Error metric / validation' column and remove the raw numeric column, or explicitly refuse cross-simulator error comparisons. A minimal check: compare DistIR's 0.071 (1−Pearson) with FlexFlow's 0.3 (maximum error) after converting both to the same definition; if the rank order changes, the current column is misleading.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central contribution is a structured comparison that enables informed simulator selection (abstract and Section III.C). Table VI is the main quantitative support for that claim, but its Error column is not a single commensurable quantity. The table's own footnotes admit this: DistIR's 0.071 is 'computed as (1−Pearson correlation); not a direct error metric,' while entries marked with the double-dagger footnote are maximum errors (e.g., FlexFlow, AMPeD, ATLAHS, Multiverse), and the unmarked entries appear to be average errors. These statistics are not interchangeable: a maximum error is always at least as large as an average error, and a correlation-derived score is not even in the same units. The column therefore cannot be used to compare simulators or to support statements such as 'Calculon has error 0.0365' versus 'DistSim has error 0.04.' Additionally, because no column records the validation workload or scale, even two entries using the same error metric are validated on different models and hardware, so the numbers are not directly comparable without additional context. Since the abstract promises 'comprehensive comparison tables' that 'support informed decision-making,' this is a load-bearing flaw, not a cosmetic one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews distributed DNN training simulators and TCO/emissions models. It proposes a taxonomy of analytical, profiling-based, and execution-driven simulation, cross-cut by workload representation granularity (configuration-based, operator/layer-level IR, machine-level). Section II reviews workload representations, Section III surveys simulators and presents a detailed comparison in Table VI, and Section IV reviews TCO/emissions models with a comparison in Table VII. The paper distills the results into a series of takeaways and claims that the structured comparison supports informed tool selection and identifies research gaps.","tokens_in":27360,"tokens_out":7391,"duration_ms":70131,"significance":"If the comparison tables and taxonomy are accurate, the survey would be a useful reference for researchers in ML systems, computer architecture, and sustainable computing. Its strengths are broad coverage of very recent 2023–2025 systems, a clearly presented taxonomy with a workload-fidelity dimension, explicit side-by-side comparison of simulator attributes, and the attempt to connect distributed training simulation with TCO/emissions modeling. The paper does not ship machine-checked proofs or reproducible code, so its value rests on the correctness and interpretability of its comparative tables. The inconsistencies identified below affect exactly that basis, so the contribution is currently only partially reliable.","major_comments":[{"comment":"The Error column of Table VI mixes validation statistics that are not commensurable. The footnotes disclose that DistIR's 0.071 is computed as (1 − Pearson correlation) and is not a direct error metric, while entries marked with the double-dagger footnote (FlexFlow, AMPeD, ATLAHS, Multiverse) are maximum errors and the remaining entries appear to be average errors. A maximum error is always at least as large as an average error, and a correlation-derived score is in different units, so entries such as Calculon's 0.0365 and DistSim's 0.04 cannot be meaningfully compared across rows. Additionally, the table does not record the validation workload or hardware scale per row, so even rows using the same metric are not readily comparable. Since the abstract promises 'comprehensive comparison tables' that 'support informed decision-making,' this column undercuts the table's central purpose. Please split the Error column by metric type, add validation conditions, or remove the column and discuss accuracy narratively.","section":"Table VI and its footnotes"},{"comment":"Section III-C, first bullet, groups Calculon with DistIR and Deepflow as simulators that 'employ IR-based inputs,' but Table VI lists Calculon's Input Format as 'cfg-based,' and Figure 5 places Calculon under configuration-based workload fidelity. This is a direct text-table contradiction on a taxonomy assignment, which is one of the paper's main contributions. The authors should determine which classification is correct and align Section III-C, Figure 5, and Table VI.","section":"Section III-C vs. Table VI"},{"comment":"Takeaway III.4 states that SimAI and MultiVerse 'have validated their results on 1024-node A100 GPU clusters,' but Table VI lists MultiVerse's Validation as '1024xH100' (not A100) and SimAI's Validation as '1024xA100, 1024xH100'; the notation appears to denote GPU counts rather than node counts. This takeaway is used to single out the strongest large-scale validation in the field, so the discrepancy and the GPU/node distinction should be corrected.","section":"Takeaway III.4 vs. Table VI"}],"minor_comments":[{"comment":"Table II lists 'XLA HLO' with reference [21], but Figure 3 attributes XLA HLO to reference [93], and reference [21] is StableHLO; the reference should be corrected.","section":"Table II and Figure 3"},{"comment":"The phrase 'profiling-based simulators 5' contains a stray superscript '5' with no corresponding footnote.","section":"Section II.B.1"},{"comment":"The text 'Y AML' appears twice in Figure 3 and once in Section II.B.1; it should read 'YAML.'","section":"Figure 3 and Section II.B.1"},{"comment":"Tool names are used inconsistently: 'LLM Compass' in Table VI versus 'LLMCompass' in the text, and 'vtrain' in Table VI versus 'vTrain' in the text.","section":"Table VI and text"},{"comment":"The first bullet of Section III-C reads 'Profiling-based simulators, a balance between fidelity and speed,' which is missing a verb; it should read 'strike a balance' or similar.","section":"Section III-C"},{"comment":"Table V's Workload row for Execution-driven lists Machine-level IR, but neither execution-driven simulator in Table VI (LLMCompass, LLMServingSim) uses a machine-level workload format; if Table V is an idealization rather than a property of the surveyed tools, the caption should say so.","section":"Table V"}],"recommendation":"major_revision","confidential_remarks":"The reviewed set includes works by the authors themselves, including AMPeD [87] and imec.netzero [56], as well as other imec references [10], [12]. The survey does not disclose this self-citation. I do not see the survey's conclusions depending on favorable treatment of these works, but a conflict-of-interest note would be appropriate. The main technical issue is the non-comparable Error column in Table VI and the taxonomy contradiction around Calculon; these are fixable but currently undermine the paper's central artifact."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for the report. I read the paper and largely share your verdict, but I'd put the error-column issue a bit lower: it's a real flaw, not fatal, and it's fixable.\n\nWhat's actually new here: it's a systematic map of distributed DNN training simulators plus workload representations plus TCO/emissions models in one place. The Guo et al. survey covered analytical models and a limited set of IR simulators; this one widens the net, adds the TCO dimension, and organizes everything into a three-way taxonomy (analytical/profiling/execution). The takeaways in Sections II-D and III-D are sensible: LLM focus, validation weakness, NVIDIA dominance, missing energy modeling, local/remote memory gaps, and the isolation of TCO models from simulators. The gap analysis is the strongest part.\n\nThe tables are the main artifact, and that's where the trouble is. Your weakest assumption is correct: Table VI's Error column mixes average error, max error, and a 1-Pearson correlation. The footnotes disclose this, but a reader scanning the table will see numbers like 0.0365 versus 0.04 and think they're comparable. They're not. That undercuts the 'informed decision-making' promise, but it's not a deep conceptual flaw; the authors should split the column, mark metrics explicitly, or drop the column and keep validation discussion in text.\n\nThere are also a few internal contradictions that any careful reader will trip on. Section III-C classifies Calculon as IR-based; Table VI lists it as cfg-based. Takeaway III.4 says SimAI and MultiVerse validated on 1024-node A100 clusters; Table VI shows Multiverse on 1024xH100 and SimAI on both A100 and H100. vTrain is listed with no TCO aspects in Table VI, but Takeaway III.8 says it accounts for operating cost via AWS pricing. And references [26] and [27] are identical. These are all minor fixes, but they matter because a survey is only as good as its tables.\n\nThe self-citation point: authors' own tools (AMPeD, STCO, imec.netzero) appear, but they're cited in context and the synthesis doesn't hinge on them. Not a problem.\n\nBottom line: this is a useful secondary source for anyone choosing a simulator or looking for research gaps. It doesn't open a new line of work, but it's a competent synthesis. With a careful revision of Table VI and the flagged contradictions, it deserves publication. I'd send it to peer review rather than desk reject, with instructions to the referees to check table accuracy against cited sources.","headline":"Useful survey with a solid taxonomy and TCO coverage, but Table VI's incomparable error metrics and a few internal contradictions keep it from being a reliable reference until revised.","tokens_in":27905,"tokens_out":2961,"would_cite":false,"duration_ms":30177,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that the scattered field of distributed DNN training simulators can be mapped by a fidelity-based taxonomy, and that the map exposes consistent gaps in validation, energy modeling, and cost integration.","keywords":["distributed DNN training","simulators","workload representation","performance modeling","total cost of ownership","carbon emissions","taxonomy","intermediate representation"],"falsifier":"A reader can settle the central comparison claim by returning to the papers behind Table VI's Error entries: if the values are indeed a mix of average error, maximum error, and one minus Pearson correlation, then the column cannot rank simulators, and the survey's stated goal of informed tool selection would need a standardized error metric instead.","tokens_in":26953,"feed_emoji":"⚙️","tokens_out":6465,"duration_ms":65906,"temperature":0.7,"pith_summary":"This survey tries to establish that distributed DNN training simulators can be usefully organized along two axes: simulation fidelity (analytical, profiling-based, or execution-based) and workload representation (configuration-based versus operator- or layer-level intermediate representations). It argues that this organization supports informed tool selection and makes visible where the field is underdeveloped, especially validation at scale, operating-energy modeling, and integration of total cost of ownership (TCO) and carbon-emissions models. The practical stakes are real: full-scale training systems are too expensive to prototype, so simulation is the main lever for early design exploration, and a reliable map of simulator capabilities changes which tools designers trust and where they invest. The paper's contribution is the structured comparison itself, not a new simulator.","feed_headline":"Survey: DNN training simulators skip energy and scale validation","feed_subtitle":"A fidelity-based taxonomy of 20+ tools shows profiling-based simulation rising while cost models stay separate.","key_machinery":"The central object is the survey's taxonomy in Figure 5 and its companion tables (Table VI for simulators, Table VII for TCO/emissions models). The taxonomy classifies simulators on a fidelity axis as analytical, profiling-based, or execution-based, and crosses that with a workload-granularity axis (configuration-based, operator-level IR, layer-level IR); the tables then carry the argument by listing each tool's input format, error, target hardware, network model, and scalability, and each cost model's coverage of fabrication, technology, network, storage, node architecture, and workload.","core_discovery":"On the survey's own terms, the discovery is that the field has converged on a common architecture—a workload graph plus separate compute, network, and scheduler models—while diverging sharply in fidelity and abstraction, and that no surveyed simulator covers the whole stack. The comparison tables show profiling-based simulation emerging as the balance point between speed and accuracy, operator-level IRs displacing configuration-based inputs, NVIDIA GPUs dominating target hardware, network modeling remaining mostly analytical with only a few congestion-aware exceptions, and energy and TCO modeling almost entirely absent from simulators while TCO/emissions models stay decoupled in a separate literature. Synthesizing those patterns into one taxonomy is what the survey claims to add.","pith_inferences":["A natural next step the authors leave implicit is a shared benchmark suite with one workload, one hardware target, and one error definition; without it, the Error column cannot support quantitative ranking.","Because the survey shows profiling-based simulators already collect execution traces, pairing those traces with component power estimators (which the survey cites as available building blocks) is the shortest path to closing the operating-energy gap.","The finding that simulators are NVIDIA-centric implies a high-value extension would be porting one profiling-based simulator to a non-NVIDIA backend; the taxonomy predicts the trace-collection layer, not the core scheduler, is where most of the work sits.","If the taxonomy is correct, the field may converge on hybrid simulators that combine analytical compute models with congestion-aware network simulators, since that combination is what the profiling-based category already points toward."],"forward_implications":["Tool selection can be structured by fidelity class: analytical simulators for fast trend exploration, profiling-based for balanced studies, and execution-based for detailed bottleneck analysis.","The shift to operator-level IRs means new simulators should expect to consume or produce graphs like Chakra rather than configuration files.","Profiling-based simulators are the natural place to add operating-energy estimation, because their traces already capture compute and communication activity.","Validation at scale is the field's main bottleneck, so any simulator claiming large-scale accuracy should be tested against published cluster results.","TCO and emissions models should be integrated with simulators; none currently model environmental impact, so a coupled tool would open a new design dimension."],"supporting_citations":[{"why":"Prior focused survey of distributed DNN training simulators that this survey positions itself against and complements.","marker":"[45]"},{"why":"ASTRA-sim appears throughout as the modular reference simulator and the basis for the profiling-based category and follow-up extensions.","marker":"[128]"},{"why":"Chakra is the framework-agnostic IR whose adoption beyond its origin supports the IR discussion in Table II.","marker":"[115]"},{"why":"DistIR supplies both an IR and an analytical simulator example used in the taxonomy and Table VI.","marker":"[106]"},{"why":"Megatron-LM is the source of configuration-based workload representations and tensor-parallel communication patterns.","marker":"[110]"},{"why":"MLIR is the compiler infrastructure whose dialects anchor the framework-agnostic IR discussion.","marker":"[76]"},{"why":"ACT is the architectural carbon modeling tool that grounds the emissions-model comparison in Table VII.","marker":"[46]"},{"why":"Google TPU life-cycle assessment supplies the CCI metric and geographic variability findings used in the emissions section.","marker":"[107]"},{"why":"Paleo is the early analytical simulator example whose layer-level IR and error entry shape Table VI.","marker":"[97]"},{"why":"DistSim is a profiling-based simulator example used for network modeling via NCCL benchmark data.","marker":"[82]"}],"fun_headline_variants":["No DNN training simulator models the full stack yet","Energy and cost models remain separate from DNN sims","Profiling-based simulators become middle ground in DNN training","Survey: operator-level IRs replace config inputs in DNN sims","NVIDIA dominates DNN training simulator targets in survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Error column in the main simulator comparison is a single comparable quantity, but its footnotes show the column mixes average errors, maximum errors, and a correlation-based value, so the tables can only guide tool selection if those discrepancies are ignored.","fun_headline_variants_meta":{"raw":{"variants":["No DNN training simulator models the full stack yet","Energy and cost models remain separate from DNN sims","Profiling-based simulators become middle ground in DNN training","Survey: operator-level IRs replace config inputs in DNN sims","NVIDIA dominates DNN training simulator targets in survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000476,"raw_usage":{"total_tokens":2334,"prompt_tokens":889,"completion_tokens":1445,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":1362}},"tokens_in":505,"tokens_out":1445,"duration_ms":11005,"temperature":1.0,"reasoning_tokens":1362,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:53:11.976794+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader can settle the central comparison claim by returning to the papers behind Table VI's Error entries: if the values are indeed a mix of average error, maximum error, and one minus Pearson correlation, then the column cannot rank simulators, and the survey's stated goal of informed tool selection would need a standardized error metric instead.","supporting_citations":[{"cited_title":"Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model train- ing at scale","cited_arxiv_id":null,"evidence_quote":"ASTRA-sim appears throughout as the modular reference simulator and the basis for the profiling-based category and follow-up extensions."},{"cited_title":"Distir: An intermediate representation for optimizing distributed neural networks","cited_arxiv_id":null,"evidence_quote":"DistIR supplies both an IR and an analytical simulator example used in the taxonomy and Table VI."},{"cited_title":"Mlir: Scaling compiler infrastructure for domain specific computation","cited_arxiv_id":null,"evidence_quote":"MLIR is the compiler infrastructure whose dialects anchor the framework-agnostic IR discussion."},{"cited_title":"Paleo: A performance model for deep neural networks","cited_arxiv_id":null,"evidence_quote":"Paleo is the early analytical simulator example whose layer-level IR and error entry shape Table VI."},{"cited_title":"Distsim: A performance model of large-scale hybrid distributed dnn training","cited_arxiv_id":null,"evidence_quote":"DistSim is a profiling-based simulator example used for network modeling via NCCL benchmark data."}],"review_version":1}