{"id":"e148f5bf-4678-48d7-9a7f-d93e7beca81f","arxiv_id":"2506.17739","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"In short data center simulations, a linear battery model with inefficiency and power limits reproduces physics-based battery behavior almost as well as electrochemical models, at a fraction of the runtime.","lead":"This paper compares four battery models for simulating data center microgrids, from a simple lossless model to detailed electrochemistry simulations. It finds that a medium-complexity linear model matches the detailed models closely in short experiments while running far faster.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CLCBattery's parameters are fitted to PyBaMM and then judged by PyBaMM's own SoC traces, which the paper concedes are only approximate; the 'closely match' claim lacks quantitative error bounds.","rationale":"The reader's conditional verdict is appropriate. I read Sections 4 and 5 carefully. The positive claim is plausible and the runtime measurements are credible; as a systems integration paper, the contribution has merit. My concern is not that PyBaMM is an unreasonable reference per se, but that the paper's own limitation statements weaken the evidence for 'closely match': CLC parameters are derived from PyBaMM, SoC is the comparison quantity, and the authors state they cannot quantify SoC differences because the complex models' SoC estimates are inaccurate. This internal inconsistency is more specific than the reader's broader concern about PyBaMM's fidelity to real cells, though the two are related. In addition, no code or data artifacts are provided, so the traces cannot be independently recomputed. None of this warrants rejection, because the paper's main contribution is the integration and the qualitative comparison, and the central claim is explicitly scoped to short-term experiments. The proposed test would either corroborate the claim or cap its strength with quantitative bounds.","tokens_in":15119,"tokens_out":3954,"duration_ms":41933,"concrete_test":"Recompute the Section 5 experiments with a quantitative SoC error metric (e.g., RMSE or MAE over time between CLCBattery and PyBaMM for each C-rate), and separately validate CLCBattery against experimental data for the M50 cell, e.g., Chen et al.'s measured voltage and current curves, using only parameters fitted at 0.2C/0.3C. If the mean SoC error exceeds a pre-stated tolerance such as 5%, or if CLCBattery's error against experimental data is comparable to or larger than its error against PyBaMM, the 'closely match' claim should be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CLCBattery closely matches physics-based models rests on a fit-and-compare loop within a single simulator. In Section 4.2, the CLC inefficiency factors and linear energy limits are determined from PyBaMM simulations of the INR21700 M50 cell. In Section 5, the same PyBaMM-based models are the reference against which CLCBattery is visually judged. This does not by itself invalidate the claim that CLCBattery can stand in for PyBaMM, but it makes the surrogate's good fit partly a construction of the chosen reference. The problem is compounded by the evaluation metric: the claimed agreement is in SoC traces, yet Section 5.1 explicitly states that 'because of inaccurate SoC estimations of the complex models, we cannot exactly quantify the SoC differences.' Section 4.3 also says the OCV-SoC lookup is only 'a reasonable estimation' and that SoC tends to be overestimated around half-full. Thus the headline 'closely match' has no quantitative error bound attached to the very quantity used to demonstrate it. The only quantitative comparisons are aggregate grid-energy differences (1.9% vs PyBaMM, 8.5% vs LiionBatteryPack in Section 5.2), which are not the claimed behavioral match and are still limited to a single two-day scenario, one battery chemistry, and one pack configuration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper extends the Vessim co-simulation framework with a common battery interface and implements four battery models: SimpleBattery, CLCBattery (a linear C-L-C model with inefficiencies and power limits), PybammBattery (a single-particle model with electrolyte via PyBaMM), and LiionBatteryPack (PyBaMM plus liionpack circuit solving). The authors evaluate SoC trajectories under constant charge/discharge at various C-rates and in a two-day data center microgrid scenario, and measure per-step runtime. They argue that CLCBattery closely matches the physics-based models in short-term experiments with much lower runtime, while SimpleBattery is inaccurate and offers no runtime advantage.","tokens_in":15406,"tokens_out":5727,"duration_ms":54845,"significance":"The practical question addressed—which battery model to use in data center co-simulations—is timely, and the modular interface is a useful engineering contribution. The runtime measurements are clean, obtained on a single node with comparable methodology, and show a clear 500x gap between linear and pack-level circuit models. The paper also honestly discloses limitations of physics-based SoC estimation. However, the accuracy claim is weakened by a calibration-and-comparison loop within PyBaMM and by the absence of quantitative SoC error metrics; only aggregate grid-energy differences (1.9%, 8.5%, 41.5% in Section 5.2) are numeric. These issues are local, not fundamental, and can be addressed with additional analysis.","major_comments":[{"comment":"The central claim that CLCBattery 'closely matches' the behavior of physics-based models is not quantitatively supported for the SoC trajectories. Section 5.1 explicitly states that 'we cannot exactly quantify the SoC differences,' and Figure 8 shows a single deterministic run without error bars or uncertainty bounds. To make the headline claim measurable, the authors should report quantitative error metrics for the SoC traces (e.g., RMSE and maximum absolute deviation over time for each scenario and C-rate), together with a description of how SoC estimation error in the complex models is accounted for. Without such metrics, the claim remains an impression based on visual inspection.","section":"Sections 5.1 and 5.2"},{"comment":"The close agreement between CLCBattery and the physics-based models is partly a calibration artifact: the CLC inefficiency factors and linear energy limits are fitted to PyBaMM simulations in Section 4.2, and PyBaMM-based models are then used as the reference in Section 5. To break this circularity, the authors should validate CLCBattery against an independent reference, such as the experimentally measured cell data from Chen et al. [11] or a separate experimental SoC dataset, for at least one scenario. Alternatively, the claim in the abstract and Section 6 should be explicitly restricted to 'CLCBattery approximates PyBaMM's single-particle model,' rather than implying general agreement with complex battery behavior.","section":"Sections 4.2 and 5"},{"comment":"The power-limit inequality as written is ambiguous and appears dimensionally inconsistent. If the energy limits a1(I) and a2(I) are expressed in Wh and I in A, then u1 has units V·h, so the term u1·V in the denominator has units V^2·h, while d_s·η_d has units h, making the two terms incompatible. The authors should define all variables with units, add explicit parentheses to the fraction, and show the derivation from the C-L-C energy limits to confirm that the implementation follows the intended constraints. This is important because an implementation error in Eq. (4) would directly affect the CLCBattery's behavior and the validity of the comparison.","section":"Section 4.2, Eq. (4)"},{"comment":"The generalization that CLCBattery 'is applicable for most use-cases utilizing microgrid simulation over a short time-frame' goes beyond the evidence presented. The evaluation uses a single cell chemistry (INR21700 M50), a single pack configuration (16S16P) for the data center scenario, and a single two-day weather trace and control policy. To support this broader claim, the authors should add at least a few variants, such as different pack sizes, different operating SoC ranges, or a multi-day scenario, or explicitly temper the generality statement to what the experiments actually cover.","section":"Section 6 (Discussion)"}],"minor_comments":[{"comment":"The phrase 'way too simple battery models' is informal and should be replaced with a more neutral formulation such as 'overly simple battery models.'","section":"Section 2 (Related Work)"},{"comment":"The sentence 'PyBaMM simulations, determined the constant inefficiency factors...' is a grammatical fragment; it should be reworded, for example, to 'PyBaMM simulations were used to determine the constant inefficiency factors...'.","section":"Section 4.2"},{"comment":"In the discussion of charging above the 0.7C limit, the text says 'resulting in an even slower discharge'; this should read 'slower charge'.","section":"Section 5.1 (Charging paragraph)"},{"comment":"The phrase 'battery pack imitated using the LiionBatteryPack model' should be 'simulated using the LiionBatteryPack model.'","section":"Section 5.3"},{"comment":"The sentence 'Each microgrid can consist of multiple simulator responsible for...' has a grammar error; 'simulator' should be 'simulators.'","section":"Section 3.1"},{"comment":"Figure 2 is dense and difficult to read at the published size; consider enlarging it or providing a simplified schematic of the 4S4P pack topology.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a workshop-style evaluation, and the authors' own text acknowledges the SoC differences cannot be exactly quantified. The runtime and architecture contributions are solid, but the central accuracy claim needs quantitative support and an independent validation to avoid the circularity concern. For a journal venue, the authors should also broaden the scenario coverage before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the integration work is real, the runtime analysis is useful, and the practical recommendation—use the C-L-C linear model instead of a lossless one for short-term data center microgrid co-simulations—is reasonable. But the headline 'closely match' is not as solid as the abstract implies.\n\nWhat's actually new: the paper adds a storage interface to Vessim, wraps three battery models (Kazhamiaka's C-L-C, PyBaMM's SPMe, liionpack) behind it, and compares them against the existing simple model in constant-power and a two-day data center scenario. The runtime measurements are clean and meaningful: CLCBattery is ~40x faster than PybammBattery and ~500x faster than a 16S16P liionpack. The finding that pack-level circuit losses barely matter over two days is a useful negative result for anyone building co-simulations.\n\nWhere it wobbles: the CLC parameters are fitted using PyBaMM simulations (Section 4.2), and PyBaMM then serves as the reference in Section 5. That makes the good fit partly a calibration artifact. The paper is honest about this—Section 5.1 says SoC differences cannot be exactly quantified because the complex models' SoC estimates are inaccurate—but that honesty undercuts the 'closely match' phrasing. There are no error bars, no SoC error metrics, and only one chemistry, one pack configuration, and one two-day weather/load trace. The grid-energy comparison shows CLCBattery within ~6-8% of the physics models, which is fine for many planning questions, but the qualitative SoC trace comparison does not support stronger statements. Also, no code or data artifacts are provided, which is a real gap for a systems paper whose main contribution is an integration.\n\nCredit where due: the authors do not oversell PyBaMM's accuracy, they flag the SoC estimation problem, and the limitations section is candid. The related work is appropriately placed. The central engineering takeaway survives: for short-term experiments, the linear C-L-C model with inefficiencies and power limits is a good substitute for detailed electrochemical models, and SimpleBattery is not, especially when charging near full or when grid-energy accounting matters.\n\nBottom line: this deserves a serious referee. A reviewer should push for quantified error bounds (e.g., SoC RMSE or time-to-empty), at least one more scenario or chemistry, and a code/data release. With those, it would be a dependable reference for the carbon-aware computing crowd.","headline":"Worth a look for the Vessim storage interface and the runtime numbers; read the accuracy claim with a grain of salt because it is a PyBaMM-calibrated surrogate compared against PyBaMM, with no quantitative SoC error bounds.","tokens_in":15991,"tokens_out":4383,"would_cite":false,"duration_ms":38784,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A linear battery model that accounts for charging inefficiencies and power limits closely tracks physics-based battery models in short-term data center microgrid simulations, while running orders of magnitude faster.","keywords":["battery modeling","energy storage","carbon-aware computing","microgrid","co-simulation","data center","lithium-ion battery","sustainable computing"],"falsifier":"Repeat the two-day data center scenario on a physical INR21700 M50 battery pack or a testbed with a real cell, logging actual state of charge and grid energy; if the real battery's net grid energy differs from the C-L-C model by substantially more than the 8.5 percent gap the paper reports against the physics model, the central claim fails.","tokens_in":14859,"feed_emoji":"🔋","tokens_out":6059,"duration_ms":57057,"temperature":0.7,"pith_summary":"This paper tries to settle a practical choice for researchers who simulate data center microgrids with battery storage: which battery model is accurate enough without slowing the simulation to a crawl. It argues that a linear model that accounts for charging inefficiencies and power limits—called the C-L-C model—closely tracks the behavior of full electrochemical battery models over short simulated experiments, while running hundreds of times faster and needing only product-sheet parameters. The simple lossless model that many simulators default to, by contrast, diverges from the physics-based models and gains no runtime advantage. If this holds, most short-term co-simulation studies of carbon-aware computing can safely use the linear model and reserve electrochemistry for long-term questions like degradation.","feed_headline":"Linear battery model matches physics models in short runs","feed_subtitle":"Data center microgrid simulators can replace slow electrochemistry with a fast, spec-based model and keep accuracy.","key_machinery":"The load-bearing object is the C-L-C linear storage model: it tracks stored energy as $b(t)=b(t-1)+\\eta_c p_s(t)d_s(t)$ when charging and $b(t-1)+\\eta_d p_s(t)d_s(t)$ when discharging, with fixed efficiency factors $\\eta_c$ and $\\eta_d$, and it enforces (dis)charge power limits through linear energy-bound curves parameterized from cell specifications. The paper couples this model to a discrete-event co-simulation through a storage interface with a battery-management system and a microgrid policy that decides how much of the power delta the battery should absorb or supply. The argument works because the model's parameters are calibrated against the same physics-based cell model that later serves as the accuracy reference, and because the comparison is restricted to short-term behavior where a constant-voltage approximation is reasonable.","core_discovery":"The central claim is that the gap between a linear, specification-driven battery model and a detailed physics-based model is small for the time horizons that data center microgrid experiments actually use. The paper implements four models—a lossless energy counter, the C-L-C linear model with constant inefficiencies and linear power limits, a single-cell electrochemistry model, and a full battery-pack circuit-plus-chemistry model—and compares their state-of-charge traces and grid energy exchange over constant-power runs and a two-day simulated data center. The C-L-C model reproduces the physics-based models' state-of-charge progression across charging and discharging rates and lands within about 8.5 percent of their net grid energy in the two-day scenario. The lossless model misses the grid energy by over 40 percent because it ignores inefficiencies. The paper concludes that the linear model is sufficient for most short-term co-simulations, that the pack-level circuit model adds little for short experiments, and that the physics-based models are hard to justify there outside degradation studies.","pith_inferences":["A natural extension is to test whether the C-L-C model's closeness survives other cell chemistries or temperatures; the paper's calibration uses a single lithium-ion cell parameterization, so the 8 percent grid-energy figure is not yet a universal bound.","The paper's implicit recommendation—use linear models for short experiments and electrochemistry only for degradation—suggests building a hybrid that runs the linear model in the hot loop and periodically corrects it with an electrochemical model, or switches to the physics model when state-of-health questions arise.","Because the battery interface separates the policy from the storage model, controllers that work with the linear model should transfer unchanged to the physics models, making the interface itself a reusable artifact for comparing energy-management algorithms.","If real-cell validation later confirms the C-L-C approximation, default simulation stacks for carbon-aware scheduling studies could drop electrochemistry entirely, removing a major barrier to reproducing those experiments."],"forward_implications":["Researchers extending the co-simulation testbed can default to the C-L-C model for experiments lasting hours to a few days and expect its grid-energy estimate to stay within roughly 8 percent of an electrochemical model.","The lossless battery model should not be used in energy-management studies; its more than 40 percent grid-energy error in the two-day scenario is large enough to change management conclusions.","For short experiments, modeling a full battery pack's circuit adds negligible accuracy over a single scaled cell, so pack-level models can be skipped unless pack imbalances or real-time stepping are the question.","The common interface lets a simulator swap models without changing controllers, so the same experiment can be run at different fidelity levels and the model choice can be justified by runtime budget.","Runtime overhead of the C-L-C model is tiny (about 196 microseconds per step) compared with about 7.65 milliseconds for single-cell electrochemistry and about 0.095 seconds for a 16S16P pack, making large parameter sweeps feasible."],"supporting_citations":[{"why":"Supplies the C-L-C model: constant efficiencies and linear power/energy limits, which the paper recommends as the middle ground.","marker":"[22]"},{"why":"Provides the physics-based single-cell reference model used to calibrate and judge the linear models.","marker":"[42]"},{"why":"Supplies the battery-pack circuit network solver behind the LiionBatteryPack model.","marker":"[43]"},{"why":"Provides the experimentally fitted lithium-ion cell parameters used to parameterize all four models.","marker":"[11]"},{"why":"Defines the co-simulation testbed architecture that the paper extends with a standardized storage interface.","marker":"[44]"},{"why":"Provides the software-in-the-loop control and scenario pattern used in the data center experiment.","marker":"[47]"}],"fun_headline_variants":["Linear battery model matches physics in data center runs","Skip slow battery physics: linear model suffices for data center sims","Fast linear battery model keeps accuracy in short data center tests","Battery model shortcut: linear fit matches physics for data center runs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison treats a physics-based simulation of one lithium-ion cell as the truth against which the linear model is judged, so if that simulator's parameters do not reflect a real battery's behavior, the claimed closeness of the linear model is unproven for real hardware.","fun_headline_variants_meta":{"raw":{"variants":["Linear battery model matches physics in data center runs","Skip slow battery physics: linear model suffices for data center sims","Fast linear battery model keeps accuracy in short data center tests","Battery model shortcut: linear fit matches physics for data center runs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000612,"raw_usage":{"total_tokens":2840,"prompt_tokens":934,"completion_tokens":1906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":1836}},"tokens_in":550,"tokens_out":1906,"duration_ms":12209,"temperature":1.0,"reasoning_tokens":1836,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:01:59.046346+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the two-day data center scenario on a physical INR21700 M50 battery pack or a testbed with a real cell, logging actual state of charge and grid energy; if the real battery's net grid energy differs from the C-L-C model by substantially more than the 8.5 percent gap the paper reports against the physics model, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the C-L-C model: constant efficiencies and linear power/energy limits, which the paper recommends as the middle ground."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the battery-pack circuit network solver behind the LiionBatteryPack model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the co-simulation testbed architecture that the paper extends with a standardized storage interface."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the software-in-the-loop control and scenario pattern used in the data center experiment."}],"review_version":2}