{"id":"620a767b-785f-41ee-8d76-c647dd30bfba","arxiv_id":"2411.17583","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"A dual-sourcing inventory model for green hydrogen shows that accounting for both stochastic local capacity and random import yield gives an average cost benefit of about 8 percent, and that flexible fixed-order policies come within 2 percent of the optimal policy.","lead":"This paper builds a mathematical model to help the Netherlands decide how much green hydrogen to produce locally versus import, given that local output is unpredictable and imports can shrink in transit. If the model holds, dual sourcing with flexible orders could cut expected supply costs by about 8 percent compared with ignoring these uncertainties.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (3) is not a valid Bellman equation for the stated average-cost objective: it omits the gain constant, so the 'optimal' policies behind the 8% and 2% claims are not rigorously grounded.","rationale":"The most load-bearing claim is that the MDP solution is optimal, because both headline results are defined relative to it. The Reader's weakest assumption identifies exactly this point, and I agree. I considered the FOQ+ construction inconsistency (§4.2 vs §6.2) and the underspecified distributions (§5.2, §6) as alternatives; those are real reproducibility concerns and may change exact numbers, but they do not undermine the 'optimal' ground truth as directly as an invalid Bellman equation. The concern is not that the algorithm is necessarily wrong—span-based stopping can work—but that the paper provides no theorem or explicit algorithm normalization establishing that the stopped policy solves the stated objective. A focused reimplementation with the gain term made explicit settles it. If the check matches reported numbers, the conditional concern is resolved and the verdict can be upgraded after minor revisions; if not, the headline percentages need revision.","tokens_in":21755,"tokens_out":9158,"duration_ms":91767,"concrete_test":"Re-implement the MDP as explicit average-cost relative value iteration: pick a reference state S_ref, update h^{k+1}(S)=min E[CO+CI+h^k(S')]-h^k(S_ref), stop when span(h^{k+1}-h^k)<ε, and recompute long-run average costs for all scenarios in Tables 2-5 (all countries, cost ratios, storage types). Compare each recomputed optimal cost and local-supply percentage to the reported values. If any optimal cost changes by more than 0.5% or any policy's local share shifts by more than 1 percentage point, the reported 8% and 2% claims are not robust; if all values match, the omission is purely a presentation issue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper states an average-cost objective ('expected cost per period', §3) but Eq. (3) writes V(St)=min E[CO+CI+V(St+1)] with neither a discount factor nor the average-cost gain constant. For positive costs no finite V satisfies this equation, so as written the optimality equation is not well-posed. Algorithm 1's stopping rule (span of successive differences) is a known heuristic for average-cost value iteration, but the paper never says it is solving the average-cost optimality equation g+h(S)=min E[CO+CI+h(S')], nor states the conditions (finite ergodic MDP, normalization of V) under which the span rule certifies optimality. Because every number in Sections 6.1-6.2—the 8% benefit and the ~2% heuristic gaps—is computed against policies labeled 'optimal' by this algorithm, the central quantitative claims rest on an unstated optimality criterion. If the intended criterion is discounted cost, a discount factor is missing; if average cost, the gain term is missing. Either way, the label 'optimal' is not rigorously supported by the equations in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a dual-sourcing inventory problem for green hydrogen in which local production is subject to stochastic capacity and imports are subject to random yield, with general lead times. The authors formulate an infinite-horizon MDP, solve it by value iteration, and propose four heuristics (FOQ, FOQ+, TBS, TBS+) that stabilize order quantities. A Netherlands case study with Norway, Morocco, and the UAE as suppliers reports an average cost benefit of 8% for considering both uncertainties relative to models that ignore them, heuristic optimality gaps of about 2% for FOQ+, and sensitivity analyses linking local supply rates to Dutch climate scenarios.","tokens_in":21993,"tokens_out":9150,"duration_ms":85009,"significance":"The application is timely and the modeling combination of stochastic capacity and random yield in dual sourcing is a reasonable extension of the inventory literature. The case study uses externally grounded cost data, and the headline 8% and roughly 2% figures are outputs of the proposed MDP and simulations, not inputs; I see no circularity. If the numerical claims are reproducible and the optimality criterion is made rigorous, the paper would offer useful, decision-relevant insights for hydrogen trade negotiations. However, the paper provides no formal structural results, no code or data, and the current statement of the Bellman equation is not a well-posed optimality condition for the stated average-cost objective, so the quantitative claims need verification before publication.","major_comments":[{"comment":"The optimality equation is not well-posed for the stated objective of minimizing expected cost per period. Equation (3) writes V(S_t) = min E[C_O + C_I + V(S_{t+1})] with no discount factor and no average-cost gain term; for positive costs no finite V satisfies this equation. Algorithm 1's span-based stopping rule resembles average-cost value iteration, but the paper never states the normalized relative value iteration, the gain constant, or the conditions under which span-based stopping certifies optimality. Because the 8% benefit and the 2% heuristic gaps in Sections 6.1 and 6.2 are all measured against policies labeled optimal by this algorithm, the central quantitative claims rest on an unstated optimality criterion. Please either introduce a discount factor and report discounted results, or state and solve the average-cost optimality equation with the gain term and a normalization, and confirm that the reported numbers are unchanged.","section":"§3.1, Eq. (3); §4.1, Algorithm 1"},{"comment":"The exact probability distributions driving the experiments are not specified. Section 5.2 gives support sets for local capacity and demand but no probability masses; Section 6 says these are normal with VarL=0.5 but does not describe the discretization/truncation used for capacity and demand, while the corresponding procedure for random yield is only described in Section 5.3. As a result, the numerical results—including the headline 8% benefit and all heuristic gaps—cannot be reproduced or independently checked. Please provide the full probability mass functions (or a data/code repository) and state how the pmfs are derived from the normal distributions.","section":"§5.2, §5.3, §6"}],"minor_comments":[{"comment":"Please report the normalization used for the value function, the value of epsilon, and the number of iterations needed for convergence; without these details the 'optimal' label is difficult to assess.","section":"§4.1, Algorithm 1"},{"comment":"FOQ+ and TBS+ are obtained by re-running Algorithm 1 on restricted action spaces; the paper should state explicitly that these are optimal policies for the constrained MDP, so the reported gaps are constrained-optimality gaps rather than heuristic approximations.","section":"§4.2"},{"comment":"The storage-cost units should be clarified. Section 5.2 says per-day costs are multiplied by 7 to obtain weekly ch, but Figure 2's axis is labeled '€/kg' and the text refers to '4 €/kg' as a threshold; specify whether the axis and thresholds are per-day or per-week costs.","section":"§5.2, Figure 2"},{"comment":"At ρl/i=1.0 with salt-cavern storage for Norway, the 'Yes/No' deviation (9.21%) exceeds the 'No/No' deviation (7.38%), which contradicts the general claim that the Yes/No case decreases deviations; please address this exception.","section":"§6.1, Table 2"},{"comment":"Tables 5–7 report gaps such as 0.00% and 0.01%; because the heuristic parameters are selected by simulation over 100,000 periods, please provide standard errors or confidence intervals so that small values can be distinguished from simulation noise.","section":"§6.2, Tables 5–7"},{"comment":"The 'within 2% of optimal' claim should be qualified as an average; individual FOQ+ gaps in Table 5 reach 4.39% (ρl/i=1.2, compressed gas, Norway).","section":"Abstract and §7"},{"comment":"Equation (1) uses the symbol s′0_t before it is introduced in Eq. (2); please define the interim inventory level before using it.","section":"Eq. (1)–(2)"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than rejection: the modeling framework and case study are valuable, and the optimality-equation issue can in principle be fixed by rewriting the Bellman equation and re-running the numerical experiments. I would also ask the authors to deposit code and data, or at least full distribution tables, since the value of the paper rests heavily on reproducible numerical claims. The paper may fit an application-oriented OR/energy journal; the methodological novelty is incremental."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the combination: a dual-sourcing inventory MDP with stochastic local supply capacity and random import yield, which the literature review convincingly shows has not been done before. The Netherlands case study is serious, with real parameter sources and a clear policy question. The 8% average cost benefit is an output of the model, not an input, so there is no circularity. Credit where due: the decomposition into single sourcing, dual sourcing ignoring one or both uncertainties is a clean way to isolate the value of each feature, and the heuristic gap analysis is useful for practitioners. FOQ+ staying within about 2% of the claimed optimal is a genuinely useful finding.\n\nNow the soft spots, in proportion. The stress-test concern is valid. Equation (3) writes V(St) = min E[CO + CI + V(St+1)] with no discount factor and no average-cost gain constant. For the stated average-cost objective, that equation is not well-posed. Algorithm 1's span-based stopping rule is the kind of thing you use in relative value iteration, but the paper never connects it to the average-cost optimality equation. So the \"optimal\" label, which backs all the quantitative claims, is not rigorously grounded in the equations as written. This is a presentation flaw rather than evidence that the numbers are wrong, but it needs to be fixed.\n\nSecond, reproducibility: the exact discrete distributions for demand, capacity, and yield are not fully specified—just means, ranges, and a truncated normal that is \"mapped to discrete values.\" No code or data are provided. That makes the computational results hard to check. Third, a direct internal contradiction: Section 4.2 says FOQ+ bounds are based on the fixed quantities from FOQ, while Section 6.2 says they first obtain ¯ql from TBS. That is a real inconsistency and undermines the reported gaps until clarified.\n\nThese are all addressable. The core idea and the case study are solid; the missing pieces are mathematical hygiene and reproducibility. A serious referee should see it, but the authors should be told to fix the Bellman equation, specify the distributions, and reconcile the FOQ+ description before the numbers can be trusted.","headline":"Novel dual-sourcing MDP for green hydrogen with a serious case study, but the Bellman equation is sloppy and the FOQ+ description contradicts itself; worth refereeing after fixes.","tokens_in":22585,"tokens_out":2147,"would_cite":true,"duration_ms":22339,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C40","90B05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that jointly modeling stochastic local supply capacity and random import yield in a dual-sourcing Markov decision process yields an average cost benefit of about 8 percent over models that ignore both, and that flexible…","keywords":["green hydrogen","dual sourcing","Markov decision process","stochastic supply capacity","random yield","inventory control","value iteration","hydrogen import"],"falsifier":"Re-solve the Netherlands case study with the same transition law using a standard average-cost MDP algorithm that includes a gain constant in the Bellman equation; if the long-run average cost of the Algorithm 1 policy differs from that average-cost optimum by more than numerical tolerance, then the reported 8% and 2% figures are not measured against the true optimal policy.","tokens_in":21537,"feed_emoji":"💧","tokens_out":9075,"duration_ms":74212,"temperature":0.7,"pith_summary":"Green hydrogen supply is unreliable in two ways: local electrolysis output fluctuates with renewable availability, and imported hydrogen arrives only after a lead time and with a random fraction lost in transport and conversion. The paper models both uncertainties together in a Markov decision process and solves it with value iteration, arguing that a policy accounting for both is the right benchmark for dual sourcing decisions. In a case study of Dutch imports from Norway, Morocco, and the UAE, considering both stochastic supply capacity and random yield gives an average cost benefit of 8 percent compared to dual sourcing models that ignore them. The paper also proposes fixed and partially adjustable order policies; FOQ+, which lets order quantities vary inside narrow bands, stays within about 2 percent of optimal cost while giving exporters more stable order patterns. These results support feasibility assessments of Dutch climate scenarios that specify local production and import shares.","feed_headline":"Green hydrogen dual sourcing cuts cost 8% when supply risk is modeled","feed_subtitle":"Dutch case study: modeling local and import uncertainties cuts cost 8% on average; flexible orders stay near optimal.","key_machinery":"The central object is a Markov decision process whose state records on-hand inventory and the pipeline of outstanding local and import orders, capturing general lead times. At each decision epoch the firm chooses a local order quantity and an import order quantity; the stochastic local capacity $K_t^l$ is realized after ordering, so the actual local delivery is $\\min\\{K_t^l, \\hat{q}_t^l\\}$, and the import order that arrives $\\tau_i$ periods later is multiplied by a random yield factor $p_t$. The optimal policy is computed by value iteration (Algorithm 1), and four heuristic policies are derived by restricting actions: FOQ fixes both order quantities, FOQ+ allows each to vary within a two-step band around fixed levels, TBS keeps a fixed import order and uses local production as a backup triggered by an inventory threshold, and TBS+ adds a band of allowed import adjustments to TBS.","core_discovery":"The paper's central claim is that the optimal dual sourcing policy for green hydrogen must be computed with stochastic local supply capacity and random import yield modeled simultaneously, because ignoring either one biases the sourcing mix and adds cost. On the evidence of the Netherlands case study, jointly modeling both produces an average cost benefit of 8% relative to dual sourcing models that ignore both, with stochastic local capacity contributing more than random yield in these settings. The paper further claims that simple, more implementable policies are near-optimal: FOQ+ has an average optimality gap around 1.6–1.8% across the three exporting countries and cost settings, and TBS+ improves on TBS at higher cost ratios and storage costs, which agrees with the known result that tailored base-surge policies improve as slower-supplier lead times grow. These findings are put forward to guide order-structure negotiations in hydrogen trade agreements and to identify conditions under which Dutch climate scenarios' local production shares are achievable.","pith_inferences":["Editorial inference: the same MDP could be re-estimated for other intermittent local generators and other hydrogen carriers such as ammonia or liquid organic hydrogen carriers; only the lead-time, yield, and cost distributions change.","Editorial inference: the sensitivity result that higher demand variability shifts sourcing toward local production suggests a direct counterfactual test—if the import lead time were equalized with the local lead time, that shift should weaken; the paper does not run this experiment.","Editorial inference: the Bellman equation in Eq. (3) is written without a discount factor or an average-cost gain term, so certifying the 'optimal' label would require restating it as an average-cost optimality equation and checking that the value iteration stopping rule converges to that gain.","Editorial inference: in practice, the width of the FOQ+ bands is set heuristically to two steps in the action grid; an obvious refinement is to optimize band width jointly with base quantities for each counterparty."],"forward_implications":["Energy planners who neglect either stochastic local capacity or random import yield should expect to pay about 8 percent more than necessary under the Dutch 2030 parameter settings.","A fixed-order contract with a narrow adjustment band (FOQ+) is a strong candidate for hydrogen trade agreements: it stabilizes order quantities and stays within roughly 2 percent of the optimal policy across countries and cost ratios.","Tailored base-surge contracts become more attractive as the import lead time and the local-to-import cost ratio grow, which guides when exporters should push for fixed base volumes.","Achieving the National Drivers climate scenario, with high local production, requires low local-to-import production costs and low variability in local capacity; the International Ambition scenario becomes feasible when local production costs exceed import costs by about 20 percent or more."],"supporting_citations":[{"why":"Introduces the tailored base-surge policy that TBS and TBS+ build on.","marker":"Allon and Van Mieghem (2010)"},{"why":"Provides the analysis of tailored base-surge policies used to motivate TBS-style heuristics.","marker":"Janakiraman et al. (2015)"},{"why":"Shows asymptotic optimality of tailored base-surge as slow-supplier lead time grows, which the paper cites to explain TBS performance.","marker":"Xin and Goldberg (2018)"},{"why":"Establishes that optimal dual sourcing with general lead times needs a state vector of net inventory plus outstanding orders, the basis for the MDP state.","marker":"Whittemore and Saunders (1977)"},{"why":"Studies dual sourcing with general lead times and supply capacity uncertainty, the closest methodological predecessor that the paper extends with random yield.","marker":"Chen and Yang (2019)"},{"why":"Treats dual sourcing with stochastic supply capacity at the faster supplier and provides a myopic heuristic, a benchmark for treating capacity uncertainty.","marker":"Jakšić and Fransoo (2018)"},{"why":"Supplies the value iteration algorithm and MDP theory used to compute optimal policies.","marker":"Puterman (2014)"},{"why":"Provides the mean 17.5% transport and conversion loss and cost data that calibrate the random import yield in the case study.","marker":"International Renewable Energy Agency (2022c)"}],"fun_headline_variants":["Green hydrogen: modeling supply risk cuts costs 8%","Dual sourcing with uncertain supply saves 8% on hydrogen","Stochastic supply modeling: 8% cheaper green hydrogen","Hydrogen import uncertainty: optimal mix saves 8%","Green H2: joint modeling of local and import risks pays off"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Algorithm 1's value iteration really produces the optimal policy for the stated infinite-horizon problem, but Eq. (3) writes the Bellman equation with no discount factor and no average-cost gain term, so the optimality criterion behind all percentage gaps is not pinned down by the equations.","fun_headline_variants_meta":{"raw":{"variants":["Green hydrogen: modeling supply risk cuts costs 8%","Dual sourcing with uncertain supply saves 8% on hydrogen","Stochastic supply modeling: 8% cheaper green hydrogen","Hydrogen import uncertainty: optimal mix saves 8%","Green H2: joint modeling of local and import risks pays off"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2836,"prompt_tokens":1009,"completion_tokens":1827,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":1743}},"tokens_in":625,"tokens_out":1827,"duration_ms":13815,"temperature":1.0,"reasoning_tokens":1743,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:57:11.520059+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-solve the Netherlands case study with the same transition law using a standard average-cost MDP algorithm that includes a gain constant in the Bellman equation; if the long-run average cost of the Algorithm 1 policy differs from that average-cost optimum by more than numerical tolerance, then the reported 8% and 2% figures are not measured against the true optimal policy.","supporting_citations":[{"cited_title":"and Van Mieghem, J","cited_arxiv_id":null,"evidence_quote":"Introduces the tailored base-surge policy that TBS and TBS+ build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the analysis of tailored base-surge policies used to motivate TBS-style heuristics."},{"cited_title":"and Goldberg, D","cited_arxiv_id":null,"evidence_quote":"Shows asymptotic optimality of tailored base-surge as slow-supplier lead time grows, which the paper cites to explain TBS performance."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that optimal dual sourcing with general lead times needs a state vector of net inventory plus outstanding orders, the basis for the MDP state."},{"cited_title":"and Yang, H","cited_arxiv_id":null,"evidence_quote":"Studies dual sourcing with general lead times and supply capacity uncertainty, the closest methodological predecessor that the paper extends with random yield."}],"review_version":1}