REVIEW 4 major objections 4 minor 34 references
Seamless Execution of Malleable Applications in Controlled and Production HPC Environments
T0 review · 4 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read This paper shows that malleable MPI applications can resize themselves on vanilla production Slurm clusters by using user-level expander jobs and process-manager reconfiguration, cutting node-hour consumption by 25–74%.
desk verdict A genuinely new orchestration for MPI malleability on vanilla Slurm, demonstrated on three production systems, but the headline savings claims overstate what the data shows and the in-memory shrink path is left half-described. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the expander-job pattern: a user-level Slurm job submitted at runtime to acquire extra nodes, combined with the PRRTE process manager's ability to grow the distributed virtual machine via MPI_Comm_spawn. Shrinkage uses either direct Slurm job-size reduction or whole-job termination of previously created expander jobs. The DMRv2 API wraps these operations behind dmr_init, dmr_check, dmr_reconfigure, and dmr_finalize, with a DMR_AUTO macro that lets applications supply data-redistribution callbacks without changing the rest of their code.
What would settle it
Run the described shrink operation on a vanilla Slurm cluster with the same Open MPI/PRRTE versions and observe whether an MPI application continues in memory after one of its nodes is released; if the communicator fails or the application must restart from checkpoint even once, the claim of seamless in-memory shrink on production systems collapses.
Extended reading notes
Core claim
DMRv2 decouples MPI malleability from RMS-level job resizing. On expansion, the parent job keeps running while DMR submits an expander job with a matching wallclock; once the resources arrive, the application is suspended, the MPI communicator grows via MPI_Comm_spawn or checkpoint/restart, data is redistributed, and execution resumes. On shrink, resources are released by terminating expander jobs or by directly reducing the parent job's node count where Slurm allows it. For in-memory shrink with Open MPI, the paper notes that PRRTE cannot currently shrink a deployed distributed virtual machine, so it relies on a wrapper-script restart or on a PRRTE configuration that tolerates daemon failur
Load-bearing premise
The load-bearing premise is that a running Open MPI application can survive having nodes removed from its process manager's virtual machine, because PRRTE cannot currently shrink the DVM and the paper relies on either daemon-failure tolerance or a wrapper-script restart to make in-memory shrinking work.
Editorial extensions
If this is right
- The same malleable application code runs in a controlled testbed and on production systems without recompilation, since the choice between Slurm4DMR and DMR@Jobs is made at deployment time.
- Both checkpoint/restart and in-memory data redistribution are viable for production malleability, so applications with an existing robust C/R mechanism can adopt DMRv2 without rewriting their data movement.
- Waiting for expansion resources does not negate the benefit of malleability: the application continues computing while the expander job is queued, so only the actual reconfiguration step stalls.
- Dynamic right-sizing can reduce node-hour consumption by 25–74% for the tested workloads while converging to efficiency targets set by communication-efficiency metrics.
- The method generalizes across CPU-only and heterogeneous CPU–GPU systems, suggesting it is not tied to a particular hardware partition.
Reading between the lines
- If PRRTE gains native DVM shrinking, the wrapper-script restart fallback can be retired, making in-memory shrink seamless and likely improving the measured savings further.
- The expander-job pattern should transfer to other schedulers that expose user-level job submission APIs, so the same DMR@Jobs approach could be hosted on Flux, OAR, or similar systems without changing the application.
- The workload-level experiment suggests reconfiguration overhead dominates when inhibition periods are very short, implying that production deployments need to tune the inhibition period to the reconfiguration cost; this is a testable knob for adopters.
- A communication-efficiency target could serve as a simple autonomous right-sizing rule for other iterative solver codes beyond the two evaluated, since it only requires per-step communication measurements.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents DMRv2, a user-level framework that aims to provide MPI malleability on unmodified (vanilla) Slurm clusters. Instead of relying on scheduler-level job resizing, DMRv2 orchestrates separate Slurm jobs and uses PRRTE/Open MPI process management to expand or shrink the MPI process layout, with data redistribution supported either through checkpoint/restart or in-memory MPI communication. The authors integrate DMRv2 into two scientific applications, Alya and MPDATA, and report experiments on three production supercomputers: Leonardo, MareNostrum 5 ACC, and MareNostrum 5 GPP. They report that DMR@Jobs (production mode) converges to similar configurations as the controlled Slurm4DMR mode and reduces node-hour consumption by 74%, 25.10%, and 55.15% in the three main comparisons.
Significance. If fully substantiated, the paper's central claim is significant: it would provide a credible path to running malleable MPI applications on production HPC systems without scheduler modifications, using only user-level mechanisms. The evaluation covers two real applications, both C/R and in-memory redistribution, and three TOP500 systems, which is a strong empirical footprint not common in the malleability literature. The dual-mode API design (Slurm4DMR for controlled development and DMR@Jobs for production) is a practical contribution. However, the quantitative claims currently rest on a baseline that reserves the maximum footprint plus a controller node, on single executions per configuration, and on an in-memory shrink mechanism that is described only partially and with an unfinished sentence that explicitly marks it as anticipated rather than demonstrated.
major comments (4)
- [Section III, DVM shrink discussion; Section V-C, MPDATA on MN5 ACC] The in-memory shrink mechanism is load-bearing but not demonstrated. Section III states that PRRTE 'currently cannot shrink the DVM once it has been deployed on a node' and offers two workarounds: a wrapper-script restart, or PRRTE 'configured to tolerate daemon failures.' The paragraph then ends with the incomplete sentence 'We anticipate that, under a PRRTE configuration that natively supports shrinking the DVM.' The paper never identifies which path the MPDATA production run in Section V-C actually used, how many shrink transitions occurred, or whether the alpha PRRTE 5.0.0a1 failure-tolerance behavior was reliable. If the wrapper restart was used, the 'in-memory' and 'seamless' characterization is misleading; if the failure-tolerance path was used, configuration and evidence are missing. This uncertainty directly affects the 3.0 node-hour result and the claim of being the first pract
- [Section V-C, Table II, abstract] The advertised reductions against 'static baselines' are not computed against a static fixed-allocation baseline. In Table II the controlled baseline is Slurm4DMR reserving 14+1 or 32+1 nodes for the entire run, while DMR@Jobs uses variable ranges [5–14] and [12–32]. Slurm4DMR is a malleability-capable controlled environment, not a conventional static allocation; it also includes an extra management/controller node. A standard static baseline would be a fixed-size Slurm job with, e.g., 14 or 32 nodes and no controller. The reported 25.10%, 55.15%, and 74% reductions are therefore partly an artifact of comparing against a maximum-reservation-plus-one baseline. The authors should add explicit static fixed-allocation comparisons or, if 'static baseline' is meant to refer to Slurm4DMR, state this unambiguously and adjust the abstract's wording.
- [Section V.A, 'single execution per configuration'; Table II] The quantitative claims are based on a single execution per configuration, as acknowledged in Section V.A. The paper reports precise percentages for node-hour reductions (25.10%, 55.15%, 74%) and compares execution times (e.g., 2.68 h vs. 2.80 h) without any measure of run-to-run variability. Given that the authors also state that production contention and node placement vary between executions, the reported numbers cannot be read as robust estimates. The authors should either provide repeated runs and confidence intervals for the central comparisons, or explicitly present the figures as illustrative single-run observations and temper the corresponding claims in the abstract and conclusions.
- [Section V.E, workload-level experiment] The workload-level experiment does not currently support any workload-level productivity claim. It uses intentionally unrealistic inhibition periods of 10–100 time steps, and reports that the average RECONF state lasts 107.14 s, so reconfiguration time dominates execution. This is a useful stress test of the mechanism, but it is not compared to a rigid-job workload or evaluated in terms of completed jobs per unit time, so any reader inference about workload-level benefit is unsupported. The text should be reframed as a feasibility/stress study rather than a quantitative demonstration.
minor comments (4)
- [Section III, last sentence of DVM shrink paragraph] The sentence beginning 'We anticipate that, under a PRRTE configuration that natively supports shrinking the DVM' is syntactically incomplete; it needs a consequent clause. More importantly, its incompleteness signals the absence of a demonstrated implementation path, which should be resolved in the main text.
- [Figure 4] The caption says 'indicating the number of time steps needed for each reconfiguration until resources are available,' while the y-axis appears to show the number of allocated GPUs. Please clarify the axes and what the plotted quantity represents, since the current description is ambiguous.
- [Reference [34]] Reference [34] (the layered approach for DMR in HPC) contains a URL that appears to be the same as reference [33] (the Flux paper). Please verify and correct the URL for the Euro-Par 2024 workshop paper.
- [Section V.B and V.D, figure labels] The captions for Figures 3 and 5 refer to 'low' and 'high' jobs. These terms are defined in the text, but the figures themselves could benefit from explicit labels such as 'starting at 5 nodes' and 'starting at 16/32 nodes' to make the comparison immediately readable.
Circularity Check
No circular steps: the paper is an empirical systems evaluation, and its self-citations are contextual rather than load-bearing definitions or uniqueness arguments.
full rationale
The manuscript contains no formal derivation and no equation whose output reduces to an input; the central claims are measured results from runs on Leonardo, MN5 ACC and MN5 GPP. The CE target (70%), inhibition periods (500, 5,000, or random 10-100), and node ranges (e.g. [2-16]) are user-set control parameters, not fitted values later reported as independent predictions; therefore the fitted-input-called-prediction pattern does not apply. The node-hour savings (74% for MPDATA, 25.10% and 55.15% for Alya) are arithmetic comparisons between a Slurm4DMR baseline that reserves maximum nodes plus a controller and a DMR@Jobs run whose allocation varies; this is a benchmark construction that can be debated on fairness, but it is not circular because the production run's node-hour count is directly accounted, not derived from the baseline. Self-citations [7], [21], [22], [23], [27], [31] provide the prior DMR API, TALP integration, and earlier malleable Alya/MPDATA work; none is used as a uniqueness theorem, none forbids alternatives, and the present experiments are independently reported. The main weakness is an evidence gap, not circularity: Section III states that PRRTE 'currently cannot shrink the DVM once it has been deployed on a node' and the proposed in-memory shrink path trails off as 'We anticipate that, under a PRRTE configuration that natively supports shrinking the DVM.' This leaves the seamless in-memory production shrink under-supported, but under-support is not definitional equivalence, so it is correctly assigned to correctness risk and does not raise the circularity score.
Assumptions & free parameters
free parameters (5)
- CE_POLICY target communication efficiency =
70% (Alya); 75% (workload)
- CE_POLICY tolerance/aggressiveness parameter =
not specified
- Inhibition period =
500 (Leonardo), 5,000 (MN5 ACC), random 10-100 (workload)
- Initial and allowed node counts =
Alya low=5, high=16/32; MPDATA range 2-16; workload range 2-32
- Slurm4DMR baseline reservation =
14+1 and 32+1 nodes
assumptions (5)
- domain assumption PRRTE can expand the DVM via MPI_Comm_spawn and can be configured to tolerate daemon failures so processes continue when nodes are removed.
- domain assumption Production Slurm permits user-level job submission/termination and SSH bootstrapping between compute nodes from inside a job.
- domain assumption Alya's checkpoint/restart is independent of MPI process count and can restart on a different process layout.
- ad hoc to paper Single executions per configuration are representative despite acknowledged production variability.
- domain assumption Alpha builds of Open MPI 5.1.0a1 and PRRTE 5.0.0a1 behave reliably at production scale.
Cite this review
Pith. "Pith review of Seamless Execution of Malleable Applications in Controlled and Production HPC Environments." pith.science (2026). https://pith.science/paper/UYW65QHW
@misc{pith2026260613266,
author = {Pith},
title = {Pith review of: Seamless Execution of Malleable Applications in Controlled and Production HPC Environments},
year = {2026},
howpublished = {\url{https://pith.science/paper/UYW65QHW}},
note = {Machine review of arXiv:2606.13266}
}
read the original abstract
Many large-scale scientific applications exhibit time-varying behavior, yet production HPC clusters still rely on rigid, fixed-size allocations, and most dynamic techniques remain confined to laboratory prototypes. This work presents a practical MPI malleability methodology that integrates with state-of-the-art high-performance computing (HPC) software stacks and operational practices. The methodology is implemented in the Dynamic Management of Resources (DMR) framework and is designed to ease adoption by existing applications without requiring intrusive code changes or scheduler modifications. We evaluate our approach by integrating the DMR API into two large-scale scientific applications and deploying them on three TOP500 supercomputers under realistic production configurations. Our non-invasive malleability solution achieves performance comparable to static baselines in controlled environments while substantially reducing node-hour consumption for identical workloads. These results show that malleability can be effectively exploited on production systems using vanilla resource managers, lowering the barrier to adoption of dynamic resource management in HPC.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Exascale workload characterization and architecture implications,
P. Balaprakash, D. Buntinas, A. Chan, A. Guha, R. Gupta, S. H. K. Narayanan, A. A. Chien, P. Hovland, and B. Norris, “Exascale workload characterization and architecture implications,” in2013 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2013, pp. 120–121. [Online]. Available: https://doi.org/10.1109/ISPASS.2013.6557153
arXiv 2013
-
[2]
Invasive computing: An overview,
J. Teich, J. Henkel, A. Herkersdorf, D. Schmitt-Landsiedel, W. Schr ¨oder- Preikschat, and G. Snelting, “Invasive computing: An overview,” Multiprocessor System-on-Chip: Hardware Design and Tool Integration, pp. 241–268, 2011. [Online]. Available: https://doi.org/10.1007/978-1 -4419-6460-1 11
doi:10.1007/978-1 2011
-
[3]
Packing schemes for gang scheduling,
D. G. Feitelson, “Packing schemes for gang scheduling,” inJob Scheduling Strategies for Parallel Processing, 1996, pp. 89–110. [Online]. Available: https://doi.org/10.1007/BFb0022289
-
[4]
A Survey on Malleability Solutions for High-Performance Distributed Computing,
J. I. Aliaga, M. Castillo, S. Iserte, I. Mart ´ın- ´Alvarez, and R. Mayo, “A Survey on Malleability Solutions for High-Performance Distributed Computing,”Applied Science, vol. 12, pp. 1–32, May 2022. [Online]. Available: https://doi.org/10.3390/app12105231
-
[5]
Malleability in modern HPC systems: Current experiences, challenges, and future opportunities,
A. Tarraf, M. Schreiber, A. Cascajo, J.-B. Besnard, M.-A. Vef, D. Huber, S. Happ, A. Brinkmann, D. E. Singh, H.-C. Hoppe, A. Miranda, A. J. Pe ˜na, R. Machado, M. G. Gasulla, M. Schulz, P. Carpenter, S. Pickartz, T. Rotaru, S. Iserte, V . Lopez, J. Ejarque, H. Sirwani, and F. Wolf, “Malleability in modern HPC systems: Current experiences, challenges, and ...
arXiv 2024
-
[6]
Slurm: Simple linux utility for resource management,
A. B. Yoo, M. A. Jette, and M. Grondona, “Slurm: Simple linux utility for resource management,” inJob Scheduling Strategies for Parallel Processing, 2003, pp. 44–60. [Online]. Available: https://doi.org/10.1007/10968987 3
doi:10.1007/10968987 2003
-
[7]
High-throughput computation through efficient resource management,
S. Iserte, “High-throughput computation through efficient resource management,” Ph.D. Thesis, Universitat Jaume I (UJI), Spain, Nov
-
[9]
Autonomic malleability in iterative mpi applications,
A. C. Sena, F. S. Ribeiro, V . E. Rebello, A. P. Nascimento, and C. Boeres, “Autonomic malleability in iterative mpi applications,” in 2013 25th International Symposium on Computer Architecture and High Performance Computing, 2013, pp. 192–199. [Online]. Available: https://doi.org/10.1109/SBAC-PAD.2013.4
Show all 34 references
-
[10]
Reshape: A framework for dynamic resizing and scheduling of homogeneous applications in a parallel environment,
R. Sudarsan and C. J. Ribbens, “Reshape: A framework for dynamic resizing and scheduling of homogeneous applications in a parallel environment,” in2007 International Conference on Parallel Processing (ICPP 2007), 2007, pp. 44–44. [Online]. Available: https://doi.org/10.1109/IC...
2007 doi
-
[11]
Dynamic malleability in iterative mpi applications,
K. E. Maghraoui, T. Desell, B. K. Szyma ´nski, and C. A. Varela, “Dynamic malleability in iterative mpi applications,”Seventh IEEE International Symposium on Cluster Computing and the Grid (CCGrid ’07), pp. 591–598, 2007. [Online]. Available: https: //api.semanticscholar.org/C...
2007
-
[12]
Towards realizing the potential of malleable jobs,
A. Gupta, B. Acun, O. Sarood, and L. V . Kal ´e, “Towards realizing the potential of malleable jobs,” in2014 21st International Conference on High Performance Computing (HiPC), 2014, pp. 1–10. [Online]. Available: https://doi.org/10.1109/HiPC.2014.7116905
2014
-
[13]
Flex-mpi: an mpi extension for supporting dynamic load balancing on heterogeneous non-dedicated systems,
G. Mart ´ın, M.-C. Marinescu, D. E. Singh, and J. Carretero, “Flex-mpi: an mpi extension for supporting dynamic load balancing on heterogeneous non-dedicated systems,” inProceedings of the 19th International Conference on Parallel Processing, ser. Euro-Par’13, 2013, p. 138–149...
2013 doi
-
[14]
Infrastructure and api extensions for elastic execution of mpi applications,
I. Compr ´es, A. Mo-Hellenbrand, M. Gerndt, and H.-J. Bungartz, “Infrastructure and api extensions for elastic execution of mpi applications,” inProceedings of the 23rd European MPI Users’ Group Meeting, ser. EuroMPI ’16, 2016, p. 82–97. [Online]. Available: https://doi.org/10...
2016
-
[15]
Extending SLURM for dynamic resource-aware adaptive batch scheduling,
M. Chadha, J. John, and M. Gerndt, “Extending SLURM for dynamic resource-aware adaptive batch scheduling,” inIEEE 27th International Conference on High Performance Computing, Data, and Analytics (HiPC), 2020, pp. 223–232. [Online]. Available: https://doi.org/10.1109/HiPC50609....
2020
-
[16]
Dynamic resource management for elastic scientific workflows using PMIx,
R. Bhattarai, H. Pritchard, and S. Ghafoor, “Dynamic resource management for elastic scientific workflows using PMIx,” inIEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), 2024. [Online]. Available: https://doi.ieeecomputersociety. org/10.1109...
2024
-
[17]
Dynamic resource allocation for efficient parallel CFD simulations,
G. Houzeaux, R. M. Badia, R. Borrell, D. Dosimont, J. Ejarque, M. Garcia-Gasulla, and V . L ´opez, “Dynamic resource allocation for efficient parallel CFD simulations,”Computers & Fluids, vol. 245, p. 105577, Sep. 2022. [Online]. Available: https://doi.org/10.1016/j.comp fluid...
2022
-
[18]
Resource optimization with MPI process malleability for dynamic workloads in HPC clusters,
S. Iserte, I. Mart ´ın- ´Alvarez, K. Rojek, J. I. Aliaga, M. Castillo, W. Folwarska, and A. J. Pe ˜na, “Resource optimization with MPI process malleability for dynamic workloads in HPC clusters,”Future Generation Computer Systems, p. 107949, 2025. [Online]. Available: https://...
2025
-
[19]
Dynamic spawning of mpi processes applied to malleability,
I. Mart ´ın- ´Alvarez, J. I. Aliaga, M. Castillo, S. Iserte, and R. Mayo, “Dynamic spawning of mpi processes applied to malleability,” International Journal of High Performance Computing Applications, vol. 38, no. 2, p. 69–93, Mar. 2024. [Online]. Available: https: //doi.org/1...
2024 doi
-
[20]
Design Principles of Dynamic Resource Management for High- Performance Parallel Programming Models,
D. Huber, M. Schreiber, M. Schulz, H. Pritchard, and D. Holmes, “Design Principles of Dynamic Resource Management for High- Performance Parallel Programming Models,” 2024. [Online]. Available: https://arxiv.org/abs/2403.17107
2024 arXiv
-
[21]
DMR API: Improving cluster productivity by turning applications into malleable,
S. Iserte, R. Mayo, E. S. Quintana-Ort ´ı, V . Beltran, and A. J. Pe ˜na, “DMR API: Improving cluster productivity by turning applications into malleable,”Parallel Computing, vol. 78, pp. 54–66, 2018. [Online]. Available: https://doi.org/10.1016/j.parco.2018.07.006
2018 doi
-
[22]
Talp: A lightweight tool to unveil parallel efficiency of large-scale executions,
V . Lopez, G. Ramirez Miranda, and M. Garcia-Gasulla, “Talp: A lightweight tool to unveil parallel efficiency of large-scale executions,” inProceedings of the 2021 on Performance EngineeRing, Modelling, Analysis, and VisualizatiOn STrategy, ser. PERMA VOST ’21, 2021, p. 3–10. ...
2021
-
[23]
MPI Malleability Validation under Replayed Real-World HPC Conditions,
S. Iserte, M. Madon, G. Da Costa, J.-M. Pierson, and A. J. Pe ˜na, “MPI Malleability Validation under Replayed Real-World HPC Conditions,” Future Generation Computer Systems, p. 108305, Dec. 2025. [Online]. Available: https://doi.org/10.1016/j.future.2025.108305
2025
-
[24]
Using an adaptive hpc runtime system to reconfigure the cache hierarchy,
E. Totoni, J. Torrellas, and L. V . Kale, “Using an adaptive hpc runtime system to reconfigure the cache hierarchy,” inSC ’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2014, pp. 1047–1058. [Online]. Available...
2014 doi
-
[25]
Predicting output performance of a petascale supercomputer,
B. Xie, Y . Huang, J. S. Chase, J. Y . Choi, S. Klasky, J. Lofstead, and S. Oral, “Predicting output performance of a petascale supercomputer,” inProceedings of the 26th International Symposium on High- Performance Parallel and Distributed Computing, 2017, p. 181–192. [Online]...
2017
-
[26]
Boosting performance of iterative applications on gpus: Kernel batching with cuda graphs,
J. Ekelund, S. Markidis, and I. Peng, “Boosting performance of iterative applications on gpus: Kernel batching with cuda graphs,” in33rd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (PDP), 2025, pp. 70–77. [Online]. Available: https...
2025
-
[27]
Malleable computational fluid dynamics simulations,
S. Iserte, G. Houzeaux, P. Sand ˚as, A. J. Pe ˜na, and M. Garcia-Gasulla, “Malleable computational fluid dynamics simulations,” inProceedings of the 36th Parallel CFD International Conference, Merida, Yucatan, Mexico, Nov. 2025, [In-press]
2025
-
[28]
Alya: Multiphysics Engineering Simulation Towards Exascale,
M. V ´azquez, G. Houzeaux, S. Koric, A. Artigues, J. Aguado- Sierra, R. Ar ´ıs, D. Mira, H. Calmet, F. Cucchietti, H. Owen, A. Taha, E. D. Burness, J. M. Cela, and M. Valero, “Alya: Multiphysics Engineering Simulation Towards Exascale,”Journal of Computational Sciences, vol. 1...
2016 doi
-
[29]
A low-dissipation finite element scheme for scale resolving simulations of turbulent flows,
O. Lehmkuhl, G. Houzeaux, H. Owen, G. Chrysokentis, and I. Rodriguez, “A low-dissipation finite element scheme for scale resolving simulations of turbulent flows,”Journal of Computational Physics, vol. 390, pp. 51–65, 2019. [Online]. Available: https: //doi.org/10.1016/j.jcp.2...
2019 doi
-
[30]
Parallelization of 3d mpdata algorithm using many graphics processors,
K. Rojek and R. Wyrzykowski, “Parallelization of 3d mpdata algorithm using many graphics processors,” inParallel Computing Technologies, 2015, pp. 445–457. [Online]. Available: https://doi.org/10.1007/978-3 -319-21909-7 43
2015 doi
-
[31]
A study of the effect of process malleability in the energy efficiency on GPU-based clusters,
S. Iserte and K. Rojek, “A study of the effect of process malleability in the energy efficiency on GPU-based clusters,”Journal of Supercomputing, vol. 76, pp. 255–274, Oct. 2020. [Online]. Available: https://doi.org/10.1007/s11227-019-03034-x
2020 doi
-
[32]
Dynamic resource management in HPC systems using dynamic processes with PSets,
D. Huber, S. Iserte, M. Schreiber, P.-F. Dutot, O. Richard, A. J. Pe ˜na, K. Gaddameedi, and T. Neckel, “Dynamic resource management in HPC systems using dynamic processes with PSets,” inProceedings of the 32nd IEEE International Conference on High Performance Computing, Data,...
2025
-
[34]
A Layered Approach for Dynamic Resource Management in HPC,
H.-J. Bungartz, P.-F. Dutot, J. Fecht, K. Gaddameedi, D. Huber, S. Iserte, M. Minion, T. Neckel, A. Pe ˜na, O. Richard, M. Schreiber, M. Schulz, and V . Sch ¨uller, “A Layered Approach for Dynamic Resource Management in HPC,” inEuro-Par 2024: Parallel Processing Workshops: Eur...
2024 doi
-
[35]
Preparing mpich for exascale,
Y . Guo, K. Raffenetti, H. Zhou, P. Balaji, M. Si, A. Amer, S. Iwasaki, S. Seo, G. Congiu, R. Latham, L. Oden, T. Gillis, R. Zambre, K. Ouyang, C. Archer, W. Bland, J. Jose, S. Sur, H. Fujita, D. Durnov, M. Chuvelev, G. Zheng, A. Brooks, S. Thapaliya, T. Doodi, M. Garazan, S. ...
2025 doi
-
[2018]
Available: http://dx.doi.org/10.6035/14101.2018.176272
[Online]. Available: http://dx.doi.org/10.6035/14101.2018.176272
2018
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.