REVIEW 4 major objections 5 minor 26 references
How scarce research compute is allocated is policy, not just engineering—and three gaps decide whether federation helps or harms.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 12:07 UTC pith:KBWOTALA
load-bearing objection Solid UK-facing roadmap: measure occupancy vs utilisation, tune Slurm before replacing it, and don’t open unmetered federation without placement rules—strongest on the first two, sandbox-thin on Braess-style local harm. the 4 major comments →
FAIR-Compute: A Roadmap for Fair and Efficient Allocation of Federated Digital Research Infrastructure
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across landscape evidence, a strategic-reporting model, and simulations, three results recur: allocation data measure occupancy not utilisation, so productive value is unanswerable; tuned transparent heuristics approach an offline full-information welfare benchmark, so tune before replacing; and unmanaged federation minimises average delay like uncoordinated routing while concentrating load—and potential harm—on receiving systems, requiring coordinated placement rules.
What carries the argument
A descriptive Slurm-style priority scheduler paired with an offline full-information welfare benchmark (weighted flow time / value) and a federation-as-congestion-game lens: placement policies from naive lowest-completion routing through switching-cost tolls, minimum-gain thresholds, and utilisation protection, scored on efficiency, delay, and fairness under actual versus requested walltime occupation.
Load-bearing premise
That synthetic high-contention traces built from low-contention public US months, on two abstract systems without real data movement, are enough to ground UK federation policy—including the warning about local harm from selfish placement.
What would settle it
Replay the same scheduling and federation policies on a high-contention, pseudonymised real UK multi-site job log and check whether tuned heuristics still sit near the offline benchmark and whether uncoordinated placement still concentrates load without protecting local users.
If this is right
- A mandatory shared job-telemetry schema is the foundation: without it, efficiency KPIs, ex-post verification, and value-aware pilots cannot be evaluated.
- Operators should publish multi-objective Slurm weight guidance rather than mandate new schedulers.
- Portable allocation needs domain-workload exchange rates before open cross-system mobility.
- Federated job movement should use minimum-gain, utilisation ceilings, or switching tolls instead of pure greedy routing.
- Occupancy-versus-utilisation reporting would expose over-reservation and support energy and carbon tracking per unit of research output.
Where Pith is reading between the lines
- If measurement is truly foundational, delaying a shared schema will quietly block every later mechanism the roadmap proposes, including value-aware priority and carbon accounting.
- The Braess-style analogy implies that adding more free cross-site channels without tolls could worsen local service even as global average wait falls—testable once real multi-site traces exist.
- Incentive-compatible urgency reporting is the unsolved hinge for any production value-aware pilot; without it, self-declared deadlines will recreate walltime padding in a new form.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FAIR-Compute frames UK federated DRI allocation as a joint mechanism-design, transport-economics, and scheduling problem. Combining a landscape review (EuroHPC, WLCG, ACCESS, JASMIN, DiRAC), stakeholder interviews, an N=10 demand-side survey, a strategic-reporting model, and simulations on Fresco/Anvil plus synthetic stress tests, it advances three recurring claims: (i) current records measure occupancy not utilisation, so productive value cannot be assessed; (ii) multi-objective-tuned transparent Slurm-like heuristics approach an offline full-information welfare ILP, so near-term gains come from tuning and instrumentation rather than replacement; (iii) unmanaged federation behaves like uncoordinated selfish routing—cutting average delay while concentrating load on receivers—so coordinated placement (tolls, minimum-gain, utilisation protection) is needed before unmetered mobility. Eight FAIRC roadmap recommendations follow, with measurement (FAIRC-1/2) as the foundation.
Significance. If the three recurring results hold under real UK contention, the paper supplies a concrete, sequenced NFCS roadmap that correctly prioritises measurement and low-risk scheduler tuning over speculative scheduler replacement, and that imports a useful congestion-game lens into federated HPC policy. Strengths include explicit triangulation across operator testimony and simulation, honest labelling of synthetic sandbox results and ILP optimality gaps, a public federation simulator (Appendix F), and recommendations written in implementable Must/Should/Could form with FTE/cost estimates. The occupancy-versus-utilisation and “tune before replace” strands are the most transferable contributions; the federation-coordination imperative is the highest-stakes and least externally validated claim.
major comments (4)
- [§5.5–5.6, FAIRC-5, Tables 11–13, Appendix D] §5.5–5.6 and FAIRC-5: The load-bearing claim that unmanaged federation requires coordinated placement before unmetered mobility rests on Experiment 4 alone—two abstract 300-node systems, 30-day synthetic traces, a fixed switching-cost stand-in, no inter-site data movement or authentication, and single-window point estimates. The text itself states that local-user harm is not directly quantified and that naive currently looks best on aggregate delay (Table 11–13; Appendix D). Either add a direct local-user slowdown/Jain metric on the receiving system’s home jobs under naive vs coordinated rules, or downgrade FAIRC-5 from a coordination imperative to a pilot contingent on FAIRC-1 UK-trace validation. As written, the Braess analogy is directional evidence, not yet a policy proof.
- [§4.3, §5.3, Table 7] §4.3 and Table 7 (Experiment 2): The offline ILP is the welfare benchmark against which “heuristics are good enough” is judged, yet the solver often terminates with a nonzero optimality gap and heuristics sometimes match or beat it on raw wait. The caveat in §4.3 is acknowledged but under-used in the executive claim. Report the optimality gap (or best-bound gap) beside every offline row, and restate the claim as “tuned heuristics are competitive with the best feasible offline solution found,” not as proximity to a proven full-information optimum. Without that, the “tune don’t replace” headline over-reaches the computation.
- [§5.1, §5.6, Tables 5–8] §5.1 and §5.6: All contention results that drive Experiments 1–3 and the federation stress tests use synthetic bursty-fill traces derived from low-contention Anvil months (real-trace utilisation pinned ≈0.25 in Table 8). Volume multipliers and fill injection are free parameters. Sensitivity of the Pareto weight sets (Table 5–6), the heuristic–ILP gap (Table 7), and transfer volumes (Table 16) to the fill multiplier and to alternative cluster counts should be reported; otherwise the policy levers and “close to offline” conclusions are tied to one sandbox calibration. This is the external-validity hinge the authors already flag as motivating FAIRC-1—make the dependence quantitative in the main text.
- [§4.2, §5.4, Appendix A, FAIRC-8/3] §4.2 and Experiment 3: The strategic-reporting utility (Eq. 2) and penalty mechanisms (Appendix A) are presented as the analytical spine for FAIRC-8 and FAIRC-3, but “the strategic-reporting mechanism itself is not exercised in the present simulations.” Experiment 3 shows runtime externalities and value loss under value-blind rules, which is related but not a test of incentive-compatible walltime or urgency reporting. Either run a minimal strategic-padding experiment (users best-responding to kill-at-walltime / fairshare penalties) or clearly separate the mechanism-design programme as future work so that FAIRC-8 is not read as simulation-backed when it is practice-codification plus untested theory.
minor comments (5)
- [Executive Summary, §2.2] N=10 survey (§2.2) is correctly labelled exploratory, but the Executive Summary’s “three strands of evidence” phrasing can be read as equal weight. Soften survey citations in the Exec. Summary to “corroborative operator-facing themes.”
- [Figures 5, 7; Tables 9–13] Figure 5 and Figure 7 need axis units and a one-line reading guide in the caption; several tables report point estimates without stating the job-window size in the table note (only in §5.6).
- [Appendix C] QoS-tier proxy cut-points (Appendix C, Table 14) are ad hoc; a one-sentence robustness note that conclusions are insensitive to cut-points would help, as claimed in the appendix but not shown.
- [§6] Recommendation numbering jumps FAIRC-1,2,4,7,6,8,3,5 in implementation order—fine for a roadmap, but add a one-line dependency diagram or explicit “depends on” column so readers do not assume numeric order is priority order.
- [§1, throughout] Minor typos/consistency: “becomespolicy” spacing in the Intro; “use-it-or-lose-it” hyphenation varies; arXiv date and “Submitted July 2026” are fine but ensure reference access dates are uniform.
Circularity Check
No significant circularity: empirical/simulational comparisons against independent benchmarks, not definitions restated as predictions.
full rationale
FAIR-Compute’s load-bearing claims are landscape triangulation plus controlled simulation comparisons, not a first-principles derivation that collapses into its inputs. The offline welfare benchmark (weighted flow-time/value ILP, §4.3 Eq. 3) is an independent full-information reference; FIFO/Slurm-like/tuned heuristics are scored against it rather than fitted to reproduce it (§5.3). Multi-objective weight search (NSGA-II) maps trade-offs among utilisation, wait, slowdown, and fairness; it does not claim to predict a quantity already used as the fit target (§5.2). QoS tiers are explicitly labelled inferred ACCESS-style proxies from annualised usage, with sensitivity disclaimed (Appendix C). Federation results compare placement policies under synthetic contention and invoke Braess’s paradox only as an interpretive congestion-game analogy, not as a uniqueness theorem or self-cited forced result (§5.5–5.6). Strategic-reporting utility (Eq. 2, Appendix A) is mechanism-design framing and is not exercised as a circular ‘prediction’ in the simulations. External citations (Slurm, ACCESS, WLCG/HEPScore, pymoo, PyJobShop) are operational or methodological scaffolding, not author-overlapping uniqueness imports that forbid alternatives. Weak external validity of the synthetic sandbox is a correctness/generalisation issue, not circularity by construction.
Axiom & Free-Parameter Ledger
free parameters (6)
- Slurm-like priority weights (age, fairshare, size, QoS, ...) =
Pareto set on synthetic filled workload (Table 5); single-objective runs in Table 6
- Federation switching-cost toll =
fixed + 10% of runtime (Appendix D)
- Minimum-gain threshold θ and utilisation ceiling U_max
- ACCESS-style QoS-tier cut-points from annualised usage =
≤4e5; (4e5,1.5e6]; (1.5e6,3e6]; >3e6 core-hour proxy (Table 14)
- Synthetic workload volume multipliers and bursty-fill injection
- Linear waiting penalty κ_j and job values V_j in welfare/utility model =
Experiment 3: 2388 value units over 500 jobs
axioms (6)
- domain assumption Production schedulers rank jobs by a weighted linear priority score over age, fairshare, QoS, size, partition (Eq. 1), with backfill of smaller jobs that do not delay the head job.
- domain assumption Users face uncertain runtimes and choose reported walltime as self-selected insurance trading kill/preempt risk against queue delay and ex-post penalties (Eq. 2).
- domain assumption Offline full-information scheduling maximising sum (V_j − κ_j W_j) is a meaningful welfare benchmark for practical rules even when the solver returns a small optimality gap.
- ad hoc to paper Uncoordinated cross-system placement is analogous to unpriced selfish routing in congestion games, so Braess-like concentration risks motivate coordinated tolls/thresholds.
- ad hoc to paper Synthetic traces that preserve Anvil job-structure clusters but amplify contention are valid instruments for evaluating allocation policy where real months are low-contention.
- domain assumption Occupancy-based accounting systematically diverges from productive utilisation and is the right efficiency gap for national KPIs.
invented entities (3)
-
FAIRC-1..8 recommendation bundle (unified schema, occ/util KPI, PVR pilot, etc.)
no independent evidence
-
Projected Value Remaining (PVR) prioritisation measure
no independent evidence
-
Inferred ACCESS-style QoS-tier proxy per account
no independent evidence
read the original abstract
As demand for high-performance computing (HPC), high-throughput computing and data storage grows, the way scarce compute is allocated -- not just how much exists -- has become a decisive factor in the productivity of UK research. FAIR-Compute studies allocation in a federated Digital Research Infrastructure (DRI) as a problem at the intersection of algorithmic game theory, transport economics and HPC scheduling. Our central observation is simple: once a shared system must decide who runs, when and under what evidence, its scheduler settings cease to be a purely technical matter and become policy. The project combined three strands of evidence: a landscape review of UK and international allocation practice (including EuroHPC, WLCG, ACCESS, JASMIN and DiRAC) supported by a stakeholder survey; a mechanism-design model of allocation under strategic and uncertain user reports; and a simulation study built on the public Fresco/Anvil workload trace and controlled synthetic stress tests. Three results recur across all three strands. First, allocation records measure occupancy (resources reserved) rather than utilisation (useful work done), so the system cannot currently answer the question UKRI and DSIT most want answered -- whether resources deliver productive value. Second, simple, transparent scheduling heuristics perform close to an offline full-information benchmark once tuned, so the near-term opportunity is to tune and instrument existing schedulers rather than replace them. Third, federation is genuinely valuable but behaves like uncoordinated routing when left unmanaged: it can minimise average delay while quietly concentrating load -- and potentially harm -- on the receiving system. This mirrors a classic effect from road-traffic economics (Braess's paradox), where letting everyone independently pick the fastest route can leave the whole network worse off.
Figures
Reference graph
Works this paper leans on
-
[1]
ACCESS Project Types.https://allocations.access-ci.org/ project-types, 2026
ACCESS. ACCESS Project Types.https://allocations.access-ci.org/ project-types, 2026. Accessed: 2026-07-01
2026
-
[2]
ACCESS Resources.https://allocations.access-ci.org/resources, 2026
ACCESS. ACCESS Resources.https://allocations.access-ci.org/resources, 2026. Accessed: 2026-07-01
2026
-
[3]
Pymoo: Multi-objective optimization in python.Ieee access, 8:89497–89509, 2020
Julian Blank and Kalyanmoy Deb. Pymoo: Multi-objective optimization in python.Ieee access, 8:89497–89509, 2020
2020
-
[4]
White paper on China’s computing power development index
China Academy of Information and Communications Technology (CAICT). White paper on China’s computing power development index. English translation (2024), Center for Security and Emerging Technology, 2022.https://cset.georgetown.edu/wp-content/ uploads/t0581_china_compute_index_2022_EN.pdf
2024
-
[5]
A fast and elitist multiobjective genetic algorithm: Nsga-ii.IEEE transactions on evolutionary computation, 6(2):182–197, 2002
Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and TAMT Meyarivan. A fast and elitist multiobjective genetic algorithm: Nsga-ii.IEEE transactions on evolutionary computation, 6(2):182–197, 2002
2002
-
[6]
The impact of epsrc’s investments in high performance comput- ing infrastructure: Final report
EPSRC. The impact of epsrc’s investments in high performance comput- ing infrastructure: Final report. Technical report, UK Research and In- novation, 2019.https://www.ukri.org/wp-content/uploads/2022/07/ EPSRC-050722-ImpactEPSRCInvestmentsHighPerformanceComputingInfrastructure. pdf
2019
-
[7]
Federation of computing infrastructures
ETP4HPC. Federation of computing infrastructures. Technical report, European Technol- ogy Platform for High-Performance Computing, 2025
2025
-
[8]
Eurohpc federation platform.https://www.eurohpc-ju
EuroHPC Joint Undertaking. Eurohpc federation platform.https://www.eurohpc-ju. europa.eu/supercomputers/eurohpc-federation-platform_en, 2026. Accessed: 2026- 06-30
2026
-
[9]
FRESCO: Open repository and analysis of system usage data.https: //www.frescodata.xyz/, 2026
FRESCO Project. FRESCO: Open repository and analysis of system usage data.https: //www.frescodata.xyz/, 2026. Accessed: 2026-05-04
2026
-
[10]
The economic impact of NVIDIA cambridge-1, 2021
Frontier Economics. The economic impact of NVIDIA cambridge-1, 2021. Estimated economic value approx. £600M (approx. $831M) over ten years
2021
-
[11]
HEPScore: A new CPU benchmark for the WLCG
Domenico Giordano et al. HEPScore: A new CPU benchmark for the WLCG. InEPJ Web of Conferences (CHEP), 2023. Adopted by WLCG April 2023, replacing HEP-SPEC06; reproducible to better than 1% across x86 and ARM.https://arxiv.org/abs/2306. 08118. 29 FAIR-Compute NFCS Flexible Fund — Final Report
2023
-
[12]
Overview of HPCI.https://www
High Performance Computing Infrastructure (HPCI). Overview of HPCI.https://www. hpci-office.jp/en/about_hpci/what_is_hpci, 2024. Accessed: 2026-07-17
2024
-
[13]
Open ondemand: A web-based client portal for hpc centers
David Hudak, Doug Johnson, Alan Chalker, Jeremy Nicklas, Eric Franz, Trey Dockendorf, and Basil L McMichael. Open ondemand: A web-based client portal for hpc centers. Journal of Open Source Software, 3(25):622, 2018
2018
-
[14]
Leon Lan and Joost Berkhout. Pyjobshop: Solving scheduling problems with constraint programming in python.arXiv preprint arXiv:2502.13483, 2025
Pith/arXiv arXiv 2025
-
[15]
SHAP: Shapley additive explanations documen- tation.https://shap.readthedocs.io/en/latest/, 2026
Scott Lundberg and SHAP contributors. SHAP: Shapley additive explanations documen- tation.https://shap.readthedocs.io/en/latest/, 2026. Accessed: 2026-05-12
2026
-
[16]
Fresco: A public multi-institutional dataset for understanding hpc system behavior and depend- ability
Joshua McKerracher, Preeti Mukherjee, Rajesh Kalyanam, and Saurabh Bagchi. Fresco: A public multi-institutional dataset for understanding hpc system behavior and depend- ability. InPractice and Experience in Advanced Research Computing 2025: The Power of Collaboration, pages 1–6. 2025
2025
-
[17]
PyJobShop Documentation.https://pyjobshop.org/stable/ index.html, 2026
PyJobShop Developers. PyJobShop Documentation.https://pyjobshop.org/stable/ index.html, 2026. Accessed: 2026-07-01
2026
-
[18]
pymoo: Multi-objective Optimization in Python.https://pymoo
pymoo Developers. pymoo: Multi-objective Optimization in Python.https://pymoo. org/, 2026. Accessed: 2026-07-01
2026
-
[19]
Depei Qian and Zhongzhi Luan. High performance computing development in China: A brief review and perspectives.Computing in Science & Engineering, 21(1):6–16, 2018. doi:10.1109/MCSE.2018.2875367
arXiv 2018
-
[20]
Slurm workload manager documentation, 2026
SchedMD. Slurm workload manager documentation, 2026. URLhttps://slurm.schedmd. com/documentation.html. Accessed: 2026-03-14
2026
-
[21]
Anvil-system architec- ture and experiences from deployment and early user operations
X Carol Song, Preston Smith, Rajesh Kalyanam, Xiao Zhu, Eric Adams, Kevin Colby, Patrick Finnegan, Erik Gough, Elizabett Hillery, Rick Irvine, et al. Anvil-system architec- ture and experiences from deployment and early user operations. InPractice and experience in advanced research computing 2022: Revolutionary: Computing, connections, you, pages 1–9. 2022
2022
-
[22]
Met office supercomputing 2020+ programme: Accounting officer assess- ment
UK Government. Met office supercomputing 2020+ programme: Accounting officer assess- ment. Technical report, Government Major Projects Portfolio, GOV.UK, 2022. Net present social value £13.74bn (25% optimism bias); 9:1 cost–benefit ratio against the do-nothing option
2020
-
[23]
National compute ecosystem town hall
UK Research and Innovation. National compute ecosystem town hall. UKRI Digital Re- search Infrastructure programme, 2026. June 2026
2026
-
[24]
Waldur: Cloud and hpc management platform.https://waldur.com/, 2026
Waldur. Waldur: Cloud and hpc management platform.https://waldur.com/, 2026. Accessed: 2026-06-30
2026
-
[25]
Memorandum of understanding for the worldwide lhc comput- ing grid.https://wlcg.web.cern.ch/organisation-mou/sample-mou, 2020
WLCG Collaboration. Memorandum of understanding for the worldwide lhc comput- ing grid.https://wlcg.web.cern.ch/organisation-mou/sample-mou, 2020. Accessed: 2026-06-30
2020
-
[26]
Toward dynamically controlling slurm’s classic fairshare algorithm
Jason Yalim. Toward dynamically controlling slurm’s classic fairshare algorithm. InPrac- tice and Experience in Advanced Research Computing 2020: Catch the Wave, pages 538–
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.