{"id":"3dc06745-283e-45bf-95ba-90f68db7e6df","arxiv_id":"2506.15418","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RISC-V remains impractical for large-scale HPC mainly because of memory bandwidth limits, absent high-performance networking, and tooling gaps, while application portability and Fortran support look adequate.","lead":"This conference position paper from the RISC-V HPC SIG reports that most HPC codes build and run on RISC-V CPUs, but the real blockers are memory subsystem limits, missing high-performance networking, and thin performance tooling. A generalist would read it for a concrete, community-written checklist of what must improve before RISC-V can challenge x86-plus-GPU systems in supercomputing.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evidence base for 'not yet realistic for large-scale HPC' rests on a single CPU and a two-row table; generalization needs a stated sampling frame.","rationale":"The reader's UNVERDICTED verdict already captures the two weaknesses I identify: generalization from the SG2042 alone and a thin, under-specified Fortran experiment. My stress-test pass does not find an additional load-bearing flaw; the paper is an extended abstract/position statement, not a hypothesis-testing research paper, and the action points (networking, tooling, memory subsystem) are plausible community-prioritization items even on the thin evidence. The most decision-relevant risk is that the headline 'not yet realistic for large-scale HPC' is an inductive generalization from one CPU generation and one benchmark table, and the manuscript explicitly defers the broader survey to the talk. That is a verifiability gap, not an internal inconsistency, so I agree with the reader's UNVERDICTED rather than moving to REJECT or ACCEPT. I also confirm the paper itself flags limitations (missing performance counters, early networking work, deferred survey data), and those flags are consistent with my reading rather than contradictory. A concrete settlement test would be to widen the hardware sample and report the survey data; if those land, the assessment could be upgraded to an evidence-backed position. No machine-checked proofs or shipped artifacts are claimed, so this is a documentation of 'insufficient evidence in the written record' rather than an identified technical error.","tokens_in":2928,"tokens_out":1569,"duration_ms":17247,"concrete_test":"Obtain the deferred application-survey data and expand the hardware sample: compile and run NPB (and one memory-bound proxy such as STREAM or GUPS) on at least two additional server-class RISC-V parts—e.g., SG2380 or another current high-core-count CPU—under identical GCC v13 flags, and report whether the SG2042 memory-bandwidth/latency gap and the Fortran-vs-C ratios reproduce. If a second part shows a materially better memory subsystem, the priority ordering and the 'not yet realistic' claim need to be re-scoped.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central action-point ordering—memory subsystem, networking, tooling—is presented as an ecosystem assessment, but the load-bearing evidence for the strongest limitation claim is [2], a characterization of the Sophon SG2042, plus Table 1's single-core GCC v13 NPB class C Fortran-vs-C comparison against AMD Rome. The interpretive step is 'SG2042 memory subsystem limits ⇒ RISC-V is not yet realistic for large-scale HPC.' That step requires that SG2042 is representative of the RISC-V server-class parts the community would deploy, or at least that no current part with a materially better memory subsystem exists. The paper states the SG2042 'is a potential contender' but does not report any attempt to establish representativeness, nor does it list which other CPUs were considered. Networking is asserted to be a 'major limitation' with one prototype Infiniband driver mentioned; no survey scope, deployment status of other interconnects, or reference is given. Fortran support is inferred from five benchmark ratios with no methodology: no error bars, no repetition count, no compiler flags, no baseline justification for why a per-benchmark ratio rather than a compiler or library check captures maturity. These gaps matter because the paper's own logic ('until we address this RISC-V will always be limited') depends on the generality. The conclusion may be right, and the talk may contain the deferred survey data, but the written artifact does not yet make the generalization testable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This extended abstract, submitted on behalf of the RISC-V HPC SIG, argues that RISC-V is not yet ready for large-scale HPC deployment. It identifies three main blockers: memory-subsystem performance on current RISC-V CPUs (specifically the Sophon SG2042), lack of mature high-performance networking, and sparse performance-analysis tooling. It reports that common HPC applications port without difficulty and that Fortran compiler support is not noticeably worse than on x86, using a small NPB benchmark comparison (Table 1). The paper concludes with action points for the community, prioritizing high-performance networking and memory-subsystem improvements.","tokens_in":3108,"tokens_out":4419,"duration_ms":46110,"significance":"The paper's value is that it frames an ecosystem-level assessment with specific, actionable recommendations and is honest about current limitations. If the SG2042 is representative of RISC-V server-class hardware, the identified blockers would indeed be the right priorities for the HPC community. The paper also contributes a useful data point in Table 1 to an ongoing debate about Fortran support. However, the evidence presented in the written artifact is narrow: the core generalizations rest on one CPU and a single small benchmark table with no experimental protocol, so the significance of the claims is currently limited by lack of a transparent sampling frame and methodology.","major_comments":[{"comment":"The central generalization that RISC-V is 'not yet a realistic proposition for large-scale HPC' is inferred from a characterization of a single CPU, the Sophon SG2042, cited from the author's own prior paper [2]. No evidence is given that the SG2042 is representative of the RISC-V server-class parts that the HPC community would deploy, nor is any list provided of other RISC-V CPUs that were considered and rejected. This matters because the priority ordering of the action points (memory subsystem before networking, tooling) depends on the generality of the SG2042 results. Please either provide a survey of current RISC-V HPC-capable CPUs with a comparison of their memory subsystems, or explicitly scope the claims to the SG2042.","section":"The state of RISC-V for HPC"},{"comment":"Table 1 is the only quantitative support for the claim that Fortran compiler support on RISC-V is not a concern, but it is presented without any experimental protocol: there are no repetition counts, no error bars or other variance measures, no statement of compiler flags or optimization levels, and only one baseline (AMD Rome) and one problem class (NPB class C) on a single core. As a result, the per-benchmark ratios (e.g., BT 3.35 on RISC-V versus 1.30 on x86) cannot be distinguished from noise or compiler-flag artifacts. The paper should either provide a documented methodology with enough repetitions to support the ratios, or soften the conclusion from 'Fortran support is comparable' to 'preliminary indications do not show a large gap'.","section":"Software tooling, Table 1"},{"comment":"The claim that high-performance networking is a 'major limitation' and that 'until we address this RISC-V will always be limited when it comes to large-scale HPC deployment' is asserted with only a single example (a prototype Infiniband driver) and no references to the current state of RISC-V networking support across vendors or interconnect technologies. This is a strong, falsifiable claim, but the written artifact provides no survey scope, deployment status, or citations. Please provide a documented overview of what networking support exists (e.g., Ethernet, Infiniband, Omni-Path, custom interconnects) on RISC-V hardware and in Linux drivers, and state how this was assessed.","section":"Infrastructure"},{"comment":"The statement that 'the majority of these built without issue on RISC-V CPUs and could run common use-cases' is a load-bearing part of the paper's claim that application portability is adequate, yet no data are shown: no list of applications, libraries, or benchmarks considered, no build results, and no description of what 'common use-cases' were run. This evidence would need to be at least summarized in the written artifact (or in an appendix) for the claim to be assessable, given that it is one of the few positive conclusions in the abstract.","section":"Application support"}],"minor_comments":[{"comment":"Typographical issues: 'maximimise' should be 'maximise', 'belif' should be 'belief', and 'and-so' appears with missing spaces in the Infrastructure section; 'T able 1' has a stray space.","section":"Throughout"},{"comment":"The caption should specify that times are wall-clock, the exact GCC v13 version, and the optimization flags used, and should indicate whether the ratios are means over repetitions or single-run values.","section":"Table 1 caption"},{"comment":"The references for Extrae and the claim about performance tooling are missing; please add citations so readers can verify the state of RISC-V tooling.","section":"Software tooling"},{"comment":"The paper would benefit from a short statement of methodology for the HPC SIG analysis, even in an appendix, so that the reader understands how the conclusions were assembled (e.g., number of systems surveyed, time period, criteria for inclusion).","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is an extended abstract, and some brevity is expected, but several of its central claims go beyond what the written evidence supports. The reliance on the author's own prior measurements ([1], [2]) is not itself a problem, but it means that independent verification or a broader survey is especially important. If the full talk contains the deferred survey data, the authors should be encouraged to make that material available in a companion technical report or supplement. I see the manuscript as a useful position statement that can be revised to match its evidence base."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nick's extended abstract is best read as a SIG position statement, not a research preprint. The one genuinely new piece is Table 1: a single-core NPB class C Fortran-versus-C comparison on the SG2042 against AMD Rome. The rest is a fair synthesis of his earlier SG2042 characterisation, the Grayskull stencil study, and widely known community gaps. That's not a criticism; the paper is upfront about being an update on behalf of the RISC-V HPC SIG, and the action points are sensible.\n\nWhat the paper does well: it states the central blockers clearly—memory subsystem on current CPUs, missing high-performance networking, and sparse performance tooling—and it ranks them in order of importance. The networking point is the strongest: the absence of deployable high-performance interconnect is a hard blocker for distributed-memory HPC, and the paper is right to make it the headline. The Table 1 experiment is a reasonable quick check on the claim that RISC-V Fortran tooling is a problem; the numbers suggest it is not dramatically worse than on x86. Credit where due: the author's cited prior measurements are legitimately reproduced or verifiable in the references, and self-citation here is not a sin.\n\nThe soft spots are exactly where the stress-test note hits. The conclusion that 'RISC-V is not yet realistic for large-scale HPC' is inferred from characterising one CPU, the SG2042, and there is no attempt to show that the SG2042 is representative of current or near-term server-class RISC-V parts. If a newer part has a materially better memory subsystem, the priority ordering changes. The paper needs a stated sampling frame: which CPUs were considered, why the SG2042 stands in for the ecosystem, and what else is on the horizon. Table 1 also lacks methodology: no repetition counts, no compiler flags, no error bars, and the choice of a per-benchmark Fortran/C ratio as a maturity proxy is questionable—it conflates compiler codegen with library alignment. The application survey is deferred to the talk, so the written artifact cannot be checked. These are real weaknesses, but proportionate to the genre: this is a two-page extended abstract, not a measurement paper.\n\nThe reader's UNVERDICTED tag is fair. I would not call this a research contribution, but it is a useful assessment document for the RISC-V HPC community, vendors, and system integrators. It deserves a serious referee as a position statement: the referee should ask for the survey data and a broader hardware baseline, not for a new theory. I'd accept it for peer review at a workshop or conference, and would welcome it in a reading group as a prompt for discussion about what evidence would be needed to make the ecosystem claim testable.","headline":"A sensible SIG position statement with one small new measurement; the load-bearing generalization from a single CPU is the main soft spot.","tokens_in":3713,"tokens_out":2388,"would_cite":false,"duration_ms":25404,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims RISC-V is not yet realistic for large-scale HPC, identifies missing high-performance networking as the top priority, and argues that application portability is not the main problem.","keywords":["RISC-V","high performance computing","SG2042","memory subsystem","high-performance networking","performance tooling","Fortran","NPB benchmarks"],"falsifier":"Run NPB class C memory-bound benchmarks on a newer RISC-V CPU with a redesigned memory subsystem; if it matches or beats x86 at comparable core count, the paper's memory-subsystem blocker would no longer hold. Separately, deploy a small RISC-V cluster with a working high-performance interconnect and run distributed-memory benchmarks; if communication bandwidth and latency reach x86 levels, the claim that networking makes RISC-V unrealistic for large-scale HPC would need revision.","tokens_in":2639,"feed_emoji":"🖥️","tokens_out":8877,"duration_ms":96724,"temperature":0.7,"pith_summary":"The paper, an extended abstract prepared by the RISC-V high-performance-computing community group, argues that RISC-V has made real progress but is not yet a realistic proposition for large-scale HPC. Its central claim is that the blockers are not application portability or Fortran compiler maturity: most common HPC codes build and run on RISC-V, and a Fortran-versus-C comparison on the NPB benchmarks gives ratios comparable to x86. Instead, the main limits are the memory subsystem of current RISC-V CPUs, sparse performance-analysis tooling, and, most urgently, the absence of deployable high-performance networking. The paper's action points follow directly: treat networking as the top priority, improve memory-subsystem performance, encourage porting of performance tools, and mature a RISC-V vectorisation dialect. A reader should care because these are concrete, addressable gaps rather than a fundamental portability wall.","feed_headline":"RISC-V for HPC: networking, memory, and tooling are the blockers","feed_subtitle":"Applications port easily, but missing high-speed networking and memory performance keep RISC-V out of large HPC.","key_machinery":"The load-bearing mechanism is a gap analysis built on two comparisons. First, the previously published performance characterisation of the 64-core SG2042 separates compute-bound from memory-bound workloads and shows that the CPU matches server-class performance on the former and loses on the latter; this turns 'RISC-V is slow' into a specific memory-subsystem deficit. Second, Table 1 compares Fortran and C versions of the NPB benchmarks (class C, GCC v13, single core) on the SG2042 and an x86 baseline; the comparable ratios are used to rule out Fortran compiler maturity as a blocker. The paper's prioritisation of action points is then inferred by matching each observed gap to a community effort: networking, memory subsystems, tooling ports, and a vectorisation dialect.","core_discovery":"The author's core discovery is that RISC-V's path to HPC is blocked by ecosystem infrastructure, not by the ability to compile and run scientific code. In testing driven by the most popular HPC applications, libraries, and benchmarks, the majority built and ran without issue on RISC-V CPUs, and the NPB Fortran-versus-C speedup ratios on a single core of a current RISC-V CPU were broadly comparable to an x86 server, with the BT benchmark actually substantially better on RISC-V. The same body of work found that the 64-core SG2042 processor performs well on compute-bound workloads, matching a previous-generation server CPU core-for-core and beating lower-core-count x86 chips at full core count, but falls far behind on memory-bandwidth- and latency-bound workloads. The paper also identifies performance-analysis tooling as nearly missing and hardware performance counters as incomplete. Taken together, the diagnosis is that until high-performance networking matures, RISC-V cannot realistically support distributed-memory parallelism at scale, and that memory-subsystem and tooling improvements are the other high-value action points.","pith_inferences":["Editorial inference: the paper's conclusions rest on a single current-generation CPU, the SG2042, and one compiler-benchmark comparison; a broader sweep across newer RISC-V chips with different memory subsystems could change the priority ordering.","Editorial inference: Fortran-versus-C speedup ratios on a handful of NPB benchmarks measure compiler differences but not absolute Fortran maturity; a fuller test suite with different compiler flags and optimisation levels would be needed to confirm the 'adequate support' conclusion.","Editorial inference: the paper implies, without testing, that accelerating RISC-V adoption via pragma-based porting will follow the GPU route; this is testable by porting a real GPU-oriented HPC code to a RISC-V accelerator and measuring the developer cost.","Editorial inference: if networking becomes mature, the next bottleneck could shift to parallel filesystems and system-administration tooling, which the paper lists as secondary but does not measure."],"forward_implications":["If the paper's diagnosis is right, adding mature high-performance networking support is a necessary condition for RISC-V to enter large-scale HPC, and no amount of CPU performance or application portability can substitute for it.","Next-generation RISC-V CPUs that improve memory bandwidth and latency will matter more for HPC than increased vector width alone, because the current bottleneck is memory, not vectorisation.","The 'Fortran support is behind' assumption should drop out of the community's priority list; efforts are better spent on tooling and infrastructure.","Porting existing performance-analysis tools and adding complete hardware performance counters would immediately help HPC developers evaluate RISC-V systems.","A RISC-V vectorisation dialect in the compiler-framework stack would enable optimisations analogous to those already available for other target architectures."],"supporting_citations":[{"why":"Supplies the comparison of a RISC-V accelerator against a 24-core x86 CPU on a scientific workload, used to argue that RISC-V specialisation can exceed incumbents.","marker":"[1]"},{"why":"Supplies the SG2042 performance characterisation (compute-bound versus memory-bound) from which the paper draws its memory-subsystem limitation conclusion.","marker":"[2]"}],"fun_headline_variants":["RISC-V for HPC: not CPUs, but networking, memory, and tools","RISC-V HPC: apps port fine, but networking and memory lag","RISC-V's HPC bottleneck: missing networking, memory, and tooling","RISC-V for HPC held back by ecosystem, not compute","RISC-V HPC: compute is fine, but networking and memory are not"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that RISC-V is not yet ready for large-scale HPC rests on treating the SG2042 as representative of current RISC-V hardware and on reading Fortran-versus-C speedup ratios as a measure of compiler maturity.","fun_headline_variants_meta":{"raw":{"variants":["RISC-V for HPC: not CPUs, but networking, memory, and tools","RISC-V HPC: apps port fine, but networking and memory lag","RISC-V's HPC bottleneck: missing networking, memory, and tooling","RISC-V for HPC held back by ecosystem, not compute","RISC-V HPC: compute is fine, but networking and memory are not"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1449,"prompt_tokens":812,"completion_tokens":637,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":548}},"tokens_in":428,"tokens_out":637,"duration_ms":6269,"temperature":1.0,"reasoning_tokens":548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:56:29.276103+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run NPB class C memory-bound benchmarks on a newer RISC-V CPU with a redesigned memory subsystem; if it matches or beats x86 at comparable core count, the paper's memory-subsystem blocker would no longer hold. Separately, deploy a small RISC-V cluster with a working high-performance interconnect and run distributed-memory benchmarks; if communication bandwidth and latency reach x86 levels, the claim that networking makes RISC-V unrealistic for large-scale HPC would need revision.","supporting_citations":[{"cited_title":"Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator","cited_arxiv_id":null,"evidence_quote":"Supplies the comparison of a RISC-V accelerator against a 24-core x86 CPU on a scientific workload, used to argue that RISC-V specialisation can exceed incumbents."},{"cited_title":"Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC","cited_arxiv_id":null,"evidence_quote":"Supplies the SG2042 performance characterisation (compute-bound versus memory-bound) from which the paper draws its memory-subsystem limitation conclusion."}],"review_version":1}