{"id":"25b1441f-ea08-4149-823a-b8f84906ce87","arxiv_id":"2603.06060","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":3.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"An update survey of stochastic rounding (2022–2026) that centers limited-precision SR, commercial hardware, probabilistic error bounds, and applications in ML and scientific computing.","lead":"This paper surveys nearly four years of progress on stochastic rounding after a 2022 review, with emphasis on limited-precision variants that fix the random-bit width. It matters because major AI hardware vendors now ship SR and the survey maps the theory, implementations, and next standardization steps.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies the paper as a timely, well-structured survey whose value is synthesis rather than original proof or measurement. The strongest claim is modest and well-supported by the cited error analyses (Sections 4, [16]) and the public ISA/patent literature (Section 5, Table 1). The only material caveat—the secondary character of the hardware sources—is already noted by the reader and does not rise to a load-bearing flaw for a survey of this type. No deeper technical soft spot (e.g., an unstated assumption in the limited-precision error model, or a contradiction between theory and the reported hardware) is present. Consequently the ACCEPT verdict with high confidence remains appropriate; no adjustment is warranted.","tokens_in":12905,"tokens_out":458,"duration_ms":35175,"concrete_test":"Spot-check three random-bit widths in Table 1 against the cited primary documents (Graphcore Tile Vertex ISA Sec. 2.14.6, AMD MI300 ISA pp. 362-363, NVIDIA PTX ISA Sec. 9.7.9.21); if any entry is materially wrong the hardware summary weakens, otherwise the survey’s organizational claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is an update survey whose central claim is organizational: limited-precision SR (SR_p^r = SR_p(fl_{p+r})) is the practically relevant new variant, with r ~ ceil((log2 n)/2) recovering the O(sqrt(n)u) probabilistic bounds of exact SR for summation and related algorithms (citing El Arar et al. [16,17] and the martingale/variance analyses), while commercial hardware already ships concrete realizations (Graphcore, AMD MI300, NVIDIA PTX, etc., Table 1). As a literature synthesis it does not assert original theorems or new silicon measurements; the secondary-source nature of the hardware descriptions is already flagged by the reader and is inherent to the genre. No internal inconsistency, hidden assumption, or unsupported leap undermines the claim that limited-precision SR is the focus of recent theory and industrial activity.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This manuscript is an update survey of stochastic rounding (SR) covering roughly 2022–2026. It organizes recent progress around limited-precision SR, defined as SR_p^r(x) = SR_p(fl_{p+r}(x)), which approximates the ideal probabilities of exact SR while remaining implementable. The paper reviews probabilistic error analyses that recover O(√n u) bounds for summation and related algorithms when r is chosen near ⌈(log2 n)/2⌉, summarizes commercial hardware realizations (Graphcore, AMD MI300, NVIDIA PTX/Blackwell, Tesla, Google TPU) and research/patent activity, and surveys applications in mixed-precision ML training, neuromorphic computing, climate simulation, and scientific computing. It also covers the IEEE P3109 interim variants (StochasticA/B/C) and software emulation packages.","tokens_in":13135,"tokens_out":802,"duration_ms":10432,"significance":"As a timely literature synthesis the paper is valuable. Limited-precision SR is the practically relevant variant now appearing in vendor ISAs and ML training stacks; collecting the random-bit widths (Table 1), the P3109 bias-complexity trade-offs, the martingale/variance analyses, and the application evidence into one place is useful for both numerical analysts and hardware designers. The survey does not claim original theorems or silicon measurements; its contribution is organizational and bibliographic, which is appropriate for the genre and for an update to Croci et al. (2022). Strengths include the clear exact-vs-limited-precision distinction (Fig. 1 and Eq. (1)), the concrete hardware table, and the breadth of application coverage.","major_comments":[],"minor_comments":[{"comment":"Section 5 / Table 1: the caption and body correctly note that the bit-widths come from public ISA documents and patents. A single explicit sentence that these are secondary sources (no independent silicon measurements) would make the inherent limitation of the genre fully transparent to readers.","section":null},{"comment":"Section 4: the heuristic r ≈ ⌈(log2 n)/2⌉ is attributed to El Arar et al. [16,17]. A one-sentence pointer to the precise statement (or theorem number) in those papers would help readers locate the supporting analysis without hunting.","section":null},{"comment":"Scattered typographical issues: missing spaces after commas/periods in several places (e.g., “round-to-nearest(RN)”, “fl +r(x)” rendering), and a few incomplete sentences near the end of Section 5 (Huawei/Google paragraphs). These are purely presentational.","section":null},{"comment":"References [17] is listed as “in preparation”; if it remains unpublished at acceptance, consider citing the arXiv version or noting the status more prominently so readers know the supporting analysis is not yet peer-reviewed.","section":null},{"comment":"Section 9 (ML): the discussion of 1-D vs 2-D scaling and double quantization is dense. A short clarifying sentence on why SR is preferred only for gradients (and not forward activations) would improve accessibility for non-ML readers.","section":null}],"recommendation":"accept","confidential_remarks":"The manuscript is a clean, useful update survey. The secondary-source nature of the hardware claims is inherent to the genre and already visible in the text; it does not warrant major revision. Fit for a numerical-analysis or computer-arithmetic venue is good. No novelty or citation-pattern concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean, useful update to Croci et al. 2022. The one thing worth knowing is that the authors put limited-precision SR (SR_p^r = SR_p(fl_{p+r})) front and center, show why r near ceil((log2 n)/2) recovers the O(sqrt(n)u) probabilistic bounds for summation and related algorithms, and then map that idea onto what Graphcore, AMD MI300, NVIDIA PTX, IEEE P3109, Tesla, and a handful of patents actually ship or propose. That is the practical story of the last four years, and they tell it clearly.\n\nWhat is new is the synthesis and the emphasis, not a theorem or a silicon measurement. They pull together the martingale and variance-based error analyses (including their own published work on limited-precision SR), the three P3109 variants, the concrete random-bit widths in Table 1, the research FPGA designs, the patent activity, the software emulators, and the application threads in LLM mixed-precision training, neuromorphic, climate, and scientific computing. Definitions are careful, citations are dense and mostly external, and the hardware table is the kind of reference people will actually keep open. As a survey it is well-organized and fair about what is exact versus limited-precision.\n\nSoft spots are minor and inherent to the genre. Hardware claims rest on public ISAs, white-papers, and patents; no independent silicon measurements. Self-citations to the limited-precision analyses exist, but those papers are already out and the survey covers a broad external literature. No original result is claimed, so there is nothing to over-sell. The “next steps for hardware” close is light, which is fine for an update.\n\nThis is for people who need a current map of SR theory, hardware, and ML/climate use cases. Numerical analysts, hardware architects, and anyone implementing low-precision training will get value from it. It deserves a serious referee; a desk reject would be a mistake. I would engage with it, cite the hardware table and the limited-precision framing, and bring it to reading group if we are talking low-precision arithmetic.","headline":"Solid, timely update survey that organizes limited-precision SR and the 2022–2026 hardware/theory wave; useful synthesis, not a new theorem paper.","tokens_in":13680,"tokens_out":532,"would_cite":true,"duration_ms":5757,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65G50","65Y04","68M07"],"pacs":[],"model":"grok-4.5","headline":"Limited-precision stochastic rounding recovers the sqrt(n) error growth of exact SR while matching what commercial chips already ship.","keywords":["stochastic rounding","limited-precision SR","probabilistic error bounds","floating-point hardware","mixed-precision training","stagnation","IEEE P3109","low-precision arithmetic"],"falsifier":"Measure the actual random-bit widths and rounding bias of the SR conversion instructions on AMD MI300, NVIDIA Blackwell and Graphcore devices; if the observed bit counts or bias diverge systematically from the ISA tables, the claim that commercial hardware realises the analysed limited-precision model fails.","tokens_in":13838,"feed_emoji":"🎲","tokens_out":774,"duration_ms":8990,"temperature":0.7,"pith_summary":"This survey updates the state of stochastic rounding (SR) since 2022, with primary attention on limited-precision SR: first round the exact value to a modest number of extra bits, then apply the usual probabilistic rounding. Exact SR makes the expected value of a rounded number equal the original and yields summation error that grows only like the square root of the number of terms with high probability, rather than linearly; limited-precision SR approximates those probabilities closely enough that the same probabilistic bounds remain valid when the extra bit-width is chosen near half the log of problem size. The paper shows that this is the form already appearing in Graphcore, AMD, NVIDIA, Tesla and Google hardware, in IEEE P3109 drafts, and in recent error analyses of summation, inner products, Horner evaluation, gradient descent and LLM training. It further documents patents, FPGA designs and software libraries that implement the same idea, and applications that use it to fight stagnation in climate models, neuromorphic plasticity and mixed-precision training. The practical message is that limited-precision SR is no longer a theoretical curiosity; it is the concrete variant that industrial hardware and numerical analysis have converged upon, and the remaining work is standardisation of random-bit widths and wider availability of the instructions.","feed_headline":"Limited-precision SR already matches commercial chips","feed_subtitle":"With roughly (log n)/2 random bits it keeps the sqrt(n) error bound while matching real silicon","key_machinery":"Limited-precision SR: SR_p^r(x) = SR_p(fl_{p+r}(x)). The outer SR uses the usual distance-based probabilities; the inner deterministic rounding fl_{p+r} fixes the random-number precision to r bits. The construction is what lets analysis recover the concentration bounds of exact SR and what matches the conversion instructions already documented by major vendors.","core_discovery":"Limited-precision stochastic rounding—defined by first rounding an exact value to precision p+r and then applying ordinary SR—is the practically relevant new form of the method. When r is chosen near ceil((log2 n)/2), the probabilistic O(sqrt(n) u) error bounds that hold for exact SR continue to hold for recursive summation, inner products, Horner’s method and pairwise summation, while the random-bit cost stays modest enough for real hardware. Commercial devices already implement concrete instances of this limited-precision rule, and the survey collates their bit-widths, standardisation proposals and application successes.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Limited-precision SR matches chips with (log n)/2 random bits","Modest random bits keep SR's sqrt(n) error bounds in hardware","Commercial silicon already uses limited-precision stochastic rounding","Survey updates: limited SR preserves probabilistic accuracy at low cost","Real devices implement limited-precision SR with proven error control"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hardware claims rest on public ISA documents, patents and white-papers rather than independent measurements of the actual silicon behaviour.","fun_headline_variants_meta":{"raw":{"variants":["Limited-precision SR matches chips with (log n)/2 random bits","Modest random bits keep SR's sqrt(n) error bounds in hardware","Commercial silicon already uses limited-precision stochastic rounding","Survey updates: limited SR preserves probabilistic accuracy at low cost","Real devices implement limited-precision SR with proven error control"]},"model":"grok-4.5","effort":"low","cost_usd":0.003808,"raw_usage":{"total_tokens":1220,"prompt_tokens":830,"num_sources_used":0,"completion_tokens":87,"cost_in_usd_ticks":38080000,"prompt_tokens_details":{"text_tokens":830,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":303,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":830,"tokens_out":87,"duration_ms":3636,"temperature":1.0,"reasoning_tokens":303,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T14:01:28.184277+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Measure the actual random-bit widths and rounding bias of the SR conversion instructions on AMD MI300, NVIDIA Blackwell and Graphcore devices; if the observed bit counts or bias diverge systematically from the ISA tables, the claim that commercial hardware realises the analysed limited-precision model fails.","supporting_citations":[],"review_version":1}