{"id":"90613575-4566-4d7a-9d42-bd8b880d781a","arxiv_id":"2507.10463","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A coalition of academic and industry researchers argues that chips exploiting natural physical dynamics, rather than enforcing digital abstractions, could dramatically cut AI computing costs.","lead":"A large team of researchers and industry affiliates proposes a new class of chips, physics-based ASICs, that would use natural physical dynamics instead of enforcing digital abstractions. The paper is a roadmap and advocacy document, not a demonstration, so its promise rests on future benchmarks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"End-to-end I/O and Amdahl fraction are never quantified; cited prototype speedups evaporate at larger problem sizes per Sec. V, so the 'necessary evolution' claim is unsupported.","rationale":"The reader's verdict is UNVERDICTED because the paper is a vision/roadmap document rather than a testable research result, and I agree with that classification. My stress test targets the strongest form of the paper's claim—'not only a viable alternative but a necessary evolution'—and finds that the paper's own Section V, Phase 1 provides direct evidence against the scalability of demonstrated prototypes due to I/O. This is precisely the same weakness the reader identified as the most fragile premise: the fraction x of a workload that can be accelerated, plus the data-movement cost, must be large enough that Amdahl's Law does not erase the raw efficiency gain. I find no internal inconsistency in the roadmap's qualitative description, but the absence of any end-to-end metric means the 'substantial gains' and 'necessary evolution' statements are not currently supported. Since the reader already marked the paper UNVERDICTED rather than ACCEPT or REJECT, my concern does not change the verdict; it reinforces it. A concrete system-level benchmark would settle whether the concern lands, and until such a benchmark exists, the paper should not be treated as a demonstrated solution to the compute crisis.","tokens_in":18101,"tokens_out":3071,"duration_ms":42757,"concrete_test":"Build a system-level model for one cited workload—e.g., the 1440-spin latch-based Ising machine of [12] or the 48-node coupled-oscillator Ising machine of [13]—that includes input encoding time/energy, spin-state readout, chip-to-host transfer (PCIe/NoC), and any unaccelerated preprocessing, and sweep problem size N from the demonstrated size up to 10^4–10^6 variables. If the end-to-end speedup or energy ratio versus the claimed CPU/GPU baseline drops below 1 before the target application size, the central system-level claim fails; if it remains above 1, the concern is answered. Report the crossover size N* and the dominant overhead at N*.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that physics-based ASICs offer substantial system-level gains and are a 'necessary evolution'—requires that after input encoding, output readout, memory transfer, and unaccelerated serial steps, the total efficiency still beats state-of-the-art digital hardware. The paper never supplies such a system-level model. Section III.B defines RT(ℓ) and RE(ℓ) in Eqs. (1)–(2) as kernel-level ratios of runtime/energy on SOTA versus the ASIC, and Section III.C invokes Amdahl's Law for the un-accelerated fraction x, but Eq. (1) does not include the time and energy needed to load data onto the physical substrate and read results off it. Section V, Phase 1 then concedes that for larger problems the latch-based Ising prototype 'often cannot achieve the same speed-up, due to the cost of loading data onto and reading data off of the physics-based ASIC' [12]. Since no cited demonstration reports an end-to-end advantage at scale, the 'substantial gains' claim rests on an unverified transfer from primitive speedups to system speedups. The weakest point is therefore not device physics but the missing accounting of I/O and memory bandwidth, which the paper's own evidence suggests can erase the raw advantage in exactly the large-problem regime that motivates the 'compute crisis'.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a perspective/roadmap paper arguing that physics-based ASICs—chips that exploit natural physical dynamics rather than enforcing digital abstractions such as statelessness, unidirectionality, determinism, and synchronization—could provide order-of-magnitude gains in energy efficiency and throughput for AI and scientific workloads. It defines the paradigm, surveys platforms (memristors, Ising machines, p-bits, thermodynamic computers, photonic systems), proposes a top-down/bottom-up co-design strategy, introduces kernel-level speedup and energy ratios RT(ℓ) and RE(ℓ), invokes Amdahl's law, discusses physical machine learning and physical learning, lists target applications (neural networks, diffusion models, sampling, optimization, scientific simulation, analog data analysis), and lays out a three-phase roadmap from proof-of-concept to system integration. The paper contains no new measurements, derivations, or system-level models; its argument is qualitative and vision-oriented.","tokens_in":18390,"tokens_out":2536,"duration_ms":32222,"significance":"If the central claim were established, the paper would identify a high-impact direction for post-CMOS computing and address real concerns about AI energy and hardware scaling. The manuscript is useful as a synthesis and a call to action: it explicitly acknowledges Amdahl's law, memory-bandwidth limitations, the difficulty of physical machine learning optimization, barren plateaus, and sim2real gaps, and it names concrete prototype results. Its main weaknesses are that the load-bearing evidence consists largely of kernel-level or small-scale demonstrations, several from the authors' own prior work, and that no end-to-end system-level accounting is provided. As a perspective, the paper is coherent and readable, but the strength of the conclusions currently exceeds the strength of the evidence.","major_comments":[{"comment":"The performance framework in Eqs. (1) and (2) defines RT(ℓ) and RE(ℓ) at the kernel level, and the Amdahl discussion in Section III.C uses only an unaccelerated runtime fraction x. Neither accounts for the time and energy of loading data onto the physics-based ASIC and reading results off it. Section V, Phase 1 then concedes that for larger problems the latch-based Ising prototype 'often cannot achieve the same speed-up, due to the cost of loading data onto and reading data off of the physics-based ASIC' [12]. Since the cited evidence does not include an end-to-end system-level demonstration, the paper's central claim of substantial system-level gains is not supported by the presented material. The manuscript should either add a quantitative system-level model that includes I/O and memory bandwidth or explicitly scope its claims to kernel-level advantages.","section":"III.B–III.C and V, Phase 1"},{"comment":"The scalability discussion asserts that tile-based designs and reconfigurable coupling 'could potentially get as large as GPUs' but provides no quantitative treatment of the costs that dominate at large scale: inter-tile communication bandwidth, minor-embedding overhead for arbitrary sparse graphs, analog noise accumulation, and on-chip routing of mixed-signal blocks. Because the 'compute crisis' motivating the paper is specifically a large-scale problem, these omissions are load-bearing. The passage should be labeled as conjecture unless accompanied by scaling estimates or references to validated large-scale designs.","section":"V, Phase 2"},{"comment":"The conclusion that physics-based ASICs are 'not only a viable alternative but a necessary evolution in how we compute' is an evaluative claim that outruns the evidence in the manuscript. The paper itself lists major open challenges—PML optimization difficulty, sim2real gaps, bandwidth-limited prototypes, and the absence of demonstrated large-scale integration—so the evidence supports a promising research direction rather than a demonstrated necessity. The conclusion should be rephrased to match the evidentiary level, or the authors should supply the missing system-level demonstration or model.","section":"VI.A"}],"minor_comments":[{"comment":"'Miniturization effects' should be 'miniaturization effects'.","section":"I"},{"comment":"In the sentence after Eq. (2), 'either RT(ℓ) or RT(ℓ) is greater than one' should read 'either RT(ℓ) or RE(ℓ) is greater than one'.","section":"III.B, Eqs. (1)–(2)"},{"comment":"'quadratic unconstained binary optimization' should be 'quadratic unconstrained binary optimization'.","section":"IV.A.4"},{"comment":"'paralellism' should be 'parallelism'.","section":"VI.B"},{"comment":"Reference [2] lists the year as 2014, but the cited arXiv paper (2405.21015) appeared in 2024.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a vision/perspective paper rather than a technical contribution. Its main value is as a field-level synthesis and roadmap, and its central claim can be made defensible by systematically softening the language and adding or explicitly deferring system-level analysis. The authors should also be asked to ensure that self-citations are used as illustrations of feasibility rather than as independent validation. Given the journal context, major revision with a request to recalibrate claims seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kavita,\n\nQuick take: this is a perspective/roadmap paper, not a research result, and it reads like one. What is actually new is modest but real: the term “physics-based ASICs” and the top-down/bottom-up co-design framing in Sections III.A and III.D give the community a useful shared vocabulary, and the paper is unusually candid about its own weak spots. It explicitly concedes in Section V, Phase 1 that latch-based Ising prototypes lose their speedup at larger problem sizes due to I/O costs, which is exactly the caveat that matters.\n\nThe paper does several things well. It catalogs existing platforms—memristors, Ising machines, p-bits, thermodynamic computers, optical neural networks—under a coherent set of relaxed abstractions (statelessness, unidirectionality, determinism, synchronization). That is a genuinely helpful synthesis. It also engages seriously with Amdahl’s Law, memory bandwidth, PML optimization difficulty, and the sim2real gap, which is more self-awareness than most vision papers show. The writing is clear and the citation pattern is broad; self-citation is heavy in the thermodynamic computing sections, but those papers are real and the text leans on them honestly rather than hiding the dependence.\n\nThe soft spots are proportionate to the genre. The central claim of “substantial gains” and “necessary evolution” is not demonstrated, and cannot be from this kind of paper. The stress-test concern holds up: Eqs. (1) and (2) define kernel-level runtime and energy ratios, but the system-level accounting—input encoding, output readout, memory transfer, Amdahl fraction x—is never quantified. The paper’s own Phase 1 text admits the I/O cost can erase the raw advantage at scale, so the “necessary evolution” conclusion rests on an unverified extrapolation from primitive speedups to end-to-end system speedups. That is a real gap, but it is the standard gap of a roadmap paper, not a fatal flaw. The paper never pretends these prototypes are production-ready; it frames them as Phase 1 demonstrations to be scaled.\n\nWho gets value from this: researchers entering the field, program managers, and anyone who needs a single reference for the landscape of physics-based computing. It is a good review-style anchor, and the co-design framework could shape how people pitch and compare work. It deserves a serious referee—not because the evidence is conclusive, but because the synthesis and framing are useful and the field needs a paper like this to argue against.\n\nMy recommendation: send it to peer review as a perspective/roadmap, with a request that the authors either soften the “necessary evolution” conclusion or add a worked system-level estimate showing a plausible end-to-end gain for at least one concrete workload.\n\nBest,\nNadia","headline":"A clear, honest roadmap paper that reframes existing physics-based computing work under one umbrella, but the 'necessary evolution' claim outruns the evidence it cites.","tokens_in":18931,"tokens_out":683,"would_cite":true,"duration_ms":10309,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that application-specific chips which compute with natural physical dynamics, rather than enforcing digital abstractions, can overcome the AI energy and scaling crisis.","keywords":["physics-based ASICs","compute crisis","hardware-software co-design","energy efficiency","probabilistic computing","thermodynamic computing","Ising machines","analog neural networks"],"falsifier":"A concrete test is to build a tile-based physics-based ASIC scaled to a production-size problem in diffusion sampling or combinatorial optimization, co-design an algorithm for it, and measure end-to-end runtime and energy against a state-of-the-art GPU; the central claim fails if the measured speedup and energy ratios are not both above 1 at that scale, or if the Amdahl fraction $x$ falls below $1 - 1/S$ for the claimed speedup $S$.","tokens_in":17951,"feed_emoji":"⚡","tokens_out":8001,"duration_ms":89941,"temperature":0.7,"pith_summary":"This paper argues that the path out of the AI compute crisis—unsustainable data-center energy use, billion-dollar training runs, and faltering CMOS scaling—is to stop fighting the physics of silicon. It proposes physics-based ASICs, application-specific integrated circuits that deliberately use the natural dynamics of circuits, including stochasticity, memory, feedback, and asynchrony, as the computation itself, rather than spending energy to enforce idealized digital behavior such as statelessness, unidirectionality, determinism, and synchronization. The paper lays out a design strategy in which algorithms and hardware are co-designed so that applications such as diffusion models, sampling, optimization, neural inference, and molecular simulation land in the set of operations a physical device can do natively, and it quantifies the attainable gain with runtime and energy ratios under Amdahl's Law. If the argument holds, specialized physics-based chips could deliver substantial energy savings and new capabilities that general-purpose digital hardware cannot reach, shifting future high-performance computing toward heterogeneous systems of specialized physical processors.","feed_headline":"Chips that use raw physics as the computer target the AI energy crisis","feed_subtitle":"By relaxing clocks and determinism, these chips run sampling, optimization, and inference as natural physical processes.","key_machinery":"The central object is the physics-based ASIC itself, defined by relaxing the four conventional abstractions of statelessness, unidirectionality, determinism, and synchronization so that computation is realized directly as a physical process. The argument's quantitative engine is algorithmic co-design: applications define a set of possible algorithms, physical structures define a set of algorithms they can run efficiently, and the design goal is to maximize their overlap while respecting Amdahl's Law, which limits the attainable speedup to a factor of $1/(1-x)$ where $x$ is the fraction of runtime movable onto the physics-based ASIC. Performance is assessed by runtime and energy ratios relative to state-of-the-art digital hardware, and the paper also proposes physical machine learning, in which the hardware itself learns its parameters from physical dynamics, as a route to co-design without a separate digital training loop.","core_discovery":"The paper's central claim is that the compute crisis can be addressed by a class of chips it calls physics-based ASICs: circuits that let computation be carried out by the natural physical dynamics of the device instead of forcing the device to imitate an ideal digital machine. Conventional ASICs spend energy, time, and complexity to approximate four abstractions that are not exactly realizable in physics: statelessness, unidirectionality, determinism, and synchronization. The paper argues that when these constraints are relaxed, components can be stateful, bidirectionally coupled, stochastic, and asynchronous, and a computed result becomes the outcome of a real physical process—for example, thermodynamic relaxation performing sampling or linear algebra. It claims this can fuse many operations, save power, and in several demonstrated or predicted cases beat CPU and GPU solvers in speed or energy, citing a latch-based Ising machine solving 1440-spin problems over 1000 times faster than a CPU solver and optical neural networks operating below one photon per scalar multiplication. The paper concludes that physics-based ASICs are not just an alternative but a necessary evolution of computing, deployed in heterogeneous systems alongside CPUs and GPUs.","pith_inferences":["If the hardware-lottery dynamic the paper describes is real, the largest effect of a successful physics-based ASIC may be to change which algorithms researchers choose to work on: algorithms that map naturally to physical dynamics would begin to win over algorithms tuned for GPUs.","A testable extension of the roadmap is to measure, for each candidate workload, the Amdahl fraction $x$ and the data-transfer energy at production scale; the paper's own Phase-1 caveat suggests that these, not the raw on-chip gain, will separate the workloads that benefit from those that do not.","The energy-time-accuracy tradeoff the paper notes could be developed into a quantitative design criterion: for a given physical device, the optimal operating point in voltage, clock rate, and noise level would be chosen by optimizing a three-way tradeoff rather than by maximizing determinism.","If physical machine learning matures, the same hardware that computes could also learn its own parameters without a digital co-processor, which would invert the usual assumption that hardware must be designed to be trainable by software."],"forward_implications":["If the central claim is right, the first practical demonstrations will be domain-specific: small physics-based ASICs will beat CPU or GPU solvers on particular workloads such as Ising optimization, sampling, or linear algebra, before they compete on general-purpose tasks.","Co-design becomes the main engineering lever: pushing algorithmic complexity into subroutines that run on the ASIC increases the fraction $x$ that Amdahl's Law rewards, so new algorithms will be judged by how much of their runtime can be made physical.","Heterogeneous systems will emerge in which CPUs, GPUs, and multiple physics-based ASICs share a workload, with a compiler mapping each step of a high-level program to the best device.","Energy efficiency is expected to arrive before raw speed: early physics-based ASICs may match a GPU's output while consuming far less power, especially for diffusion, sampling, and stochastic neural-network workloads.","Once scaled and integrated, the technology could enable computations that digital machines cannot do affordably, such as approximation-free Bayesian inference for reliable AI predictions and large-scale molecular dynamics."],"supporting_citations":[{"why":"Supplies the 'hardware lottery' observation that algorithms are implicitly co-designed to whatever hardware wins, which motivates the paper's intentional co-design strategy.","marker":"[5]"},{"why":"Provides the Phase-1 speedup example of a latch-based Ising machine solving 1440-spin problems over 1000 times faster than a CPU solver.","marker":"[12]"},{"why":"Provides the coupled-oscillator Ising machine result used to argue that analog physics-based ASICs can overtake GPU solvers beyond roughly 150 spins.","marker":"[13]"},{"why":"Describes a thermodynamic computing system for AI applications, serving as an existing physics-based ASIC platform and a baseline for the paper's applications discussion.","marker":"[19]"},{"why":"Gives the thermodynamic linear algebra result that the paper cites as an asymptotic complexity advantage over digital methods.","marker":"[34]"},{"why":"Gives the thermodynamic Bayesian inference result, cited as an example of a physics-based ASIC application with better asymptotic scaling.","marker":"[36]"},{"why":"Reports an optical neural network operating below one photon per scalar multiplication, used as key evidence for the energy-efficiency advantage.","marker":"[74]"},{"why":"Describes self-adjusting resistor networks that learn and compute physically, cited as evidence of machine learning without a processor and of potentially million-fold energy savings.","marker":"[10, 11]"}],"fun_headline_variants":["Physics-based chips compute by being, not simulating","Relaxing chip abstractions to harness physics for AI","Letting physics do the math: ASICs for the compute crunch","From digital mimicry to physical computation in ASICs","Compute crisis fix: ASICs that compute via natural dynamics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The roadmap holds only if, after algorithmic co-design, a large enough fraction of each real workload can be run on the physics-based ASIC, and the cost of moving data on and off the chip stays small enough, that Amdahl's Law does not erase the raw efficiency gain; the paper itself notes in its Phase-1 discussion that larger prototypes often lose their speed-up to exactly these data-transfer costs.","fun_headline_variants_meta":{"raw":{"variants":["Physics-based chips compute by being, not simulating","Relaxing chip abstractions to harness physics for AI","Letting physics do the math: ASICs for the compute crunch","From digital mimicry to physical computation in ASICs","Compute crisis fix: ASICs that compute via natural dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1244,"prompt_tokens":954,"completion_tokens":290,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":210}},"tokens_in":570,"tokens_out":290,"duration_ms":3538,"temperature":1.0,"reasoning_tokens":210,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:29:45.577621+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test is to build a tile-based physics-based ASIC scaled to a production-size problem in diffusion sampling or combinatorial optimization, co-design an algorithm for it, and measure end-to-end runtime and energy against a state-of-the-art GPU; the central claim fails if the measured speedup and energy ratios are not both above 1 at that scale, or if the Amdahl fraction $x$ falls below $1 - 1/S$ for the claimed speedup $S$.","supporting_citations":[{"cited_title":"Melanson, M","cited_arxiv_id":null,"evidence_quote":"Gives the thermodynamic linear algebra result that the paper cites as an asymptotic complexity advantage over digital methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports an optical neural network operating below one photon per scalar multiplication, used as key evidence for the energy-efficiency advantage."}],"review_version":1}