REVIEW 3 major objections 6 minor 7 references
PETAT -- An ASIC for Simple and Efficient Readout of Large PET Scanners
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A readout chip that daisy-chains PET detector data could drop per-module FPGAs and cut supply losses.
desk verdict A useful, honest engineering proposal for a PET readout ASIC; the time-sorted daisy-chain with timeout events is new, but the wrap-around proof and the silicon validation are lighter than the claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two load-bearing mechanisms are the time-sorted daisy chain and serial powering. In the daisy chain, each chip's serial input links are buffered in FIFOs, a merger compares the oldest time stamp of every FIFO and always forwards the smallest, and periodically injected timeout events keep FIFOs non-empty and resolve the wrap-around of the 20-bit, roughly 52 microsecond time-stamp counter. In serial powering, chips are connected in a current chain, and each chip is shunted by a regulator that compares its supply voltage to an external reference, so the whole chain draws a single current and trace voltage drops scale with that one current instead of the sum of all chip currents.
What would settle it
Feed one PETAT input link a silent gap longer than a quarter of the 52 microsecond wrap period while the other input carries a sustained full-rate stream; if the timeout injection cannot keep the merger's FIFOs non-empty, hits will overflow or the time order will be corrupted. The paper's uniform, idealized simulation does not cover this overload case.
Extended reading notes
Core claim
The paper claims that by moving data aggregation into the front-end ASIC itself, so that each PETAT chip receives serial hit streams from neighbouring chips, merges them by time stamp, and forwards a single time-sorted stream, the per-module FPGA can be eliminated. It further claims that the resulting low pin count makes serial powering practical: chips stacked in series with shunt regulators draw one common current, so resistive losses in supply traces and cables fall far below the parallel-powering case, at the cost of a higher total supply voltage. The authors' goal is to show that these two changes together make large PET scanners simpler, less power-hungry, and easier to build.
Load-bearing premise
The readout stays correct only if hits arriving on each serial link are already time-sorted and if timeout events are frequent enough that no time-stamp gap ever exceeds half the wrap-around period; the paper's simulation of idealized gammas does not bound the worst-case burst that could fill a FIFO.
Editorial extensions
If this is right
- A large PET scanner built from PETAT chips needs no FPGA on each detector module; one downstream FPGA or computer can ingest several time-sorted links.
- Because the output stream is time-sorted, coincidence filtering, epoch extension, and cluster building can be done in a single downstream processor without re-sorting.
- Serial powering reduces supply current and cable cross-section; in the paper's example, trace loss drops from 1.2 W to 0.08 W for five 1 A, 2 V chips.
- Cycle-accurate simulation of a 64-chip tree indicates that the link can be filled to its bandwidth limit of about 3.9 million hits per second per link up to roughly 2.1 MBq activity before hits are lost.
Reading between the lines
- If the time-sorted merge holds at scale, the same daisy-chain idea could apply to other time-stamped detector arrays, such as Cherenkov or time-of-flight systems, not just PET.
- The serial-powering benefit grows quadratically with the number of chips in a parallel chain, so this design is most attractive exactly where PET scanners are largest; a total-body scanner with thousands of chips is the natural stress test.
- A realistic validation would need a simulation with Compton scattering, attenuation, and nonuniform activity to confirm that the chosen timeout-event rate and FIFO depths cover worst-case load imbalance; the paper's idealized back-to-back gammas leave this open.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PETAT, a new ASIC generation in the PETA family for SiPM-based PET readout, with two main architectural novelties: (i) a hierarchical serial readout in which front-end chips daisy-chain or tree-merge hit data without local module FPGAs, producing a time-sorted output stream via FIFO merging and periodically injected 'timeout events'; and (ii) a serial powering scheme using shunt regulators to reduce supply currents and cable losses. The digital readout architecture is validated by a cycle-accurate HDL simulation of 64 and 1024 chips in several topologies, and a first test chip (PETAT1) demonstrates the shunt regulators, a 3-chip serial chain, and data transfer across voltage domains. The full analogue readout path with a standard crystal array was not tested because PETAT1's analogue section was configured for a different SiPM readout concept; a follow-up chip PETAT2 has been submitted but its test results are not yet available.
Significance. If the readout concept holds, it would simplify detector construction for large-area PET scanners by removing per-module FPGAs, reducing auxiliary power distribution circuitry, and enabling a time-sorted data stream that allows early coincidence filtering and data reduction in a single downstream FPGA. The paper's strengths include cycle-accurate HDL simulation rather than abstract modelling, performance numbers derived from the clock frequency and packet size rather than fitted parameters, and an honest, clearly separated report of which features were silicon-verified (shunt regulators, 3-chip chain, inter-voltage-domain links) versus simulation-only or still pending (full PET readout on PETAT2). The serial-powering analysis is a useful quantitative comparison even if the exact trace-resistance accounting needs clarification.
major comments (3)
- [Section 2.1] The correctness of the time-sorted merger is load-bearing, but the paper only asserts that injecting timeout events every quarter period 'guarantees' that the difference between two time stamps never exceeds half the wrap-around period. No proof or invariant is given. Please provide a concise formal argument showing that, if each input stream (including the local hit FIFO) delivers packets whose timestamps are spaced by at most T/4, then at every merge step the head timestamps of all FIFOs lie within a window smaller than T/2, and that a temporarily empty FIFO cannot invalidate this bound. Also specify where timeout-event removal takes place: if TEs are dropped inside intermediate chips before the final merge, the per-stream packet-rate guarantee used for the wrap-around correction is violated.
- [Section 2.2] The cycle-accurate simulation uses an idealized phantom with uniform back-to-back gamma pairs and no Compton scattering or absorption, so it does not exercise the pathological load patterns that stress the time-sorted merger: for example, a hot branch with sustained high hit rate and a cold branch that contributes only timeout events, or bursts that cause one FIFO to fill while another is empty. Since the paper claims loss-free operation up to a well-defined activity limit, please add stress-test simulations with asymmetric per-region rates and temporal bursts, and report the FIFO depths and any overflows. This would substantiate the claim that the half-period wrap guarantee and the chosen FIFO sizes are sufficient in realistic conditions.
- [Section 4] The paper's title and abstract present PETAT as an ASIC for PET readout, but the only silicon results reported are for the shunt regulators, a 3-chip serial chain, and digital link operation across voltage domains. The full readout path with a standard crystal array was not tested because PETAT1's analogue section could not read such an array, and PETAT2 has not yet been tested. Please state this limitation prominently in the abstract and conclusions, and clearly differentiate the 'validated on silicon' claims from the 'simulated' and 'planned' claims, so that readers do not mistake the simulated performance for measured end-to-end performance.
minor comments (6)
- [Section 3.1 / Figure 8] The serial-powering voltage-drop calculation appears to count only one inter-chip resistance per segment (0.02 Ω total drop per segment), whereas the parallel-powering example counts both power and ground traces (2 × 0.02 Ω per segment). Please state explicitly whether the serial-chain resistance includes both conductors and, if so, recompute the total drop (which would then be 0.16 V rather than 0.08 V in the example).
- [Section 2.1] The sentence 'we add sufficient TEs to ensure a data packet at least once every ¼ of the period' should specify whether this guarantee applies to each serial input link individually and to the local hit FIFO, and whether a 'packet' includes real hits as well as timeout events.
- [Section 2.3] The proposed FPGA processing pipeline in Figure 7 is described as 'not yet implemented nor simulated in detail'. Since this pipeline is part of the claimed readout chain, please clarify in the text whether the time-sorted end-to-end coincidence filtering is conceptual only or is supported by simulation.
- [Section 3.2] The statement 'This is still work in progress' at the end of Section 3.2 is in tension with the earlier statement that 'the shunt regulators work as expected'. Please clarify which aspects of the shunt regulation remain unverified.
- [References] Reference [4] gives 'DOI: 1908.05878', which is an arXiv identifier rather than a DOI; please correct the reference format.
- [Section 3.3] The text says 'The single-ended cells are uses for a standard JTAG'; this should be 'are used for'.
Circularity Check
No significant circularity: the time-sorted readout is an implemented protocol and serial powering is standard, with only an unproven wrap-around bound as a correctness risk.
full rationale
The paper does not fit parameters to a target result, nor does it invoke a uniqueness theorem from its own prior work. The central time-sorted readout claim is an inductive invariant: each chip sorts its local hits by age ('a selection logic transfers the asynchronously occurring hits according to their age into a further (shorter) FIFO'), the merger preserves order when FIFOs are non-empty, and timeout events keep links non-empty. The sentence 'The time-ordered output stream is the input for further chips, so that the assumption of time-ordered input data streams is self-consistent' is a constructive self-consistency argument, not a circular derivation: the base case is the chip's own age-sorted selection logic, and the design is validated by a cycle-accurate HDL simulation. The only genuine gap is that the half-period wrap-around guarantee is asserted ('we add sufficient TEs to ensure a data packet at least once every 1/4 of the period') without a formal bound on merger stall times under worst-case FIFO imbalance; this is a correctness risk, not a circularity. The serial powering section is standard Ohm's-law arithmetic and cites independent prior work [5,6] for the technique. The PETA4 self-citation [3] is background about inherited circuit blocks and does not carry the argument. Therefore no circular step can be exhibited.
Assumptions & free parameters
free parameters (1)
- Timeout-event injection interval =
At least one data packet per quarter of the time-stamp wrap-around period (~13 us for the 52 us period)
assumptions (4)
- domain assumption Events on each incoming link are time-sorted, with lower time stamps first.
- domain assumption Time-stamp difference between any two hits never exceeds half the wrap-around period, guaranteed by injecting timeout events.
- domain assumption The system simulation uses idealized back-to-back gamma pairs with no Compton scattering or gamma absorption.
- domain assumption Clock counters in all chips can be synchronized or corrected so that time stamps are comparable after per-chip correction.
invented entities (1)
-
Timeout event (TE)
independent evidence
Cite this review
Pith. "Pith review of PETAT -- An ASIC for Simple and Efficient Readout of Large PET Scanners." pith.science (2026). https://pith.science/paper/5WG4365L
@misc{pith2026241202394,
author = {Pith},
title = {Pith review of: PETAT -- An ASIC for Simple and Efficient Readout of Large PET Scanners},
year = {2026},
howpublished = {\url{https://pith.science/paper/5WG4365L}},
note = {Machine review of arXiv:2412.02394}
}
abstract
Modern PET scanners based on scintillating crystals use solid state photo detectors for light readout. The small area of these devices is beneficial for spatial resolution, but also leads to a large number of electronic channels to be read out, mostly by application specific integrated circuits (ASICs) containing amplification, noise reduction, hit finding, time stamping and amplitude measurement. Although each ASIC provides up to $\approx 64$ channels, a large number of chips is required with the need for auxiliary electronic components like voltage regulators or FPGAs for control and data readout. The FPGAs in turn often require multiple supply voltages and configuration infrastructure, so that PCBs get complicated, cumbersome and power-hungry, in addition to the significant power requirement of the front-end ASICs. We address this issue in the latest generation of our PETA readout ASIC for SiPMs by a simplified control scheme and, in particular, by a hierarchical serial data readout which does not require any additional FPGA. In addition, it provides a time-sorted stream of hit data, allowing early on-detector data reduction and hit pre-processing like the removal of hits with no coincident partner. The simplicity of this readout facilitates a supply scheme where power/ground of multiple ASICs are connected in series instead of the standard parallel connection. This 'serial-powering' approach can reduce supply current (while increasing overall supply voltage) so that voltage drop issues in the supply are alleviated.
Reference graph
Works this paper leans on
-
[1]
7, Article number: 35 (2020), DOI: 10.1186/s40658-020-00290-2
For instance: S.Vandenberghe, P.Moskal and J.S.Karp: State of the art in total body PET, EJNMMI Physics vol. 7, Article number: 35 (2020), DOI: 10.1186/s40658-020-00290-2
-
[2]
al.: Compact MR-compatible DC-DC converter module, Journal of Instrumentation, Vol
C.Ritzer et. al.: Compact MR-compatible DC-DC converter module, Journal of Instrumentation, Vol. 14, September 2019, DOI 10.1088/1748-0221/14/09/P09016
-
[3]
8, December 2013, DOI: 10.1088/1748-0221/8/12/c12013
I.Sacco, P.Fischer and M.Ritzert: PETA4: a multi-channel TDC/ADC ASIC for SiPM Readout, Journal of Instrumentation, Vol. 8, December 2013, DOI: 10.1088/1748-0221/8/12/c12013
-
[4]
V.Nadig, B.Weissler, H.Radermacher, V.Schulz and D.Schug: Investigation of the Power Consumption of the PETsys TOFPET2 ASIC, IEEE Trans. on Rad. and Plasma Medical Sciences, Vol. 4, No. 3, May 2020, DOI: 1908.05878
work page Pith review arXiv 2020
-
[5]
D.B.Ta et. al.: Serial powering: Proof of principle demonstration of a scheme for the operation of a large pixel detector at the LHC, Nucl. Instr. and Methods A, Vol. 557, Issue 2, 15 February 2006, pp. 445-459, DOI: https://doi.org/10.1016/j.nima.2005.11.115
-
[6]
R.Ceccarelli et. al.: Serial powering characterisation for the CMS Inner Tracker at the High Luminosity LHC, Journal of Instrumentation, Vol. 19, January 2024, DOI: 10.1088/1748-0221/19/01/C01056
-
[7]
J.Debus et. al.: Design and first performance tests of the SAFIR-II PET-MR Scanner, IEEE NSS, MIC and RTSD, November 2023, DOI: 10.1109/nssmicrtsd49126.2023.10338699 – 10 –
arXiv 2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.