{"id":"67574d5f-fa75-4e22-988c-1e82f669ad6d","arxiv_id":"2607.25174","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Compact BDT waveform classifiers deployed in the Belle II CDC front-end FPGA suppress cross-talk noise by ~2x at the wire level and reduce fake L1 trigger rates by up to ~72% with roughly 5-11% signal-acceptance loss.","lead":"Belle II physicists trained small decision-tree models to recognize cross-talk noise in the central drift chamber's front-end electronics, and deployed them on the detector's FPGAs to filter noise before the trigger. In test runs, trigger rates from fake tracks dropped by up to ~70% while keeping real track acceptance nearly unchanged, showing that machine learning can run inside front-end boards.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Offline wire-level performance lacks a documented held-out evaluation; the factor-of-two cross-talk rejection and >98% efficiency may reflect training-data memorization rather than generalization.","rationale":"The reader's weakest assumption concerned representativeness to future high-luminosity runs, which is a legitimate limitation but is acknowledged by the authors and does not invalidate the feasibility claim under the tested conditions. The more fundamental issue is that the offline wire-level performance—the quantitative basis for the factor-of-two and >98% efficiency—appears to lack a documented held-out evaluation. This is a correctness risk because the reported numbers could be inflated by training-data memorization, especially with 96 channel-specific models. The online validation provides end-to-end evidence of trigger-rate reduction but does not directly measure wire-level classification quality, so it cannot substitute for an independent offline test. The paper should either provide a clear train/test split or explicitly state that the reported numbers are on training data, which would materially lower their evidentiary value. The acceptance inconsistency for the Neural 3D Tracker (11% loss vs. 'within 10%') is a secondary issue that can be fixed by rewording. The recommendation remains conditional acceptance pending the holdout analysis, so the reader's verdict is unchanged.","tokens_in":10996,"tokens_out":6153,"duration_ms":61326,"concrete_test":"Re-train the 96 BDT models on a randomly selected subset (e.g., 80%) of the 2024 waveform dataset, reserving the remaining 20% as a held-out test set. Ensure that the same wire hit does not appear in both sets (split by event or by wire channel). Recompute Table II using only the held-out test set for signal efficiency, background rejection, and cross-talk rejection at cut levels 1–4. If the cross-talk rejection at cut level 2 decreases by more than 0.10 absolute or the signal efficiency drops below 98%, the offline claims in the abstract are not supported. Also recompute the ROC curve on the held-out set and compare the AUC with Fig. 8 to quantify any overfitting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims of 'cross-talk noise reduced by approximately a factor of two' and 'signal efficiency above 98%' rest on Table II, which reports wire-level performance of the quantized BDT models. Section III-B describes training using approximately 10,000 signal and 10,000 background samples, but Section IV-A states that 'the performance was evaluated using all waveform samples' and does not state that the training samples were excluded. The thresholds for cuts 1–4 are derived from ROC curves (Fig. 8), which are presumably computed on the same data used for training. With 96 channel-specific BDT models, each trained and evaluated on overlapping data, the reported background and cross-talk rejection rates are at risk of optimistic bias from overfitting. The online validation in Section IV-B is an end-to-end test using new data (March 2026), but it measures only trigger rates and acceptance, not wire-level classification efficiency or rejection. It therefore does not independently confirm the specific wire-level numbers in Table II. Without a clear train/test split, the headline offline performance—the basis for the factor-of-two claim—is not established. This is a more direct threat to the central claim than the extrapolation to full luminosity, because it affects the quantitative result even under the tested conditions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the development and deployment of compact BDT-based waveform classifiers in the front-end FPGAs of the Belle II CDC to suppress cross-talk noise before it reaches the Level-1 trigger. The authors train 96 channel-specific quantized BDTs using five ADC samples around each TDC hit, implement a two-cycle custom VHDL inference engine using ~4.5% of the Virtex-5 resources, and validate it in dedicated March 2026 single-beam calibration runs. Offline processing (Table II) reports signal efficiencies above ~98% with roughly half of cross-talk hits rejected; the online runs show 48–72% reductions in the 2D Finder trigger rate and ~50% reductions in outer-super-layer track-segment rates, while event-level acceptance for events containing IP tracks is preserved at roughly the 90% level.","tokens_in":11272,"tokens_out":13530,"duration_ms":136882,"significance":"If the offline wire-level numbers survive a proper holdout evaluation, this is a valuable, experimentally demonstrated use of ML in detector front-end electronics under tight resource and latency constraints. The main strengths are the real FPGA deployment in an operating experiment, the measured online trigger-rate reduction, the explicit statement of resource and power usage, the channel-by-channel quantization study, and the very small degradation from quantization. The main concern is that the wire-level offline table is the quantitative basis for the headline 'factor of two / >98%' claim, but its train/test separation is not documented. The acceptance statement for one trigger also needs a small correction relative to Table V.","major_comments":[{"comment":"The offline wire-level performance is the basis for the factor-of-two cross-talk rejection and the >98% signal efficiency, but the text does not state that the evaluation set excludes the samples used to train the 96 BDT models. Section IV-A says 'the performance was evaluated using all waveform samples' and the cut thresholds are 'determined from their ROC curves' without indicating whether those curves come from the same data. Given ~10k training samples per class and 96 channel-specific models, in-sample evaluation could inflate the reported rejection rates. Please add an explicit train/test split (e.g., a held-out run, time-ordered split, per-channel holdout, or k-fold CV) and report Table II on the held-out data, including the number of samples in the evaluation set.","section":"§IV-A, Table II"},{"comment":"The claim that track-trigger acceptance for events with tracks is preserved 'within 10%' is not supported for both triggers at cut level 2. Table V shows the Neural 3D Tracker acceptance for events with IP tracks changing from 0.790±0.032 (no cut) to 0.702±0.027 (cut 2), a relative loss of about 11.1%; only the 2D Finder (5.1% relative loss) satisfies the strict 10% statement. Please qualify the claim (e.g., 'approximately 10%'), combine both triggers into a single metric, or discuss the uncertainty and why the 11% loss is acceptable.","section":"Abstract/Conclusion vs. Table V"}],"minor_comments":[{"comment":"The paper states 14,336 anode wires and 48 channels per FEE board, then mentions 292 FEE boards. 292×48 = 14,016, leaving 320 wires unaccounted for; please check the number of boards or the per-board channel count.","section":"§II-A/II-B/II-C"},{"comment":"The 'consecutive increments' defining cut levels 2–4 are not specified. Please give the actual threshold values or the exact rule used to move from one cut level to the next, so the deployed cuts are reproducible.","section":"§III-B/IV-A"},{"comment":"For waveform samples with multiple TDC hits, the paper states that a wire passes if at least one corresponding classifier output exceeds the threshold. This OR rule makes the selection more permissive and should be justified; it also affects how the reported signal and background efficiencies should be interpreted.","section":"§IV-A"},{"comment":"No uncertainties are reported for the efficiencies/rejection rates in Table II. With 96 models and finite samples, per-channel spread or binomial confidence intervals would help assess whether the differences between cut levels and between SL0 and SL1–SL8 are meaningful.","section":"Table II"},{"comment":"The Conifer HLS comparison is obtained on a Kintex-7 device because HLS does not support Virtex-5. The caption notes this, but the text should explicitly warn that the resource numbers for Conifer HLS are not directly comparable to the Virtex-5 implementation.","section":"§III-C/Table I"},{"comment":"The sentence 'a loss of less than 10% loss' contains a typo; should be 'a loss of less than 10%'.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"This is a solid engineering paper with a convincing hardware deployment and valuable online measurements. The central issue is the missing documentation of train/test separation for Table II, which directly underpins the factor-of-two claim; this is fixable with a proper holdout analysis and does not require re-running the hardware. The acceptance statement discrepancy is minor but should be corrected. I would not reject; after the offline evaluation is clarified and the wording fixed, the paper would be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: this is a real paper about a real deployment. The group put a BDT waveform classifier inside the CDC front-end FPGA of a running experiment, with a custom two-cycle VHDL implementation on a Virtex-5, and then measured the effect during dedicated calibration runs. That is new. Previous ML trigger work (Conifer, GNN hit filtering) sits in the back-end; moving the intelligence into the FEE, where the full waveform exists, is a legitimate step, and the engineering constraints—97k LUTs, 15.7 ns added latency, ~2% power increase—are presented concretely. The online validation is the strongest part: at cut 2 the 2D Finder rate drops from 82.6 to 23.3 kHz, and track acceptance for IP-track events stays within about 5%; segment-rate reductions are plausibly flat in azimuth. That is the kind of evidence that convinces me the thing works.\n\nThe soft spot is the offline wire-level analysis that supports the headline 'factor of two / >98% efficiency' claim. Section IV-A says performance was evaluated using all waveform samples, and I do not see a statement that the ~10k signal and ~10k background training samples were excluded, nor any cross-validation. The cut thresholds are chosen from ROC curves that appear to be computed on the same data. With 96 channel-specific models, that is a textbook setup for optimistic in-sample numbers. The online validation uses new data, but it measures trigger and segment rates, not per-wire classification, so it does not independently confirm Table II. In my view this is the one thing a referee should pin down. It is not a fatal flaw—the practical noise suppression is demonstrated—but the specific wire-level efficiency/rejection numbers are not yet established as generalization.\n\nTwo smaller points. The abstract's 'within 10%' acceptance wording sits awkwardly next to Table V, where the Neural 3D Tracker acceptance drops from 0.790 to 0.702 at cut 2—about 11%. And the calibration runs used a single electron beam, not collision conditions, so the extrapolation to future high-luminosity running is an expectation, not a measurement. The authors say the cross-talk origin is not fully understood, so that caveat matters, but the feasibility claim survives it.\n\nWho should read this: anyone designing front-end or trigger systems for high-rate experiments, and people working on FPGA ML inference. It deserves a serious referee. The review should ask for an explicit train/test split or cross-validated ROC, a rewording of the acceptance claim, and a comment on how the single-beam results map to collision conditions. I would not block acceptance on the missing split if the authors can quickly produce one or if they clarify that the evaluation set was indeed disjoint.","headline":"First real ML-in-FEE deployment for Belle II CDC noise suppression; online numbers are convincing, but the offline factor-of-two claim needs a documented train/test split.","tokens_in":11816,"tokens_out":2270,"would_cite":true,"duration_ms":23698,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Compact decision-tree classifiers in the Belle II drift chamber front-end FPGA cut cross-talk noise by about half while keeping signal efficiency above 98 percent, and in calibration runs reduced track-trigger rates by up to around 50 perce","keywords":["boosted decision trees","cross-talk noise","FPGA inference","front-end electronics","waveform discrimination","Belle II central drift chamber","Level-1 trigger"],"falsifier":"Record full ADC waveforms during regular beam-beam operation at design luminosity and measure the deployed BDT's cross-talk rejection rate and 2D Finder trigger rate; if the factor-of-two rejection and roughly 50 percent trigger-rate reduction do not reproduce, or if track acceptance drops by more than about 10 percent, the central claim would be weakened.","tokens_in":10859,"feed_emoji":"⚛️","tokens_out":4623,"duration_ms":41883,"temperature":0.7,"pith_summary":"This paper tries to prove that cross-talk noise in the Belle II central drift chamber can be stopped at its source, inside the front-end electronics, by classifying each wire's ADC waveform with a tiny machine-learning model. The classifiers are boosted decision trees small enough to fit into the Virtex-5 FPGA on each front-end board, where they run in a fully pipelined way with a latency of two clock cycles. Offline, the models reject roughly half of cross-talk hits while keeping signal efficiency above 98 percent. In dedicated calibration runs with the actual firmware, track-segment and track-trigger rates fell by up to about 50 percent, and trigger acceptance for events containing real tracks stayed within 10 percent of the baseline. If this holds at full luminosity, it provides a template for moving intelligent waveform processing into the front-end electronics of future detectors.","feed_headline":"Tiny decision trees halve fake triggers at Belle II","feed_subtitle":"Waveform classification inside the drift chamber front-end chips cuts cross-talk noise in half while keeping over 98% of real signals.","key_machinery":"The central mechanism is a per-channel Boosted Decision Tree (BDT) classifier, an ensemble of four shallow decision trees whose outputs are summed, converted to a custom two-stage pipelined VHDL implementation in the FEE's Virtex-5 FPGA. All decision-node conditions are evaluated in parallel in one clock cycle and leaf outputs combined with AND gates in the next, giving 15.7 ns inference. Five ADC samples around each TDC hit, truncated to eight bits, are the input features; independent models are trained per wire channel and for the two chamber geometries (SL0 vs SL1–SL8). The design keeps TDC timing flowing to the trigger immediately and attaches the BDT classification later at the Track Se","core_discovery":"The central claim is that the full ADC waveform, available only in the drift chamber's front-end electronics, contains enough information to distinguish real charged-track hits from FEE cross-talk noise, and that a suitably compact boosted-decision-tree classifier can exploit that information in real time on a limited FPGA. The paper reports that five waveform samples centered on the TDC timing, quantized to eight bits, feed four depth-four decision trees per wire channel; 96 such models cover the 48 channels in the two chamber geometries. Offline this configuration cuts cross-talk noise by roughly a factor of two at a signal efficiency above 98 percent. When deployed in firmware, it reduced","pith_inferences":["Editorial inference: The training and validation waveforms came from a 2024 calibration run and a March 2026 run with only an electron beam at 1 A; the paper does not measure performance under full beam-beam background, so the deployed thresholds may need retraining if cross-talk shapes change with luminosity.","Editorial inference: Because the BDT uses only five samples per hit, its discrimination is effectively a template on pulse shape and timing; a natural extension the paper itself identifies is using neighboring-channel information via CNN or GNN architectures.","Editorial inference: The approach could generalize beyond cross-talk: any front-end electronics with access to raw waveforms could use compact BDTs for pile-up separation, baseline correction, or other preprocessing, provided the FPGA resources are comparable."],"forward_implications":["Cross-talk wire hits can be cut by roughly half at the front end while retaining more than 98 percent of genuine signal hits.","Reduced wire-hit rates translate into track-segment and Level-1 trigger rate reductions of up to about 50 percent in calibration runs, with the 2D Finder rate falling from 82.6 kHz to 23.3 kHz at the stricter cut.","Track trigger acceptance for events with reconstructed IP tracks is preserved, with losses below about 10 percent.","The added power consumption on the front-end board is about 2 percent, and inference latency is two clock cycles (15.7 ns), so the approach fits the existing 4.4 microsecond trigger budget.","Channel-wise independent waveform classification is sufficient, so no inter-channel data sharing is needed in the FPGA."],"fun_headline_variants":["ML in Belle II front-end halves cross-talk noise","Belle II drift chamber chips use ML to cut noise","On-chip decision trees halve Belle II fake triggers","Waveform ML on Belle II FEE reduces noise by half","Belle II FEE ML: 98% signal, half cross-talk"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Training and validation waveforms came from a 2024 calibration run and a March 2026 run with only an electron beam at 1 A; the paper assumes these faithfully represent cross-talk noise under future high-luminosity beam-beam collisions, and it states the cross-talk origin is not yet fully understood.","fun_headline_variants_meta":{"raw":{"variants":["ML in Belle II front-end halves cross-talk noise","Belle II drift chamber chips use ML to cut noise","On-chip decision trees halve Belle II fake triggers","Waveform ML on Belle II FEE reduces noise by half","Belle II FEE ML: 98% signal, half cross-talk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000905,"raw_usage":{"total_tokens":3785,"prompt_tokens":858,"completion_tokens":2927,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":2843}},"tokens_in":602,"tokens_out":2927,"duration_ms":19498,"temperature":1.0,"reasoning_tokens":2843,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:11:08.141216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record full ADC waveforms during regular beam-beam operation at design luminosity and measure the deployed BDT's cross-talk rejection rate and 2D Finder trigger rate; if the factor-of-two rejection and roughly 50 percent trigger-rate reduction do not reproduce, or if track acceptance drops by more than about 10 percent, the central claim would be weakened.","supporting_citations":[],"review_version":1}