{"id":"442606e3-3066-407a-8d84-b139acb2b63c","arxiv_id":"2606.20926","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A new RL-LS BP decoder for QLDPC codes combines learned variable-node scheduling with list-based search and cumulative path metrics, showing improved performance over standard BP on the depolarizing channel.","lead":"The paper proposes a reinforcement learning-based list sequential belief propagation decoder for quantum LDPC codes that extends prior RL sequential scheduling with list exploration of alternative error paths. A smart generalist might read it because efficient decoding remains a central bottleneck for making quantum error correction practical in fault-tolerant quantum computers.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flagged generalization as the weakest link, but once the full manuscript is examined the claim itself is narrowly scoped to the reported benchmarks and does not assert cross-code or cross-channel universality. No further load-bearing gap is visible.","tokens_in":1735,"tokens_out":246,"duration_ms":13072,"concrete_test":"Reproduce the reported frame-error-rate curves for the largest benchmark code at the highest simulated physical error rate using the exact training and list-pruning hyperparameters stated in the experimental section; if the gap versus plain RL-S disappears under identical random seeds and Monte-Carlo sample size, the headline numerical claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical performance improvement on representative QLDPC codes under depolarizing noise, obtained by combining learned sequential scheduling with a specific list-exploration rule (soft bias to second-most-likely symbol plus cumulative path metric). The paper supplies the numerical evidence for that claim on the stated benchmarks; no internal inconsistency, missing derivation, or unstated assumption that would invalidate the reported gains on those instances is apparent from the provided text.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes a reinforcement learning-based list sequential belief propagation (RL-LS BP) decoder for quantum LDPC codes. It extends the prior RL-S framework by incorporating list-based search: a learned policy selects the next variable node to update, the decoder retains the RL-S trajectory while exploring a competing branch via soft biasing of the post-update LLR pair toward the second-most likely Pauli symbol, and trajectories are ranked and pruned using a proposed cumulative path metric. Numerical results on representative QLDPC benchmark codes over the depolarizing channel are reported to show improved decoding performance relative to the underlying decoder and favorable comparisons with existing BP-based methods.","tokens_in":1804,"tokens_out":347,"duration_ms":23096,"significance":"If the reported empirical gains are reproducible, the work offers a practical contribution to decoding QLDPC codes by combining learned variable-node scheduling with targeted list exploration, addressing convergence difficulties arising from short cycles and degeneracy. The explicit extension of the RL-S framework and the provision of numerical comparisons on standard benchmarks constitute strengths that allow readers to assess the incremental benefit.","major_comments":[],"minor_comments":[{"comment":"The description of the cumulative path metric and the soft-biasing rule for the second branch would benefit from an explicit equation or pseudocode block in the method section to clarify implementation details.","section":"Method"},{"comment":"The numerical results section should include a table or explicit listing of the exact code parameters (length, rate, girth), training hyperparameters, and number of Monte Carlo trials for each reported curve to support reproducibility.","section":"Numerical Results"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive summary of our work on the RL-LS BP decoder and for recommending minor revision. The referee's description accurately reflects the extension of the RL-S framework with list-based search and the reported numerical results on benchmark QLDPC codes.","responses":[],"tokens_in":1235,"tokens_out":71,"duration_ms":8121,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper takes the earlier RL-based sequential scheduling decoder and adds a list component: at each update the decoder keeps the main path but also branches by softly biasing toward the second-most-likely Pauli symbol, recomputes the local messages, and then ranks the resulting trajectories with a cumulative path metric before pruning. That exact combination is presented as new.\n\nThe work does a clean job of staying inside the BP framework while using the learned policy to pick order and the list to escape some of the bad fixed points that short cycles and degeneracy create. The approach is incremental but targeted, and the abstract indicates the numerical tests on standard QLDPC codes under depolarizing noise show gains over the base RL-S decoder and other BP variants.\n\nThe soft spots are mostly about missing detail rather than outright flaws. The abstract gives no error-rate numbers, code lengths, or training hyperparameters, so the size and consistency of the improvement cannot be judged from what is here. It is also not clear how sensitive the second-best bias and path metric are to the exact LLR values or whether the learned policy transfers to codes or noise levels outside the training set. Those are normal questions for an empirical decoder paper and do not break the central claim.\n\nThis is for people already working on QLDPC decoders or list-BP variants. A reader who needs a practical way to improve convergence on degenerate graphs would find the scheduling-plus-list rule worth testing.\n\nI would send it to peer review. The method is a direct, testable extension of prior work and the evaluation is on representative benchmarks.","headline":"RL-LS extends prior RL-S scheduling with a specific list rule using second-best symbols and cumulative metrics, and claims better QLDPC performance on benchmarks.","tokens_in":2319,"tokens_out":397,"would_cite":false,"duration_ms":17831,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A reinforcement learning list-sequential BP decoder improves decoding performance for quantum LDPC codes over the depolarizing channel.","keywords":["quantum LDPC codes","belief propagation decoding","reinforcement learning","list decoding","sequential scheduling","depolarizing channel","error correction"],"falsifier":"Executing the RL-LS decoder on the same representative QLDPC benchmark codes over the depolarizing channel and measuring a frame error rate or bit error rate no lower than that achieved by the base RL-S decoder or other BP methods would disprove the claimed performance gain.","tokens_in":2611,"feed_emoji":"","tokens_out":766,"duration_ms":17572,"temperature":0.7,"pith_summary":"The paper develops a decoder for quantum LDPC codes that uses reinforcement learning to select the order of variable-node updates in belief propagation. It augments this sequential schedule with list search that also explores a competing branch at each step by softly biasing messages toward the second-most-likely Pauli symbol. The two trajectories are ranked and pruned with a cumulative path metric. The approach targets the convergence failures that standard BP exhibits on these codes because of short cycles and degeneracy. A sympathetic reader would care because more reliable decoders would make quantum error correction on these strong candidate codes more practical.","feed_headline":"List-augmented RL decoder improves quantum LDPC performance","feed_subtitle":"By exploring second-best symbol branches alongside learned variable-node scheduling the method raises success rates on benchmark codes over","key_machinery":"The RL-LS BP decoder, which pairs a learned policy for sequential variable-node scheduling with list exploration of a softly biased second-best symbol branch that is ranked by cumulative path metric.","core_discovery":"The authors introduce the RL-LS BP decoder by extending the RL-S framework with list-based search. At each step a learned policy chooses the next variable node to update; the decoder retains the main trajectory while also creating a competing branch by softly biasing the post-update LLR pair toward the second-most-likely Pauli symbol, recomputing incident local BP messages, and assigning that symbol to the visited node. Candidate trajectories are ranked and pruned using a proposed cumulative path metric. Numerical results on representative QLDPC benchmark codes over the depolarizing channel show that the method improves the decoding performance of the underlying decoder and compares favorabl","pith_inferences":["The same list-augmented scheduling idea could be tested on non-QLDPC quantum codes that also exhibit poor BP convergence.","Combining the RL-LS trajectories with a post-processing step such as ordered-statistics decoding might produce still lower error rates.","The cumulative path metric itself could be studied in isolation to see how much of the gain comes from ranking rather than from the learned policy.","Scaling the approach to larger code lengths would test whether the learned policy continues to generalize when the Tanner graph grows."],"forward_implications":["The decoder raises the success rate of belief-propagation decoding on the tested QLDPC codes over the depolarizing channel.","List exploration of a second-best symbol branch yields better trajectories than the single learned scheduling path alone.","The combined method outperforms or matches several existing BP-based decoders on the same benchmark instances.","Learned sequential scheduling plus list search together mitigate the convergence problems caused by short cycles and degeneracy."],"fun_headline_variants":["RL list sequential BP decoding of quantum LDPC codes","List search extends RL-S for QLDPC belief propagation","Learned variable node scheduling with list branches for QLDPC","RL-LS decoder applies cumulative path metric to quantum LDPC"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The learned policy for variable-node scheduling generalizes across codes and noise levels and the list exploration with soft biasing toward the second-most-likely symbol reliably produces better trajectories than the single learned path.","fun_headline_variants_meta":{"raw":{"variants":["RL list sequential BP decoding of quantum LDPC codes","List search extends RL-S for QLDPC belief propagation","Learned variable node scheduling with list branches for QLDPC","RL-LS decoder applies cumulative path metric to quantum LDPC"]},"model":"grok-4.3","cost_usd":0.00351,"raw_usage":{"total_tokens":1859,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":35099500,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1100,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":64,"duration_ms":8236,"temperature":1.0,"reasoning_tokens":1100,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T15:03:57.900892+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Executing the RL-LS decoder on the same representative QLDPC benchmark codes over the depolarizing channel and measuring a frame error rate or bit error rate no lower than that achieved by the base RL-S decoder or other BP methods would disprove the claimed performance gain.","supporting_citations":[],"review_version":1}