{"id":"d6f7cbaa-0f5b-45e0-b40f-d7b494a3c6ac","arxiv_id":"2501.15588","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The TDSC-ABUS 2023 challenge delivers a public 200-case ABUS benchmark for three tumor analysis tasks and summarizes the top 16 participating algorithms.","lead":"This paper reports the first challenge for tumor detection, segmentation, and classification in automated 3D breast ultrasound (ABUS), built on 200 annotated scans from one hospital. It documents the winning algorithms and makes the benchmark platform publicly available for future AI development in breast ultrasound screening.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Detection leaderboard is invalid as a FROC benchmark because the test set has no healthy/tumor-free ABUS volumes, so false positives on normal scans are never measured.","rationale":"The strongest claim is that the challenge provides a pioneering benchmark for all three ABUS tasks. For that to hold, each task's primary metric must measure what it claims. The detection metric is FROC, which is designed for screening populations containing both positive and negative scans. The challenge test set is 70 tumor-positive cases only; false positives per scan can only be counted on scans that already contain a lesion. Consequently, the detection leaderboard measures sensitivity and 'extra detections per tumor-positive scan,' not the clinically critical false-positive rate on normal tissue. Because the conclusion presents the benchmark as definitive (Abstract; Section 5), this is a load-bearing gap, not a stylistic limitation. The authors themselves flag the missing healthy tissue in Section 4.5, but do not revise the benchmark claim. This concern is distinct from the reader's broader case-mix concern: it is not primarily about demographic diversity or label noise, but about an incorrect match between the FROC definition and the evaluation set. The proposed test is feasible if the challenge archives its Docker submissions, and it would settle whether the detection ranking changes when normal cases are included. Other issues, such as the ad hoc infinite-HD penalty and team count mismatches, are real but secondary; they affect specific table entries without changing the overall winner, whereas the missing negative cases changes the interpretation of the entire detection leaderboard. The paper contributes a genuinely useful challenge platform and dataset; the concern is about the validity of the detection metric for the stated benchmark purpose, not about the quality of the submitted methods. Therefore the conditional verdict remains appropriate pending this check.","tokens_in":22964,"tokens_out":10827,"duration_ms":99849,"concrete_test":"Acquire 30–50 healthy/normal ABUS volumes from the same scanner or a public ABUS source, run the 17 archived Docker submissions (or at least the top-ranked detection teams) on these negative cases, and recompute each team's FROC at the standard false-positive operating points. If any team shows per-scan false-positive rates on normal volumes that are comparable to the inter-team gaps in Table 6, the current detection leaderboard overestimates specificity and the 'definitive benchmark' claim fails for detection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that TDSC-ABUS offers a pioneering benchmark for ABUS CAD assessment, including detection. For detection, the ranking is based solely on FROC (Section 2.3), with sensitivity averaged at false-positive rates of 0.125–8 per scan. However, the entire dataset (Section 2.1, Table 1) consists of tumor-bearing cases; the 70-case test split contains no healthy or tumor-free ABUS volumes. FROC counts false positives per scan, so with no negative scans the metric cannot penalize a detector that fires indiscriminately on normal breast tissue. A non-specific detector that finds the true lesions would receive the same sensitivity as a highly specific one, because no normal scans exist to expose the false positives. Section 4.5 concedes the dataset 'lacked representation of other abnormalities or healthy tissue,' yet the conclusion still presents the benchmark as definitive. This is a mismatch between the FROC definition and the evaluation set, and it directly undermines the detection component of the trifecta benchmark claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the organization, results, and analysis of the TDSC-ABUS 2023 challenge, a MICCAI-affiliated competition for tumor detection, segmentation, and classification in automated 3D breast ultrasound (ABUS). The authors describe a dataset of 200 ABUS volumes with radiologist-derived annotations, the challenge protocol (training, validation, and Docker-based test phases), the evaluation metrics (FROC for detection, Dice/HD for segmentation, ACC/AUC for classification), and the ranking scheme. The paper summarizes the algorithms of the participating teams and discusses common strategies and limitations. The central contribution claimed is an inaugural public benchmark combining the three tasks for ABUS CAD assessment.","tokens_in":23147,"tokens_out":5073,"duration_ms":47024,"significance":"If the inconsistencies and evaluation limitations are properly addressed, this is a useful community resource: the released dataset addresses a genuine lack of publicly available labeled ABUS data, and the challenge framework with held-out test evaluation and per-case disclosure is a sound basis for future benchmarking. The paper includes a large international participation (106 approved teams, 17 qualified submissions) and documents a broad range of technical approaches. However, the strength of the benchmark claims is currently undercut by the absence of tumor-free volumes in the detection evaluation and by several internal numerical and naming inconsistencies that must be resolved before the results can be considered reliable.","major_comments":[{"comment":"The reported number of participating teams is internally inconsistent. Section 3.1 states that 17 qualified submissions were obtained; Section 3.2 says \"We have 17 teams that produced valid results, but one of them did not submit short paper... Thus we summarize key point of other 16 teams,\" and yet the text immediately proceeds to describe 18 distinct teams (T1 through T18). Table 2 is titled \"Summary of the benchmark methods of top ten teams.\" These discrepancies make it impossible for the reader to know how many teams actually contributed to each leaderboard and which algorithms are included in each summary. This is load-bearing because the manuscript's contribution is a documented benchmark and leaderboard, not just a set of algorithm descriptions.","section":"Section 3.1, Section 3.2, Table 2"},{"comment":"The detection task is evaluated with FROC, which averages sensitivity over false-positive rates per scan, but the dataset contains no tumor-free ABUS volumes (Table 1 shows only malignant and benign cases; Section 4.5 concedes the dataset \"lacked representation of other abnormalities or healthy tissue\"). With all test scans containing at least one lesion, the FROC false-positive rate can only reflect false detections within tumor-bearing volumes, and cannot measure the detector's tendency to fire on healthy tissue. For a screening-oriented CAD benchmark, this omission is consequential: a non-specific detector that always outputs a candidate region would not be penalized for false positives on normal scans because no such scans exist. The limitation is acknowledged in Section 4.5, but the central claim in Section 1 of a \"pioneering benchmark for ABUS CAD algorithm assessment\" is not tempered accordingly, and the detection leaderboard is presented without this caveat in Section 3.3.3. The authors should either include tumor-free test volumes (or explicitly state that the detection benchmark is limited to lesion-positive volumes) and revise the wording of the benchmark claim.","section":"Section 2.1, Section 2.3, Section 4.5"},{"comment":"The classification results text says \"with Team Shiontao clearly leading, particularly in AUC,\" but no team with this name appears in Table 5 or anywhere else in the manuscript. The leading team in Table 5 is T1 (SZU). This appears to be a leftover from another challenge or an editing artefact, and it directly confuses the reported ranking. The sentence should be corrected or the team name mapped to the table entry.","section":"Section 3.3.2"},{"comment":"The relationship between the unpenalized and penalized segmentation leaderboards is not explained clearly. Table 3 lists 10 teams with finite HD values; Table 4, after the 105% penalty for infinite HD, lists 14 teams, adding T8, T10, T1, and T4. The text does not state which teams had \"inf\" HD scores, why some teams appear only in Table 4, or how the per-case replacement of \"inf\" with 105% of the worst valid HD score affects the normalization step. Without this information, the segmentation ranking—which is a core result—is not reproducible from the tables as presented.","section":"Section 3.3.1, Tables 3 and 4"}],"minor_comments":[{"comment":"The keywords block in the article front matter reads \"Segmentation, Pulmonary artery, Multi-level, Efficiency,\" which is unrelated to the paper's content on breast ultrasound and appears to be a template leftover. The keywords should be replaced with appropriate ABUS/challenge-related terms.","section":"Article header / Front matter"},{"comment":"Figure 10 includes teams T8 and T9 in the scatter plot, but Table 5 lists only 8 classification teams and does not include T8 or T9. The figure and table should be reconciled, or the figure legend should clarify which teams are shown.","section":"Figure 10"},{"comment":"The abstract states that ABUS has advantages \"over handheld mammography,\" which appears to be a typo for \"handheld ultrasound\" (or should be phrased differently). Mammography and ultrasound are distinct modalities, and the comparison as written is implausible.","section":"Abstract"},{"comment":"The text says label noise in the training dataset \"was corrected in the test set to ensure fairness,\" but the paper does not describe how this correction was performed or how the test set labels were verified. A brief explanation (e.g., re-annotation protocol, number of cases) would improve transparency.","section":"Section 4.5"},{"comment":"The x-axis label of Figure 11 says \"False Positives per Image (FPPI),\" while the text consistently refers to false positives per scan. The terminology should be aligned to avoid ambiguity about whether a \"scan\" and an \"image\" are the same unit in the FROC computation.","section":"Figure 11"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a thorough technical editing pass: the template keyword leftovers, the appearance of a nonexistent team name, and the inconsistent team counts suggest that the final version was not carefully checked. The authors should also be asked to align the benchmark claims with the acknowledged limitation of lacking healthy cases; this is a substantive issue for a challenge paper whose primary contribution is a benchmark. With these corrections, the paper could become a solid contribution to the ABUS CAD literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper is a legitimate benchmark contribution, not just a methods collection. The TDSC-ABUS challenge is the first public ABUS dataset I know of that includes detection, segmentation, and classification labels in one place, and the organizers followed the BIAS reporting guidelines. The test split is held out, runs were done in Docker, and the per-case scores are public. For a field that has been stuck on method papers with private data, this is useful infrastructure. The team method summaries are mostly standard adaptations—nnU-Net, YOLO, MedSAM—but that's fine; the value is in the benchmark, and the paper doesn't oversell the algorithms. The soft spots are real but not fatal. The front matter still has a leftover 'Pulmonary artery' keyword block from another paper. The team accounting is inconsistent: 17 valid submissions, text says 16 were summarized, but I count 18 team descriptions (T1–T18), and 'Team Shiontao' appears without definition. The evaluation also leans on min-max normalization across the participant pool, which makes rankings relative to whoever shows up, and the 105% infinity-HD penalty is ad hoc. No confidence intervals or significance tests. These are all correctable and should be fixed. On the stress-test concern: I don't think it lands. FROC counts false positives per scan on all scans, and every ABUS volume here contains plenty of normal breast tissue even though every case has a lesion. A detector that fires indiscriminately on normal tissue would accumulate FPs on these scans exactly as it would on healthy scans. The lack of entirely tumor-free scans is a limitation for estimating the false-positive rate per scan at the whole-volume level, but it doesn't invalidate the detection ranking as claimed. The bigger generalization problem is single-center, single-protocol data and the absence of healthy cases, which the paper acknowledges. Bottom line: this is a useful resource with sloppy presentation. Send it to a serious referee; the issues are fixable. I would not cite the leaderboard numbers as definitive truth without reading the limitations, but I would cite the dataset.","headline":"A genuinely useful ABUS benchmark that needs editorial cleanup and a more careful limitations section before its leaderboard should be treated as definitive.","tokens_in":23849,"tokens_out":3669,"would_cite":true,"duration_ms":31992,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The TDSC-ABUS challenge supplies the first open benchmark that jointly scores tumor detection, segmentation, and classification on 200 automated 3D breast ultrasound volumes, and it reports rankings of 17 validated algorithm submissions.","keywords":["automated breast ultrasound","ABUS","tumor detection","tumor segmentation","tumor classification","benchmark challenge","computer-aided diagnosis","deep learning"],"falsifier":"Take the same 17 validated algorithm containers to an independent multi-center ABUS dataset with second-opinion radiologist labels and recompute all scores; if the leaderboard order changes substantially or the Dice-versus-Hausdorff ordering flips, the benchmark's claim to be a definitive yardstick is refuted.","tokens_in":1564,"feed_emoji":"🩺","tokens_out":13253,"duration_ms":142351,"temperature":0.7,"pith_summary":"This paper reports the first public benchmark that covers all three core computer-aided diagnosis tasks - tumor detection, segmentation, and classification - on automated 3D breast ultrasound (ABUS) images, and it uses that benchmark to compare 17 algorithms that completed a standardized evaluation. The organizers built the benchmark from 200 ABUS volumes with expert-drawn tumor boundaries and benign/malignant labels, split into 100 training, 30 validation, and 70 test cases, and ran it as an international challenge with independent scoring per task plus a combined leaderboard. The result the authors are trying to establish is that a fixed, openly accessible ABUS dataset and evaluation protocol can serve as a common yardstick for future research. A sympathetic reading sees the contribution as the benchmark itself: the paper's descriptions of winning strategies are secondary to the public platform they document.","feed_headline":"First 3D breast-ultrasound benchmark ranks 17 AI pipelines","feed_subtitle":"A 200-scan, three-task scoring protocol gives future algorithms one public yardstick to beat.","key_machinery":"The carrying object is the challenge protocol itself: 200 three-dimensional ABUS volumes with known voxel spacings, radiologist-drawn tumor masks, and benign/malignant labels, with detection bounding boxes derived from those masks. Three independent metrics are normalized by min-max scaling and combined by fixed formulas; for instance, the segmentation score is (1 + normalized Dice - normalized Hausdorff)/2, and the overall score sums the three task scores, with infinite Hausdorff values replaced by 105% of the worst valid value. Each final submission ran inside a Docker container in a standardized environment, so the leaderboard reflects the algorithms rather than their host machines. This protocol is what turns a collection of images into a benchmark: it fixes the data split, the hit criterion (a detection counts when its box has IoU over 0.3 with ground truth), and the ranking rule.","core_discovery":"The central claim is that the TDSC-ABUS challenge provides the first standard benchmark for assessing ABUS computer-aided diagnosis algorithms on detection, segmentation, and classification together. The authors argue this matters because ABUS images are hard for algorithms - tumors are small relative to the volume, boundaries are blurred, lesion-to-background similarity is high, and no public well-labeled dataset existed to compare methods fairly. On this dataset the best combined score was achieved by a pipeline of three specialized networks, one per task, while several leading segmentation entries built on U-Net-style architectures with patch-based training and ensembles; detection was frequently handled by deriving bounding boxes from segmentation output. The evaluation protocol scores segmentation with Dice and Hausdorff distance, classification with accuracy and area under the ROC curve, and detection with the free-response ROC curve averaged over false-positive levels from 0.125 to 8 per scan. The paper's thesis is that this combination of public data, metrics, and containerized testing makes algorithm rankings reproducible and gives future work a concrete baseline to beat.","pith_inferences":["If the organizers add healthy cases and multi-center data, as the paper says they plan to, the current leaderboard should be re-run; rankings that hold up would be robust, while shifting rankings would show the 2023 results were partly artifacts of a single scanner and reader pool.","The paper's 105%-of-worst penalty for infinite Hausdorff distances is one of several defensible choices, so the leaderboard should be read as a comparison under that specific rule rather than an absolute ordering of methods.","The finding that the best combined performance came from three specialized models rather than one end-to-end network suggests that, at this data size, task-specific inductive biases still beat joint training; a future multi-task architecture might challenge that conclusion.","A future challenge could score runtime and resource use; the paper notes it did not, so reported performance does not yet distinguish clinically deployable methods from heavy ensembles."],"forward_implications":["Future ABUS algorithms can be compared against a public, fixed test set instead of private data, making published gains auditable.","Because the combined leaderboard ranks only teams that solve all three tasks, the benchmark creates pressure to build integrated pipelines rather than single-task solutions.","Segmentation quality must be reported with both overlap and boundary accuracy; the results show that a high Dice score does not imply a low Hausdorff distance.","Detection performance is better summarized by the FROC curve across false-positive rates than by a single accuracy number, since lesions occupy a small fraction of each volume.","Segmentation-guided detection and patch-based training are the strategies that most consistently worked across leading entries."],"supporting_citations":[{"why":"Provides the transparent-challenge reporting standard (BIAS) that the organizers say they followed when designing TDSC-ABUS.","marker":"Maier-Hein et al. (2020)"},{"why":"Supplies the caution that biomedical challenge rankings must be interpreted carefully, motivating the paper's ranking design.","marker":"Maier-Hein et al. (2018)"},{"why":"The nnU-Net self-configuring segmentation method used by multiple leading teams as their backbone.","marker":"Isensee et al. (2021)"},{"why":"nnDetection, the self-configuring detection framework that one leading team adapted for tumor detection.","marker":"Baumgartner et al. (2021)"},{"why":"Segment Anything, used as the base model for a team's fine-tuned medical segmentation approach.","marker":"Kirillov et al. (2023)"},{"why":"U-Net, the basic architecture several teams chose for segmentation and classification branches.","marker":"Ronneberger et al. (2015)"}],"fun_headline_variants":["First open 3D breast-ultrasound benchmark ranks 17 AI pipelines","New benchmark gives 3D breast ultrasound AI a public test to beat","17 AI pipelines face off on first public 3D breast ultrasound benchmark","Three-task benchmark for 3D breast ultrasound AI: 17 pipelines ranked"],"cache_read_input_tokens":25856,"weakest_assumption_plain":"The whole benchmark rests on 200 scans from one hospital with one imaging protocol and labels drawn by radiologists from that setting, so if the case mix, scanner, or boundaries are not representative of other clinics, the rankings will not transfer there.","fun_headline_variants_meta":{"raw":{"variants":["First open 3D breast-ultrasound benchmark ranks 17 AI pipelines","New benchmark gives 3D breast ultrasound AI a public test to beat","17 AI pipelines face off on first public 3D breast ultrasound benchmark","Three-task benchmark for 3D breast ultrasound AI: 17 pipelines ranked"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000788,"raw_usage":{"total_tokens":3503,"prompt_tokens":999,"completion_tokens":2504,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":2422}},"tokens_in":615,"tokens_out":2504,"duration_ms":17959,"temperature":1.0,"reasoning_tokens":2422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:08:15.958045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 17 validated algorithm containers to an independent multi-center ABUS dataset with second-opinion radiologist labels and recompute all scores; if the leaderboard order changes substantially or the Dice-versus-Hausdorff ordering flips, the benchmark's claim to be a definitive yardstick is refuted.","supporting_citations":[{"cited_title":", author Reinke, A","cited_arxiv_id":null,"evidence_quote":"Provides the transparent-challenge reporting standard (BIAS) that the organizers say they followed when designing TDSC-ABUS."},{"cited_title":", author Eisenmann, M","cited_arxiv_id":null,"evidence_quote":"Supplies the caution that biomedical challenge rankings must be interpreted carefully, motivating the paper's ranking design."},{"cited_title":", author Jaeger, P.F","cited_arxiv_id":null,"evidence_quote":"The nnU-Net self-configuring segmentation method used by multiple leading teams as their backbone."},{"cited_title":", author J \\\"a ger, P.F","cited_arxiv_id":null,"evidence_quote":"nnDetection, the self-configuring detection framework that one leading team adapted for tumor detection."}],"review_version":1}