{"id":"03cc94d9-ce8b-49b2-b476-a7976da70495","arxiv_id":"1908.04767","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A deep learning pipeline can grade hemosiderin-laden macrophages in equine lung cytology slides at human-expert-level concordance, with whole-slide scoring in under two minutes.","lead":"This paper trains deep learning models to automatically count and grade iron-laden macrophages in horse lung fluid slides, and compares them to nine human experts. The automated system matches or slightly exceeds the average human expert at single-cell grading, and can score a whole slide in under two minutes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The single reference pathologist's annotations are used as ground truth; with inter-observer kappa 0.67, the reported 0.85 concordance and 0.66 mAP measure agreement with one expert, not diagnostic accuracy.","rationale":"The reader identified the single pathologist ground truth as the weakest assumption; I agree. The paper is honest about this limitation in the Discussion, but the abstract and conclusion nevertheless make a strong accuracy claim ('accurate, reproducible and quick EIPH scoring') that is not supported without an external reference. The dataset is large and the code release is a real contribution, and the method may still be valuable for reproducibility even if it only matches one expert. But the central quantitative claims—0.85 concordance and 0.66 mAP—are computed against that same expert's labels, so they are agreement scores, not accuracy scores. The proposed check is feasible because the nine-expert ratings already exist; no new data collection is required. If the check passes, the concern is largely resolved; if it fails, the verdict should become REJECT for the accuracy claim or the claim should be narrowed to 'reproduces the reference expert.' Since the reader's verdict was already CONDITIONAL and this is the same condition, no verdict change is needed.","tokens_in":13105,"tokens_out":3687,"duration_ms":40190,"concrete_test":"Using the already-collected ratings of the 2,000 cells from all nine experts, construct a consensus reference (majority vote or adjudicated label) and recompute the DL concordance and each expert's concordance against this consensus, plus pairwise DL-vs-each-expert concordance. If the DL model's consensus concordance is at or above the mean expert consensus concordance, the ground-truth concern is mitigated; if it drops materially below the expert mean or is only high against the single reference, the headline accuracy claim is an artifact of training to one pathologist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of an accurate automated EIPH scoring pipeline rests on treating one veterinary pathologist's annotations as ground truth. The paper explicitly states in the Discussion that 'all annotations were made by a single veterinary pathologist' and 'there is currently no true gold standard methods such as chemical measurement of iron content.' Human experts themselves only reach Fleiss' kappa = 0.67 inter-observer and 0.68–0.88 intra-observer concordance, and their concordance with the reference ranges 0.68–0.86. Thus the reference label is one person's subjective reading, not a stable truth. Against that reference, the DL classifier's 0.85 concordance and the detector's mAP of 0.66 quantify how well the model reproduces that particular expert's judgements. The conclusion that the pipeline enables 'accurate' scoring is therefore not established; it could simply be an idiosyncratic expert emulator. The paper's Fig. 1 even shows regions the human expert missed but RetinaNet detected, which is direct evidence that the ground truth has omissions. This does not invalidate the methods as reproducibility tools, but it does invalidate the accuracy framing and the 'partially exceeding human concordance' headline as a statement about diagnostic quality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript addresses automated scoring of exercise-induced pulmonary hemorrhage (EIPH) in equine bronchoalveolar lavage cytology whole slide images. The authors introduce a dataset of 17 fully annotated WSIs with 78,047 hemosiderophages annotated by one veterinary pathologist, evaluate nine human experts on single-cell classification with repeated sessions, and compare them with deep-learning classifiers and regressors. They also train RetinaNet-based object detectors with a novel quad-tree sampling strategy to detect and grade cells directly on gigapixel WSIs. The reported single-cell concordance is 0.85 for the deep-learning methods versus 0.68 to 0.86 for human experts, inter-observer Fleiss' kappa is 0.67, and the object detector reaches mAP 0.66 with a mean score error of 9 and inference under two minutes per slide. The paper concludes that the pipeline enables accurate, reproducible, and quick EIPH scoring.","tokens_in":13379,"tokens_out":6863,"duration_ms":65616,"significance":"If the central claim were appropriately bounded, the paper would be a useful contribution to veterinary digital pathology: it provides the largest fully annotated EIPH cytology dataset of which I am aware, a systematic human-variability study for this task, a novel quad-tree sampling strategy, and an open, fast object-detection pipeline. The reproducibility of the code and trained model, together with the quantitative documentation of inter- and intra-observer variability, are clear strengths. The main value is as a benchmark and as evidence that automated systems can match a specific expert's grading style; the claim of diagnostic accuracy is not supported by the evidence and needs revision.","major_comments":[{"comment":"The reference standard is a single veterinary pathologist's annotation, and the paper itself states that there is no true gold standard such as chemical measurement of iron content. Because the nine experts reach only moderate inter-observer agreement (Fleiss' kappa = 0.67) and intra-observer concordance of 0.68 to 0.88, the reported deep-learning concordance of 0.85 and the detection mAP of 0.66 quantify agreement with that one expert's grading style rather than diagnostic accuracy. The abstract and conclusion claim that the pipeline enables 'accurate' EIPH scoring, which is not supported by the present evidence; this claim must be reframed as agreement with an expert reference, or the model must be validated against an independent gold standard such as clinical outcome or chemical iron quantification.","section":"Material; Discussion and Outlook"},{"comment":"The human comparison is restricted to single-cell classification on pre-extracted cells, whereas the object-detection pipeline performs detection and classification on whole slides. The abstract's statement that the deep-learning approach 'partially exceeding human expert concordance' refers only to the single-cell task, not to the end-to-end WSI scoring task; no human expert scored whole slides, so the mAP of 0.66 and the score error of 9 have no human WSI-level comparator. The paper should either provide such a comparison, for example by having experts score the three test slides at the whole-slide level, or explicitly restrict the superiority claim to single-cell classification.","section":"Human Performance Evaluation; Results (Object Detection)"},{"comment":"The key mAP and concordance values are reported without confidence intervals, and the mAP of 0.66 is computed from only three test slides. Given that the upper human concordance is 0.86 and the deep-learning concordance is 0.85, the claim of partially exceeding human performance depends on a difference well within plausible sampling variability. The paper should report per-slide results, confidence intervals, and ideally a statistical comparison with the human observers.","section":"Results (Object Detection)"},{"comment":"The ground truth is known to be incomplete: Fig. 1 shows a region missed by the human annotator but detected by RetinaNet. Consequently the reported mAP of 0.66 is a conservative lower bound, because some false-positive detections may be true hemosiderophages absent from the reference. The paper acknowledges this qualitatively but should quantify its potential impact, for example by estimating missed-cell prevalence on a small re-annotated subset or by reporting the sensitivity of the conclusions to label noise.","section":"Object Detection Evaluation; Fig. 1"},{"comment":"The derived 'hypothetical mAP' of 0.57 to 0.74 for human experts assumes perfect detection and converts single-cell classification concordance into a detection metric. This is not a measured human detection baseline and should not be presented as one; the conversion needs a derivation or should be removed.","section":"Results (Cell Classification)"}],"minor_comments":[{"comment":"The word 'Resultsf' is a typo and should read 'Results'.","section":"Abstract"},{"comment":"The text contains typos including 'Unfortunatelly', 'Scince', and 'variablity'; these should be corrected.","section":"Discussion and Outlook"},{"comment":"The acronym for the single-cell task is inconsistent: 'CoSH' appears in the abstract while 'CoCH' appears in Methods and Discussion; use one acronym throughout.","section":"Abstract; Methods"},{"comment":"Equation (1) appears to contain the regression loss term twice with identical notation; clarify that one term corresponds to the cell regression head and the other to the patch regression head, and define c_i and ĉ_i separately for each head.","section":"Methods, Eq. (1)"},{"comment":"The table caption says 'results per WSI', but the table rows are architectures rather than slides; provide a per-slide breakdown or change the caption accordingly.","section":"Table 2"},{"comment":"The phrase 'a convexity value of 0.1' is ambiguous; this is presumably the RBF cost parameter C and should be named and defined.","section":"Methods (Support vector machine)"},{"comment":"The term 'maximal learning rate schedule' should be explained, for example by specifying whether a one-cycle policy or a similar schedule was used, so that the training protocol is reproducible.","section":"Methods (Single Cell Classification)"},{"comment":"The concordance metric should be defined explicitly, and results should be reported separately for test set I and test set II because the two sets have different class distributions.","section":"Results (Cell Classification)"}],"recommendation":"major_revision","confidential_remarks":"I am not recommending rejection because the flaws are primarily in the interpretation and statistical reporting rather than in the engineering itself, and the necessary caveats are partially acknowledged in the Discussion. The authors should be asked to revise the accuracy claim, add confidence intervals and per-slide results, and clarify that the human comparison is single-cell only."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a good applied paper with a genuinely new dataset, and the results are honestly presented. The main thing to know is that the 'accuracy' claim rests on a single pathologist's annotations, so treat it as reproducibility of one expert's judgment rather than diagnostic ground truth.\n\nWhat's actually new: 17 completely annotated whole-slide images of equine BAL cytology with 78,047 labeled hemosiderophages—that's a real contribution. Applying RetinaNet-based detection to the five-grade EIPH scoring is new, and the quad-tree sampling strategy is a practical trick that works. The human performance study with nine experts and two sessions is careful and gives useful variability numbers, including the low inter-observer kappa. Code and trained model are available, which matters.\n\nThe paper does well on its own terms. The single-cell classification matches human-level concordance (0.85 vs 0.68–0.86 for humans), and the WSI detection runs in under two minutes. That is genuinely useful for standardizing a monotonous task.\n\nSoft spots: the ground truth problem is real and the paper admits it. All annotations came from one veterinary pathologist; there is no independent gold standard. With Fleiss kappa at 0.67, experts often disagree with each other, so the model may be emulating one person's style rather than measuring true hemosiderin load. The paper even shows that the human annotator missed cells that the detector found—good honesty, but it makes the 0.85 concordance and mAP 0.66 harder to interpret as diagnostic accuracy. Also, the human baseline is only on single cells, not on whole-slide scoring, so the comparison is not apples-to-apples. Minor: no confidence intervals, and the mAP is moderate.\n\nNone of this kills the paper. The dataset alone deserves referee time, and the limitations are openly stated. What is needed is a tempering of the 'accurate, reproducible' conclusion: it is accurate relative to one expert.\n\nWho should read it: anyone doing applied DL in cytology or WSI cell detection, and people who want a concrete example of how to handle sparse class distributions. I would cite it for the dataset and the sampling strategy.\n\nYes, send it to peer review. A serious referee can push for a more careful interpretation and maybe external validation.","headline":"A useful applied DL paper with a new dataset and honest limitations, but the 'accuracy' claim is really agreement with one pathologist.","tokens_in":13968,"tokens_out":2177,"would_cite":true,"duration_ms":21388,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automated deep learning can grade pulmonary hemosiderophages on whole slide cytology images with 85% concordance, matching or exceeding the average human expert while taking under two minutes per slide.","keywords":["exercise-induced pulmonary hemorrhage","hemosiderophages","whole slide image analysis","deep learning object detection","RetinaNet","bronchoalveolar lavage cytology","observer variability","cell grading"],"falsifier":"Measure the iron content of bronchoalveolar lavage fluid from the same 17 cases by atomic absorption spectroscopy or a similar chemical assay and compare it with the automated Total Hemosiderin Score. A correlation near zero, or a case where the algorithm grades a high-iron sample as low-iron, would show the system measures visual staining patterns rather than the underlying hemorrhage.","tokens_in":12931,"feed_emoji":"🐴","tokens_out":7425,"duration_ms":74086,"temperature":0.7,"pith_summary":"Exercise-induced pulmonary hemorrhage in racehorses is diagnosed by grading iron-laden macrophages in bronchoalveolar lavage cytology, a manual task that is slow and subjective. This paper claims that a deep learning pipeline can perform the grading automatically on whole slide images, reaching 85% concordance with a pathologist's ground truth and running in under two minutes per slide. That performance sits at the upper end of nine human experts, whose concordance ranged from 68% to 86% with a mean of 73%, and whose intra-observer agreement was also variable. The authors argue that automation therefore offers a reproducible alternative to a scoring system that even experts apply inconsistently.","feed_headline":"Deep learning scores horse lung bleeding as well as experts","feed_subtitle":"Model agrees with the reference pathologist on 85% of cells and processes whole slides in under two minutes.","key_machinery":"The carrying mechanism is a modified RetinaNet detector: a feature pyramid network built on a ResNet-18 backbone predicts object boxes and classes at multiple scales, and an added regression head predicts a continuous hemosiderin score per cell, while a patch regression head estimates the slide score directly. To train on rare high-grade cells, the paper introduces a quad-tree based patch sampler that allocates sampling probability according to cell density and grade, so grade-3 and grade-4 cells are not starved of training examples. The detector processes whole slides in about 101 seconds on a modern GPU.","core_discovery":"The central claim is that a single end-to-end object detection system, built on RetinaNet with a ResNet-18 backbone and an extra regression head, can both locate and grade hemosiderophages across gigapixel whole slide images. On a new, fully annotated dataset of 17 slides containing 78,047 cells, the system achieves a mean average precision of 0.66 over the five Golde grades and a cell-level concordance of 0.85 with the reference annotation, exceeding the mean human concordance of 0.73. The same pipeline estimates the whole-slide Total Hemosiderin Score with a mean error of about 9 points on a 0-400 scale, compared with 19 for deep regression and 21 for an SVM baseline. The paper also establishes that human grading is highly variable: inter-observer Fleiss' kappa is 0.67 and intra-observer concordance ranges from 0.68 to 0.88, which is why the authors propose automation as a more reliable route to EIPH scoring.","pith_inferences":["If the reference labels are treated as noisy rather than as truth, the model's 0.85 concordance could understate how well it matches an expert consensus; a panel majority-vote benchmark would be a stronger test.","The missing independent validation the paper itself notes could be supplied by measuring iron content chemically in the same lavage samples; a strong correlation would confirm that the visual grades track actual hemorrhage.","The same quad-tree sampling and regression-head detector should transfer to other rare-cell grading tasks, such as human pulmonary hemorrhage cytology, provided external datasets are used.","Because human baseline variability is high, future comparisons should report agreement against a consensus ground truth rather than a single pathologist's labels."],"forward_implications":["EIPH scoring can be fully automated on whole slide images, removing the need for a human to select and grade individual cells; this would make the test faster and cheaper in routine equine practice.","Because the algorithm is deterministic, repeated runs give identical scores, directly addressing the observed 0.68-0.88 intra-observer variability and 0.67 inter-observer kappa of human raters.","The system's mAP of 0.66 approaches the estimated human upper bound of 0.74 mAP given the measured concordance levels, so further gains are limited more by label noise than by detector architecture.","Regression-based grading can reveal within-grade differences in iron load, offering a continuous measure of hemorrhage severity instead of the five discrete Golde grades.","The released dataset and model provide a basis for building interactive annotation tools that flag regions missed by human experts."],"supporting_citations":[{"why":"Defines the five-grade hemosiderin scoring system used to label every cell and to compute EIPH scores.","marker":"4"},{"why":"Adapts the Golde score to horses and defines the Total Hemosiderin Score with the diagnostic threshold of 75.","marker":"34"},{"why":"Supplies the RetinaNet detector with focal loss that the paper modifies for whole-slide analysis.","marker":"23"},{"why":"Faster R-CNN serves as a two-stage object detection baseline for comparison.","marker":"21"},{"why":"SSD serves as a one-stage detection baseline for comparison.","marker":"22"},{"why":"Provides the annotation tool and database used to build the fully labeled whole-slide dataset.","marker":"33"},{"why":"Provides the residual backbone used for feature extraction in classification and detection.","marker":"35"},{"why":"Provides the feature pyramid network for multi-scale feature extraction in the detector.","marker":"37"},{"why":"Documents speed and accuracy trade-offs of modern detectors and motivates the RetinaNet choice.","marker":"32"},{"why":"Earlier application of the same detection approach to fully annotated multiclass whole-slide images; the closest prior work.","marker":"30"}],"fun_headline_variants":["AI matches human experts at grading horse lung bleeding","Deep learning rivals cytology experts for EIPH scoring","Automated horse lung bleed scoring: as good as experts","AI lung bleed grading in horses matches expert level","Horse EIPH scoring: AI equals human expertise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that the single veterinary pathologist's annotations are a valid ground truth for grading, even though the paper reports only moderate inter-observer agreement (Fleiss' kappa = 0.67) and no independent gold standard such as a chemical iron measurement exists.","fun_headline_variants_meta":{"raw":{"variants":["AI matches human experts at grading horse lung bleeding","Deep learning rivals cytology experts for EIPH scoring","Automated horse lung bleed scoring: as good as experts","AI lung bleed grading in horses matches expert level","Horse EIPH scoring: AI equals human expertise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001238,"raw_usage":{"total_tokens":5133,"prompt_tokens":1044,"completion_tokens":4089,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":4012}},"tokens_in":660,"tokens_out":4089,"duration_ms":30354,"temperature":1.0,"reasoning_tokens":4012,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:48:11.576176+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the iron content of bronchoalveolar lavage fluid from the same 17 cases by atomic absorption spectroscopy or a similar chemical assay and compare it with the automated Total Hemosiderin Score. A correlation near zero, or a case where the algorithm grades a high-iron sample as low-iron, would show the system measures visual staining patterns rather than the underlying hemorrhage.","supporting_citations":[{"cited_title":"W., Drew, W","cited_arxiv_id":null,"evidence_quote":"Defines the five-grade hemosiderin scoring system used to label every cell and to compute EIPH scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Adapts the Golde score to horses and defines the Total Hemosiderin Score with the diagnostic threshold of 75."},{"cited_title":"& Dollár, P","cited_arxiv_id":null,"evidence_quote":"Supplies the RetinaNet detector with focal loss that the paper modifies for whole-slide analysis."},{"cited_title":"& Sun, J","cited_arxiv_id":null,"evidence_quote":"Faster R-CNN serves as a two-stage object detection baseline for comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SSD serves as a one-stage detection baseline for comparison."},{"cited_title":"& Maier, A","cited_arxiv_id":null,"evidence_quote":"Provides the annotation tool and database used to build the fully labeled whole-slide dataset."},{"cited_title":"& Sun, J","cited_arxiv_id":null,"evidence_quote":"Provides the residual backbone used for feature extraction in classification and detection."},{"cited_title":"Feature pyramid networks for object detection","cited_arxiv_id":null,"evidence_quote":"Provides the feature pyramid network for multi-scale feature extraction in the detector."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents speed and accuracy trade-offs of modern detectors and motivates the RetinaNet choice."},{"cited_title":"Towards Deep Cellular Phenotyping in Placental Histology","cited_arxiv_id":"1804.03270","evidence_quote":"Earlier application of the same detection approach to fully annotated multiclass whole-slide images; the closest prior work."}],"review_version":1}