{"id":"1cb16e7c-b312-45c4-b55b-da5ca0cab20f","arxiv_id":"2606.05458","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"YOLOv12, optical flow thresholding, and fine-tuned VideoMAE achieve macro-F1 of 0.898 for blink classification and 0.926 for binary detection on a public horse video dataset.","lead":"This paper develops three computer vision methods to automatically detect and classify eye blinks in horse videos. If effective, it could support automated monitoring of pain and stress for better equine welfare.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Blink categories' alignment with affective states lacks independent validation","rationale":"The reader's weakest_assumption matches the load-bearing link between technical performance and the paper's stated application. No other internal inconsistency in the reported scores is identifiable from the given abstract and claim.","tokens_in":1678,"tokens_out":218,"duration_ms":33640,"concrete_test":"Locate and list all citations in §1 or §2 of the full manuscript that validate half/full blinks as affective state indicators; if none exist or they are only general references, re-run the reported F1 evaluation after consulting domain experts on category alignment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim reports high F1 scores for blink classification/detection to support equine affective state assessment. This requires that the chosen categories (half/full blinks) align with established pain/stress indicators. The abstract states they are 'recognised indicators' but provides no citations, equine welfare studies, or validation for the specific classification scheme used in the public dataset evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript develops and evaluates three methods for automated detection and classification of half and full eye blinks in horse videos (YOLOv12 frame-based detector, optical flow magnitude thresholding, and fine-tuned VideoMAE) on a publicly available dataset. It reports macro-F1 scores of 0.898 for blink classification and 0.926 for binary detection, positioning the work as a step toward automated equine affective state assessment via facial action units.","tokens_in":1703,"tokens_out":449,"duration_ms":27455,"significance":"If the evaluation details and category validation were provided, the results could offer a practical contribution to fine-grained AU detection in animal videos. The work highlights challenges in micro-expression detection but its significance for welfare applications is limited by the unvalidated link between the specific blink categories and pain/stress indicators.","major_comments":[{"comment":"Abstract: the claim that half and full blinks are 'recognised indicators of pain and stress' is stated without any citations to equine welfare literature or validation studies, which is load-bearing for the paper's application to affective state assessment.","section":"Abstract"},{"comment":"Methods/Results: the abstract reports F1 scores of 0.898 and 0.926 but supplies no dataset size, class distribution, train/test split, cross-validation procedure, or per-class error analysis, preventing assessment of whether the scores are supported by rigorous evaluation.","section":"Methods"},{"comment":"Introduction/Results: the central application claim requires that the chosen blink categories align with established pain/stress indicators, yet the manuscript provides no independent validation or references for this alignment on the public dataset.","section":"Introduction"}],"minor_comments":[{"comment":"Abstract: the three methods are listed but it is unclear which achieves the headline scores or how they compare in the reported results.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's scope (technical CV methods for a veterinary application) may be a better fit for an interdisciplinary or animal science venue than a core computer vision journal; the lack of methodological transparency suggests it is not yet ready for review without major additions."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments. We address each major comment point by point below, indicating where revisions will be made.","responses":[{"response":"We agree that citations are required to support this claim. We will add relevant references from equine welfare literature in the revised abstract and introduction to substantiate the statement.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the claim that half and full blinks are 'recognised indicators of pain and stress' is stated without any citations to equine welfare literature or validation studies, which is load-bearing for the paper's application to affective state assessment."},{"response":"The full manuscript provides these details in the Methods and Results sections. To address the concern, we will revise the abstract to include key information on dataset size, split, and evaluation procedure while maintaining conciseness.","revision_made":"yes","referee_comment":"[Methods] Methods/Results: the abstract reports F1 scores of 0.898 and 0.926 but supplies no dataset size, class distribution, train/test split, cross-validation procedure, or per-class error analysis, preventing assessment of whether the scores are supported by rigorous evaluation."},{"response":"We will add references supporting the alignment of half and full blinks with established pain/stress indicators. Independent validation on the public dataset is not included, as the work focuses on detection methods rather than clinical validation.","revision_made":"partial","referee_comment":"[Introduction] Introduction/Results: the central application claim requires that the chosen blink categories align with established pain/stress indicators, yet the manuscript provides no independent validation or references for this alignment on the public dataset."}],"tokens_in":1293,"tokens_out":418,"duration_ms":33155,"standing_objections":["Independent validation of the blink categories as pain/stress indicators on the public dataset (beyond adding references), as this is outside the scope of the detection-focused study."]},"desk_editor":{"model":"grok-4.3","letter":"The paper applies three off-the-shelf video methods to detect half and full blinks in horses and reports macro-F1 of 0.898 for classification and 0.926 for binary detection on a public dataset. That is the core result.\n\nThe new part is the domain. Prior work on equine facial action units for welfare is thin, so running YOLOv12, optical flow thresholding, and a fine-tuned VideoMAE on horse blink footage counts as a legitimate first application rather than a routine repeat. Comparing the three approaches on the same data is useful and shows the frame-based and learned models both reach usable performance.\n\nThe soft spot is the link to affective state assessment. The abstract states that half and full blinks are recognised indicators of pain and stress, yet supplies no citations to equine welfare studies and no validation that the dataset categories match those indicators. The stress-test concern holds on the evidence given; the high scores are for blink detection, not for inferring pain or stress. Dataset size, split details, and error analysis are also missing from the abstract, which leaves the numbers hard to assess without the full methods.\n\nThis is for applied CV people or animal science groups who need a concrete starting point for horse video tools. It is narrow in scope but the use of public data makes the numbers checkable.\n\nThe work shows clear thinking on the technical side and honest engagement with an under-served application. It deserves a serious referee to check the methods and ask for the missing validation on the blink categories.","headline":"Solid F1 numbers on public horse video data for blink detection, but the affective state claim rests on an unbacked assertion about what the blinks mean.","tokens_in":2210,"tokens_out":384,"would_cite":false,"duration_ms":21215,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Three computer vision methods detect and classify horse eye blinks at macro-F1 scores of 0.898 and 0.926 from video.","keywords":["horse blink detection","equine welfare","computer vision","YOLOv12","VideoMAE","optical flow","facial action units","pain assessment"],"falsifier":"Expert re-annotation of the same videos for affective state under controlled conditions that would show whether the automated scores align with actual welfare outcomes.","tokens_in":2548,"feed_emoji":"🐴","tokens_out":557,"duration_ms":29399,"temperature":0.7,"pith_summary":"The paper tests three approaches to automatically spot half and full eye blinks in horse videos, which serve as subtle signs of pain and stress. These movements are too fine-grained for easy human observation and require frame-by-frame review. The methods reach strong performance numbers on a public dataset, indicating that video analysis could support ongoing equine welfare checks. Results also note remaining difficulties in handling the fine details of these facial actions.","feed_headline":"Horse blink detection reaches 0.926 F1 on video","feed_subtitle":"Three methods classify half and full blinks as indicators for equine pain and stress monitoring.","key_machinery":"Evaluation of YOLOv12, optical flow thresholding, and VideoMAE on subtle half and full eye blinks treated as facial action units for affective state monitoring.","core_discovery":"A frame-based YOLOv12 detector, an optical flow magnitude thresholding approach, and a fine-tuned VideoMAE model were applied to horse videos, producing a macro-F1 score of 0.898 for classifying blinks and 0.926 for binary blink detection on a publicly available dataset.","pith_inferences":["Combining blink detection with other facial action units could build a broader automated system for horse state assessment.","Real-time implementation on farm cameras would allow immediate alerts for potential pain or stress.","The same video techniques might adapt to detect similar subtle expressions in other large animals."],"forward_implications":["Automated detection reduces the need for manual frame-by-frame inspection in equine welfare monitoring.","Binary detection outperforms multi-class classification, indicating simpler tasks may be more immediately practical.","The results point to both feasibility and ongoing challenges for fine-grained facial action unit detection in horses.","Such methods could extend to continuous monitoring in veterinary or farm environments."],"fun_headline_variants":["YOLOv12 scores 0.926 F1 on horse blink detection","VideoMAE scores 0.898 macro-F1 for blink classification","Three methods score 0.926 F1 for equine blink detection","0.898 macro-F1 on automated horse blink classification"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The public dataset reflects real-world horse video variation and the blink categories match established pain and stress indicators without extra validation.","fun_headline_variants_meta":{"raw":{"variants":["YOLOv12 scores 0.926 F1 on horse blink detection","VideoMAE scores 0.898 macro-F1 for blink classification","Three methods score 0.926 F1 for equine blink detection","0.898 macro-F1 on automated horse blink classification"]},"model":"grok-4.3","cost_usd":0.006921,"raw_usage":{"total_tokens":3163,"prompt_tokens":574,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":69212000,"prompt_tokens_details":{"text_tokens":574,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2516,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":574,"tokens_out":73,"duration_ms":18150,"temperature":1.0,"reasoning_tokens":2516,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T06:27:17.904440+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Expert re-annotation of the same videos for affective state under controlled conditions that would show whether the automated scores align with actual welfare outcomes.","supporting_citations":[],"review_version":1}