{"id":"ff78f4e2-28dd-41d2-b584-2952313b89c8","arxiv_id":"2412.04945","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HOLa, a HoloLens 2 app using SAM-Track, automatically labels a single object in recorded video with quality close to human annotators and a 500x speedup.","lead":"Researchers built HOLa, a tool that combines HoloLens 2 recording with the SAM-Track segmentation algorithm to label objects in augmented reality video after a single user click. In tests on liver surgery scenes and phantoms, it labels about 500 times faster than manual annotation with Dice scores close to human agreement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline Dice/speed claims rest on a post-hoc, non-random 90-frame selection and a reference rater who recorded the data; a blinded all-frame/random-frame re-evaluation is required before the quality numbers can be taken at face value.","rationale":"The paper's engineering contribution is concrete: a Unity/Python HoloLens pipeline built on SAM-Track, with code released, and the phantom results are plausibly strong. The weakness is not internal inconsistency but an evaluation protocol that can inflate agreement. The reader's weakest assumption already identifies this; my reading agrees. Since the issue is empirical and fixable by re-evaluation, it supports conditional acceptance rather than rejection. I would not change the reader's verdict: UNCHANGED. The concern is the most load-bearing because the headline numbers are the main evidence for the central claim, and without an unbiased frame sample and an independent reference, the 'comparable to human annotators' conclusion is not secured.","tokens_in":4811,"tokens_out":4922,"duration_ms":53437,"concrete_test":"For each of the five experiments, re-run HOLa on a fixed random sample of frames stratified by time (or on the full recorded sequence), and have an independent radiologist or surgeon who did not record the videos produce blinded reference masks. Report per-frame Dice distributions, tracking-failure counts, and the mean Dice on the unbiased sample. If the unbiased mean falls below the published 90-frame mean by more than the inter-rater standard deviation (roughly 0.01-0.03), the headline quality claim should be revised to be conditional on frame selection. Publish the frame indices or the selection criterion to make the original choice reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('Dice scores between 0.875 and 0.982, comparable to human annotators') is supported only by Section 3's evaluation protocol. For each of five recordings, 90 frames were selected by the authors as 'best represent[ing] the variations during recording,' and the ground-truth reference (HA 1) is the person who recorded the data and placed the initial seed point. The selection is retrospective and subjective: frames where the tracker drifts, the object leaves the view, or contrast is briefly unfavorable can be excluded without this being visible in the reported aggregate. The sentence 'No frames were excluded in advance' does not cure this, because frames were excluded post hoc by the selection step. The Discussion's argument that the 10-frame subset differs by less than 0.013 Dice from the 90-frame subset only shows consistency between two non-random subsets; it does not establish representativeness. If the selection oversampled easy frames, the reported quality range is overoptimistic, especially for Experiment 5, where the paper itself notes low-contrast boundaries. The 500x speed figure is also end-to-end only for the automatic pass, not for the human quality-control pass that the Discussion says is always needed, but the frame-selection bias is the more direct threat to the quantitative headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents HOLa, an integrated Unity/Python application for the Microsoft HoloLens 2 that combines the Segment Anything Model (SAM) with the DeAOT tracker (SAM-Track) to annotate a single object of interest across recorded video frames. The user selects a seed point via an AR sphere cursor and initiates recording by voice command; the labeling mode then produces pixel-wise masks for every frame. The method is evaluated on three phantom experiments and two open-liver-surgery recordings. The authors report mean Dice scores between 0.875 and 0.982, close to human inter-rater concordance on a 10-frame subset, and an annotation speed of 5 fps versus roughly 0.008–0.010 fps for humans, yielding the advertised 'more than 500 times' speedup. The code is publicly available.","tokens_in":5088,"tokens_out":2645,"duration_ms":33051,"significance":"If the reported quality and speed hold, HOLa would be a practically useful tool for building annotated datasets for medical AR tracking research, with a meaningful reduction in manual labeling cost. The paper is, to the authors' knowledge, the first evaluation of SAM-based tracking on HoloLens RGB data, and the integration with a HoloLens recording application is a tangible contribution. There is no obvious circularity in the evaluation: HOLa is not fit to the test data, and the Dice scores compare its outputs with independent human annotations (HA 1). The main strength is the concrete, open-source system that others can reuse. The main weakness is the evaluation protocol: only five short recordings are used, the 90-frame subsets are chosen retrospectively and subjectively by the authors, and the reference annotator is the person who recorded the data. These issues directly affect the credibility of the headline quantitative claims, although they are fixable with additional experiments and analysis.","major_comments":[{"comment":"The evaluation protocol does not support the claim that the reported Dice scores are representative of HOLa's performance on typical HoloLens usage. The 90 frames per experiment are selected by the authors 'across the entire sequence that best represent the variations during recording,' and the statement 'No frames were excluded in advance' does not address the fact that frames were excluded post hoc by this subjective selection step. If the selection oversamples easy, high-contrast frames and undersamples tracker drift or boundary-ambiguity frames, the mean Dice scores, especially the 0.875 for Experiment 5, will be optimistically biased. The validity of the central quality claim requires an evaluation on all recorded frames or on a pre-registered random/blinded subset, reported with frame-level distributions and worst-case scores.","section":"Section 3 (Experiments)"},{"comment":"The defense of the evaluation in the Discussion is insufficient. The sentence 'the metrics differ by less than 0.013 Dice compared to the results based on 90 frames, suggesting that the selection is representative for the total set' only compares two non-random subsets (90 vs. 10 frames) that were both hand-picked by the authors; consistency between two subjective selections does not establish representativeness of the full recording. Furthermore, HA 1 is 'the same person who recorded the data,' so the reference annotations may benefit from a familiarity with the scenes that an independent oracle would not have. The authors should provide either a blinded re-annotation study or an all-frame/random-frame evaluation to rule out this bias.","section":"Section 5 (Discussion)"},{"comment":"The headline 'more than 500 times' speedup is computed for the automatic labeling pass only, but the Discussion states that 'this will never replace a human cross-check' and that future work will analyze quality control more closely. As stated, the end-to-end labeling workflow includes a human QC step whose cost is not included in the reported 5 fps. The speed claim should be qualified to the automatic pass, or the authors should estimate the full time including QC to support the practical workload-reduction claim.","section":"Table 1 and Section 5"}],"minor_comments":[{"comment":"The phrase 'fully automatic single object annotation ... while requiring minimal human participation' is almost contradictory; consider rephrasing to 'automatic labeling after a single user-provided seed point' to match the actual workflow.","section":"Abstract"},{"comment":"The formatting of the fps values ('0 .008 𝑓𝑝𝑠') contains stray spaces and uses italic text; the inconsistent spacing should be corrected.","section":"Tables 1 and 2"},{"comment":"The sentence 'We transform all recorded PV frames to a video prior to frame-wise labeling' is slightly unclear because the video is then processed frame-wise; rephrase to clarify that the frames are assembled into a video for input to SAM-Track.","section":"Section 2 (Methods)"},{"comment":"The caption describes 'distortions in labeling' while the text describes 'incomplete segmentation' of a multi-segment liver; align the terminology and explain in the text what type of distortion occurs (e.g., background leakage vs. missing segments).","section":"Figure 4 caption and Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a brief systems evaluation with a small dataset. The core idea is sound and the open-source contribution is useful, but the current evaluation protocol leaves the quantitative claims insufficiently supported. If the authors can report results on all frames (or a blinded random sample) and clarify the speed claim with QC cost, the paper would be acceptable. I see no scope or novelty issues that would justify rejection. The manuscript appears to be an early-stage report; a revised version with the requested evaluation would strengthen it considerably."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nShort version: this is a legitimately useful tool paper, and the evaluation is mostly sound, but the headline Dice numbers sit on a non-random frame-selection step that needs rechecking. I'd send it out for review, with a request for a more robust validation.\n\nWhat's actually new: HOLa is an open-source HoloLens 2 recording and labeling app that wires HL2SS streaming to SAM-Track, modified for single-object, seed-point-prompted tracking. The integration is straightforward—no new architecture—but the package is real. The paper reports five experiments (three phantoms, two in-situ liver surgeries) with Dice between 0.875 and 0.982, and a ~500x speedup over manual labeling. The public code and the clinical data with ethics approval are concrete assets. The authors also state the limitations honestly: single-object only, SAM fails on low contrast, seed placement matters, and human QC is always required. That last point undercuts the 'fully automatic' label in the abstract, but they don't hide it.\n\nWhere it gets soft: the 90 frame evaluation. You pick frames 'that best represent the variations during recording'—that's post-hoc selection, not sampling. The authors say no frames were excluded 'in advance,' but that doesn't address frames excluded by the selection itself. The reference rater is the person who recorded the data and placed the seed point, which is a familiarity bias even if two medical experts revised the labels. The 10-frame subset used for inter-rater comparison is also non-random, and the 0.013 Dice difference between 10 and 90 frames only shows two non-random subsets agree with each other. So the reported quality range is probably optimistic for arbitrary frames, especially in low-contrast Experiment 5. This is fixable: re-annotate random or all frames without the 'best represent' filter, or at least report per-frame Dice distributions and failure cases. The speedup is also only for the automatic pass; the human QC pass burns time, though that's a minor framing issue.\n\nThe center of the paper—that a SAM-Track pipeline seeded by a HoloLens cursor can produce useful masks—holds up. The weak spot is the external validity of the exact numbers, not the tool itself.\n\nWho it's for: anyone building HoloLens medical-AR datasets who wants to skip manual pixel annotation. It's a solid engineering contribution, and the evaluation is good enough to be a serious referee assignment, not a desk reject. I'd ask for the random-frame re-evaluation before acceptance, but I wouldn't block on anything else.","headline":"HOLa is a useful, honest tool paper for HoloLens annotation, but the headline Dice numbers need a random-frame re-evaluation before I'd trust them.","tokens_in":5574,"tokens_out":3015,"would_cite":false,"duration_ms":32888,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HOLa: a single seed point plus a tracking model labels HoloLens surgery video at human-level quality, more than 500 times faster than manual annotation.","keywords":["HoloLens 2","object labeling","data annotation","Segment Anything Model","SAM-Track","augmented reality surgery","video object segmentation","medical AR"],"falsifier":"Re-run HOLa on the complete, unedited recordings from the five experiments, not on the author-selected 90 frames, and have annotators who were not involved in recording label all frames; if the mean Dice on the full sequences falls below the human inter-rater concordance by more than the small margins reported here, the claim of human-level automatic labeling on HoloLens recordings is refuted.","tokens_in":4656,"feed_emoji":"🩺","tokens_out":5162,"duration_ms":51693,"temperature":0.7,"pith_summary":"The paper introduces HOLa, a HoloLens 2 recording and labeling pipeline that turns a single center-of-object seed point into pixel-wise masks for every RGB frame in a recording. Its central claim is that this fully automatic annotation reaches human-level quality: in five liver-surgery and phantom experiments the mean Dice score against a human reference lies between 0.875 and 0.982, close to the agreement among human annotators, while annotation throughput rises from about 0.008–0.010 frames per second to 5 frames per second. The authors argue this removes the main bottleneck in collecting training data for AR-guided surgery tracking models. Because the segmentation backbone is a foundation model, the pipeline is not tuned to a specific image appearance and is offered as a general solution for AR object labeling.","feed_headline":"HOLa labels HoloLens surgery footage 500 times faster","feed_subtitle":"Automatic SAM-based labeling matches human annotators, with Dice scores from 0.875 to 0.982.","key_machinery":"The load-bearing mechanism has three parts: the seed-point-prompted Segment Anything Model (SAM), which turns one click at the frame center into an initial object mask by picking the highest-IoU proposal among three candidates; the DeAOT tracker from SAM-Track, which propagates that mask across the video; and the HL2SS sensor-streaming plugin, which brings the HoloLens RGB camera, depth, point cloud, and poses into the labeling pipeline. The authors replace SAM-Track's 'Segment Everything' mode with the single-seed SAM prompt so that exactly one object is followed. This combination is what lets a single initialization label an entire recording.","core_discovery":"HOLa's core discovery is that a promptable segmentation foundation model combined with a video object tracker can replace frame-by-frame manual labeling on HoloLens 2 recordings without loss of quality. The user points a sphere cursor at the object and speaks a command; SAM is prompted at the first frame with that seed point, the highest-IoU mask among its proposals initializes the DeAOT tracker, and the tracker propagates the mask through all subsequent frames. On the reported five experiments the mean Dice scores of HOLa versus the human reference are 0.982, 0.966, 0.981, 0.925, and 0.875, and on a 10-frame subset the HOLa-versus-human concordance tracks the human-versus-human concordance closely (for example 0.887 versus 0.917 on the most difficult surgery scene). The same experiments show a speedup of more than 500 times relative to manual annotation.","pith_inferences":["The reported 500x speedup compares automated post-processing on a high-end GPU with manual labeling at a desk; an end-to-end count that includes recording time, seed placement, and quality control would be smaller, though likely still large.","Because the reference annotator was the person who recorded and selected the frames, the human-level scores may partly reflect familiarity with the scenes; an evaluation with independent annotators selecting frames at random would give a stricter estimate.","A natural extension is to feed the synchronized depth stream and point cloud into the tracker; the paper's own examples show boundary errors from shadows and low color contrast, and depth cues are precisely the kind of signal that could correct those.","If the same seed-point-and-track recipe is applied to newer promptable segmentation models, the ranking of results across the five scenes would probably track the model's boundary quality on low-contrast imagery rather than anything specific to HoloLens."],"forward_implications":["Researchers collecting HoloLens training data for organ tracking can generate pixel-wise labels for entire recordings by marking one seed point, cutting annotation labor by orders of magnitude.","The labeling quality on clearly separated organs (Dice above 0.96 in phantoms, 0.925 in the first surgery scene) means the output can serve as training masks with only light quality control.","In low-contrast scenes where an organ blends into surrounding tissue, automatic labels degrade to about 0.875 Dice, so those recordings need human review or additional seed points.","Because the method tracks one object only, complex multi-segment structures require placing extra seed points during quality control, and frames where the object leaves the view are not labeled.","The approach transfers without appearance-specific tuning, so the same pipeline applies to non-medical HoloLens AR labeling tasks."],"supporting_citations":[{"why":"Supplies the promptable segmentation foundation model whose first-frame mask initializes the tracking in HOLa.","marker":"[2]"},{"why":"Provides the SAM-Track pipeline that HOLa adapts for single-object label propagation.","marker":"[3]"},{"why":"Provides the DeAOT video object segmenter used by SAM-Track to propagate the mask across frames.","marker":"[4]"},{"why":"Provides the HL2SS plugin that streams HoloLens RGB, depth, point cloud, and poses to the PC for recording.","marker":"[5]"}],"fun_headline_variants":["HOLa labels surgery footage 500x faster than humans","HOLa: 500x faster labeling with human-level Dice scores","SAM-powered HOLa labels HoloLens video 500x faster","Auto-labeling for HoloLens: 500x speedup, human-grade masks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline quality numbers rest on the assumption that the 90 frames selected per experiment as best representing the recording are a representative, unbiased sample; because the same person who made the recordings chose those frames and served as the reference annotator, agreement on arbitrary unselected HoloLens frames could be lower.","fun_headline_variants_meta":{"raw":{"variants":["HOLa labels surgery footage 500x faster than humans","HOLa: 500x faster labeling with human-level Dice scores","SAM-powered HOLa labels HoloLens video 500x faster","Auto-labeling for HoloLens: 500x speedup, human-grade masks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000698,"raw_usage":{"total_tokens":3141,"prompt_tokens":918,"completion_tokens":2223,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":2144}},"tokens_in":534,"tokens_out":2223,"duration_ms":15885,"temperature":1.0,"reasoning_tokens":2144,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:05:58.425258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run HOLa on the complete, unedited recordings from the five experiments, not on the author-selected 90 frames, and have annotators who were not involved in recording label all frames; if the mean Dice on the full sequences falls below the human inter-rater concordance by more than the small margins reported here, the claim of human-level automatic labeling on HoloLens recordings is refuted.","supporting_citations":[{"cited_title":"Segment anything","cited_arxiv_id":null,"evidence_quote":"Supplies the promptable segmentation foundation model whose first-frame mask initializes the tracking in HOLa."},{"cited_title":"Decoupling features in hierarchical propa- gation for video object segmentation","cited_arxiv_id":null,"evidence_quote":"Provides the DeAOT video object segmenter used by SAM-Track to propagate the mask across frames."}],"review_version":1}