{"id":"002ca8d1-d247-4e1b-a0da-7ae9b69d431c","arxiv_id":"1908.08914","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A level-set region tracker based on RGB Kullback-Leibler divergence with area and length penalties locates the iris and pupil in test image sequences, reporting 82% desired-region coverage and 20% undesired-region coverage.","lead":"This engineering thesis builds a variational image-tracking algorithm that locates a driver's iris and pupil by minimizing a color-based cost function with level sets. It reports 82% desired-region coverage on a small set of test images, but the method is not real-time, needs per-eye tuning, and is not compared with existing eye trackers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 82% coverage claim rests on per-eye multiplier customization, so the reported accuracy may be a tuning artifact rather than a property of the fixed final functional.","rationale":"The reader's weakest assumption identifies the fragility of static RGB intensity statistics across eyes and illuminations, citing Section 5.5.1. My stress-test converges on the related but more operational issue: the reported accuracy was obtained with per-eye multiplier customization, so there is no evidence that the fixed Design #4B functional generalizes across image sequences. This is the single most load-bearing concern because the central claim is empirical: the algorithm is said to provide approximate eye-location data across tested sequences. The paper is a clearly written undergraduate capstone with acknowledged limitations, and its mathematical derivation is standard variational level-set methodology. The weakness is not internal inconsistency; it is that the evaluation protocol cannot support the headline accuracy claim. A fixed-parameter held-out test would settle whether the result is a property of the algorithm or of the tuning. Since the reader already recommended REJECT and my concern reinforces rather than redirects that verdict, no verdict change is needed.","tokens_in":13085,"tokens_out":2079,"duration_ms":23331,"concrete_test":"Select one image sequence as a training set, fix lambda_1..lambda_4 to values chosen only on that sequence, and then run Design #4B on held-out sequences from at least five different subjects with different eye colors, without any re-tuning. Report Desired Region Coverage and Undesired Region Coverage per sequence and as a fixed-parameter average, and state whether the level-set curve was initialized per frame or propagated from the previous frame. If the fixed-parameter average falls materially below 82% or the undesired coverage rises substantially, the headline accuracy is a tuning artifact rather than a property of the final functional.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that Design #4B, minimized by gradient descent on a level-set representation, yields 82% desired-region coverage with 20% undesired-region coverage (Section 8). For this claim to support the stated objective, the accuracy must be attributable to the fixed algorithm, not to per-sequence or per-subject adjustment of the functional. Section 5.5.1 states that 'this final design also requires customization of all multipliers in order to function of different eyes,' and Section 6 repeats that 'the algorithm currently requires tooling to each individual eye that it operates on' and calls this 'not acceptable for an implementable solution.' The reported averages appear after such customization, but no information is given about how many sequences were used, whether the same parameters were held fixed across all frames, how ground-truth regions were defined, or whether the metrics are per-frame or per-sequence. Without a fixed-parameter, held-out evaluation, the 82%/20% numbers cannot be distinguished from the result of fitting four multipliers to a small set of images. The functional design is coherent and the limitations are honestly stated, but the load-bearing claim about accuracy is not supported by controlled evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a variational level-set method for tracking the iris and pupil in an image sequence, motivated by a driver-inattention warning system. The authors propose a series of functionals, culminating in Design #4B, which combines per-channel RGB Kullback-Leibler divergences with an area-preservation term and a length regularization term, minimized by gradient descent on a level-set representation. They report average accuracy metrics of 82% desired-region coverage and 20% undesired-region coverage on a small set of synthetic and real eye images, and they discuss next steps and broader societal impacts. The paper is an MTHE 493 undergraduate thesis posted to arXiv, not a peer-reviewed archival contribution.","tokens_in":13331,"tokens_out":4613,"duration_ms":49127,"significance":"The paper provides a clear pedagogical exposition of the variational level-set tracking pipeline and a sensible design progression from intensity-based to RGB-distribution-based functionals. The observation that color information helps disambiguate shadowed regions from the eye is a useful qualitative insight. However, the central empirical claim--the 82%/20% accuracy figures--is not supported by controlled evidence. The authors explicitly state in Sections 5.5.1 and 6 that the final design requires per-eye customization of all multipliers and that this is 'not acceptable for an implementable solution.' The evaluation set is unspecified, no fixed-parameter protocol is described, and no error bars or statistical details are given. As a controlled validation, the paper falls short of the standard required for a journal publication; its value is primarily as a proof-of-concept description of a functional design.","major_comments":[{"comment":"The reported averages of 82% desired-region coverage and 20% undesired-region coverage are obtained after customizing all multipliers for each eye, as stated in Section 5.5.1 ('This final design also requires customization of all multipliers in order to function of different eyes') and repeated in Section 6 ('The algorithm currently requires tooling to each individual eye that it operates on'). Because the parameters are tuned on the same small set of images that is used to compute the averages, the headline result may be an artifact of parameter fitting rather than a property of the fixed functional. No evidence is provided that a single parameter vector was held fixed across frames or subjects, or that the reported numbers are robust to parameter choices. The central accuracy claim is therefore not supported by controlled experimentation.","section":"Section 5.5.1 and Section 6"},{"comment":"The evaluation protocol is underspecified to the point of non-reproducibility. The two metrics are described only informally around Figure 9, with no exact formulas, no statement of how ground-truth desired regions were obtained, no count of image sequences or frames, no per-frame versus per-sequence aggregation rule, and no error bars or variance estimates. To justify the claim that the final functional 'on average' yields 82%/20%, the authors must report the data set, the annotation procedure, the exact parameter vector(s) used, and the statistical summary. Without this information, the reader cannot assess the reliability of the empirical result.","section":"Section 4.2 and Section 5.5.1"},{"comment":"The conclusion that the algorithm 'gives an approximation of the location of a subject's iris and pupil' is in tension with the manuscript's own statement that the algorithm is not acceptable for an implementable solution because it must be tooled to each individual eye. Since the stated application is a driver-inattention system intended to work for all drivers, the reported accuracy is not sufficient to support the claimed practical relevance. The paper should either demonstrate that the customization requirement can be removed or reframe the contribution as an initial design study with clearly limited generalizability.","section":"Section 6 and Section 8"}],"minor_comments":[{"comment":"There are typographical errors in the text, e.g., 'efficent' should be 'efficient' and 'computiation' should be 'computation.'","section":"Section 2.4"},{"comment":"In the paragraph describing Figure 23, the phrase 'covered covered' should read 'covered.'","section":"Section 5.5.1"},{"comment":"The term 'Kullback-Liebler' should be spelled 'Kullback-Leibler.'","section":"Section 5.3"},{"comment":"The sentence 'With 10% accuracy, the system could save up to 183 lives per year' introduces a '10% accuracy' figure that is not derived from the 82%/20% reported metrics; the connection should be explained or the statement should be removed.","section":"Section 7.4"},{"comment":"The finite-difference notation is inconsistent, with δx used in some definitions of D+x and D−x and ∆x used in others; this should be harmonized for clarity.","section":"Section 2.3"}],"recommendation":"reject","confidential_remarks":"This manuscript is an undergraduate thesis, and it reads like one: the exposition is clear but the empirical evaluation is not at the level of a peer-reviewed archival paper. The central quantitative claim is admitted by the authors themselves to depend on per-eye tuning of all multipliers, and the evaluation set is not described. The novelty is also limited relative to the existing level-set region-tracking literature (e.g., Mansouri 2002). I recommend rejection. The authors should be encouraged to either substantially expand the evaluation with a fixed-parameter protocol on a public dataset or reposition the work as a technical report rather than a research article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a clearly written capstone that builds a level-set region tracker for eyes; the math is standard and the limitations are honestly stated, but the 82% figure is not a reliable empirical claim because the functional's multipliers are tuned per eye and the test set is undocumented.\n\nWhat's new: a specific functional combining per-channel RGB Kullback-Leibler divergences with an area-preservation term and length regularization, inside the level-set tracking framework of Mansouri [11]. That's a legitimate extension, not a new framework. The paper does a good job explaining the variational derivation, the level-set discretization, and the design iteration, and it is honest that Designs #1-#3 fail on shadows and that Design #4B still needs per-eye tuning.\n\nSoft spots: the main result is Section 8's average of 82% desired-region coverage with 20% undesired coverage. That number is reported without the number of images, the definition of ground truth, per-frame vs per-sequence averaging, or any error bar. More importantly, Section 5.5.1 says the final design \"requires customization of all multipliers in order to function of different eyes\" and Section 6 repeats that it \"requires tooling to each individual eye\" and calls it \"not acceptable for an implementable solution.\" So the reported average is produced after hand-fitting four multipliers to a small, unspecified set. Without a fixed-parameter, held-out evaluation, the claim cannot be separated from a tuning artifact. The abstract overreaches by saying the algorithm has \"sufficient computational efficiency and accuracy\" when the paper itself says real-time and robustness remain future work. Section 7's broad impact analysis is tangential and doesn't add technical support. There is no code or data released, and no comparison to any baseline.\n\nThe math and derivation appear sound: the gradient descent, level-set representation, and discretization follow standard references. The citation pattern is appropriate—it builds directly on the supervisor's region-tracking work and cites Mumford-Shah, Sethian, and eye-tracking surveys. So the softness is in the evidence, not the theory.\n\nWho it's for: a reader who wants to see how a level-set region tracker can be assembled for eye images in a course project, or an instructor looking for an example of an honest limitations section. As a research preprint it doesn't support its headline quantitative claim. I'd desk reject it. If the authors came back with a fixed-parameter evaluation on a public dataset, with code and error bars, it could be a modest workshop contribution.","headline":"Coherent undergraduate thesis with a standard level-set functional, but the headline 82% accuracy is a tuning artifact on an undocumented small set.","tokens_in":13849,"tokens_out":2824,"would_cite":false,"duration_ms":25907,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The eye region is tracked by minimizing a designed functional, with the final design reaching 82% desired-region coverage","keywords":["eye tracking","region tracking","level set methods","variational methods","gradient descent","driver inattention","relative entropy","image sequence analysis"],"falsifier":"Run the final Design #4B functional on a frame in which a shadowed patch of skin or clothing has the same per-channel RGB distribution and approximate area as the iris-pupil region; if the minimized curve locks onto that patch instead of the eye, the central claim that the functional distinguishes the eye from similar-looking regions fails.","tokens_in":12895,"feed_emoji":"👁️","tokens_out":6178,"duration_ms":59009,"temperature":0.7,"pith_summary":"The paper designs an eye-tracking algorithm for a driver-alert system and argues it can localize the iris and pupil by minimizing a specially constructed functional over the image domain. The central claim is that the minimizer of this functional, evolved by gradient descent on a level-set representation, gives an approximation of the eye region across a sequence of frames. On the tested synthetic, grayscale, and color images, the final design reports 82% desired-region coverage, with 20% of the tracked region falling outside the iris and pupil. The authors present this as enough to provide approximate eye-location data to a larger system that could warn a distracted driver, while noting that real-time speed, robustness across eye colors, and per-driver parameter tuning remain open.","feed_headline":"Minimizing one functional tracks the eye across image sequences","feed_subtitle":"A level-set gradient-descent method locates the iris and pupil with 82% desired-region coverage for a driver-alert system.","key_machinery":"The load-bearing object is the energy functional $E_{4B} = E_4 + \\lambda_4(\\text{area}(R_1) - \\text{area}(R_0))^2$, where $E_4$ is the sum of three per-channel distribution-divergence terms (red, green, blue) plus a curve-length penalty. The curve is represented implicitly as the zero level set of a function $u$ on the image domain, and the minimization is carried out by gradient descent following the Euler–Lagrange equations, with the level-set update $\\partial u/\\partial t = F\\|\\nabla u\\|$. This representation gives numerical stability and lets the curve change topology freely, and the divergence terms let the functional compare regional appearance without tracking feature points or requiring a static background. The functional's role is to make the eye region a minimum that survives across frames despite shadows and illumination changes.","core_discovery":"The paper's discovery is a variational formulation of eye tracking: given a region R0 chosen in the first image, find the curve in a later image whose interior best matches R0 according to a functional that rewards similar appearance and penalizes boundary length. The final functional, Design #4B, compares the per-channel RGB intensity distributions of the tracked region with those of the original region using divergence terms, adds a squared difference in area, and includes a length-regularization term. Minimizing this functional by gradient descent yields a level-set curve whose zero level set approximates the iris and pupil boundary. In the paper's reported experiments this functional outperformed earlier intensity-only, relative-entropy, and gradient-based versions, giving an average desired-region coverage of 82% and undesired coverage of 20%.","pith_inferences":["A natural extension is to replace per-channel divergences with the joint RGB distribution; the paper chose the faster per-channel sum for speed, so the accuracy cost of that trade-off is untested.","Because the multipliers must be customized for each eye, practical deployment would likely need a calibration step at ignition, such as using detected eye color to select parameters.","The same variational machinery could be tested on other colored anatomical regions or on driver gaze direction rather than just iris and pupil location.","The 20% undesired coverage suggests a downstream distraction detector should be trained on the tracker's actual output rather than on perfectly labeled eye regions, since the errors are systematic."],"forward_implications":["With this functional, eye position can be estimated without computing motion or requiring a fixed background, which suits a moving driver's head.","The RGB-distribution terms are the main source of accuracy; intensity-only versions repeatedly failed on shadowed regions.","The area term prevents the optimizer from shrinking to a subset of the eye region with the same average appearance.","The reported 82%/20% figures mean the output is a coarse eye-location estimate, not precise gaze, so the larger alert system must tolerate some false region.","Computation time remains too high for real-time deployment; parallelizing per-pixel updates is suggested as a path forward."],"supporting_citations":[{"why":"Introduces the variational image-segmentation approach that motivates designing functionals whose minimizer is the desired region.","marker":"[7]"},{"why":"Supplies the calculus-of-variations and gradient-descent machinery used to derive the curve evolution.","marker":"[8]"},{"why":"Provides the level-set discretization and numerical schemes used in the implementation.","marker":"[9]"},{"why":"Gives the region-tracking-via-level-set-PDE method without motion computation that this design extends.","marker":"[11]"}],"fun_headline_variants":["Level-set gradient descent tracks driver's eye in video","Variational eye tracker hits 82% region coverage","Minimize functional, track eye in image sequences","Gradient descent on level-set functional locates iris and pupil","Eye tracker for drivers: 82% coverage via energy minimization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the iris and pupil are the only region in the image whose RGB distribution and area match the initial eye region closely enough, for every driver and lighting condition; the paper's own note that all multipliers must be customized per eye shows this invariance does not hold automatically.","fun_headline_variants_meta":{"raw":{"variants":["Level-set gradient descent tracks driver's eye in video","Variational eye tracker hits 82% region coverage","Minimize functional, track eye in image sequences","Gradient descent on level-set functional locates iris and pupil","Eye tracker for drivers: 82% coverage via energy minimization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000473,"raw_usage":{"total_tokens":2289,"prompt_tokens":822,"completion_tokens":1467,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":438,"completion_tokens_details":{"reasoning_tokens":1388}},"tokens_in":438,"tokens_out":1467,"duration_ms":12469,"temperature":1.0,"reasoning_tokens":1388,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:26:10.008420+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the final Design #4B functional on a frame in which a shadowed patch of skin or clothing has the same per-channel RGB distribution and approximate area as the iris-pupil region; if the minimized curve locks onto that patch instead of the eye, the central claim that the functional distinguishes the eye from similar-looking regions fails.","supporting_citations":[{"cited_title":"Optimal approximations by piecewise smooth functions and associated variational problems,","cited_arxiv_id":null,"evidence_quote":"Introduces the variational image-segmentation approach that motivates designing functionals whose minimizer is the desired region."},{"cited_title":"Gelfand and S","cited_arxiv_id":null,"evidence_quote":"Supplies the calculus-of-variations and gradient-descent machinery used to derive the curve evolution."},{"cited_title":"Sethian, Level Set Methods and Fast Marching Methods","cited_arxiv_id":null,"evidence_quote":"Provides the level-set discretization and numerical schemes used in the implementation."},{"cited_title":"Region tracking via level set pdes without motion computation,","cited_arxiv_id":null,"evidence_quote":"Gives the region-tracking-via-level-set-PDE method without motion computation that this design extends."}],"review_version":1}