{"id":"2e75fb8d-ab26-48fe-8265-db88def69615","arxiv_id":"2505.16228","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Shape-aware focus selection chooses per-camera focus distances to maximize the in-focus body surface, reporting about 85 to 95 percent in-focus coverage and resolution near 0.06 mm per pixel for total body photography.","lead":"This paper describes a total body photography system that automatically chooses each camera's focus distance based on the estimated 3D shape of a person, so more of the body surface is captured sharply. It reports about 85 percent of the surface in focus in simulation and 95 percent on a mannequin, with resolution around 0.06 mm per pixel.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline in-focus percentages are computed from the same DoF frustum model the optimizer maximizes, leaving measured sharpness unvalidated.","rationale":"The reader's weakest-assumption analysis correctly identifies the load-bearing issue: the in-focus surface metric is the same DoF frustum model used as the optimization objective, with no independent sharpness validation. My stress-test pass found no additional concern more severe than this. The algorithmic contribution—EM-based focus-distance selection over estimated body geometry—is internally coherent, and the robustness analyses in Section V-D are reasonable. The paper's relative comparisons could be meaningful even under model error, provided the DoF model is monotonically related to true sharpness, but the absolute coverage percentages and the claim of outperforming auto-focus need external sharpness evidence. A focus-sweep experiment with a measured sharpness metric would directly test whether the frustum model tracks real image quality. Since the reader already assigned CONDITIONAL for essentially this reason, my assessment leaves the verdict unchanged.","tokens_in":18150,"tokens_out":2309,"duration_ms":19941,"concrete_test":"On the real mannequin, run a focus sweep at each camera pose (sampling focus distances around the predicted optimum and around baseline distances), and compute a measured sharpness map from the captured images—e.g., variance of Laplacian or MTF from Siemens-star targets placed at known surface points. Assign each surface point to its best camera as in Eq. (8), and label it sharp when the measured metric exceeds a threshold calibrated against human-judged acceptable sharpness. Then compare the measured in-focus area for shape-aware versus closest, average, and auto-focus. If shape-aware's measured margin over average or auto-focus shrinks substantially relative to the reported 15.1–16.4 percentage points, or if sharpness classification agrees poorly with membership in V_s^c, the headline coverage numbers are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim—approximately 85% and 95% of surface area in-focus, and shape-aware focus outperforming baselines—is evaluated with the same depth-of-field model that the optimizer directly maximizes. In Eq. (1), the proximity-to-focal-plane cost term is 1(p∉V_s^c), where V_s^c is the frustum bounded by thin-lens near/far DoF limits, with the hyperfocal distance manually doubled in Section IV. The minimization step in Eq. (10) reduces exactly to maximizing the number of assigned points inside V_s^c, and the Section V-C1 metric of in-focus surface area is P(p∈V_s^c). Thus the reported coverage measures how well the algorithm places points inside its own geometric frusta, not whether those frusta correspond to real image sharpness. The qualitative crops in Fig. 7 provide some support, but no independent sharpness metric (MTF, contrast, blur estimation) is reported over the surface for any method, and the auto-focus comparison is qualitative only. If the thin-lens model, particularly the doubled hyperfocal factor, overstates true depth of field, the absolute 85%/95% figures overstate achieved sharpness; if it is too conservative, the method may underuse available sharpness. The relative ranking between methods could survive, but the load-bearing absolute coverage claim is not externally validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a shape-aware Total Body Photography (TBP) system that combines depth and RGB cameras on a rotating beam with 3D body-shape estimation and an EM-based optimization of per-camera focus distances. The cost function in Eq. (1) penalizes poor projected area, large optical-axis deviation, and points lying outside the depth-of-field frustum V_s^c; the minimization step in Eq. (10) reduces to maximizing, for each camera, the number of assigned sample points inside that frustum. The system is calibrated with incremental SfM and a ChArUco cuboid, and 3D shape is obtained by SMPL-NICP or TSDF reconstruction. Evaluation on 400 3DBodyTex meshes and one real mannequin scan reports roughly 85% and 95% in-focus surface area for simulation and real scan respectively, with average resolutions of 0.068 mm/pixel and 0.0566 mm/pixel. The paper also reports robustness to calibration noise, shape-estimation error, and postural sway, and compares against closest-focus, average-focus, and, qualitatively, auto-focus protocols.","tokens_in":18444,"tokens_out":4743,"duration_ms":39675,"significance":"If the central claims hold, the paper makes a useful contribution: it gives a clean formulation of focus-distance selection as a surface-coverage optimization problem, an efficient EM solver with a simple assignment step and a piecewise-constant minimization step, and a complete system pipeline that is evaluated on a large simulated population and in a real prototype. The robustness experiments in Section V-D are thoughtful and directly address the practical concerns of calibration error, shape-estimation error, and patient sway. The authors also provide a concrete resolution target (0.075 mm/pixel) motivated by lesion-detection requirements, which helps put the reported system resolution in context. However, the headline in-focus coverage percentages are computed with the same geometric depth-of-field model that the optimizer maximizes, and the only full-surface evidence against real image sharpness is qualitative. The real-scan evidence is a single mannequin acquisition, and the auto-focus comparison is not quantitative. These gaps limit the strength of the absolute coverage and superiority claims as currently stated.","major_comments":[{"comment":"The headline in-focus surface percentages in Table I are computed with the same geometric DoF-frustum model that the shape-aware focus optimizer maximizes. Specifically, Eq. (10) selects S_c to maximize the sum of 1(p in V_s^c) over assigned points, and the Section V-C1 metric is the proportion of sampled points satisfying p in V_s^c for some camera. The only image-based support is the qualitative crops in Fig. 7; no independent sharpness metric (MTF, contrast, blur estimation) is reported over the surface for any method. Because the near and far DoF limits come from a thin-lens model with a manually doubled hyperfocal distance (Section IV), the absolute 84.9% and 95.1% coverage figures are not externally validated: if the model overstates true depth of field, the numbers overstate achieved sharpness, and if it is too conservative, the method may underuse available sharpness. Please add an independent validation of sharpness on real images (e.g., MTF or estimated blur over the assigned surface, or a calibration of the DoF model against the actual lens), or present the coverage values explicitly as predictions of the geometric model rather than as achieved sharpness.","section":"V-C1, Eq. (1), Eq. (10)"},{"comment":"The claim that shape-aware focus outperforms auto-focus is not supported by quantitative evaluation. Table I compares only closest focus and average focus; auto-focus appears only as qualitative image crops in Fig. 7, where the precise AF focus point is unknown and the comparison is not repeated or scored. Since 'outperforms existing focus protocols (e.g. auto-focus)' is stated in the Abstract and Section V-C4, this should either be supported by a quantitative comparison on real captures (e.g., the same in-focus percentage or a full-surface sharpness metric for AF), or the claim should be softened to what the current evidence supports.","section":"Abstract; Section V-C4; Fig. 7"},{"comment":"The real-scan evidence consists of a single mannequin acquisition with no repeated trials and no variation in pose or body shape, so the real-row entries in Table I carry no uncertainty and cannot support a general claim about real-scan performance. The mannequin also has painted texture dots and lines that may make auto-focus easier, which is acknowledged in Fig. 7 but further complicates the qualitative AF comparison. At minimum, repeated scans of the same mannequin are needed; ideally the real validation should include multiple subjects or mannequins in different poses before the 95.1% coverage figure and the reported superiority over baseline protocols are presented as established results.","section":"V-C3"}],"minor_comments":[{"comment":"The percentage improvements stated in the text do not match Table I. For example, in the simulation row, K(S) changes from 4710 to 2986, which is a 36.6% reduction, not the stated 17%, and the in-focus area change from 31.6% to 84.9% is a 53.3 percentage-point increase, not a 53.3% increase. The same inconsistency appears in the real row. Please report changes consistently as relative changes or percentage-point changes.","section":"V-C2, V-C3"},{"comment":"The standard deviations in Section V-D3 are reported as percentages ('sigma = 246%' and 'sigma = 257%') although the quantities are total costs; these should be expressed in the same units as the costs or as a coefficient of variation. Also, the caption of Fig. 8 contains the typo '3DBdoyTex' for '3DBodyTex'.","section":"V-D3, Fig. 8 caption"},{"comment":"The algorithm is called 'Expectation-Minimization,' but the standard name for this alternating assignment/minimization procedure is Expectation-Maximization. In addition, the stopping-rule tolerance epsilon in Algorithm 1 is never assigned a numerical value; please state the value used in the experiments.","section":"Algorithm 1"},{"comment":"The indicator function 1(·) is used before it is defined; consider adding a brief definition in the list of notation, or using bold/blackboard notation to distinguish it from the scalar 1.","section":"Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the core idea is genuinely new: it frames multi-view TBP focus selection as a per-camera focus distance optimization over an estimated 3D body surface, and solves it with an EM assignment that reduces to maximizing in-frustum point count, which is a piecewise-constant problem. That is clean, implementable, and not in the cited TBP or extended-DoF literature. Second, the headline numbers—85% and 95% in-focus surface area, and the claimed superiority over autofocus—are only partially supported. The in-focus percentage is computed as membership in the same depth-of-field frustum that the optimizer maximizes. So the absolute coverage figures are model-based, not measured sharpness.\n\nWhat the paper does well: the system integration is substantial. Calibration with SfM and ChArUco cuboid, TSDF-based shape estimation, a rotary beam with seven 48-MP cameras and depth sensors, and a real mannequin scan. The robustness experiments are thoughtful: noise in camera calibration, shape estimation error, and simulated postural sway all get quantitative treatment. The qualitative image crops in Fig. 7 show visibly sharper anatomy for the shape-aware method than for closest or average focus, which gives real support for the relative claim. The limitations section is honest about focus breathing, pose stability, contrast/skin tone, and occlusion.\n\nSoft spots, in proportion. The metric circularity is real and load-bearing for the absolute claim. The doubled hyperfocal distance is a manual design decision; it could under- or over-state true depth of field. The real scan is a single mannequin with no repeated trials, and the autofocus comparison is qualitative only—no focus distance, no sharpness metric, no quantitative comparison. These are fixable. The relative ranking among closest/average/shape-aware is likely robust because it is corroborated by the crops, but the absolute 85%/95% should be read as \"in-focus according to the model,\" not independently verified sharpness.\n\nAlso, the paper has minor typos (e.g., 3DBdoyTex, \"Qualititative\") and some places where standard deviations for the pose-change cost are reported as 246% and 257%, which look like percent-of-mean or are mislabeled; worth a close read.\n\nWho is this for: anyone building TBP systems or working on multi-view focus selection. It deserves a serious referee. At review, ask for an independent sharpness metric (MTF or blur estimate), a quantitative autofocus baseline, and at least a few repeated real scans. Even if those weaken the absolute numbers, the algorithmic contribution and the system proof-of-concept stand.","headline":"A genuinely new EM-based focus selection for TBP, with a solid system proof-of-concept; the headline coverage numbers are model-relative and need an independent sharpness check.","tokens_in":18953,"tokens_out":2225,"would_cite":true,"duration_ms":17172,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A total-body photography scanner that sets each camera's focus from the patient's estimated 3D body shape keeps 85–95% of the skin surface in focus at sub-0.075 mm/pixel resolution.","keywords":["total body photography","skin cancer screening","focus distance optimization","depth of field","multi-view scanning","expectation-maximization","3D body shape estimation","extended depth of field"],"falsifier":"Run one full mannequin scan with slant-edge or contrast targets placed at known depths, measure local sharpness across all surface patches, and compare the independently measured out-of-focus area with the area predicted by membership in $V_s^c$; a substantial mismatch would show that the reported 85–95% in-focus coverage is an artifact of the depth-of-field model rather than true sharpness.","tokens_in":17976,"feed_emoji":"📸","tokens_out":10077,"duration_ms":83216,"temperature":0.7,"pith_summary":"Total-body photography (TBP) is used to screen people at high risk for skin cancer, but when cameras are close enough to resolve small lesions their depth of field is shallow, so curved body surfaces produce many blurred images. This paper claims that blur can be largely removed by solving a focus-selection problem: given the patient's 3D body shape and the calibrated camera poses, choose a focus distance per camera so that the total surface area lying inside at least one camera's sharp zone is maximized. The authors build a rotary rig with seven RGB and three depth cameras, estimate the body shape before scanning, and solve the focus selection with an alternating assignment-and-update (EM) procedure. They report 84.9% of the simulated surface in focus at 0.068 mm/pixel and 95.1% of a real mannequin in focus at 0.0566 mm/pixel, improvements over closest-distance, average-distance, and auto-focus protocols. If the results hold, whole-body imaging can reach the sharpness needed for automated lesion detection without pausing the scan to focus.","feed_headline":"Focus from 3D body shape lifts in-focus skin to 85–95%","feed_subtitle":"Per-camera focus distances are chosen from an estimated 3D body model, beating fixed and auto-focus in tests.","key_machinery":"The machinery is the depth-of-field frustum $V_s^c$ paired with a pointwise cost. For camera $c$ focused at distance $s$, $V_s^c$ is the set of points whose images are deemed acceptably sharp, with near and far limits from a thin-lens depth-of-field model and a hyperfocal distance doubled for safety. The cost $\\kappa_s^c(p)$ in Eq. (1) is a weighted sum of a projected-area term, an optical-axis deviation term, and a binary term that is 1 outside $V_s^c$; the total objective $K(S)$ integrates the pointwise minimum over cameras. Because the binary term is piecewise constant in $s$, each EM update reduces to finding the focus interval containing the largest number of assigned points, making the solve fast enough to run in about five seconds. Importantly, the reported in-focus percentage is computed by exactly the same frustum membership that the optimization maximizes, so the metric and the objective coincide.","core_discovery":"The central claim is that per-camera focus distance is the main controllable lever for sharp whole-body coverage, and that the right setting can be computed, not guessed. The paper defines a per-point cost $\\kappa_s^c(p)$ that penalizes large projected surface area, distance from the optical axis, and points outside the depth-of-field frustum $V_s^c$, then minimizes the surface integral of the pointwise minimum over all cameras. The optimization is solved with the EM procedure: surface points are assigned to the camera that images them most cheaply, and each camera's focus distance is updated to the depth interval whose frustum contains the most assigned points. On 400 simulated body meshes this raises in-focus surface from 31.6% for closest focus and 68.5% for average focus to 84.9%; on a real mannequin the shape-aware method reaches 95.1% versus 38.4% and 80%. The authors also report that the gains are stable under realistic calibration noise, estimated-shape errors, and simulated postural sway, and qualitative crop comparisons show sharper hands, knees, inner arms, and thighs than auto-focus.","pith_inferences":["The same cost function could be inverted to optimize camera poses or angular camera density for a given body shape, a direction the paper lists as future work but does not implement.","The per-point camera assignment $\\phi(p)$ could serve as a correspondence prior for longitudinal tracking of lesions across scans, since it already identifies which camera best images each surface point.","A clinical validation on human subjects with an independent sharpness metric (such as MTF or local contrast) would test whether the geometric in-focus proxy used here matches perceived image quality; the current evidence is one mannequin scan plus simulation.","Extending the cost function with a contrast or illumination term would let the same EM framework jointly optimize focus and lighting, directly addressing the paper's stated limitation on skin-tone contrast."],"forward_implications":["A scanner using this method needs no per-patient manual focus tuning: one 8-second depth capture and a 5-second solve determine all focus distances before the 80-second image capture.","At the reported resolutions, roughly 73% of the simulated body surface meets the 0.075 mm/pixel threshold often cited for reliable automated detection of small lesions.","The assignment map produced by the EM solve can be used directly for image navigation, letting a clinician select any surface point and jump to the camera that images it best.","Because the focus solve is robust to the expected real-world errors, the protocol should transfer to new subjects without recalibration beyond the one-time system calibration."],"supporting_citations":[{"why":"Fixed-focus total-body system whose limited depth of field motivates the optimization approach.","marker":"[6]"},{"why":"Autofocus-based whole-body imaging baseline used for resolution and protocol comparison.","marker":"[8]"},{"why":"Supports the claim that autofocus is greedy in multi-view capture because it ignores already-covered surface.","marker":"[15]"},{"why":"Supplies the 400 textured body meshes used in the simulation evaluation.","marker":"[30]"},{"why":"Provides the camera intrinsic calibration method used in the system calibration pipeline.","marker":"[36]"},{"why":"Provides the template-based body registration alternative tested for 3D shape estimation.","marker":"[43]"},{"why":"Provides the volumetric signed-distance reconstruction used for the deployed 3D shape estimation.","marker":"[45]"},{"why":"Supplies the postural-sway range used to model pose changes in the robustness analysis.","marker":"[50]"}],"fun_headline_variants":["Shape-aware focus lifts in-focus skin to 85–95%","3D body model guides focus for sharper total-body scans","Body-shape focus optimization boosts in-focus skin area","Total-body camera focus set by 3D shape estimation","Per-camera focus from body shape beats auto-focus"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything depends on the assumption that the depth-of-field frustum, computed from a thin-lens model with a doubled hyperfocal distance, correctly identifies which surface points will actually look sharp in the captured photos.","fun_headline_variants_meta":{"raw":{"variants":["Shape-aware focus lifts in-focus skin to 85–95%","3D body model guides focus for sharper total-body scans","Body-shape focus optimization boosts in-focus skin area","Total-body camera focus set by 3D shape estimation","Per-camera focus from body shape beats auto-focus"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001065,"raw_usage":{"total_tokens":4514,"prompt_tokens":1045,"completion_tokens":3469,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":3385}},"tokens_in":661,"tokens_out":3469,"duration_ms":24264,"temperature":1.0,"reasoning_tokens":3385,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:04:50.249705+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one full mannequin scan with slant-edge or contrast targets placed at known depths, measure local sharpness across all surface patches, and compare the independently measured out-of-focus area with the area predicted by membership in $V_s^c$; a substantial mismatch would show that the reported 85–95% in-focus coverage is an artifact of the depth-of-field model rather than true sharpness.","supporting_citations":[{"cited_title":"A new total body scanning system for automatic change detection in multiple pigmented skin lesions,","cited_arxiv_id":null,"evidence_quote":"Fixed-focus total-body system whose limited depth of field motivates the optimization approach."},{"cited_title":"Monitoring of pigmented skin lesions using 3d whole body imaging,","cited_arxiv_id":null,"evidence_quote":"Autofocus-based whole-body imaging baseline used for resolution and protocol comparison."},{"cited_title":"Revisiting autofocus for smartphone cameras,","cited_arxiv_id":null,"evidence_quote":"Supports the claim that autofocus is greedy in multi-view capture because it ignores already-covered surface."},{"cited_title":"3dbodytex: Textured 3d body dataset,","cited_arxiv_id":null,"evidence_quote":"Supplies the 400 textured body meshes used in the simulation evaluation."},{"cited_title":"A flexible new technique for camera calibration,","cited_arxiv_id":null,"evidence_quote":"Provides the camera intrinsic calibration method used in the system calibration pipeline."},{"cited_title":"Nicp: Neural icp for 3d human registration at scale,","cited_arxiv_id":null,"evidence_quote":"Provides the template-based body registration alternative tested for 3D shape estimation."},{"cited_title":"A volumetric method for building complex models from range images,","cited_arxiv_id":null,"evidence_quote":"Provides the volumetric signed-distance reconstruction used for the deployed 3D shape estimation."},{"cited_title":"Feedforward ankle strategy of balance during quiet stance in adults,","cited_arxiv_id":null,"evidence_quote":"Supplies the postural-sway range used to model pose changes in the robustness analysis."}],"review_version":1}