{"id":"b24b9dc5-c482-489a-935b-f3ef13e00f6b","arxiv_id":"2607.24746","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Humanoid robots systematically lack biological ISO 7250 landmarks but retain measurable joint-level geometry, with four recurring deviation patterns (additive, subtractive, exaggerative, speculative) driving ergonomic risk.","lead":"This paper benchmarks six humanoid robots against the ISO 7250 human body measurement standard using photos and videos, finding that many anatomical landmarks don't exist on robots while joint-level measurements mostly do. It proposes four ways robot bodies deviate from human proportions and argues these deviations, not just size, drive ergonomic risk in human-robot collaboration.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Risk inference step is unvalidated: deviations are catalogued, but the central claim that 'ergonomic risk arises' from them is asserted without task, posture, force, or validated risk metrics; the paper's own §2 and §3.4 hedge with 'may.'","rationale":"The reader's weakest assumption identified two linked issues: (1) reliability of visual landmark classification, and (2) the unsupported step from geometric deviation to ergonomic risk. This stress-test selects the risk-inference step as the single most load-bearing concern because the paper's novel and headline-level contribution is precisely the ergonomic-risk conclusion; even if landmark classification were made fully reliable with annotated images and inter-rater checks, the causal claim would still lack support without a demonstrated link between deviation patterns and worker strain. The paper explicitly hedges with 'may' in §2 and §3.4, yet the abstract and conclusions state the risk claim categorically. This is an internal-evidence gap, not a disagreement with external consensus. The verdict remains CONDITIONAL: the descriptive taxonomy is plausible and reproducible in principle, but the paper must either temper its central claim or supply task/posture/force evidence. The proposed simulation is a concrete, low-cost test that directly targets the disputed inference. Agreement is 'partial' because the reader centered visual reliability, while this pass centers the downstream risk inference, though both appear in the reader's weakest_assumption.","tokens_in":4653,"tokens_out":3296,"duration_ms":37794,"concrete_test":"Conduct a digital human simulation with an ergonomic risk model (e.g., Siemens Jack, AnyBody, or 3DSSPP) for a standardized collaborative task—for instance, a handover from a human worker to a robot at two work heights (e.g., 100 cm and 140 cm). Simulate two robot morphologies matched for overall stature but differing in a documented deviation pattern: one with human-like limb proportions and one with Digit-like backward-bent legs and altered joint heights. For each condition, fix the human worker's posture and compute RULA/REBA scores or L5/S1 compression. If risk scores are indistinguishable across the two morphologies after controlling for stature, the central claim 'risk arises from how/where bodies diverge rather than size alone' fails; if scores differ systematically, the claim receives direct support. Also report inter-rater reliability of landmark coding on a subsample of the six","verdict_should_be":"UNCHANGED","load_bearing_attack":"Even granting the author's visual landmark classifications, the paper's central claim—'ergonomic risk arises not from size alone, but from how and where humanoid bodies diverge from human anthropometric logic'—is not supported by the evidence presented. The study catalogs geometric deviations (additive, subtractive, exaggerative, speculative) and shows that some ISO 7250 landmarks are unidentifiable, but it never measures or models an actual ergonomic risk outcome. §2 explicitly states that risks are 'infer[red] ... despite the absence of task-specific posture data'; §3.4 uses hedged language: deviations 'may affect reach alignment, clearance, and posture' and 'may increase ... ergonomic strain.' RULA/REBA are mentioned only as 'analytical anchors,' not applied to any worker–robot task, and no biomechanical or human-factors data are reported. Thus the load-bearing inference from 'humanoid bodies diverge from human proportions' to 'ergonomic risk arises' rests on plausibility rather than evidence. The four deviation patterns are defined by examples, not by quantitative thresholds or a demonstrated link to physical strain. If geometry alone does not predict ergonomic risk, the headline conclusion overreaches; the paper itself contains the admission that task-specific data are absent. This is not an external-consensus objection but an internal-evidence gap in the causal chain.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks six contemporary humanoid robots against the ISO 7250 human body measurement standard using publicly available images and videos. It classifies ISO landmarks as identifiable, proxy-identifiable, or non-identifiable; assesses which body measurements are feasible; and describes four recurring patterns of anthropometric deviation (additive, subtractive, exaggerative, speculative). The headline claim is that ergonomic risk in human–humanoid collaboration arises not from scale alone but from how humanoid bodies diverge from human anthropometric logic, and that ISO 7250 remains useful as a geometric reference but not as a full anthropometric standard for robots.","tokens_in":4942,"tokens_out":2295,"duration_ms":25800,"significance":"If the central claim were fully supported, the paper would make a useful conceptual contribution by giving practitioners a taxonomy for thinking about humanoid–human physical compatibility and by highlighting the limits of transferring human anthropometric standards to non-biological embodiments. Strengths of the work include the use of an external benchmark (ISO 7250), the absence of fitted free parameters, a qualitatively reproducible benchmarking protocol, and a clearly stated, testable thesis. However, as it stands the evidence is largely a single-author visual classification exercise. The quantitative statements in §3.4 are not backed by raw data or tables, and the leap from geometric deviation to ergonomic risk is asserted rather than demonstrated. These gaps prevent the paper from supporting its headline conclusion in its current form.","major_comments":[{"comment":"The quantitative claims (e.g., 70–85% identifiable/proxy-identifiable landmarks, CV 10–15% for key proportions, 100% non-identifiable for biological landmarks, ≈17% facial landmark coverage, 'at least two deviation types' for all robots) appear with no supporting tables, raw measurements, or confidence intervals. The classifications are based on the author's visual inspection of unspecified public images and videos, with no inter-rater reliability check. These numbers are load-bearing for the cross-platform claims and must be either backed by an appendix containing a per-robot landmark/measurement matrix and the underlying normalized measurements, or removed from the paper.","section":"§3.4, §2"},{"comment":"The central inference from geometric deviation to ergonomic risk is not validated. The paper's own §2 states that risks are inferred 'despite the absence of task-specific posture data,' and §3.4 hedges with 'may affect reach alignment, clearance, and posture' and 'may increase ... ergonomic strain.' No RULA/REBA scores, task simulations, biomechanical models, or human-factors outcome data are reported. The headline conclusion—'ergonomic risk arises not from size alone, but from how and where humanoid bodies diverge from human anthropometric logic'—therefore overreaches the evidence. The authors should either reframe the conclusion as a testable hypothesis or add a concrete ergonomic analysis, e.g., applying RULA/REBA to documented collaboration scenarios.","section":"§3.3, §4"},{"comment":"Reproducibility is not yet achieved as written. The methodology says robots were analyzed from 'publicly available images and videos,' but no data sources, URLs, frame identifiers, or measurement procedures are listed. The selection of which images/videos were used and how landmarks were localized is not documented. Without this, another researcher cannot reproduce the classifications. Please provide a data/evidence appendix with sources and, where feasible, annotated images or screenshots.","section":"§2, §3.1"}],"minor_comments":[{"comment":"The height range '187.00 cm to 5.77 cm' appears to be a typo; 5.77 cm is not physically plausible for a humanoid robot. Please verify the intended value (e.g., 57.7 cm) and correct it.","section":"Abstract / §1"},{"comment":"Figures 1–7 are referenced but not included in the preprint text examined. Ensure that final version contains all figures with clear annotations, and that figure captions fully explain what is being classified (e.g., which robot, which landmark, which view).","section":"Figures"},{"comment":"Reference [7] is a self-citation about human–automated system conflict; its relevance to the ergonomic-challenges sentence in §1 is not obvious. Clarify the connection or replace with a more directly relevant citation.","section":"Reference [7]"}],"recommendation":"major_revision","confidential_remarks":"The paper reads as a conference-scale conceptual note rather than a fully evidenced empirical study. The taxonomy and the ISO 7250 transferability argument are plausible and could be a useful contribution to the human–robot interaction and ergonomics community. However, the current manuscript's central claim is broader than its evidence base, and the quantitative statements are unverifiable without raw data. For a journal version, I would expect either a substantial empirical supplement (data tables, inter-rater checks, or an applied RULA/REBA example) or a carefully de-scoped conclusion that frames the taxonomy as a hypothesis-generating framework. The paper's own hedged language in §3.4 closely tracks my assessment: the evidence supports 'may affect,' not 'arises.'"},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a short, honest paper that does something genuinely new — benchmarking six humanoids against ISO 7250 landmarks and naming four deviation patterns (additive, subtractive, exaggerative, speculative). The descriptive core holds together. The problem is the load-bearing claim that these deviations produce ergonomic risk; that's asserted from geometry, not measured. The paper itself admits \"absence of task-specific posture data\" (§2) and uses \"may\" in §3.4. So treat the risk language as overreach, not as the main contribution.\n\nWhat's good: Applying ISO 7250 to robots is a legitimate new application. The distinction between identifiable joint landmarks and non-identifiable soft-tissue landmarks is clear and plausible. The taxonomy is simple but usable. The method is externally reproducible in principle — no internal specs, no fitted parameters, and the author's self-citation [7] is not a problem. For a conference paper, the scope is reasonable.\n\nWhere it's soft: The evidence is qualitative visual inspection with no raw annotation tables, no inter-rater reliability, no confidence intervals. The quantitative line (70–85% identifiability, CV 10–15% vs 3–5%) appears without supporting data; those numbers are not load-bearing, but they shouldn't be presented as results. There's also a glaring typo: \"height varies from 187.00 cm to 5.77 cm\" — that 5.77 cm is impossible and needs a fix. The more substantive weakness is the inference from geometric deviation to ergonomic risk. You'd need task-specific posture, force, or RULA/REBA application to substantiate it; mentioning RULA/REBA as \"analytical anchors\" doesn't do the job. So the conclusion should be reframed as an hypothesis or a screening rationale, not a finding.\n\nWho it's for: people working on human-robot workspace design, anthropometry standards, and robot disclosure practices. It's a conference-level contribution, and with the risk claims scaled back, it's publishable.\n\nMy recommendation: send it to peer review, not desk reject. The flaws are fixable, the core observation is useful, and the framework can be tested by other researchers. But the referee should require the raw data or a clear data-sharing statement, and a more careful separation of description from risk inference.","headline":"Useful descriptive framework for benchmarking humanoids against ISO 7250, but the ergonomic-risk conclusion outruns the evidence; deserves a serious referee.","tokens_in":5401,"tokens_out":2222,"would_cite":false,"duration_ms":21375,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A study of six humanoid robots finds that ergonomic risk in human–robot collaboration comes not from size but from how robot bodies diverge from human anthropometric logic, making ISO 7250 only a geometric reference.","keywords":["human-robot collaboration","ergonomics","anthropometry","ISO 7250","humanoid robots","landmark identifiability","measurement feasibility","anthropometric deviation"],"falsifier":"Measure the six robots—from CAD models, 3D scans, or manufacturer data—and see whether the paper's landmark classifications hold; then run a standardized collaborative task and check whether measured RULA/REBA scores correlate with the four deviation patterns. If biological landmarks turn out to be locatable or deviations do not predict strain, the central claim collapses.","tokens_in":4513,"feed_emoji":"🤖","tokens_out":5411,"duration_ms":44431,"temperature":0.7,"pith_summary":"The paper tries to establish that the human-centered measurement standard ISO 7250 cannot be carried over to humanoid robots as a full anthropometric standard. Benchmarking six humanoid robots against the standard's landmarks and measurements using only externally observable geometry, it finds that landmarks tied to biological anatomy (such as the hip bone, nipple, shin, and ear notch) are not identifiable on any robot, while joint-based and geometric dimensions are usually measurable. The author identifies four recurring patterns of deviation from human proportions—additive, subtractive, exaggerative, and speculative—and argues that these patterns, not height alone, are what create ergonomic risk by changing clearance, reach alignment, motion predictability, and risk perception. If correct, the result matters because standard ergonomic risk tools assume human anatomical landmarks, so humanoid-specific anthropometry is needed. A sympathetic reader would see a testable benchmarking procedure and a useful taxonomy for robot–human physical compatibility.","feed_headline":"Humanoid bodies fail ISO's human measurement standard","feed_subtitle":"Biological landmarks are unmeasurable on machines, so ergonomic risk is a geometry problem, not a size problem.","key_machinery":"The carrying mechanism is the ISO 7250-1:2017 body measurement framework applied as a geometric benchmark. Landmarks are classified as identifiable, proxy-identifiable, or not identifiable, and measurements as feasible or infeasible, based only on externally observable geometry. The four-part deviation taxonomy—additive, subtractive, exaggerative, speculative—is the analytical device that converts landmark and measurement gaps into ergonomic-risk claims. The author anchors these claims with established ergonomic principles such as RULA and REBA to infer strain from geometry in the absence of task-specific posture data.","core_discovery":"Using ISO 7250 as a reference framework, the author benchmarked six contemporary humanoid robots—Ameca, Optimus Gen 3, Figure 03, Atlas Electric, Unitree G1, and Digit—from publicly visible geometry. The central discovery is a structural mismatch: landmarks that mark rigid body extremes or joints (top of head, shoulder point, elbow) are identifiable or proxy-identifiable across all robots, while landmarks tied to skeletal or soft-tissue anatomy (ASIS, thelion, tibiale, tragion) are non-identifiable on every platform, making the associated measurements infeasible. Across platforms the author observes four deviation patterns that often co-occur: additive (extra parts such as Digit's leg suppor","pith_inferences":["Inference: the deviation taxonomy could be mapped to specific failure modes—additive protrusions to contact hazards, subtractive faces to reduced motion legibility, exaggerative features to biased risk perception—but the paper does not test these mappings.","Inference: a direct test would compute RULA/REBA scores for a standardized task using each robot's measured joint geometry and compare them with human-human baselines; if scores do not track deviation patterns, the geometry-to-risk link would weaken.","Inference: the visual-inspection method could be hardened with CAD models or 3D scans of the same robots, turning qualitative deviation categories into quantified distance and proportion measures.","Inference: if humanoid design continues toward function-optimized bodies, the field may need a parallel robot-anthropometry standard built from joint centers and contact surfaces rather than anatomical landmarks."],"forward_implications":["If the paper is right, standard ergonomic assessment tools like RULA and REBA cannot be directly applied to humanoid robots, because they rely on anatomical landmarks that are absent.","If the paper is right, robot manufacturers should disclose dimensional and joint data in anthropometrically compatible form, and designers should treat each humanoid as a fixed configuration rather than a scaled human.","If the paper is right, the four deviation patterns give engineers a practical checklist for anticipating clearance, reach, and perception problems before deployment.","If the paper is right, the variability among humanoid platforms (about 10–15% CV for key ratios, exceeding human 3–5%) means a workplace tuned for one robot may not transfer to another."],"fun_headline_variants":["Humanoid robots miss key human landmarks for safety","Biological landmarks unmeasurable on humanoid robots","Ergonomic risk from robot geometry, not just size","ISO human scale standards don't apply to humanoids","Humanoid bodies diverge from human anthropometric norms"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that visual inspection of publicly available photos and videos by a single rater reliably identifies which ISO landmarks exist on a robot, and that geometric deviation predicts ergonomic strain without task or posture measurements.","fun_headline_variants_meta":{"raw":{"variants":["Humanoid robots miss key human landmarks for safety","Biological landmarks unmeasurable on humanoid robots","Ergonomic risk from robot geometry, not just size","ISO human scale standards don't apply to humanoids","Humanoid bodies diverge from human anthropometric norms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000355,"raw_usage":{"total_tokens":1762,"prompt_tokens":734,"completion_tokens":1028,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":952}},"tokens_in":478,"tokens_out":1028,"duration_ms":9568,"temperature":1.0,"reasoning_tokens":952,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T13:40:36.174190+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the six robots—from CAD models, 3D scans, or manufacturer data—and see whether the paper's landmark classifications hold; then run a standardized collaborative task and check whether measured RULA/REBA scores correlate with the four deviation patterns. If biological landmarks turn out to be locatable or deviations do not predict strain, the central claim collapses.","supporting_citations":[],"review_version":1}