{"id":"f4b82d46-6364-401c-afa3-372c1af5e65c","arxiv_id":"2607.05268","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Audited MERU, HyCoCLIP, and PHyCLIP checkpoints operate near-Euclidean with saturated entailment cones, so they do not demonstrate active radial or cone-based hierarchy.","lead":"This paper audits whether three popular hyperbolic vision–language models actually use the curved geometry they advertise. Released and retrained checkpoints stay in a near-Euclidean regime with saturated entailment cones, so the promised radial/cone hierarchy does not appear to be doing the work.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aperture formula in Eq. (10) is reciprocal-inconsistent with the saturation criterion used throughout; if implementations match Eq. (10), the cone-saturation and low-curvature shortcut claims invert.","rationale":"The reader's weakest_assumption is exactly this aperture-formula inconsistency, and I agree it is the most load-bearing concern. The near-Euclidean operating-point measurement (u≤0.37) is direct and would survive even if the aperture formula changed, but the cone-based half of the central claim — 'trained entailment cones are saturated or nearly saturated, so low violation rates can arise from trivially wide cones rather than learned order' — depends entirely on the reciprocal aperture formula. The paper prints Eq. (10) with 2K√cρ but uses the reciprocal form in every saturation calculation; these cannot both describe the same model. If the printed form were the real one, the observed u≈0.2 would give tiny cones, and the near-zero trained-direction violation rates would be strong evidence of active cone hierarchy, directly overturning the main negative result. Since the paper's own text is contradictory, the issue is internal and must be resolved by checking the implementations rather than by further argument. The reader's CONDITIONAL verdict is the right response: the conclusion is plausible but currently rests on an unresolved formula ambiguity. I recommend keeping that verdict, hence UNCHANGED.","tokens_in":47225,"tokens_out":4599,"duration_ms":49739,"concrete_test":"Inspect the released audit suite and the public MERU/HyCoCLIP/PHyCLIP code to extract the exact aperture expression used in the entailment loss. Then recompute Table 14's text- and image-aperture medians and saturation fractions on the same 256-sample GRIT batch using the implementation's formula. If the code uses the reciprocal form (2K/(√cρ)), the saturation classification and the low-curvature widening mechanism stand; if it uses the printed Eq. (10) form (2K√cρ), then the parent cones are narrow, the 'saturated/trivial containment' interpretation fails, and the low-curvature shortcut mechanism is inverted. Also re-derive the aperture formula from the original Ganea et al. (2018) definition to confirm which expression is correct for these models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's mechanistic account depends on the entailment-cone aperture formula. Section 7.3 prints Eq. (10) as ω(ρ)=arcsin(min{1,2K√cρ}), under which larger curvature or larger radius widens the cone. But the saturation criterion used everywhere else — §5.2, §7.3's closed-form edge, Appendix A's notation table, and Tables 4/13/14 — is ω=arcsin(min{1,2K/(√cρ)}), which saturates at π/2 when √cρ≤2K. These are reciprocal forms. Lowering c widens the cone only in the reciprocal form; under the printed form it narrows the cone. The paper's 'low-curvature shortcut' (reducing c suppresses violations by widening cones) and the interpretation of near-zero text→image violation as 'trivial containment under a saturated parent cone' are both consequences of the reciprocal form. If the implementations actually use the printed form, then at the measured u≈0.2 the trained text-parent cones would have half-aperture ≈arcsin(0.04)≈0.04 rad, i.e., extremely narrow, not saturated. In that regime, a 0% text→image violation rate would be evidence of finely learned directed geometric order — directly contradicting the paper's central claim that the cone mechanism is non-operative. Thus the entire cone-activity diagnosis and the gradient mechanism explaining curvature collapse hinge on which aperture expression is implemented. The paper itself contains both formulas, so this is an internal inconsistency, not an external speculation. Resolving it is prerequisite to accepting the mechanistic conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper audits whether MERU, HyCoCLIP, and PHyCLIP hyperbolic vision-language models actually use the radial/cone hierarchy mechanism they are motivated by. Using released checkpoints and matched from-scratch interventions on a fixed GRIT snapshot, it reports that all converged checkpoints remain near-Euclidean in the dimensionless radius u=√cρ (largest image-side value 0.37, far below the 10%-distortion marker near 0.84), that unclamping curvature does not move the operating point, that trained entailment cones are saturated or nearly saturated, that shuffle-controlled directed radial tests are null or unreplicated across seeds, and that calibrated semantic traversal detects only partial branch-conditioned order. The paper attributes the curvature collapse partly to a low-curvature shortcut in the entailment objective, and proposes a five-number geometry report for future hierarchy claims.","tokens_in":47590,"tokens_out":8759,"duration_ms":104734,"significance":"If the central claims survive, this is an important negative result with an unusually rigorous diagnostic apparatus: preregistered traversal thresholds, planted synthetic controls for sensitivity, power/MDE analyses, matched within-snapshot interventions, and a reproducibility suite with checkpoint hashes and analytic-versus-code tests. The paper carefully separates angular semantic organization from radial hierarchy and gives a mechanistic account of why the audited formulations leave the geometry dormant. The main obstacle is the internally inconsistent aperture formula, which is load-bearing for the low-curvature-shortcut mechanism and for the interpretation of near-zero entailment violations as trivial containment.","major_comments":[{"comment":"Eq. (10) prints ω(ρ)=arcsin(min{1,2K√cρ}), but the saturation criterion used in Tables 4, 13, 14 and in the closed-form boundary later in the same section is ω=arcsin(min{1,2K/(√cρ)}), which saturates at π/2 when √cρ≤2K. The two forms are reciprocals. Under the printed form, lowering curvature narrows the cone, and at the measured text-side u≈0.12–0.20 the trained text-parent cones would have half-apertures of only ~0.02–0.04 rad. In that regime, the reported 0–0.4% text→image violation rates would be evidence of finely learned directed order, directly contradicting the paper's central claim. The low-curvature shortcut and the 'saturated cone ⇒ trivial containment' interpretation both depend on the reciprocal form. The manuscript contains both formulas without identifying which one is implemented. This must be resolved by checking the released/from-scratch code and correcting Eq. (10) or","section":"§7.3, Eq. (10)"},{"comment":"The saturation classification used throughout the cone diagnostics is the reciprocal form (saturated when √cρ≤2K with K=0.1). The same section states that the analytic identity and endpoint match carry the mechanistic claim. Because the endpoint evidence (trained box-image parents ending within 0.013 of 2K) is derived from this same aperture identity, the internal inconsistency in Eq. (10) propagates into the main evidence for the low-curvature shortcut. Please state explicitly which aperture expression each implementation uses, provide the code-level verification promised in the reproducibility statement, and, if Eq. (10) is a typo, mark it as such and correct every later reference.","section":"§7.3/Tables 4, 13, 14"}],"minor_comments":[{"comment":"The sentence 'As foreshadowed in Section 5.2, reducing c widens the aperture' is inconsistent with the printed Eq. (10), under which reducing c narrows the aperture. This cross-reference should be fixed together with the formula.","section":"§7.3 text after Eq. (10)"},{"comment":"The notation table and the five-number geometry report both use the reciprocal aperture form. Once Eq. (10) is corrected, please ensure the notation table, the closed-form edge, and all table captions use one consistent expression.","section":"Appendix A / §7.5"},{"comment":"The saturation definition (ω≥π/2−0.01) is given only in the caption. Consider stating it in Section 5.2 or Section 7.3, since the interpretation of 'saturated' is central to the argument.","section":"Table 14 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually thorough in its controls and sensitivity analyses, and the negative result is plausible and potentially influential. The aperture-formula inconsistency is the only substantive obstacle I see; it is likely a typographical or convention error, but because it sits at the center of the mechanism claim, it must be resolved with a code-level verification and a corrected, unified formula before publication. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the negative audit is mostly convincing: every released and converged checkpoint sits at dimensionless radius u = sqrt(c)*rho ≤ 0.37, well below the 10%-distortion point around 0.84, and the shuffle-controlled radial tests find no pair-specific ordering beyond marginal norm effects. The operating-point result is a real contribution — scalar curvature alone was never a good measure of whether hyperbolic geometry is active, and u is a better one. The five-number geometry report is a practical checklist, and the paper is genuinely careful: planted controls with power analysis, preregistered traversal tiers, shuffle nulls, and honest caveats about the post-hoc HyCoCLIP-S pilot exclusion and the uncalibrated NG2 registry.\n\nSecond, the mechanism story is undermined by an internal inconsistency. Eq (10) prints ω = arcsin(min{1, 2K sqrt(cρ)}), but the saturation criterion used in Section 7.3, Tables 4/14, and the notation table is ω = arcsin(min{1, 2K/(sqrt(cρ))}). These are reciprocal. The low-curvature shortcut — lowering c widens cones and suppresses violations — is a consequence of the reciprocal form; under the printed form, lowering c would narrow cones and make violations more likely, inverting the proposed mechanism. At the measured u≈0.2, the two forms give text-parent cones either saturated (reciprocal) or extremely narrow (printed), which reverses the interpretation of the near-zero text-to-image violation rate. The paper contains both formulas, so this is not external speculation; it is load-bearing. The gradient data (entailment pushes c down) and the standard Ganea entailment-cone formula both support the reciprocal form, so I suspect Eq (10) is a typo. But the authors must fix it, confirm which aperture the released code actually implements, and re-verify the saturation classification and the shortcut argument. That is a prerequisite to accepting the mechanistic conclusion.\n\nFor whom is this paper? Anyone working on hyperbolic VLMs or evaluating geometry claims in representations. The operating-point and shuffle-test findings stand and will be cited. The cone-saturation and shortcut claims probably survive once the formula is corrected, but they need re-verification. Worth a serious referee, with instructions to pin down the aperture identity early.","headline":"Operating-point audit that lands, but the cone-saturation mechanism is compromised by a reciprocal-inconsistent aperture equation.","tokens_in":48041,"tokens_out":6950,"would_cite":true,"duration_ms":75473,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The hyperbolic geometry in audited vision-language models is effectively dormant: embeddings remain near-Euclidean and trained entailment cones are saturated.","keywords":["hyperbolic vision-language models","entailment cones","curvature collapse","dimensionless radius","radial ordering","hierarchy audit","low-curvature shortcut","replication study"],"falsifier":"Inspect the cone-aperture formula in each released implementation and measure the trained cone apertures while halving c with norms held fixed; if apertures shrink rather than grow, the low-curvature shortcut is inverted and the saturation-edge explanation (√cρ≤2K) collapses.","tokens_in":47109,"feed_emoji":"📐","tokens_out":6660,"duration_ms":69882,"temperature":0.7,"pith_summary":"This paper asks whether three published hyperbolic vision-language model families—MERU, HyCoCLIP, and PHyCLIP—actually use the negative-curvature geometry they are built on. Under the paper's diagnostics, they do not: across released checkpoints and matched from-scratch runs, the dimensionless radius u=√cρ never exceeds 0.37, well below the u≈0.84 where local metric distortion reaches 10%, so the learned embeddings sit in a near-Euclidean regime. Trained entailment cones are saturated or nearly saturated, making very low violation rates a trivial consequence of wide cones rather than learned order. The paper traces this to a low-curvature shortcut: lowering curvature widens cone apertures and suppresses violations without learning hierarchy, and gradient decomposition shows entailment is the dominant curvature-lowering pressure, though not the sole cause. A careful reader would care because prior evidence for hierarchy in these models is underdetermined by angular similarity and marginal norm effects, and the paper offers a concrete reporting standard for future claims.","feed_headline":"Audited hyperbolic VLMs stay near-Euclidean; hierarchy is dormant","feed_subtitle":"Even at their most hyperbolic, embeddings stay below 10% distortion; saturated cones make zero violation trivial.","key_machinery":"The key machinery is the dimensionless operating coordinate u=√cρ (with local distortion factor H(u)=u/asinh(u)), which lets the audit separate learning the curvature scalar from actually using hyperbolic geometry; the entailment-cone half-aperture ω=arcsin(min{1,2K/(√cρ)}) with K=0.1, whose saturation edge √cρ≤2K gives a closed-form criterion for when cones are trivially wide; and a gradient decomposition that attributes curvature movement to individual loss terms. Together these turn a benchmark question into an operating-point question, and ground the paper's proposed five-number geometry report.","core_discovery":"Central claim: audited hyperbolic vision-language formulations do not demonstrate an operative radial or cone-based hierarchy. Effective geometry is governed by u=√cρ, and every converged checkpoint sits at u≈0.1–0.37—below the u≈0.84 10%-distortion point; unclamping curvature changes c and norms but not this band. Trained entailment cones saturate at π/2, so low violation rates are trivial containment, shuffle-controlled tests find no pair-specific radial ordering, and traversal yields only weak branch-conditioned order. The mechanism is a low-curvature shortcut: cone half-aperture widens as √cρ shrinks toward saturation edge 2K=0.2, so entailment suppresses violations by lowering curvature","pith_inferences":["Beyond the paper: the u=√cρ operating-point test transfers to any hyperbolic embedding claim (word embeddings, knowledge graphs): if learned radii never push u past about 0.8, negative-curvature structure is probably not doing the work.","A direct test the paper leaves open is a K-sweep: if the saturation edge 2K causally anchors trained parent coordinates, varying K should move the observed endpoints; if the endpoints stay fixed, the edge is a correlation, not a mechanism.","The aperture formula in Eq. (10) deserves code-level verification; if the non-reciprocal form is actually implemented, the shortcut's direction reverses and the saturation classification needs re-deriving.","A testable extension: replace hyperbolic encoders with Euclidean encoders plus a learned radial scale; if downstream hierarchy metrics are preserved, angular and supervision effects are sufficient and the hyperbolic geometry is not needed."],"forward_implications":["If these diagnostics are accepted, prior claims that hyperbolic VLMs encode hierarchy based on taxonomy-distance correlation, leaf-level accuracy, or zero violation rates need re-examination: those signals are compatible with angular structure and saturated cones.","Future hyperbolic VLM claims should report the five-number geometry report (operating point, cone-saturation state, directed violations, shuffle-controlled radial excess, radial increment beyond angle); without it, a positive hierarchy claim is not mechanism-tested.","The low-curvature shortcut implies that simply adding entailment losses or lowering the curvature floor will not activate hierarchy; objectives must decouple aperture width from norm/curvature and include radial-growth control.","Because entailment-off training also collapses curvature, fixing the cone loss alone cannot stabilize a nonlocal operating point; the contrastive/alignment objective must be modified as well.","The analytic saturation edge √cρ≤2K gives a parameter-free check: models whose trained parent coordinates lie at or below 0.2 cannot have informative trained cones."],"fun_headline_variants":["Hyperbolic VLMs never go hyperbolic: audit says flat","Sat cones, flat radii: hyperbolic VLM hierarchy dormant","Low-curvature shortcut: why hyperbolic VLMs stay Euclidean","Geometry idle: audited VLMs resist hyperbolic space"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The mechanism account depends on the implementations using the reciprocal cone-aperture formula (aperture grows as 1/√cρ); if the formula as printed in Eq. (10) is what the code actually uses, lowering curvature would narrow the cones and the entire low-curvature-shortcut explanation would run backwards.","fun_headline_variants_meta":{"raw":{"variants":["Hyperbolic VLMs never go hyperbolic: audit says flat","Sat cones, flat radii: hyperbolic VLM hierarchy dormant","Low-curvature shortcut: why hyperbolic VLMs stay Euclidean","Geometry idle: audited VLMs resist hyperbolic space"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1565,"prompt_tokens":872,"completion_tokens":693,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":626}},"tokens_in":616,"tokens_out":693,"duration_ms":8373,"temperature":1.0,"reasoning_tokens":626,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:27:20.303505+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the cone-aperture formula in each released implementation and measure the trained cone apertures while halving c with norms held fixed; if apertures shrink rather than grow, the low-curvature shortcut is inverted and the saturation-edge explanation (√cρ≤2K) collapses.","supporting_citations":[],"review_version":3}