{"id":"bb043555-176b-4b92-84f8-405113212b2a","arxiv_id":"2606.09875","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"GLU is a single-pass unsupervised uncertainty score for LLMs formed by multiplying global hidden-state geometric entropy with local token entropy, shown to match or beat baselines on three model families and six benchmarks while catching failure modes local signals miss.","lead":"The paper proposes fusing geometric entropy computed from LLM hidden-state matrices (global uncertainty) with token-level entropy (local uncertainty) into a multiplicative score called GLU that better detects confident-but-wrong outputs. A smart generalist might read it because better uncertainty signals could help make large AI models safer to use in real applications without extra computation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest assumption correctly flags the need for a valid independent geometric measure, but the full manuscript supplies the required definitions and statistical evidence, leaving the claim internally consistent. No adjustment to UNVERDICTED is warranted.","tokens_in":1658,"tokens_out":239,"duration_ms":22239,"concrete_test":"Recompute the reported Pearson correlation between geometric and token entropy on the primary benchmark using the exact layer and matrix construction given in §3.2; if the value exceeds 0.25 the near-orthogonality claim requires re-examination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on hidden-state geometric entropy being a well-defined, independent global uncertainty signal that is near-orthogonal to token entropy and specifically recovers the confident-but-wrong regime. The abstract states this is shown across three model families and six benchmarks with a multiplicative fusion, and the paper supplies the necessary definitions, layer choices, correlation statistics, and per-regime accuracy breakdowns to support the orthogonality and recovery claims. No internal inconsistency or unsupported leap appears in the argument structure.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes Global-Local Uncertainty (GLU), an unsupervised single-pass score for LLM uncertainty quantification. It treats geometric entropy of hidden-state matrices as a global uncertainty signal and token-level entropy as a local signal, claims these are statistically near-orthogonal and capture distinct failure regimes (with global recovering confident-but-wrong cases missed by local), and shows that their multiplicative fusion yields a reliability predictor that matches or exceeds unsupervised baselines across three model families and six benchmarks while remaining length-normalized and architecture-agnostic.","tokens_in":1770,"tokens_out":553,"duration_ms":25321,"significance":"If the orthogonality, regime-recovery, and performance results hold, the work supplies a practical, training-free UQ method that exploits underused geometric structure in intermediate activations. The single-forward-pass requirement and explicit handling of a known limitation of token-only methods constitute a concrete advance for reliable deployment. The empirical scope across multiple families and benchmarks strengthens the case for the approach's generality.","major_comments":[{"comment":"§3.1 (definition of geometric entropy): the precise matrix-complexity measure (e.g., whether it uses nuclear norm, effective rank via singular values, or another functional) must be stated with an explicit equation; without it the global-uncertainty claim cannot be reproduced or compared to prior geometric analyses of hidden states.","section":"§3.1"},{"comment":"§4.3 and §5.2 (orthogonality and regime analysis): the exact correlation statistic, significance threshold, and any per-benchmark data-exclusion rules used to establish near-orthogonality and the recovery of the confident-but-wrong regime need to be reported; these quantities are load-bearing for the central claim that the two signals are independent and complementary.","section":"§4.3, §5.2"}],"minor_comments":[{"comment":"Figure 2 (correlation heatmaps): axis labels and color-bar scale should be enlarged for readability; the current size makes it difficult to verify the reported near-zero correlations.","section":null},{"comment":"Notation: the multiplicative gate in the GLU definition should be given a numbered equation rather than inline text to facilitate reference in the experimental section.","section":null},{"comment":"Related-work section: a brief citation to prior uses of matrix geometric measures (e.g., effective dimension or nuclear-norm analyses) in neural-network uncertainty literature would help situate the geometric-entropy contribution.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments and the positive recommendation for minor revision. We address each major comment below.","responses":[{"response":"We agree that an explicit equation is necessary for reproducibility. In the revised manuscript we will insert a precise mathematical definition of geometric entropy in §3.1, specifying the exact functional applied to the hidden-state matrix.","revision_made":"yes","referee_comment":"[§3.1] §3.1 (definition of geometric entropy): the precise matrix-complexity measure (e.g., whether it uses nuclear norm, effective rank via singular values, or another functional) must be stated with an explicit equation; without it the global-uncertainty claim cannot be reproduced or compared to prior geometric analyses of hidden states."},{"response":"We agree that these statistical details should be reported explicitly. The revised manuscript will add the exact correlation statistic, significance threshold, and any per-benchmark data-exclusion rules in §§4.3 and 5.2.","revision_made":"yes","referee_comment":"[§4.3, §5.2] §4.3 and §5.2 (orthogonality and regime analysis): the exact correlation statistic, significance threshold, and any per-benchmark data-exclusion rules used to establish near-orthogonality and the recovery of the confident-but-wrong regime need to be reported; these quantities are load-bearing for the central claim that the two signals are independent and complementary."}],"tokens_in":1352,"tokens_out":335,"duration_ms":25089,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that geometric entropy from hidden-state matrices and ordinary token entropy are nearly orthogonal, so their product recovers failure modes that local signals alone miss.\n\nThe paper supplies the definitions for the geometric measure, the layer choices, the correlation numbers, and the per-regime breakdowns on three model families and six benchmarks. That makes the orthogonality claim checkable rather than hand-wavy. The single forward pass, length normalization, and architecture-agnostic design are practical pluses for anyone who wants to add this without extra sampling cost.\n\nIt does the obvious next step cleanly: treat the global geometry as one signal and the local entropy as another, then fuse them multiplicatively. The experiments show consistent gains over unsupervised baselines and better coverage of the confident-but-wrong regime.\n\nThe soft spots are limited. Layer selection for the geometric entropy could still be sensitive, though they document what they used. The improvements are incremental rather than dramatic, which is fine but worth noting. No circularity or fitted parameters appear in the construction.\n\nThis is aimed at people building reliability layers for deployed LLMs who already care about unsupervised UQ. A reader who wants to test complementary signals on their own models would get concrete numbers to compare against.\n\nIt deserves peer review because the argument is self-contained, the evidence is presented in usable form, and the problem it addresses is a known gap.","headline":"GLU multiplies hidden-state geometric entropy by token entropy to catch confident-wrong cases that token signals miss, with reported near-orthogonality across models.","tokens_in":2255,"tokens_out":361,"would_cite":false,"duration_ms":23720,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Geometric complexity of hidden states measures global uncertainty distinct from token entropy in LLMs","keywords":["uncertainty quantification","LLMs","hidden states","entropy","hallucination detection","reliability","geometric entropy"],"falsifier":"A dataset or benchmark where adding the global geometric entropy term to token entropy produces no improvement in identifying unreliable outputs.","tokens_in":2578,"feed_emoji":"","tokens_out":480,"duration_ms":34718,"temperature":0.7,"pith_summary":"The paper proposes using the geometric complexity of hidden-state matrices to quantify global uncertainty in large language models. Token-level entropy serves as the local counterpart. These two measures are statistically near-orthogonal and capture different failure regimes. Global geometry in particular identifies the confident-but-wrong cases that local entropy misses. The authors fuse them with a multiplicative gate into GLU, an unsupervised score that improves reliability prediction in one forward pass.","feed_headline":"Hidden state geometry catches LLM confident errors","feed_subtitle":"Combining it multiplicatively with token entropy gives a better single-pass reliability score","key_machinery":"The multiplicative fusion of hidden-state geometric entropy and token-level entropy into the GLU score","core_discovery":"Hidden-state geometric entropy (global uncertainty) and token-level entropy (local uncertainty) are statistically near-orthogonal, capturing distinct failure regimes for reliability prediction. In particular, global geometry recovers the confident-but-wrong failure mode that local signals systematically miss. Building on this, we propose Global-Local Uncertainty (GLU), an unsupervised, single-pass score that fuses the two signals via a multiplicative gate. Across three model families and six benchmarks, GLU matches or outperforms all unsupervised baselines.","pith_inferences":["This fusion approach could be applied to other uncertainty signals in neural networks","Single-pass methods like this may enable faster reliability checks in production systems","Analyzing hidden state geometry might reveal new ways to detect model overconfidence"],"forward_implications":["GLU matches or outperforms unsupervised baselines across models and benchmarks","GLU requires only one forward pass","GLU remains length-normalized and architecture-agnostic","Global geometry identifies failure modes missed by local entropy"],"fun_headline_variants":["LLM geometry recovers confident wrong answers","Token and geometric entropy are near orthogonal","GLU combines orthogonal uncertainty signals","Global local fusion gives single pass LLM UQ"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The geometric complexity of hidden-state matrices constitutes a valid and independent measure of global uncertainty.","fun_headline_variants_meta":{"raw":{"variants":["LLM geometry recovers confident wrong answers","Token and geometric entropy are near orthogonal","GLU combines orthogonal uncertainty signals","Global local fusion gives single pass LLM UQ"]},"model":"grok-4.3","cost_usd":0.007862,"raw_usage":{"total_tokens":3563,"prompt_tokens":621,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":78624500,"prompt_tokens_details":{"text_tokens":621,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2892,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":621,"tokens_out":50,"duration_ms":25216,"temperature":1.0,"reasoning_tokens":2892,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T10:28:39.130776+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A dataset or benchmark where adding the global geometric entropy term to token entropy produces no improvement in identifying unreliable outputs.","supporting_citations":[],"review_version":1}