{"id":"7d712907-9935-465d-ab99-128d1fbbcb79","arxiv_id":"2607.13568","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Prompt-point activations carry a graded, steerable entity-familiarity signal that is robust to Polish/English stem changes and is stronger in Polish-adapted models than in base models.","lead":"This paper asks whether a language model can sense how well it knows an entity before it starts answering, by reading its internal activations at the end of the question. It finds that a simple probe trained on Polish entities separates real from fabricated names across four model families, tracks popularity in Polish-adapted models, and that steering one activation direction makes a refusing model refuse more or less on command.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-vs-fabricated anchor is not matched on name-surface statistics, so the probe (and the L30 direction built from the same contrast) may be reading name plausibility rather than entity familiarity; the paper's own Limitations (§10) concedes this.","rationale":"The reader's weakest assumption already identifies the lexical-naturalness confound in the fabricated set, and the paper's own Limitations section concedes that character n-grams, name-component frequency, and morphological likelihood were not controlled. This is the most load-bearing concern because the real-vs-fabricated contrast is the primary evidence for the familiarity construct, and the same contrast feeds both the probe and the steering direction. If the probe is exploiting surface form, the headline separation numbers and the causal steering interpretation are both in question. The paper's other limitations—LLM judges, single steering domain, same-entity steering evaluation—are real but either affect secondary claims or are explicitly scoped by the authors. The lexical confound, by contrast, reaches the central claim. Since the reader's verdict is already CONDITIONAL and explicitly requires controlling for character-level name statistics, my read does not change the verdict; it strengthens the justification for the condition.","tokens_in":20244,"tokens_out":6784,"duration_ms":72167,"concrete_test":"Construct a new fabricated set matched to real entities on character n-gram distributions, given/surname component frequencies, and morphological plausibility (or use real entities with post-training-cutoff Wikipedia pages as unknown-real controls); then (a) recompute the probe AUROC on real-vs-fabricated (Table 3a) and (b) re-derive the L30 familiarity direction and test steering on held-out entities from the matched set. If probe AUROC falls toward the 0.786 char-n-gram ceiling and steering no longer transfers to held-out matched entities, the readout is a surface-form artifact rather than entity familiarity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central construct—entity familiarity—is operationalized primarily through a real-vs-fabricated contrast. This contrast is used both as the headline discrimination target (Table 3a) and as the training/selection contrast for the probe that drives the gradation (Table 1) and the steering direction (L30: mean(known) − mean(unknown∪fabricated)). The fabricated set was matched only on token length (§4, §10), not on character n-gram statistics, name-component frequency, or morphological likelihood. Hidden states at the final prompt token can encode surface statistics, so the probe's 0.859–0.934 AUROC may be a more flexible version of the lexical-naturalness signal whose char-n-gram ceiling is 0.786. The gap is real but not decisive, and the same potential confound propagates into the gradation and causal claims: if the probe/direction is largely a 'name plausibility' axis, then the monotone popularity correlation may reflect name-form correlations with pageviews, and the L30 steering effect could be manipulating a surface-form axis rather than entity knowledge. The paper's Limitations section explicitly states this shortcut cannot be fully excluded, making it the primary unresolved threat to construct validity. If it lands, the abstract's 'separation between representational familiarity and policy' is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies activations at the final prompt token in twelve instruction-tuned models from the Bielik, PLLuM, Gemma-4, and Qwen3 families, using a new Polish-entity dataset with Wikipedia-pageview deciles and fabricated controls. It introduces a supervised logistic-regression familiarity probe and an unsupervised dispersion metric, and reports four main results: (i) probes separate real from fabricated entities in all families and show graded popularity correlation in the Polish-adapted families; (ii) probe transfer across a Polish/English stem substitution retains 96–101% of within-language AUROC in the paired setting; (iii) a rank-one familiarity direction in Gemma-4-12B steers refusal rates monotonically in both directions; and (iv) the calibrated one-pass probe is the best pre-generation gate, while post-generation detectors have better average behavioral-error prediction. The paper is unusually transparent: it reports out-of-fold probe scores, bootstrap CIs, exact permutation tests, before/after continual-pretraining controls, and an explicit Limitations section that flags the main threats to construct validity.","tokens_in":20560,"tokens_out":4973,"duration_ms":56770,"significance":"If the core claims hold, the paper provides a valuable one-forward-pass, pre-generation entity-familiarity readout for Polish and a concrete method for routing long-tail queries to retrieval. Methodologically, the paper is strong: probe scores are out-of-fold, CIs are bootstrap-based, the permutation test has exact null resolution, and the PLLuM base-vs-adapted comparisons are a useful quasi-experimental control. The paper also ships code, dataset, and provenance artifacts. However, the central construct—entity familiarity—is operationalized through a real-vs-fabricated contrast whose surface-statistics confound the paper itself concedes cannot be fully excluded, and the behavioral labels come from a single LLM judge with no human audit. These are not presentation issues; they bear directly on whether the abstract's 'separation between representational familiarity and policy' is established.","major_comments":[{"comment":"The fabricated-anchor contrast is the training/selection contrast for the probe, the dispersion metric, and the L30 steering direction, yet the fabricated names are matched only on token length. The paper concedes in §10 that 'a lexical-naturalness shortcut cannot be fully excluded.' The character-n-gram AUROC ceiling of 0.786 bounds the real-vs-fabricated classification, but it does not bound the popularity gradation (Table 1, Fig. 1) or the steering axis (Fig. 6). A probe trained on top-deciles-vs-fabricated could exploit surface cues that correlate both with top-decile real names and with fabricated status; the per-decile gradient would then be partly a name-surface gradient rather than a familiarity gradient. Please add direct controls: (a) train the probe on top-decile vs bottom-decile real entities only (no fabricated examples) and report the popularity gradient; (b) rematch fabric","section":null},{"comment":"The familiarity direction is estimated as mean(known) − mean(unknown-real ∪ fabricated) on the same 42 athletes per condition that are then steered and evaluated. The held-out check in §7 is correlational (AUROC on saved activations), not an intervention on held-out entities. Thus the dose–response could partly reflect test-set-specific direction estimation rather than a general familiarity axis. Please report steering refusal rates on held-out entities using a direction built from a disjoint training subset, across the same alpha grid, with explicit CIs. The random-direction control is coarse, as the paper notes: it shows that arbitrary directions are degenerate at high norm, not that non-familiarity semantic directions are inert. An additional control direction built from an unrelated contrast (e.g., cities vs people) at matched norm is needed to support the causal specificity claim.","section":"§7, Fig. 6/Table 2"},{"comment":"The behavioral-prediction results (Table 3b, Fig. 8, and the abstract's 'post-generation detectors better predict behavioral error on average') inherit a target definition in which the strict LLM judge scores explicit refusals as correct ~88% of the time. For Gemma-4-12B, the only abstaining model, this inflates the apparent success of post-generation detectors. The §8 answered-only reanalysis shows that once refusals are treated as abstentions, no gate separates on Gemma-4 (AURC 0.83–0.93 against a 0.90 answered base error). Since the paper itself identifies this as 'the principal evaluation revision,' the answered-only analysis should be the primary behavioral target, or the abstract and Table 3b should be explicitly qualified. As written, the headline behavioral comparison is not robust to the paper's own preferred label treatment.","section":"§8/§10"}],"minor_comments":[{"comment":"The abstract's '96–101% within-language AUROC' should be accompanied by the entity-disjoint transfer numbers (98–100% for Bielik but 74–93% for Gemma-4). The current phrasing understates the family difference in generalization to unseen entities.","section":"§6"},{"comment":"The trichotomy of 'rising/flat/falling' AUROC curves is described as coarse; consider moving it to supplementary material and keeping the primary evidence as the per-cell Spearman correlations with CIs, which are currently only in released artifacts.","section":"§5"},{"comment":"In Figure 6, the 'random: 100% degenerate' annotation is mostly relevant at the amplitudes where familiarity steering saturates; at lower amplitudes random controls are 39–100% degenerate. Please state this in the caption to avoid implying all doses are degenerate.","section":"§7"},{"comment":"The behavioral mirror uses a single LLM judge with no human audit. The second-judge agreement (κ=0.65–0.69) is substantial but not high, particularly for Gemma-4 (κ=0.469). Please state explicitly in the main text that all behavioral conclusions are conditional on this judge reliability, rather than only in §10.","section":"§5"},{"comment":"The supervision asymmetry note is good, but the 'unknown sign' of the MIND fidelity gap is easy to miss. Consider moving this to the main text of §8, since it directly affects how the probe-vs-MIND comparison in Table 3(b) should be read.","section":"§8, Table 3 footnote"}],"recommendation":"major_revision","confidential_remarks":"This is a careful, honest paper; the main reservations are already in the Limitations section. The fabricated-anchor surface confound is the most serious unresolved threat because it propagates into the gradation, steering, and gating claims. I would like to see the suggested controls (real-only probe training, n-gram-matched fabricated names, post-cutoff entities, or equivalent) before publication. The behavioral-label issue is also load-bearing and should be resolved by making the answered-only analysis primary. I do not see grounds for rejection, provided these construct-validity points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a careful, unusually self-aware empirical paper, and the headline result—a one-pass, prompt-point familiarity probe that separates real from fabricated entities and tracks popularity in Polish-adapted models—holds up better than I expected from the abstract. The paper earns its keep. The dataset construction is transparent (QID-resolved, decile-stratified, fabricated anchors). The measurements are described with out-of-fold scores, bootstrap CIs, permutation tests, and explicit notes on which baselines are imprecise reimplementations. The PLLuM base-vs-adapted comparison is the cleanest evidence that Polish continual pretraining, not scale, drives the gradation—that is a real contribution.\n\nThe soft spots are real but mostly not fatal. The primary one is the one the authors themselves flag in Section 10: the fabricated entities are matched on token length but not on character n-grams, name-component frequency, or morphological plausibility. The probe's 0.859–0.934 AUROC clears the character-n-gram ceiling of 0.786, so it is not purely lexical, but the gap is modest, and the steering direction is built from the same real-vs-fabricated contrast. If part of that signal is name plausibility rather than entity familiarity, the L30 steering result is still interesting but less about \"familiarity\" and more about a surface-form axis. The gradation result is partly insulated because the probe trains on real-vs-fabricated and then correlates with pageviews among real entities—a lexical shortcut would need to track popularity to explain that—but the insulation is not complete.\n\nSecond soft spot: the behavioral labels come from a single LLM judge with no human audit. The second judge agreement (kappa 0.65) is substantial but not decisive, and the known issue of refusals counted as correct is handled honestly but adds uncertainty.\n\nThird, the steering is one model, one domain, with the direction estimated and evaluated on the same 42 entities per condition. The held-out-entity check is correlational only. The paper says this. For a causal claim, that is narrow.\n\nThis is a solid, credible empirical study. It deserves serious peer review. I'd recommend conditional acceptance after human labels on a long-tail sample, better lexical matching of fabricated names, and at least held-out-entity steering. I would cite it; it is more careful than most work in this area.","headline":"Careful, honest empirical study of a pre-generation familiarity probe; the main results survive scrutiny, but the surface-form confound and single-judge labels need addressing.","tokens_in":21074,"tokens_out":2331,"would_cite":true,"duration_ms":24404,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Final-prompt-token activations carry a graded entity-familiarity readout, and a single familiarity direction moves Gemma-4 refusal from 0.24 to 1.00 on known entities and from 0.73 to 0.00 on unknown ones.","keywords":["entity familiarity","prompt-point probing","activation steering","refusal behavior","hallucination detection","Polish language models","popularity gradation","selective answering"],"falsifier":"Train the same probe on fabricated names matched to real names on character n-grams, name-component frequency, and morphological likelihood, and on real entities created after the model's training cutoff; if AUROC drops to the ~0.786 lexical-ceiling level in both cases, the 'familiarity' readout is mostly surface-form detection rather than exposure. For the causal claim, rebuild the steering direction from a held-out subset of entities and measure refusal on unseen entities only; if the dose–response disappears, the layer-30 effect is an in-sample artifact.","tokens_in":20085,"feed_emoji":"🧠","tokens_out":7371,"duration_ms":69556,"temperature":0.7,"pith_summary":"The paper tries to establish that a language model estimates how familiar an entity is before it starts answering: the hidden state at the final question token contains a graded familiarity signal, not just a binary known/unknown bit. Using 1,440 Polish entities stratified by Wikipedia popularity plus fabricated controls, a supervised probe on that token separates real from fabricated entities at AUROC 0.859–0.934 in all four tested model families. The signal is monotone in popularity for Polish-adapted models (Bielik and PLLuM) but nearly flat in Gemma-4 and Qwen3, and the difference tracks Polish continual pretraining more than model size. In Gemma-4-12B, adding a single familiarity direction at layer 30 moves refusal rates monotonically in both directions, showing the readout is causally potent rather than merely correlational. If true, this gives a one-forward-pass pre-generation gate for routing or abstention, while separating what the model represents from the policy that decides whether to say 'I don't know.'","feed_headline":"Probe reads entity familiarity before a model answers","feed_subtitle":"A one-forward-pass probe on the final question token separates real from fabricated entities and can steer refusal rates.","key_machinery":"Final-token hidden states at the 'prompt point' — the last question token before any generated token — measured in one forward pass. The main readout is a supervised logistic-regression probe on the residual-stream hidden state, with the layer chosen by cross-validation; unsupervised activation-dispersion metrics and first-token entropy serve as secondary signals. The dataset is 1,440 Polish entities across four domains, stratified into ten log-pageview deciles, plus 240 token-length-matched fabricated names as unfamiliar controls. The causal tool is rank-one activation steering: a difference-of-means familiarity direction at layer 30 and a refusal direction at layer 44, added to the residua","core_discovery":"The central discovery is that entity familiarity is present, graded, and locally linear in the residual stream before any answer is generated. A logistic-regression probe trained on hidden states at the final prompt token discriminates real from fabricated entities in every family (0.859–0.934 AUROC), rising with log pageviews in Polish-adapted models (mean Spearman rho 0.28–0.57) and nearly flat in Gemma-4 and Qwen3 (at most 0.11). In a paired experiment, keeping entity names unchanged and switching only the question stem from Polish to English retains 96–101% of within-language AUROC; entity-disjoint transfer is strong for Bielik (98–100%) but weaker for Gemma-4 (74–93%). A rank-one famili","pith_inferences":["A testable extension the paper leaves implicit: apply a familiarity-derived steering direction to a never-refusing Polish model (Bielik or PLLuM) to see whether abstention can be installed; if it can, the missing ingredient in those families is policy rather than representation.","If the probe's edge over the character-n-gram ceiling (0.859–0.934 vs. 0.786) survives matching fabricated names on character n-grams, name-component frequency, and morphological likelihood, the readout is a genuine exposure signal; if not, the gap shrinks toward a surface-form artifact — the paper itself flags this as needed follow-up.","The dose–response steering result raises a safety extension the author acknowledges: the same familiarity direction that induces refusal can suppress it, so any deployed abstention gate built on this signal would be vulnerable to activation-level manipulation.","The entity-disjoint transfer gap between Bielik (98–100%) and Gemma-4 (74–93%) suggests cross-lingual robustness of the familiarity readout may be tied to Polish-adapted training rather than being a general multilingual feature — testable by adding more model families and languages."],"forward_implications":["A familiarity probe can act as a pre-generation gate, routing long-tail queries to retrieval before any answer token is produced, at the cost of one forward pass.","In Polish-adapted models the readout is graded with popularity, so a calibrated gate could abstain or retrieve proportionally to long-tail risk rather than at a single threshold.","Because the layer-30 familiarity direction flips Gemma-4-12B refusal in both directions, representational familiarity and abstention policy are separable: a model can represent familiarity without acting on it, and acting on it can be steered.","The Polish/English stem-swap result implies that, in the paired setting, the readout is not primarily a Polish-surface-form artifact, narrowing where language effects could enter.","The probe beats adapted post-generation detectors on real-vs-fabricated discrimination but not on predicting behavioral error on average, so the two targets measure different things."],"fun_headline_variants":["LM probe reveals entity familiarity pre-answer","Graded familiarity readout in LM activations","Familiarity probe steers refusals in one layer","Probe tracks entity popularity in Polish LMs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that fabricated names are a faithful operational stand-in for 'entities the model has never met,' yet they are matched to real names only in token length, not in character n-gram statistics, name-component frequency, or morphological naturalness; behavioral labels also come from a single LLM judge with no human audit, so if either gives way, the probe's edge over the 0.786 lexical ceiling and the behavioral mirror both shrink.","fun_headline_variants_meta":{"raw":{"variants":["LM probe reveals entity familiarity pre-answer","Graded familiarity readout in LM activations","Familiarity probe steers refusals in one layer","Probe tracks entity popularity in Polish LMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1526,"prompt_tokens":859,"completion_tokens":667,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":615}},"tokens_in":603,"tokens_out":667,"duration_ms":8369,"temperature":1.0,"reasoning_tokens":615,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T04:46:36.641484+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same probe on fabricated names matched to real names on character n-grams, name-component frequency, and morphological likelihood, and on real entities created after the model's training cutoff; if AUROC drops to the ~0.786 lexical-ceiling level in both cases, the 'familiarity' readout is mostly surface-form detection rather than exposure. For the causal claim, rebuild the steering direction from a held-out subset of entities and measure refusal on unseen entities only; if the dose–response disappears, the layer-30 effect is an in-sample artifact.","supporting_citations":[],"review_version":1}