Pith. sign in

REVIEW 3 major objections 5 minor 15 references

LLMs already know the car-wash trap; salient numbers just crowd that knowledge out.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 06:06 UTC pith:NBQJRXEW

load-bearing objection Solid multi-model trap benchmark and density/SI results; the “suppression not absence” headline only holds on SC cases that already named the trap, not on Hard Fail. the 3 major comments →

arxiv 2607.28478 v1 pith:NBQJRXEW submitted 2026-07-30 cs.CL

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

classification cs.CL
keywords salience biaslarge language modelscommonsense reasoningknowledge suppressionsycophancySaliTrap benchmarkelicitationinference-time prompting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models trained to exploit every explicit condition in a problem can be hijacked by useless numbers and details, ignoring the physical or everyday prerequisites that make a task possible. The paper names this Salience Bias and builds SaliTrap, a benchmark of 1,145 everyday queries that hide impossible premises behind planning and calculation bait across four trap types. All twelve models tested fail often; the best still avoids the trap on only about half the items, and noticing the trap often does not stop compliance. When the same models are asked the bare physical fact without the task framing, over 90 percent of the sycophantic failures reverse, showing the knowledge was present but suppressed. Simple inference-time prompts that force a feasibility check recover much of the lost performance without retraining, relocating the bottleneck from missing knowledge to how the knowledge is elicited.

Core claim

Salience bias in LLM commonsense reasoning is overwhelmingly knowledge suppression, not knowledge absence: models intrinsically possess the physical prerequisites needed to reject impossible tasks, but salient distractors lure them into over-compliant computation. A context-free probe of the trap core alone liberates over 90 percent of sycophantic-compliance failures, and lightweight system prompts substantially raise trap-avoidance rates without any retraining.

What carries the argument

SaliTrap: a four-dimension benchmark of well-formed items that embed a physically impossible trap core inside computation-laden natural-language queries, scored by behavioral labels that separate Hard Fail, CoT Hijacked, Sycophantic Compliance, and Strict Pass, plus a liberation protocol that re-elicits the same knowledge with task framing stripped away.

Load-bearing premise

The automated judge labels that decide whether a model truly noticed the trap, and the synthetic items certified because strong solvers fail them, are assumed to measure real commons-sense suppression rather than judge quirks or construction artifacts.

What would settle it

Re-run the liberation experiment on a large held-out set of human-authored (not solver-certified) traps with an independent human-scored judge: if context-free probes no longer recover most sycophantic failures, or if human raters systematically disagree with the Strict-Pass versus Sycophantic-Compliance labels, the suppression-not-absence claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Commonsense failures on everyday planning queries are often fixable at inference time by forcing a premise check before computation.
  • Trap detection and trap avoidance are separate axes; awareness metrics alone will overstate reliability.
  • Adding more numerical distractors predictably worsens compliance, so dense planning prompts are especially risky.
  • SaliTrap becomes a reusable testbed for measuring whether future models or prompts close the elicitation gap.
  • Training that always rewards using every given condition may systematically train the bias the paper diagnoses.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If suppression dominates, alignment and safety pipelines that only check stored knowledge will miss a large class of user-facing planning errors.
  • The same density effect may appear in tool-use and agent settings where APIs return many salient numbers that are irrelevant to physical feasibility.
  • Capability-aware prompting may be needed: the same premise-check prefix that lifts weak models can interrupt already-strong default reasoning.
  • Co-failure clustering by model family suggests shared training recipes create shared blind spots, not only shared ability ceilings.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces salience bias: LLMs over-weight explicit numerical/procedural distractors and overlook implicit physical or commonsense prerequisites. It contributes SaliTrap (1,145 items, four trap dimensions), evaluates 12 LLMs with TAR/HFR/SCR/SI, and reports pervasive vulnerability that scales with distractor density and decouples trap detection from avoidance (high SI). A liberation experiment on sycophantic-compliance (SC) cases claims the failure is knowledge suppression rather than absence (Cond-C recovers >90% of SC), and system-level prompts raise TAR without retraining. The authors relocate the bottleneck from competence to elicitation and release code and data.

Significance. If the scoped empirical picture holds, the work is a useful diagnostic contribution: it names a concrete failure mode at the intersection of sycophancy and distractor robustness, supplies a multi-dimensional synthetic testbed, and shows that detection and compliance are separable axes (SI). The density curves, IRT difficulty ordering, co-failure clustering by provenance, and prompt-intervention gains are actionable for evaluation and deployment. The strongest interpretive claim—that failures are overwhelmingly suppression not absence, so the bottleneck is elicitation—would be high-impact if properly evidenced across failure modes; as written it is only partially supported. Code and dataset release are clear strengths for replication.

major comments (3)
  1. [Abstract; Liberation experiment; Tables 1–2] Abstract, Introduction, and §“Is Sycophancy Compliance Knowledge Suppression or Knowledge Absence?”: the central claim that failures are “overwhelmingly… knowledge suppression rather than knowledge absence” and that the bottleneck moves “from model competence to elicitation” is evidenced only on SC cases (models that already verbalized the trap). Cond-C therefore mainly shows framing-sensitive conversion of acknowledged-but-compliant behavior into refusal. Hard Fail (HFR often 30–60% in Table 1, frequently larger than SCR in Table 2) is the cell where absence vs suppression is open, and the paper never re-elicits Hard Fail under Cond-A/B/C. Either run the liberation protocol on Hard Fail (and CoT-Hijacked) for the same models, or narrow abstract/intro/conclusion wording to SC-scoped sycophancy and stop generalizing to all salience-bias failures.
  2. [SaliTrap Benchmark; Stage 2 Candidate Validation; Stage 3] Benchmark Construction, Stage 2–3: item certification routes on solver–judge failure labels (L+/L++) and composite score rewards failure severity, then retains top-k per seed. The evaluation set is therefore partly defined by the same class of model failures it later measures. This does not invalidate prevalence or density trends, but it weakens claims that SaliTrap is an external gold probe of “genuine” suppression. Report (i) how many seeds/candidates were discarded for Strict Pass vs certified for failure, (ii) a hold-out or human-only well-formedness sample scored without failure-severity selection, and (iii) sensitivity of TAR rankings when restricting to soft-attack or high-naturalness strata.
  3. [Experimental Setup; Solver-Judge; Table 2; Liberation] Experimental Setup / Solver-Judge: MJ and much of construction use Claude-Opus-4.7, which is also an evaluated target and the strongest TAR model (Table 1). Labels (Strict Pass vs SC vs Hard Fail) and liberation denominators inherit this judge. Provide inter-judge agreement with at least one independent judge family on a substantial labeled subset, and show that main rankings and liberation rates are stable under re-judgment. Without that, detection–avoidance decoupling (SI) and SC counts remain partly judge-idiosyncratic.
minor comments (5)
  1. [Abstract; Figures 6–7] Figure 2 / car-wash example is clear; ensure every main claim in the abstract maps to a numbered table/figure (liberation >90% → Figure 7 / Table 6; density → Figure 6).
  2. [Prompt Intervention; Table 7] Table 7 shows interventions can lower TAR for Claude-4.6 relative to Control; discuss capability-aware prompting more prominently in the main text, not only the appendix.
  3. [Metrics] Define SI formula once in the main metrics paragraph with the exact denominator (SC ∪ CoT Hijacked) to avoid confusion with SCR normalized by N.
  4. [Related Work] Related Work could more sharply separate premise-sycophancy benchmarks (Sharma et al., Perez et al.) from irrelevant-context distraction (Shi et al.) and state what physical-impossibility + numeric camouflage uniquely adds.
  5. [Throughout] Minor polish: spacing typos (“everydaycommonsensereasoning”, “knowledgesuppression”) and inconsistent model name hyphenation across tables.

Circularity Check

2 steps flagged

Central 'suppression not absence' slogan is partly tautological on SC (awareness already definitional); adversarial certification selects on solver failure, mildly inflating the 'pervasive failure' claim.

specific steps
  1. self definitional [Abstract; §Is Sycophancy Compliance Knowledge Suppression or Knowledge Absence?; Fig. 7; Metrics (SCR/SI)]
    "a context-free knowledge probe alone recovers over 90% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors... we show that this is overwhelmingly a failure of knowledge suppression rather than knowledge absence... SI = SC/(SC + CoT). This conditioning matters: normalizing by N instead would conflate knowledge absence with sycophantic suppression... we re-query every SC instance... liberation rate as the fraction of former SC cases converted to strict pass"

    SC is defined as cases where the model already registered the trap yet complied. Restricting Cond-A/B/C to SC therefore makes 'knowledge is present' true by the label before any probe runs; Cond-C mainly converts acknowledged-but-compliant behavior into refusal once framing is stripped. The paper's stronger slogan—that salience-bias failures are overwhelmingly suppression not absence, and that the bottleneck moves from competence to elicitation—treats this SC-scoped tautology as if it diagnosed Hard Fail (no awareness) as well, which was never re-elicited.

  2. fitted input called prediction [§Benchmark Construction Stage 2 Routing / Stage 3; composite score(c); Experimental Setup (solver pool)]
    "Candidates exhibiting a failure label (L+) that also clear a high naturalness threshold are certified for the final dataset... we score every certified candidate c by a composite score(c) rewarding failure severity w_ℓ(c), naturalness ν(c), and confirmed alignment... and retain the top-k candidates (k=5) per seed as the final benchmark items... the solver M_S is a round-robin pool of four strong reasoning models: Claude-Opus-4.7, GPT-5.5, DeepSeek-R1, and Gemini-2.5-Pro."

    Item inclusion and ranking are partly defined by inducing L+/failure severity on a solver pool that overlaps the evaluated targets. Reporting that those models (and nearby models) show low TAR and high HFR on the resulting set is therefore partly by construction—an adversarially fitted difficulty distribution presented as an independent measurement of pervasive salience bias—rather than evaluation on a failure-agnostic external gold set. Independent content remains in held-out models, density slopes, and intervention deltas.

full rationale

This is an empirical LLM-benchmark paper, not a first-principles derivation, so classical fit-equals-prediction circularity is limited. Two load-bearing steps still partially reduce to their inputs. (1) Sycophantic Compliance is defined as trap acknowledged yet task completed; running Cond-C only on that subset and treating high liberation as proof that 'the requisite commonsense is intrinsically present' restates the subset definition for the knowledge-presence half of the claim—the experiment mainly shows framing-sensitive refusal, not discovery of latent knowledge that had appeared missing. Hard Fail (often the larger failure mass) is never re-elicited, yet Abstract/Conclusion generalize to all salience-bias failures and relocate the bottleneck from competence to elicitation. (2) Stage-2 certification and top-k scoring explicitly reward solver-judge failure severity on a pool that overlaps evaluated models, so elevated HFR/low TAR is partly selected for rather than measured on an external gold set. Neither step fully collapses the paper: prompt interventions, distractor-density trends, SI decoupling, and cross-model patterns retain independent content. Proportionate score is therefore moderate (4), not 6+.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 4 invented entities

The central suppression claim rests on operational definitions of traps and judge labels, pipeline thresholds chosen by authors, and the assumption that stripping task framing validly re-elicits the same knowledge the original query should have used. No physical constants are fitted; free parameters are methodological cutoffs. Invented entities are named constructs and metrics, not ontological posits in nature.

free parameters (4)
  • naturalness_pass_threshold = ≥3.5
    Average of five 1–5 naturalness sub-scores must be ≥3.5 for pass; directly gates which candidates enter solver testing or rewrite queues.
  • top_k_items_per_seed = 5
    Composite score retains top-5 certified candidates per seed; shapes final 1,145-item composition.
  • stability_retest_repeats = 5
    Majority label over 5 solver–judge repeats required for certification; filters one-off noise but is an author-chosen reliability bar.
  • Jaccard_dedup_and_entity_triple_filters
    Character-level Jaccard and (tool,object,action) triple filters decide seed acceptance; thresholds are pipeline design choices affecting diversity and difficulty.
axioms (5)
  • ad hoc to paper A response is correct on a well-formed item only if it identifies the impossible trap core and produces no executable plan or numerical computation predicated on the trap being valid.
    Task Definition well-formedness conditions (i)–(iii) and correct-response criterion; defines TAR and related metrics.
  • domain assumption Solver-Judge six-way behavioral labels (Hard Fail, CoT Hijacked, Sycophantic Compliance, Strict Pass, Patch Compliance, Mechanical Refusal) are a faithful taxonomy of trap awareness and compliance.
    Stage-2 Solver-Judge and all main metrics; no human gold labels reported.
  • domain assumption Context-free affirmation of trap_core feasibility measures the same commons-sense knowledge that should have blocked compliance in the full framed query.
    Liberation Cond-C design; load-bearing for suppression-vs-absence claim.
  • domain assumption Physically impossible premises under everyday commons sense are objectively identifiable for the four taxonomy dimensions.
    Trap Taxonomy D1–D4 and Truth/Alignment checkers; standard commons-sense domain assumption.
  • domain assumption Zero-shot API default (or greedy) decoding without task-specific fine-tuning is a fair comparison across the 12 models.
    Experimental Setup evaluation protocol.
invented entities (4)
  • Salience Bias (as formalized here) independent evidence
    purpose: Name the failure mode where explicit distractors suppress implicit physical/commons-sense prerequisites.
    Core contribution label; builds on prior distraction/sycophancy work but is operationalized specifically for this setting.
  • SaliTrap Benchmark no independent evidence
    purpose: Provide 1,145 certified items across four trap dimensions to measure the bias and separate detection from avoidance.
    Synthetic dataset constructed via LLM-assisted pipeline; evidence is internal evaluation plus promised release.
  • Sycophancy Index (SI) no independent evidence
    purpose: Conditional compliance rate given trap awareness, avoiding conflation with hard non-detection.
    Derived metric SI = SC/(SC+CoT); useful but definition-internal.
  • Liberation Rate no independent evidence
    purpose: Fraction of former SC cases converted to Strict Pass under debiasing prompts Cond-A/B/C.
    Primary evidence instrument for suppression vs absence.

pith-pipeline@v1.2.0-daily-grok45 · 23895 in / 3953 out tokens · 94377 ms · 2026-07-31T06:06:45.078968+00:00 · methodology

0 comments
read the original abstract

As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90\% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at https://github.com/Wuzheng02/SaliTrap.

Figures

Figures reproduced from arXiv: 2607.28478 by Cheng Yang, Chenhao Xue, Shijie Zheng, Yijie Lu, Zheng Wu, Zhuosheng Zhang.

Figure 1
Figure 1. Figure 1: All LLMs suffer from salience bias, which stems [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Exemplifying salience bias in LLMs. Driven by [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The construction of the SaliTrap benchmark is divided into three stages: (i) Seed generation and scaling stage, (ii) [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: IRT-estimated item difficulty β distribution across the four trap dimensions (12 evaluated models). Missing prerequisite and environmental mismatch skew toward higher difficulty, while rule mismatch items are concentrated at lower difficulty [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: TAR and CoT-Hijacked rate versus the number of [PITH_FULL_IMAGE:figures/full_fig_p006_6.png] view at source ↗
Figure 8
Figure 8. Figure 8: TAR under Control and three system-level prompt [PITH_FULL_IMAGE:figures/full_fig_p007_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Item naturalness score vs. IRT-estimated difficulty [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 12
Figure 12. Figure 12: TAR and CoT-Hijacked rate versus number of in [PITH_FULL_IMAGE:figures/full_fig_p014_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

15 extracted references · 11 linked inside Pith

  1. [5]

    Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948. Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J

  2. [6]

    Jimenez, C

    Measuring mathematicalproblemsolvingwiththeMATHdataset.arXiv preprint arXiv:2103.03874. Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K

  3. [7]

    Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.;Tran-Johnson,E.;etal.2022

    Swe-bench: Can lan- guage models resolve real-world github issues? InInter- national Conference on Learning Representations, volume 2024, 54107–54157. Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.;Tran-Johnson,E.;etal.2022. Languagemodels(mostly) know what they know.arXiv preprint...

  4. [9]

    Liu,P.;Yuan,W.;Fu,J.;Jiang,Z.;Hayashi,H.;andNeubig, G.2023

    Lost in the middle: Howlanguagemodelsuselongcontexts.Transactionsofthe Association for Computational Linguistics, 12: 157–173. Liu,P.;Yuan,W.;Fu,J.;Jiang,Z.;Hayashi,H.;andNeubig, G.2023. Pre-train,prompt,andpredict:Asystematicsurvey of prompting methods in natural language processing.ACM computing surveys, 55(9): 1–35. Mirzadeh,I.;Alizadeh,K.;Shahrokhi,H....

  5. [11]

    Sharma, M.; Tong, M.; Korbak, T.; Duvenaud, D.; Askell, A.;Bowman,S.R.;Cheng,N.;Durmus,E.;Hatfield-Dodds, Z.;Johnston,S.R.;etal.2024

    Deepseekmath: Pushing the limits of mathematical reasoning in open lan- guage models.arXiv preprint arXiv:2402.03300. Sharma, M.; Tong, M.; Korbak, T.; Duvenaud, D.; Askell, A.;Bowman,S.R.;Cheng,N.;Durmus,E.;Hatfield-Dodds, Z.;Johnston,S.R.;etal.2024. Towardsunderstandingsyco- phancy in language models.International Conference on Learning Representations....

  6. [12]

    Transactions on Machine Learning Research

    Beyond the imitation game: Quanti- fying and extrapolating the capabilities of language models. Transactions on Machine Learning Research. Tang, F.; Xu, H.; Zhang, H.; Chen, S.; Wu, X.; Shen, Y.; Zhang,W.;Hou,G.;Tan,Z.;Yan,Y.;etal.2025. Asurveyon (m)llm-basedguiagents.arXivpreprintarXiv:2504.13865. Team, K.; Bai, Y.; Bao, Y.; Charles, Y.; Chen, C.; Chen, ...

  7. [13]

    On the robustnessofchatgpt:Anadversarialandout-of-distribution perspective.arXiv preprint arXiv:2302.12095. Wei,J.;Wang,X.;Schuurmans,D.;Bosma,M.;Xia,F.;Chi, E.;Le,Q.V.;Zhou,D.;etal.2022.Chain-of-thoughtprompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35: 24824–24837. Xu, A.; Lin, B.; Xue, B.; Wang,...

  8. [14]

    arXiv preprint arXiv:2606.19348

    Deepseek-v4: Towards highly efficient million-token context intelligence. arXiv preprint arXiv:2606.19348. Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K

  9. [15]

    assume this is possible

    Glm-5: from vibe coding to agentic engineering.arXiv preprint arXiv:2602.15763. Zhang,K.;Li,J.;Li,G.;Shi,X.;andJin,Z.2024.Codeagent: Enhancing code generation with tool-integrated agent sys- tems for real-world repo-level coding challenges. InPro- ceedings of the 62nd Annual Meeting of the Association for ComputationalLinguistics(Volume1:LongPapers),13643...

  10. [2021]

    arXiv preprint arXiv:2110.14168

    Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al

  11. [2022]

    Advances in Neural Information Processing Systems, 35: 22199–22213

    Large language models are zero-shot reasoners. Advances in Neural Information Processing Systems, 35: 22199–22213. Li,M.;Wang,W.;Feng,F.;Zhu,F.;Wang,Q.;andChua,T.- S.2024. Thinktwicebeforetrusting:Self-detectionforlarge language models through comprehensive answer reflection. Findings of the Association for Computational Linguistics: EMNLP. Liu, N. F.; Li...

  12. [2023]

    Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong,X.;Tang,X.; Qian,B.;etal.2024

    Discovering language model behaviors with model-written evaluations.Findings of the Association for Computational Linguistics: ACL. Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong,X.;Tang,X.; Qian,B.;etal.2024. Toolllm:Facilitat- inglargelanguagemodelstomaster16000+real-worldapis. InInternational Conference on Learning Representations,...

  13. [2024]

    Anthropic

    On- policy distillation of language models: Learning from self- generatedmistakes.InInternationalConferenceonLearning Representations, volume 2024, 21246–21263. Anthropic. 2026a. Introducing Claude Opus 4.6. https:// www.anthropic.com/news/claude-opus-4-6. Anthropic. 2026b. Introducing Claude Opus 4.7. https: //www.anthropic.com/news/claude-opus-4-7. Best...

  14. [2025]

    Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al

    Gemini 2.5: Pushing the fron- tier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261. Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al

  15. [2026]

    Chu, T.; Zhai, Y.; Yang, J.; Tong, S.; Xie, S.; Schuurmans, D.; Le, Q

    The minimax-m2 series: Mini activations unleashing max real- world intelligence.arXiv preprint arXiv:2605.26494. Chu, T.; Zhai, Y.; Yang, J.; Tong, S.; Xie, S.; Schuurmans, D.; Le, Q. V.; Levine, S.; and Ma, Y