LaneRoPE adds an inter-sequence attention mask and extended RoPE to enable collaborative parallel sequence generation in LLMs, yielding accuracy gains on math reasoning under length limits.
Mixed citations
Metarag: Metamorphic testing for hallucination detection in rag systems
Mixed citation behavior. Most common role is background (50%).
citation-role summary
citation-polarity summary
years
2026 16verdicts
UNVERDICTED 16representative citing papers
Reversible-jump MCMC analysis of LIGO binary black hole mergers identifies three subpopulations with distinct properties and independent redshift evolution.
LGMT is a logic-grounded metamorphic testing framework that detects hidden reasoning defects in LLMs by checking consistency on semantically invariant inputs derived from FOL equivalences.
Statistical model checking on the K+S model shows macro-financial and structural parameters produce stronger transient effects on unemployment and GDP growth than heuristic-rule parameters under fixed precision policies.
uxCUA is a trained computer use agent that assesses GUI usability more accurately than larger models by learning to prioritize and execute important user interactions on labeled interface datasets.
XR Blocks supplies an LLM-optimized Reality Model and Vibe Coding XR workflow that converts high-level prompts into working physics-aware XR applications with high one-shot success.
Agentic iteration improves perceived quality of generated multiview genomics visualizations over direct LLM generation, but adding more specialist agents or a reviewer yields no further gains across 159 test cases.
KAPLAN-HR applies B-spline KANs to nonparametric hazard estimation in survival analysis, recovering GAMs in the single-layer case, capturing interactions via deeper layers, with convergence rates independent of covariate dimension for KAN-representable targets, and competitive performance on six cli
SimpleTES scales test-time evaluation in LLMs to discover state-of-the-art solutions on 21 scientific problems across six domains, outperforming frontier models and optimization pipelines with examples like 2x faster LASSO and new Erdos constructions.
An encoding of Solidity contracts and first-order Hennessy-Milner logic into Lustre enables Kind 2 model checking of complex temporal properties in smart contracts.
RLVR training raises verified Dafny pass rates from 9.7% to 31.1% on a filtered benchmark while a Lean proof scaffold lifts success from 46.2% to 69.2% on a pilot set and solves 7 of 42 prior unsolved tasks.
Introduces a taxonomy of nine LLM code smells, a static detection tool, and reports 73.5% prevalence with 91.3% precision and 71.8% recall across 692 projects.
The paper claims that alignment requires treating AI as part of the self through cognitive co-regulation, identifying risks like deskilling and automation bias while drawing on System 0 cognition theory.
Literature on system prompts for AI shows fragmented and contradictory claims that complicate policy efforts to use them as reliable governance mechanisms.
Verbalized confidence from small LMs enables cost-effective cascade routing for automated educational scoring, matching large-model accuracy at 76% lower cost when discrimination is strong.
Reviews coherent and incoherent radio emission in ordered stellar magnetospheres, links massive-star CBO mechanism to UCD luminosity trends, and predicts SKA detection of ~1000 UCDs to test the hypothesis.
citing papers explorer
-
LaneRoPE: Positional Encoding for Collaborative Parallel Reasoning and Generation
LaneRoPE adds an inter-sequence attention mask and extended RoPE to enable collaborative parallel sequence generation in LLMs, yielding accuracy gains on math reasoning under length limits.
-
Reversible-jump MCMC reveals binary black hole subpopulations with distinct redshift evolution
Reversible-jump MCMC analysis of LIGO binary black hole mergers identifies three subpopulations with distinct properties and independent redshift evolution.
-
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
LGMT is a logic-grounded metamorphic testing framework that detects hidden reasoning defects in LLMs by checking consistency on semantically invariant inputs derived from FOL equivalences.
-
Statistical Model Checking of the Keynes+Schumpeter Model: A Transient Sensitivity Analysis of a Macroeconomic ABM
Statistical model checking on the K+S model shows macro-financial and structural parameters produce stronger transient effects on unemployment and GDP growth than heuristic-rule parameters under fixed precision policies.
-
Training Computer Use Agents to Assess the Usability of Graphical User Interfaces
uxCUA is a trained computer use agent that assesses GUI usability more accurately than larger models by learning to prioritize and execute important user interactions on labeled interface datasets.
-
Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini
XR Blocks supplies an LLM-optimized Reality Model and Vibe Coding XR workflow that converts high-level prompts into working physics-aware XR applications with high one-shot success.
-
Agentic Authoring of Interactive Multiview Visualizations in Genomics
Agentic iteration improves perceived quality of generated multiview genomics visualizations over direct LLM generation, but adding more specialist agents or a reviewer yields no further gains across 159 test cases.
-
KAPLAN: Kolmogorov-Arnold Prognostic Learnable Activation Networks for Survival Analysis
KAPLAN-HR applies B-spline KANs to nonparametric hazard estimation in survival analysis, recovering GAMs in the single-layer case, capturing interactions via deeper layers, with convergence rates independent of covariate dimension for KAN-representable targets, and competitive performance on six cli
-
Evaluation-driven Scaling for Scientific Discovery
SimpleTES scales test-time evaluation in LLMs to discover state-of-the-art solutions on 21 scientific problems across six domains, outperforming frontier models and optimization pipelines with examples like 2x faster LASSO and new Erdos constructions.
-
KindHML: formal verification of smart contracts based on Hennessy-Milner logic
An encoding of Solidity contracts and first-order Hennessy-Milner logic into Lustre enables Kind 2 model checking of complex temporal properties in smart contracts.
-
Automating Formal Verification with Reinforcement Learning and Recursive Inference
RLVR training raises verified Dafny pass rates from 9.7% to 31.1% on a filtered benchmark while a Lean proof scaffold lifts success from 46.2% to 69.2% on a pilot set and solves 7 of 42 prior unsolved tasks.
-
LLM Code Smells: A Taxonomy and Detection Approach
Introduces a taxonomy of nine LLM code smells, a static detection tool, and reports 73.5% prevalence with 91.3% precision and 71.8% recall across 692 projects.
-
Position: AI as Part of Self -- Extending the Mind Requires Cognitive Co-Regulation
The paper claims that alignment requires treating AI as part of the self through cognitive co-regulation, identifying risks like deskilling and automation bias while drawing on System 0 cognition theory.
-
Prompt Governance? On Governing Technologies Governed by Natural Language
Literature on system prompts for AI shows fragmented and contradictory claims that complicate policy efforts to use them as reliable governance mechanisms.
-
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
Verbalized confidence from small LMs enables cost-effective cascade routing for automated educational scoring, matching large-model accuracy at 76% lower cost when discrimination is strong.
-
Coherent and Incoherent Emission from the Ordered Magnetospheres of Low-Mass Stars, UCDs, and Massive Stars
Reviews coherent and incoherent radio emission in ordered stellar magnetospheres, links massive-star CBO mechanism to UCD luminosity trends, and predicts SKA detection of ~1000 UCDs to test the hypothesis.