STARE uses step-wise RL to attack multimodal models, achieving 68% higher attack success rate while revealing that adversarial optimization concentrates conceptual toxicity early and detail toxicity late in the generation trajectory.
SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression
17 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 17roles
background 1polarities
background 1representative citing papers
EDEN adaptively sets branching factor proportional to next-token entropy, achieving better accuracy per expansion than fixed beam search while providing a proof that monotone entropy-based branching outperforms any fixed budget allocation.
EngiAgent deploys a fully connected multi-agent coordinator to achieve higher feasibility rates when using LLMs to solve open-ended engineering problems under physical and data constraints.
A bilevel textual-gradient optimizer that discovers competitive LLM workflow structures and prompts from data without human pipeline engineering.
CREATE is a benchmark that scores LLMs on their ability to produce many specific and diverse associative paths between concepts drawn from parametric knowledge.
A stateless, SVD-free regularizer approximates polar factors to induce low-rank weight structure during training, enabling better post-training compression of vision models and LLMs at under 8% overhead.
Heading-anchored steering vectors exert bidirectional causal control over tool-invocation behavior across five LLMs, but geometric analysis reveals diffuse, non-linear structure inconsistent with parametric concept encoding.
MBR decoding is reformulated via a noisy-channel decomposition into four weighted probabilistic terms, revealing that channel importance is metric-specific and task-agnostic, and that reweighting can improve performance.
Introduces C²R cross-sample consistency regularization that mitigates feature splitting and absorption in SAEs while preserving reconstruction fidelity.
A Judge-Aware Gated Multi-Task Learning architecture with outcome taxonomy supervision achieves SOTA accuracy on 13,937 UK Employment Tribunal decisions using an order of magnitude fewer parameters than generative SFT baselines on a 26B model.
MÖVE presents a new German-language benchmark evaluating 39 LLMs on performance and governance criteria using ten public-administration datasets.
Divergence Decoding steers LLM logits using small auxiliary models to unlearn specific data at inference time, outperforming baselines and generalizing to images.
LFQ improves generation accuracy of low-bit quantized LLMs by applying cross-entropy logit alignment specifically to the final block, addressing misalignment from omitted unembedding layers and MSE objectives in standard block-wise PTQ.
BITE applies LinUCB bandits to discover style edits that raise LLM judge scores by 1-2 points on a 9-point scale with >65% success while preserving meaning.
CL-bench Life shows frontier language models achieve only 13.8% average success on real-life context tasks, with the best model at 19.3%.
BARRED uses dimension decomposition and asymmetric multi-agent debate to generate high-fidelity synthetic data that lets small fine-tuned models outperform proprietary LLMs and existing guardrail models on custom policies.
DOVE measures LLM cultural value alignment via a rate-distortion value codebook and unbalanced optimal transport between human and model open-ended text distributions.
citing papers explorer
-
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
STARE uses step-wise RL to attack multimodal models, achieving 68% higher attack success rate while revealing that adversarial optimization concentrates conceptual toxicity early and detail toxicity late in the generation trajectory.
-
Entropy-informed Decoding: Adaptive Information-Driven Branching
EDEN adaptively sets branching factor proportional to next-token entropy, achieving better accuracy per expansion than fixed beam search while providing a proof that monotone entropy-based branching outperforms any fixed budget allocation.
-
EngiAgent: Fully Connected Coordination of LLM Agents for Solving Open-ended Engineering Problems with Feasible Solutions
EngiAgent deploys a fully connected multi-agent coordinator to achieve higher feasibility rates when using LLMs to solve open-ended engineering problems under physical and data constraints.
-
FlowBot: Inducing LLM Workflows with Bilevel Optimization and Textual Gradients
A bilevel textual-gradient optimizer that discovers competitive LLM workflow structures and prompts from data without human pipeline engineering.
-
CREATE: Testing LLMs for Associative Creativity
CREATE is a benchmark that scores LLMs on their ability to produce many specific and diverse associative paths between concepts drawn from parametric knowledge.
-
SLORR: Simple and Efficient In-Training Low-Rank Regularization
A stateless, SVD-free regularizer approximates polar factors to induce low-rank weight structure during training, enabling better post-training compression of vision models and LLMs at under 8% overhead.
-
Controlling Tool Use with Heading-Specific Activation Steering
Heading-anchored steering vectors exert bidirectional causal control over tool-invocation behavior across five LLMs, but geometric analysis reveals diffuse, non-linear structure inconsistent with parametric concept encoding.
-
Noisy-Channel Minimum Bayes Risk Decoding
MBR decoding is reformulated via a noisy-channel decomposition into four weighted probabilistic terms, revealing that channel importance is metric-specific and task-agnostic, and that reweighting can improve performance.
-
C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders
Introduces C²R cross-sample consistency regularization that mitigates feature splitting and absorption in SAEs while preserving reconstruction fidelity.
-
Towards Explainable Adjudicative Variance: Quantifying Judicial Discretion via Gated Multi-Task Learning
A Judge-Aware Gated Multi-Task Learning architecture with outcome taxonomy supervision achieves SOTA accuracy on 13,937 UK Employment Tribunal decisions using an order of magnitude fewer parameters than generative SFT baselines on a 26B model.
-
M\"OVE: A Holistic LLM Benchmark for the German Public Sector
MÖVE presents a new German-language benchmark evaluating 39 LLMs on performance and governance criteria using ten public-administration datasets.
-
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
Divergence Decoding steers LLM logits using small auxiliary models to unlearn specific data at inference time, outperforming baselines and generalizing to images.
-
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
LFQ improves generation accuracy of low-bit quantized LLMs by applying cross-entropy logit alignment specifically to the final block, addressing misalignment from omitted unembedding layers and MSE objectives in standard block-wise PTQ.
-
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
BITE applies LinUCB bandits to discover style edits that raise LLM judge scores by 1-2 points on a 9-point scale with >65% success while preserving meaning.
-
CL-bench Life: Can Language Models Learn from Real-Life Context?
CL-bench Life shows frontier language models achieve only 13.8% average success on real-life context tasks, with the best model at 19.3%.
-
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
BARRED uses dimension decomposition and asymmetric multi-agent debate to generate high-fidelity synthetic data that lets small fine-tuned models outperform proprietary LLMs and existing guardrail models on custom policies.
-
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
DOVE measures LLM cultural value alignment via a rate-distortion value codebook and unbalanced optimal transport between human and model open-ended text distributions.