Delta Attention Residuals attend over per-sublayer deltas instead of cumulative hidden states, producing higher-contrast attention weights and 1.7-8.2% validation perplexity gains over standard and attention residuals across 220M-7.6B models.
Title resolution pending
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
SocialIQA is the first large-scale benchmark with 38k crowdsourced questions testing commonsense about social interactions, where pretrained language models trail humans by over 20% but transfer to improve performance on Winograd Schemas and COPA.
River-LLM enables token-level early exit in decoder-only LLMs by routing exited tokens through 4-bit quantized copies of backbone layers that share the KV cache addressing scheme, achieving 1.53–2.16× wall-clock speedup without training.
Translation function vectors extracted from one language direction transfer to unseen target languages, indicating a language-agnostic translation signal in multilingual LLMs.
citing papers explorer
-
Delta Attention Residuals
Delta Attention Residuals attend over per-sublayer deltas instead of cumulative hidden states, producing higher-contrast attention weights and 1.7-8.2% validation perplexity gains over standard and attention residuals across 220M-7.6B models.
-
SocialIQA: Commonsense Reasoning about Social Interactions
SocialIQA is the first large-scale benchmark with 38k crowdsourced questions testing commonsense about social interactions, where pretrained language models trail humans by over 20% but transfer to improve performance on Winograd Schemas and COPA.
-
River-LLM: Large Language Model Seamless Exit Based on KV Share
River-LLM enables token-level early exit in decoder-only LLMs by routing exited tokens through 4-bit quantized copies of backbone layers that share the KV cache addressing scheme, achieving 1.53–2.16× wall-clock speedup without training.
-
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
Translation function vectors extracted from one language direction transfer to unseen target languages, indicating a language-agnostic translation signal in multilingual LLMs.