REVIEW 6 cited by
Towards Resilient and Efficient LLMs: A Comparative Study of Efficiency, Performance, and Adversarial Robustness
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the increasing demand for practical applications of Large Language Models (LLMs), many attention-efficient models have been developed to balance performance and computational cost. However, the adversarial robustness of these models remains under-explored. In this work, we design a framework to investigate the trade-off between efficiency, performance, and adversarial robustness of LLMs and conduct extensive experiments on three prominent models with varying levels of complexity and efficiency -- Transformer++, Gated Linear Attention (GLA) Transformer, and MatMul-Free LM -- utilizing the GLUE and AdvGLUE datasets. The AdvGLUE dataset extends the GLUE dataset with adversarial samples designed to challenge model robustness. Our results show that while the GLA Transformer and MatMul-Free LM achieve slightly lower accuracy on GLUE tasks, they demonstrate higher efficiency and either superior or comparative robustness on AdvGLUE tasks compared to Transformer++ across different attack levels. These findings highlight the potential of simplified architectures to achieve a compelling balance between efficiency, performance, and adversarial robustness, offering valuable insights for applications where resource constraints and resilience to adversarial attacks are critical.
Forward citations
Cited by 6 Pith papers
-
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
A new seven-language ophthalmology benchmark shows LLMs are less accurate in LMIC languages, and an agentic translation-plus-RAG pipeline reduces the gap.
-
Can Large Language Models Effectively Process and Execute Financial Trading Instructions?
On a 500-item trading instruction dataset, five LLMs produced well-formatted JSON most of the time but achieved only 5-10% full accuracy and frequently asked unnecessary follow-up questions.
-
LoRA-LiteE: A Computationally Efficient Framework for Chatbot Preference-Tuning
An ensemble of two LoRA-finetuned 8-9B models gets 80.2% accuracy on Chatbot Arena preference prediction, slightly above GPT-4's 78.3%, but with no error bars or code.
-
An Improved Dung Beetle Optimizer for Random Forest Optimization
Adding circle mapping and crossover to the Dung Beetle Optimizer yields faster convergence and better accuracy on selected benchmark functions, and improved random forest hyperparameters on a retail dataset.
-
Enhanced Recommendation Combining Collaborative Filtering and Large Language Models
A simple weighted sum of collaborative filtering scores and LLM text embeddings is claimed to improve recommendation accuracy, but the reported experiments are not reproducible.
-
A Survey: Towards Privacy and Security in Mobile Large Language Models
A survey of privacy and security challenges for mobile large language models, summarizing known attack types and defenses without introducing new results.
Discussion (0). Continue with ORCID to comment.