Pith. sign in

REVIEW 6 cited by

Towards Resilient and Efficient LLMs: A Comparative Study of Efficiency, Performance, and Adversarial Robustness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04585 v3 pith:GSPA4J3V submitted 2024-08-08 cs.CL

classification cs.CL
keywords adversarialrobustnessefficiencymodelsperformancetransformeradvglueglue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the increasing demand for practical applications of Large Language Models (LLMs), many attention-efficient models have been developed to balance performance and computational cost. However, the adversarial robustness of these models remains under-explored. In this work, we design a framework to investigate the trade-off between efficiency, performance, and adversarial robustness of LLMs and conduct extensive experiments on three prominent models with varying levels of complexity and efficiency -- Transformer++, Gated Linear Attention (GLA) Transformer, and MatMul-Free LM -- utilizing the GLUE and AdvGLUE datasets. The AdvGLUE dataset extends the GLUE dataset with adversarial samples designed to challenge model robustness. Our results show that while the GLA Transformer and MatMul-Free LM achieve slightly lower accuracy on GLUE tasks, they demonstrate higher efficiency and either superior or comparative robustness on AdvGLUE tasks compared to Transformer++ across different attack levels. These findings highlight the potential of simplified architectures to achieve a compelling balance between efficiency, performance, and adversarial robustness, offering valuable insights for applications where resource constraints and resilience to adversarial attacks are critical.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new seven-language ophthalmology benchmark shows LLMs are less accurate in LMIC languages, and an agentic translation-plus-RAG pipeline reduces the gap.

  2. Can Large Language Models Effectively Process and Execute Financial Trading Instructions?

    cs.CE 2024-12 conditional novelty 4.0 of 10

    On a 500-item trading instruction dataset, five LLMs produced well-formatted JSON most of the time but achieved only 5-10% full accuracy and frequently asked unnecessary follow-up questions.

  3. LoRA-LiteE: A Computationally Efficient Framework for Chatbot Preference-Tuning

    cs.CL 2024-11 reject novelty 4.0 of 10

    An ensemble of two LoRA-finetuned 8-9B models gets 80.2% accuracy on Chatbot Arena preference prediction, slightly above GPT-4's 78.3%, but with no error bars or code.

  4. An Improved Dung Beetle Optimizer for Random Forest Optimization

    math.OC 2024-11 reject novelty 3.0 of 10

    Adding circle mapping and crossover to the Dung Beetle Optimizer yields faster convergence and better accuracy on selected benchmark functions, and improved random forest hyperparameters on a retail dataset.

  5. Enhanced Recommendation Combining Collaborative Filtering and Large Language Models

    cs.AI 2024-12 reject novelty 1.0 of 10

    A simple weighted sum of collaborative filtering scores and LLM text embeddings is claimed to improve recommendation accuracy, but the reported experiments are not reproducible.

  6. A Survey: Towards Privacy and Security in Mobile Large Language Models

    cs.CR 2025-09 conditional

    A survey of privacy and security challenges for mobile large language models, summarizing known attack types and defenses without introducing new results.

Pith tools