Pith. sign in

REVIEW 1 cited by

TRAWL: Tensor Reduced and Approximated Weights for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17261 v3 pith:XNMDY5JW submitted 2024-06-25 cs.CL

classification cs.CL
keywords modelslanguagemodeltensortrawlweightsapproximateddecompositions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent research has shown that pruning large-scale language models for inference is an effective approach to improving model efficiency, significantly reducing model weights with minimal impact on performance. Interestingly, pruning can sometimes even enhance accuracy by removing noise that accumulates during training, particularly through matrix decompositions. However, recent work has primarily focused on single matrix decompositions or lower precision techniques, which may fail to fully capture structural patterns. To address these limitations, we introduce TRAWL (Tensor Reduced and Approximated Weights for Large Language Models), a technique that applies tensor decomposition across multiple weight matrices to effectively denoise LLMs by capturing global structural patterns. Our experiments show that TRAWL improves model performance by up to 16% over baseline models on benchmark datasets, without requiring additional data, training, or fine-tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A post-training tensor decomposition of multi-head attention weights with shared factor matrices improves reasoning accuracy on several LLM benchmarks while compressing attention parameters.

Pith tools