Pith. sign in

REVIEW 6 cited by

Designing Large Foundation Models for Efficient Training and Inference: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.01990 v5 pith:HVVTNHOK submitted 2024-09-03 cs.DC cs.LG

classification cs.DCcs.LG
keywords efficientinferencetrainingdesignfoundationmodelmodelssystem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper focuses on modern efficient training and inference technologies on foundation models and illustrates them from two perspectives: model and system design. Model and System Design optimize LLM training and inference from different aspects to save computational resources, making LLMs more efficient, affordable, and more accessible. The paper list repository is available at https://github.com/NoakLiu/Efficient-Foundation-Models-Survey.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OLIVE: Online Low-Rank Incremental Learning for Efficient Adaptive Exoskeletons

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    OLIVE decomposes exoskeleton policy adaptations into low-rank residuals updated via sensor-driven policy gradients with gating and dynamic rank scheduling, reporting gait improvements on a wearable platform.

  2. CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

    cs.CL 2025-11 reject novelty 6.0 of 10

    CSV-Decode uses cluster-radius upper bounds to skip vocabulary tokens with certified exact top-k or ε-approximate softmax, but its reported speedups are not supported by the data as presented.

  3. SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling

    cs.CL 2025-08 reject novelty 4.0 of 10

    A semantic-aware tokenizer that merges similar and low-entropy text spans cuts long-context token counts by up to 59% and inference latency by roughly 2x, with no reported quality loss.

  4. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  5. An Intelligent Fault Self-Healing Mechanism for Cloud AI Systems via Integration of Large Language Models and Deep Reinforcement Learning

    cs.AI 2025-06 reject novelty 3.0 of 10

    An LLM-plus-deep-RL hybrid is proposed for cloud fault self-healing, claiming 37% faster recovery on unknown faults with weak experimental documentation.

  6. Anomaly Detection and Early Warning Mechanism for Intelligent Monitoring Systems in Multi-Cloud Environments Based on LLM

    cs.LG 2025-06 reject novelty 3.0 of 10

    A CNN-LSTM-LLM-deep SVM hybrid is proposed for multi-cloud anomaly detection, but the evaluation is qualitative and Equation (8) is mathematically wrong.

Pith tools