Pith. sign in

REVIEW 1 cited by

Deploying Open-Source Large Language Models: A performance Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.14887 v4 pith:JNGYILPO submitted 2024-09-23 cs.PF cs.AIcs.LG

classification cs.PFcs.AIcs.LG
keywords modelsavailablelanguagelargeperformancedeploydifferentevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since the release of ChatGPT in November 2022, large language models (LLMs) have seen considerable success, including in the open-source community, with many open-weight models available. However, the requirements to deploy such a service are often unknown and difficult to evaluate in advance. To facilitate this process, we conducted numerous tests at the Centre Inria de l'Universit\'e de Bordeaux. In this article, we propose a comparison of the performance of several models of different sizes (mainly Mistral and LLaMa) depending on the available GPUs, using vLLM, a Python library designed to optimize the inference of these models. Our results provide valuable information for private and public groups wishing to deploy LLMs, allowing them to evaluate the performance of different models based on their available hardware. This study thus contributes to facilitating the adoption and use of these large language models in various application domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReservoirChat: Interactive Documentation Enhanced with LLM and Knowledge Graph for ReservoirPy

    cs.SE 2025-07 conditional novelty 4.0 of 10

    ReservoirChat, a RAG and knowledge-graph assistant for ReservoirPy, improves domain-specific question answering and code debugging over its base model, but its custom benchmark may be contaminated by its own knowledge base.

Pith tools