Pith. sign in

REVIEW 2 cited by

Do We Need Domain-Specific Embedding Models? An Empirical Investigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18511 v4 pith:7WRAHT5R submitted 2024-09-27 cs.CL cs.IR

classification cs.CLcs.IR
keywords embeddingmodelsdomain-specificperformancefinmtebmtebtextdomain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Embedding models play a crucial role in representing and retrieving information across various NLP applications. Recent advancements in Large Language Models (LLMs) have further enhanced the performance of embedding models, which are trained on massive amounts of text covering almost every domain. These models are often benchmarked on general-purpose datasets like Massive Text Embedding Benchmark (MTEB), where they demonstrate superior performance. However, a critical question arises: Is the development of domain-specific embedding models necessary when general-purpose models are trained on vast corpora that already include specialized domain texts? In this paper, we empirically investigate this question, choosing the finance domain as an example. We introduce the Finance Massive Text Embedding Benchmark (FinMTEB), a counterpart to MTEB that consists of financial domain-specific text datasets. We evaluate the performance of seven state-of-the-art embedding models on FinMTEB and observe a significant performance drop compared to their performance on MTEB. To account for the possibility that this drop is driven by FinMTEB's higher complexity, we propose four measures to quantify dataset complexity and control for this factor in our analysis. Our analysis provides compelling evidence that state-of-the-art embedding models struggle to capture domain-specific linguistic and semantic patterns. Moreover, we find that the performance of general-purpose embedding models on MTEB is not correlated with their performance on FinMTEB, indicating the need for domain-specific embedding benchmarks for domain-specific embedding models. This study sheds light on developing domain-specific embedding models in the LLM era. FinMTEB comes with open-source code at https://github.com/yixuantt/FinMTEB

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-Encoder

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PRISM produces interpretable political bias embeddings by mining controversial topics, generating left/right bias indicators, and scoring articles with a political-aware cross-encoder.

  2. FinBERT2: A Specialized Bidirectional Encoder for Bridging the Gap in Finance-Specific Deployment of Large Language Models

    cs.IR 2025-05 conditional novelty 4.0 of 10

    A 32B-token Chinese financial corpus and FinBERT2 model outperform prior FinBERTs, general BERTs, and several large LLMs on five classification and retrieval benchmarks.

Pith tools