Pith. sign in

REVIEW 7 cited by

Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17016 v1 pith:T6K26LDM submitted 2024-02-26 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords modelstextbilingualembeddinglanguagetaskslearningmulti-task
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce a novel suite of state-of-the-art bilingual text embedding models that are designed to support English and another target language. These models are capable of processing lengthy text inputs with up to 8192 tokens, making them highly versatile for a range of natural language processing tasks such as text retrieval, clustering, and semantic textual similarity (STS) calculations. By focusing on bilingual models and introducing a unique multi-task learning objective, we have significantly improved the model performance on STS tasks, which outperforms the capabilities of existing multilingual models in both target language understanding and cross-lingual evaluation tasks. Moreover, our bilingual models are more efficient, requiring fewer parameters and less memory due to their smaller vocabulary needs. Furthermore, we have expanded the Massive Text Embedding Benchmark (MTEB) to include benchmarks for German and Spanish embedding models. This integration aims to stimulate further research and advancement in text embedding technologies for these languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LMEB: Long-horizon Memory Embedding Benchmark

    cs.CL 2026-03 unverdicted novelty 7.0 of 10

    LMEB is a new benchmark that evaluates embedding models on long-horizon memory retrieval and shows this skill is largely orthogonal to traditional passage-retrieval performance.

  2. Large Concept Models: Language Modeling in a Sentence Representation Space

    cs.CL 2024-12 conditional novelty 7.0 of 10

    A sentence-level language model trained to autoregressively predict SONAR sentence embeddings can summarize, expand, and generate text in unseen languages.

  3. Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A text-centered multimodal model with attention and multi-task ordinal classification predicts classroom discourse quality scores with QWK 0.384, near human inter-rater reliability of 0.326.

  4. Nano-ESG: Extracting Corporate Sustainability Information from News Articles

    cs.IR 2024-12 conditional novelty 6.0 of 10

    Nano-ESG is a released dataset of about 51,000 ESG-relevant German corporate news summaries with sentiment, aspect, and timestamp labels, plus an evaluation showing about 80% expert agreement.

  5. Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning

    cs.CL 2025-06 conditional novelty 5.0 of 10

    State-of-the-art text embeddings lag far behind on tasks requiring pragmatic inference, stance detection, and social meaning, relative to their strong performance on surface semantic benchmarks.

  6. On the Suitability of pre-trained foundational LLMs for Analysis in German Legal Education

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Pre-trained open LLMs underperform bag-of-words baselines on German Gutachtenstil identification and legal essay grading, but retrieval-based example selection narrows the gap on simpler tasks.

  7. OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain

    cs.CL 2024-12 conditional novelty 5.0 of 10

    OmniEval is a finance-domain RAG benchmark that scores retrieval and generation across a 5-task by 16-topic matrix using rule-based and fine-tuned LLM metrics.

Pith tools