Pith. sign in

REVIEW

Japanese SimCSE Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19349 v1 pith:4SLLHLCG submitted 2023-10-30 cs.CL

classification cs.CL
keywords japanesesentencesimcseembeddingmodelsreportdatasetsbaseline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We report the development of Japanese SimCSE, Japanese sentence embedding models fine-tuned with SimCSE. Since there is a lack of sentence embedding models for Japanese that can be used as a baseline in sentence embedding research, we conducted extensive experiments on Japanese sentence embeddings involving 24 pre-trained Japanese or multilingual language models, five supervised datasets, and four unsupervised datasets. In this report, we provide the detailed training setup for Japanese SimCSE and their evaluation results.

Discussion (0). Continue with ORCID to comment.

Pith tools