REVIEW 3 major objections 4 minor 1 cited by
CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read CSE-SFP performs unsupervised contrastive sentence representation for decoder-only LLMs with a single forward pass, producing higher-quality embeddings while cutting training time and memory use.
desk verdict A genuinely useful efficiency trick for unsupervised contrastive learning on decoder-only LLMs, but the paper's quality gains are confounded with a template change; the efficiency claim is solid, the quality claim is not yet pinned down. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the two-stage prompt with two representation tokens, Rep1 and Rep2, combined with the causal attention mask of decoder-only transformers. Because the mask prevents the suffix from influencing the prefix, the same input sentence yields two embeddings computed under different attention scopes and different instruction conditions in a single forward pass: Rep1 reflects the model's encoding capability, while Rep2 reflects its generative capability through the final next-token prediction. The two vectors are both semantically close to the sentence and mutually distinct, which is what makes them usable as a positive pair for the InfoNCE loss.
What would settle it
Compute the cosine similarity between the Rep1 and Rep2 views for a large sample of sentences before and during contrastive training; if the two views are almost identical (similarity near 1) or nearly orthogonal (similarity near 0) across the sample, the single-pass positive-pair construction would fail to provide the learning signal the method claims. A controlled ablation that keeps the two templates fixed and compares one-pass CSE-SFP against two independent forward passes with the same prefix and suffix templates would also isolate whether the reported gains come from the single-pass sharing or from the template composition alone.
Extended reading notes
Core claim
The central claim is that a single forward pass suffices for effective unsupervised contrastive learning of sentence embeddings in decoder-only PLMs. CSE-SFP concatenates a prefix prompt and a suffix prompt, each terminated by a representation token; the causal attention mask ensures the prefix embedding is computed independently of the suffix, while the suffix embedding is computed at the end of the sequence and draws on the model's generative next-token prediction. These two vectors, Rep1 and Rep2, are treated as the positive pair and anchor in the InfoNCE loss. Across OPT6.7b, LLaMA2-7b, Mistral-7b, and LLaMA3-8b, the paper reports that CSE-SFP consistently outperforms two-pass baselines on seven STS benchmarks and eight MTEB IR tasks, and that it reduces training time by about 40 percent while consuming less GPU memory.
Load-bearing premise
The method assumes that the two views of the same sentence produced by the prefix token and the suffix token really are close enough to be a valid positive pair and different enough to teach the contrastive loss something, a behavioral property that the paper verifies only indirectly through downstream performance.
Editorial extensions
If this is right
- Unsupervised contrastive tuning of 7B-scale decoder-only models becomes practical on a single multi-GPU node, since training time drops to roughly 60 percent of two-pass methods and memory usage falls by about 2 to 8 GB in the reported settings.
- The two-stage prompt acts as a versatile augmentation strategy that can wrap any existing sentence-representation prompt, so future template designs can inherit the single-pass efficiency without changing the training objective.
- Because CSE-SFP does not need dropout-based augmentation, it removes a major obstacle to applying unsupervised contrastive learning to generative PLMs such as LLaMA that lack reliable dropout in the intended places.
- The ratio-based alignment-uniformity metrics give a single scalar that tracks Spearman rank performance, providing a way to compare semantic spaces that avoids the ambiguity when one encoder wins on alignment and another on uniformity.
Reading between the lines
- A testable extension is to vary the distance between the sentence and Rep1 or the length of the suffix; the causal-mask account predicts that these changes alter the diversity of the positive pair in a predictable direction, which could be measured with the ratio metrics.
- The single-pass trick may transfer to supervised contrastive settings or to asymmetric retrieval, where a query-side prefix and a document-side suffix could be encoded in one forward pass while remaining separate at inference.
- The Ratio 1 and Ratio 2 metrics, being cheap to compute, could plausibly serve as a training-time early-stopping signal or as a selection criterion for choosing among prompt templates without needing a full STS evaluation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CSE-SFP, an unsupervised contrastive sentence representation method for decoder-only LLMs. The method concatenates a two-stage prompt with two representation tokens, Rep1 and Rep2, and exploits the causal attention mask so that the prefix-stage embedding (Rep1) and the suffix-stage embedding (Rep2) of the same sentence are produced in a single forward pass and used as a positive pair for InfoNCE. Experiments on seven STS benchmarks and eight MTEB IR tasks with four LLM backbones report that CSE-SFP outperforms PromptEOL, PromptSUM, and PromptSTH in quality while reducing training time by roughly 40% and memory use by several GB. The paper also proposes two ratio-based metrics derived from alignment and uniformity to evaluate embedding spaces.
Significance. If the quality claim holds, the efficiency contribution is practically important: it makes unsupervised contrastive fine-tuning of 7B-scale LLMs substantially cheaper (about 40% less training time and 7-8 GB less GPU memory in Table 5). The method is simple, general across template families and backbones, and the authors release code and checkpoints. The efficiency measurement is well supported by the single-forward-pass design. However, the quality comparison is currently not fully controlled, and the absence of significance testing weakens the claim that CSE-SFP 'produces higher-quality embeddings.' The proposed ratio metrics are an interesting idea but receive only limited validation.
major comments (3)
- [§4.1, Tables 3 and 4] The statement in §4.1 that the comparison with PromptEOL/PromptSUM/PromptSTH 'also functions as an ablation study' is not supported, because CSE-SFP differs from these baselines in two variables at once: the prompt template (two concatenated stages with two representation tokens) and the number of forward passes used to obtain positive pairs. The higher STS and IR scores in Tables 3 and 4 could therefore be caused by the richer two-stage template or by using Rep2 as the anchor, rather than by the single-pass shared-context construction. Please add control experiments, for example a two-pass version of CSE-SFP that computes Rep1 and Rep2 from the same concatenated template in separate forward passes, and single-pass versions of the baseline templates, so that the efficiency property and the template design are disentangled.
- [Tables 3 and 4] No multiple seeds, confidence intervals, or significance tests are reported for any quality comparison. Several average differences are small (e.g., OPT6.7b: 78.54 vs. 78.26; LLaMA2: 80.12 vs. 79.20), and on individual datasets CSE-SFP sometimes trails a baseline (e.g., STS-15 for LLaMA2: 83.64 vs. 84.49). The abstract's claim that CSE-SFP 'produces higher-quality embeddings' needs repeated training runs and a paired significance test over the benchmark suite before it can be considered empirically established.
- [§3.2 and §5.1] The central premise that Rep1 and Rep2 are 'sufficiently diverse to support effective contrastive learning' while preserving semantic similarity is asserted rather than demonstrated. Table 6 reports a lower alignment value for CSE-SFP, which is indirect evidence, but no diagnostic varies the template combination (for example, swapping which template serves as prefix vs. suffix, or using the same template for both stages) to show that the quality of the positive pair, rather than the choice of anchor position or suffix instruction, drives the improvement. A small controlled study of template combinations would make the mechanism transparent.
minor comments (4)
- [§1, Figure 1] Figure 1 contains duplicated text ('The battle resulted in a Roman victory.') and a 'Copy' label that appears to be a leftover editing artifact; please clean up the figure and its caption.
- [§5.1, Eq. (9)] The definition of Ratio 2 uses a positive exponent in both numerator and denominator, unlike the standard uniformity formula which uses a negative exponent; the paper should state explicitly why this formulation is preferable and explain the conditions under which lower Ratio 1/Ratio 2 values are guaranteed to reflect a better semantic space.
- [§4.1] The temperature τ, QLoRA rank, learning rate, and number of training steps are not reported; these details are needed for reproducibility of the quality and efficiency numbers in Tables 3-5.
- [§4.4, Table 5] The memory usage values exceed the capacity of a single RTX 4090 (24 GB), so the paper should state explicitly how GPU memory is aggregated across the four GPUs and whether the reported time includes data loading and evaluation or only training steps.
Circularity Check
No significant circularity: the single-forward-pass method is evaluated on external STS/IR benchmarks; the ratio metrics are post hoc; self-citations are not load-bearing.
full rationale
Walking the derivation chain, CSE-SFP's core construction (Eq. 6) concatenates a prefix and suffix prompt so that causal attention (Eq. 4) yields Rep1 and Rep2 in one forward pass. This is a design statement, not a prediction derived from its own outputs. The claimed quality improvements are tested on external SentEval STS benchmarks and MTEB IR tasks (Tables 3-4), with no fitted parameter being relabeled as a prediction. The Ratio 1/Ratio 2 metrics (Eq. 9) are computed post hoc from alignment and uniformity and do not enter the contrastive loss, so they cannot make the method true by construction. The efficiency gain (Table 5) follows directly from replacing two forward computations with one and is externally measurable. The only self-citations that appear are to the authors' earlier PromptSUM/PromptSTH templates [42], used as components and baselines; Section 4.1's claim that the comparison 'also functions as an ablation study' is methodologically loose because the template design and pass count change together, but that is a confounding-design concern, not circularity: the central claims do not reduce to those citations or to any equation that defines the outcome in terms of the inputs. No uniqueness theorem, ansatz, or fitted value is imported from the authors' prior work as proof. Therefore no load-bearing circular step exists.
Assumptions & free parameters
assumptions (4)
- standard math A token in an autoregressive (causal) decoder cannot attend to later positions, so a representation token in the prefix is not influenced by the suffix.
- ad hoc to paper The two embeddings from the two-stage template are semantically similar yet sufficiently distinct to serve as positive pairs in InfoNCE.
- domain assumption Unsupervised contrastive learning with InfoNCE improves the alignment and uniformity of sentence embeddings.
- domain assumption The baseline methods (PromptEOL, PromptSUM, PromptSTH) are implemented in the same unsupervised setting with equivalent positive-pair construction, so the comparison isolates the single-pass template design.
Cite this review
Pith. "Pith review of CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass." pith.science (2026). https://pith.science/paper/75DZHNDD
@misc{pith2026250500389,
author = {Pith},
title = {Pith review of: CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass},
year = {2026},
howpublished = {\url{https://pith.science/paper/75DZHNDD}},
note = {Machine review of arXiv:2505.00389}
}
read the original abstract
As a fundamental task in Information Retrieval and Computational Linguistics, sentence representation has profound implications for a wide range of practical applications such as text clustering, content analysis, question-answering systems, and web search. Recent advances in pre-trained language models (PLMs) have driven remarkable progress in this field, particularly through unsupervised embedding derivation methods centered on discriminative PLMs like BERT. However, due to time and computational constraints, few efforts have attempted to integrate unsupervised sentence representation with generative PLMs, which typically possess much larger parameter sizes. Given that state-of-the-art models in both academia and industry are predominantly based on generative architectures, there is a pressing need for an efficient unsupervised text representation framework tailored to decoder-only PLMs. To address this concern, we propose CSE-SFP, an innovative method that exploits the structural characteristics of generative models. Compared to existing strategies, CSE-SFP requires only a single forward pass to perform effective unsupervised contrastive learning. Rigorous experimentation demonstrates that CSE-SFP not only produces higher-quality embeddings but also significantly reduces both training time and memory consumption. Furthermore, we introduce two ratio metrics that jointly assess alignment and uniformity, thereby providing a more robust means for evaluating the semantic spatial properties of encoding models.
Figures
Forward citations
Cited by 1 Pith paper
-
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
State-of-the-art text embeddings lag far behind on tasks requiring pragmatic inference, stance detection, and social meaning, relative to their strong performance on surface semantic benchmarks.
Reference graph
Works this paper leans on
-
[1]
Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. Llm2vec: Large language models are secretly powerful text encoders. arXiv preprint arXiv:2404.05961 (2024)
arXiv 2024
-
[2]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...
2020
-
[3]
Nuo Chen, Linjun Shou, Jian Pei, Ming Gong, Bowen Cao, Jianhui Chang, Jia Li, and Daxin Jiang. 2023. Alleviating Over-smoothing for Unsupervised Sentence Representation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, 3552–3566. ...
-
[4]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Se- bastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24, 240 (2023), 1–113
2023
-
[5]
Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, and James Glass
-
[6]
Alexis Conneau and Douwe Kiela. 2018. SentEval: An Evaluation Toolkit for Uni- versal Sentence Representations. In Proceedings of the Eleventh International Con- ference on Language Resources and Evaluation (LREC 2018). European Language Re- sources Association (ELRA), Miyazaki, Japan. https://aclanthology.org/L18-1269/
work page 2018
-
[7]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2024. QLORA: efficient finetuning of quantized LLMs. In Proceedings of the 37th In- ternational Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, Article 441, 28 pages
work page 2024
-
[8]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Comput...
Show all 49 references
-
[9]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[10]
Kawin Ethayarajh. 2019. How Contextual are Contextualized Word Represen- tations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Pro- cessing and the 9th International Joint Conference ...
2019 doi
-
[11]
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple Con- trastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Compu- tational Linguistics, Online and Punta Cana, Domini...
2021 doi
-
[12]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 770–778. doi:10.1109/CVPR.2016.90
2016 doi
-
[13]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556 (2022)
2022 arXiv
-
[14]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[15]
Ting Jiang, Shaohan Huang, Zhongzhi Luan, Deqing Wang, and Fuzhen Zhuang
-
[16]
Ting Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, and Qi Zhang. 2022. Prompt- BERT: Improving BERT Sentence Embeddings with Prompts. InProceedings of the 2022 Conference on Empirical Methods in Natural Language ...
2022 doi
-
[17]
Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the Sentence Embeddings from Pre-trained Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguisti...
2020
-
[18]
Xianming Li and Jing Li. 2023. DeeLM: Dependency-enhanced Large Language Model for Sentence Embeddings. arXiv preprint arXiv:2311.05296 (2023)
2023 arXiv
-
[19]
Jiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Kai Chen, and Rui Yan. 2023. RankCSE: Unsupervised Sen- tence Representations Learning via Learning to Rank. In Proceedings of the 61st Annual Meeting of the Association for Computational Ling...
2023 doi
-
[20]
Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692 364 (2019)
2019 arXiv
-
[21]
Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Aman- preet Singh, and Douwe Kiela. 2024. Generative representational instruction tuning. arXiv preprint arXiv:2402.09906 (2024)
2024 arXiv
-
[22]
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. MTEB: Massive Text Embedding Benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics . Association for Computational Linguistics, Dubrovnik,...
2023 doi
-
[23]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2018)
2018 arXiv
-
[24]
Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler, See-Kiong Ng, and Soujanya Poria. 2024. A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the CHATGPT Era and Beyond. In Proceedings of the 18th Conference of the European Chap...
2024
-
[25]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCN...
2019 doi
-
[26]
Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen Lee, and James T Kwok. 2022. Revisiting over-smoothing in bert from the perspective of graph. arXiv preprint arXiv:2202.08625 (2022)
2022 arXiv
-
[27]
Zhan Shi, Guoyin Wang, Ke Bai, Jiwei Li, Xiang Li, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, and Xiaodan Zhu. 2023. OssCSE: Overcoming Surface Struc- ture Bias in Contrastive Learning for Unsupervised Sentence Embedding. In Proceedings of the 2023 Conference on Empirical Me...
2023 doi
-
[28]
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language m...
2022 arXiv
-
[29]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[30]
Haochen Tan, Wei Shao, Han Wu, Ke Yang, and Linqi Song. 2022. A Sentence is Worth 128 Pseudo Tokens: A Semantic-Aware Contrastive Learning Framework for Sentence Embeddings. In Findings of the Association for Computational Lin- guistics: ACL 2022. Association for Computational...
2022 doi
-
[31]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[32]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, ...
2017
-
[33]
Hao Wang and Yong Dou. 2023. SNCSE: Contrastive Learning for Unsuper- vised Sentence Embedding with Soft Negative Samples. In Advanced Intelligent Computing Technology and Applications: 19th International Conference, ICIC 2023, Zhengzhou, China, August 10–13, 2023, Proceedings...
2023 doi
-
[34]
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368 (2023)
2023 arXiv
-
[35]
Qian Wang, Weiqi Zhang, Tianyi Lei, Yu Cao, Dezhong Peng, and Xu Wang. 2023. CLSEP: Contrastive learning of sentence embedding with prompt. Know.-Based Syst. 266, C (April 2023), 11 pages. doi:10.1016/j.knosys.2023.110381
2023
-
[36]
Tongzhou Wang and Phillip Isola. 2020. Understanding Contrastive Represen- tation Learning through Alignment and Uniformity on the Hypersphere. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119) , Hal Da...
2020
-
[37]
Tianduo Wang and Wei Lu. 2022. Differentiable Data Augmentation for Con- trastive Sentence Representation Learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computa- tional Linguistics, Abu Dhabi, United Arab E...
2022
-
[38]
Wei Wang, Liangzhu Ge, Jingqiao Zhang, and Cheng Yang. 2022. Improving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Re...
2022
-
[39]
Xing Wu, Chaochen Gao, Liangjun Zang, Jizhong Han, Zhongyuan Wang, and Songlin Hu. 2022. ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding. In Proceedings of the 29th Inter- national Conference on Computational Linguistics . I...
2022
-
[40]
Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu
-
[41]
Bowen Zhang, Kehua Chang, and Chunping Li. 2024. CoT-BERT: Enhancing Unsupervised Sentence Representation Through Chain-of-Thought. In Artificial Neural Networks and Machine Learning – ICANN 2024: 33rd International Confer- ence on Artificial Neural Networks, Lugano, Switzerla...
2024 doi
-
[42]
Bowen Zhang, Kehua Chang, and Chunping Li. 2024. Simple Techniques for Enhancing Sentence Embeddings in Generative Language Models. In Advanced Intelligent Computing Technology and Applications: 20th International Conference, ICIC 2024, Tianjin, China, August 5–8, 2024, Procee...
2024 doi
-
[43]
Bowen Zhang and Chunping Li. 2024. Advancing Semantic Textual Similar- ity Modeling: A Regression Framework with Translated ReLU and Smooth K2 Loss. In Proceedings of the 2024 Conference on Empirical Methods in Natural Lan- guage Processing. Association for Computational Lingu...
2024 doi
-
[44]
Bowen Zhang and Chunping Li. 2024. Pcc-tuning: Breaking the Contrastive Learn- ing Ceiling in Semantic Textual Similarity. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Miami, Florida, USA, ...
2024 doi
-
[45]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068 (2022)
2022 arXiv
-
[46]
Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, and Bin Cui. 2024. Retrieval-augmented generation for ai-generated content: A survey. arXiv preprint arXiv:2402.19473 (2024)
2024 arXiv
-
[2019]
In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19)
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Rep- resentations from Transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19). Association for Computing Machinery, New York, NY, US...
-
[2021]
ConSERT: A Contrastive Framework for Self-Supervised Sentence Repre- sentation Transfer. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)...
-
[2022]
In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
DiffCSE: Difference-based Contrastive Learning for Sentence Embed- dings. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Association for Computational Linguistics, Seattle, Uni...
2022 doi
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.