REVIEW 8 cited by
Text Clustering with Large Language Model Embeddings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text clustering is an important method for organising the increasing volume of digital content, aiding in the structuring and discovery of hidden patterns in uncategorised data. The effectiveness of text clustering largely depends on the selection of textual embeddings and clustering algorithms. This study argues that recent advancements in large language models (LLMs) have the potential to enhance this task. The research investigates how different textual embeddings, particularly those utilised in LLMs, and various clustering algorithms influence the clustering of text datasets. A series of experiments were conducted to evaluate the impact of embeddings on clustering results, the role of dimensionality reduction through summarisation, and the adjustment of model size. The findings indicate that LLM embeddings are superior at capturing subtleties in structured language. OpenAI's GPT-3.5 Turbo model yields better results in three out of five clustering metrics across most tested datasets. Most LLM embeddings show improvements in cluster purity and provide a more informative silhouette score, reflecting a refined structural understanding of text data compared to traditional methods. Among the more lightweight models, BERT demonstrates leading performance. Additionally, it was observed that increasing model dimensionality and employing summarisation techniques do not consistently enhance clustering efficiency, suggesting that these strategies require careful consideration for practical application. These results highlight a complex balance between the need for refined text representation and computational feasibility in text clustering applications. This study extends traditional text clustering frameworks by integrating embeddings from LLMs, offering improved methodologies and suggesting new avenues for future research in various types of textual analysis.
Forward citations
Cited by 8 Pith papers
-
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
A systematic review of 153 CHI papers shows LLMs are used across ten domains, mostly as system engines and in empirical or artifact contributions, with widespread validity and reproducibility concerns.
-
Cequel: Cost-Effective Querying of Large Language Models for Text Clustering
Cequel reduces the LLM query cost of text clustering by selecting a few informative text pairs or triples, turning LLM answers into must-link and cannot-link constraints, and clustering with a PMI-weighted constrained...
-
A Dynamic Framework for Semantic Grouping of Common Data Elements (CDE) Using Embeddings and Clustering
LLM embeddings clustered with HDBSCAN group 6,390 NIH common data elements into 118 semantic clusters, and a random forest on the same embeddings reaches 90.46% accuracy on cluster labels.
-
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
BPO balances knowledge breadth and depth in preference data by compressing prompts and dynamically augmenting the number of response pairs per prompt using gradient-based clustering, achieving stronger alignment with ...
-
AI-Driven Automation Can Become the Foundation of Next-Era Science of Science Research
The paper defines a five-level AI4SoS automation hierarchy and demonstrates a preliminary LLM multi-agent society that partially reproduces known correlations between team diversity and citation impact.
-
A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions
A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.
-
Advanced Topic Modeling Techniques for Categorizing Software Vulnerabilities
Existing embedding-based topic models produce interpretable clusters on Cisco vulnerability Threat text, but without quantitative coherence scores, baselines, or downstream prioritization metrics.
-
LLMs are Also Effective Embedding Models: An In-depth Overview
A structured survey of using decoder-only LLMs as text embedding models, covering prompting, fine-tuning, data construction, benchmarks, and open problems.
Discussion (0). Continue with ORCID to comment.