REVIEW 2 cited by
TakeLab Retriever: AI-Driven Search Engine for Articles from Croatian News Outlets
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
TakeLab Retriever is an AI-driven search engine designed to discover, collect, and semantically analyze news articles from Croatian news outlets. It offers a unique perspective on the history and current landscape of Croatian online news media, making it an essential tool for researchers seeking to uncover trends, patterns, and correlations that general-purpose search engines cannot provide. TakeLab retriever utilizes cutting-edge natural language processing (NLP) methods, enabling users to sift through articles using named entities, phrases, and topics through the web application. This technical report is divided into two parts: the first explains how TakeLab Retriever is utilized, while the second provides a detailed account of its design. In the second part, we also address the software engineering challenges involved and propose solutions for developing a microservice-based semantic search engine capable of handling over ten million news articles published over the past two decades.
Forward citations
Cited by 2 Pith papers
-
What Makes You CLIC: Detection of Croatian Clickbait Headlines
CLIC, a new 2,907-headline Croatian clickbait dataset, shows about 53% of sampled headlines are clickbait and fine-tuned BERTić (F1 0.78) beats zero/few-shot LLMs on detection.
-
Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings
Croatian news embeddings from 2000 to 2024 reveal semantic shifts tied to COVID-19, EU accession, and AI, and indicate a positivity drift in recent-period embeddings on sentiment classifiers.
Discussion (0). Continue with ORCID to comment.