LMR-BENCH measures LLM agents on reproducing masked functions from 23 NLP papers, and every tested model and agent passes under 43% of the unit tests.
Scholar Inbox: Personalized Paper Recommendations for Scientists
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Scholar Inbox is a new open-access platform designed to address the challenges researchers face in staying current with the rapidly expanding volume of scientific literature. We provide personalized recommendations, continuous updates from open-access archives (arXiv, bioRxiv, etc.), visual paper summaries, semantic search, and a range of tools to streamline research workflows and promote open research access. The platform's personalized recommendation system is trained on user ratings, ensuring that recommendations are tailored to individual researchers' interests. To further enhance the user experience, Scholar Inbox also offers a map of science that provides an overview of research across domains, enabling users to easily explore specific topics. We use this map to address the cold start problem common in recommender systems, as well as an active learning strategy that iteratively prompts users to rate a selection of papers, allowing the system to learn user preferences quickly. We evaluate the quality of our recommendation system on a novel dataset of 800k user ratings, which we make publicly available, as well as via an extensive user study. https://www.scholar-inbox.com/
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
LMR-BENCH measures LLM agents on reproducing masked functions from 23 NLP papers, and every tested model and agent passes under 43% of the unit tests.