REVIEW 11 cited by
Recipes for building an open-domain chatbot
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Building open-domain chatbots is a challenging area for machine learning research. While prior work has shown that scaling neural models in the number of parameters and the size of the data they are trained on gives improved results, we show that other ingredients are important for a high-performing chatbot. Good conversation requires a number of skills that an expert conversationalist blends in a seamless way: providing engaging talking points and listening to their partners, and displaying knowledge, empathy and personality appropriately, while maintaining a consistent persona. We show that large scale models can learn these skills when given appropriate training data and choice of generation strategy. We build variants of these recipes with 90M, 2.7B and 9.4B parameter models, and make our models and code publicly available. Human evaluations show our best models are superior to existing approaches in multi-turn dialogue in terms of engagingness and humanness measurements. We then discuss the limitations of this work by analyzing failure cases of our models.
Forward citations
Cited by 11 Pith papers
-
On the Fitness Landscape in the $NK$ Model
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
-
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
LobRA reduces GPU seconds for multi-tenant LoRA fine-tuning by 45.03%-60.67% through heterogeneous FT replicas and per-step workload-balanced dispatching.
-
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
semi-PD is an LLM serving system that disaggregates prefill and decode compute at the streaming multiprocessor level over a unified GPU memory pool, and reports 1.27-2.58x lower average latency and 1.55-1.72x higher S...
-
Precise Length Control in Large Language Models
Adding a reversed, scaled sinusoidal positional encoding to a fine-tuned decoder-only LLM lets it end responses within about three tokens of a requested length.
-
Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery
A 10-day field study found that an LLM chatbot supported eating disorder recovery through private storytelling, yet also produced unnoticed harmful responses such as praising weight loss and restriction.
-
Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction
Modern LLMs mimic child and caregiver speech at the word and sentence level but exaggerate conversational alignment and show less diversity than real parent-child dialogue.
-
MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
MMGraphRAG links scene-graph entities from images to text knowledge graph entities via SpecLink, and reports accuracy gains over naive RAG and GraphRAG on multimodal document QA.
-
Automated Feedback Loops to Protect Text Simplification with Generative AI from Information Loss
Adding all missing named entities back into simplified biomedical text yields the highest cosine similarity and ROUGE-1 to the original among five insertion strategies, but the evaluation is partly circular and lacks ...
-
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
LLM-based translation and paraphrasing can effectively augment multilingual bias testing, and low-resource languages tend to show worse bias-detection scores than high-resource ones.
-
Large Language Model guided Deep Reinforcement Learning for Decision Making in Autonomous Driving
An LLM-guided reinforcement learning framework with a JS-divergence policy constraint and intermittent safety interventions reaches 90% lane-change task success in highway-env, outperforming SAC-based baselines.
-
Real-Time Textless Dialogue Generation
A streaming, textless dialogue model predicts turn-taking actions every 160 ms and generates speech units, improving naturalness over cascaded systems at the cost of lower semantic coherence.
Discussion (0). Continue with ORCID to comment.