REVIEW 3 cited by
Simulating Task-Oriented Dialogues with State Transition Graphs and Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper explores SynTOD, a new synthetic data generation approach for developing end-to-end Task-Oriented Dialogue (TOD) Systems capable of handling complex tasks such as intent classification, slot filling, conversational question-answering, and retrieval-augmented response generation, without relying on crowdsourcing or real-world data. SynTOD utilizes a state transition graph to define the desired behavior of a TOD system and generates diverse, structured conversations through random walks and response simulation using large language models (LLMs). In our experiments, using graph-guided response simulations leads to significant improvements in intent classification, slot filling and response relevance compared to naive single-prompt simulated conversations. We also investigate the end-to-end TOD effectiveness of different base and instruction-tuned LLMs, with and without the constructed synthetic conversations. Finally, we explore how various LLMs can evaluate responses in a TOD system and how well they are correlated with human judgments. Our findings pave the path towards quick development and evaluation of domain-specific TOD systems. We release our datasets, models, and code for research purposes.
Forward citations
Cited by 3 Pith papers
-
Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures
On multi-step procedural tasks, LoRA fine-tuning underperforms full fine-tuning at every rank tested because procedural knowledge requires high-rank weight updates.
-
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
ICAT automatically scores long-form text on both factual accuracy and diverse aspect coverage, with its best variant correlating with human judgments better than standard n-gram and embedding metrics.
-
Leveraging Graph Structures and Large Language Models for End-to-End Synthetic Task-Oriented Dialogues
Users describe a conversation flow as a JSON graph and two LLM agents generate the full task-oriented dialogue from it.
Discussion (0). Continue with ORCID to comment.