Pith. sign in

REVIEW 3 cited by

Enhancing Q&A with Domain-Specific Fine-Tuning and Iterative Reasoning: A Comparative Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11792 v2 pith:PJWNFSTP submitted 2024-04-17 cs.AI

classification cs.AI
keywords reasoningdomain-specificfine-tunedmodelstechnicalcomponentsembeddingfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper investigates the impact of domain-specific model fine-tuning and of reasoning mechanisms on the performance of question-answering (Q&A) systems powered by large language models (LLMs) and Retrieval-Augmented Generation (RAG). Using the FinanceBench SEC financial filings dataset, we observe that, for RAG, combining a fine-tuned embedding model with a fine-tuned LLM achieves better accuracy than generic models, with relatively greater gains attributable to fine-tuned embedding models. Additionally, employing reasoning iterations on top of RAG delivers an even bigger jump in performance, enabling the Q&A systems to get closer to human-expert quality. We discuss the implications of such findings, propose a structured technical design space capturing major technical components of Q&A AI, and provide recommendations for making high-impact technical choices for such components. We plan to follow up on this work with actionable guides for AI teams and further investigations into the impact of domain-specific augmentation in RAG and into agentic AI capabilities such as advanced planning and reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Reviews to Dialogues: Active Synthesis for Zero-Shot LLM-based Conversational Recommender System

    cs.IR 2025-04 conditional novelty 5.0 of 10

    Active sample selection over review, metadata, and collaborative seed data plus LLM-generated synthetic dialogues improves fine-tuned conversational recommendation on ReDial and INSPIRED, though not uniformly across a...

  2. A Retrieval-Augmented Generation Framework for Academic Literature Navigation in Data Science

    cs.IR 2024-12 conditional novelty 4.0 of 10

    A five-stage enhanced RAG pipeline for data science literature is reported to improve LLM-judged context relevance, though the evaluation is self-contained and not externally validated.

  3. CAPRAG: A Large Language Model Solution for Customer Service and Automatic Reporting using Vector and Graph Retrieval-Augmented Generation

    cs.CL 2025-01 reject novelty 3.0 of 10

    CAPRAG combines vector and graph retrieval with a router and Cypher templates for banking queries, but provides no quantitative evidence of effectiveness.

Pith tools