REVIEW 11 cited by
StructuredRAG: JSON Response Formatting with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability of Large Language Models (LLMs) to generate structured outputs, such as JSON, is crucial for their use in Compound AI Systems. However, evaluating and improving this capability remains challenging. In this work, we introduce StructuredRAG, a benchmark of six tasks designed to assess LLMs' proficiency in following response format instructions. We evaluate two state-of-the-art LLMs, Gemini 1.5 Pro and Llama 3 8B-instruct with 4-bit quantization using two distinct prompting strategies. We introduce these prompting strategies as f-String and Follow the Format (FF) prompting. Across 24 experiments, we find an average success rate of 82.55%. We further find a high variance in performance across tasks, models, and prompting strategies with success rates ranging from 0 to 100%. We find that Llama 3 8B-instruct often performs competitively with Gemini 1.5 Pro. We observe that task complexity significantly influences performance, with tasks involving lists or composite object outputs proving more challenging. Our findings highlight the need for further research into improving the reliability and consistency of structured output generation in LLMs. We have open-sourced our experimental code and results at github.com/weaviate/structured-rag.
Forward citations
Cited by 11 Pith papers
-
Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
Interactive Reasoning, instantiated as Hippo, lets users view and edit an LLM's chain-of-thought as a tree, and a 16-person study reports improved perceived control, sense-making, and assumption awareness.
-
Evaluating Language Models as Synthetic Data Generators
AgoraBench shows that an LM's ability to solve problems does not predict its ability to generate useful synthetic training data.
-
CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning
A cost-aware tool planning framework using tokenized plans and offline RL lets a 7B model beat GPT-4 on a new cost-aware planning benchmark.
-
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
A single trained token pair inserted around any target text forces Qwen-2 7B and Llama-3.1 8B to output that text on 54 to 75 percent of unseen prompts.
-
Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity
A modular multimodal generative AI framework produces synthetic residential building data from public sources, with reported overlaps exceeding 65% against a national reference dataset.
-
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
Asking a reasoning model several problems at once reveals large accuracy drops and exposes differences that single-question benchmarks miss.
-
OASBuilder: Generating OpenAPI Specifications from Online API Documentation with Large Language Models
A modular pipeline of rules and LLMs converts HTML API documentation into OpenAPI specifications, evaluated on hundreds of APIs and deployed in an enterprise setting.
-
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance
On most RDF and SPARQL engineering tasks, larger open LLMs score higher, but plateau, ceiling, and occasional intra-family drops mean bigger is not always better.
-
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
LLM-KG-Bench 3.0 is an open, extensible benchmark framework with automatic scoring that compares more than 30 LLMs on RDF and SPARQL knowledge graph tasks.
-
AI Prototyper: A Figma Plugin for Decomposition-Based GUI Prototyping with LLMs
An LLM-based Figma plugin with a human-editable feature list and RAG component retrieval produced more completed prototypes and higher expert quality ratings than manual Figma in a small Thai-language pilot.
-
Querying Databases with Function Calling
A new tool definition and synthetic benchmark show top LLMs can format database query calls via function calling, with the best models scoring around 74% exact-match accuracy.
Discussion (0). Continue with ORCID to comment.