Pith. sign in

REVIEW 3 cited by

Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.03107 v1 pith:XT2A3DFC submitted 2023-08-06 cs.AI

classification cs.AI
keywords dataretrievalstructuredagriculturedocumentsembedding-basedextractpest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pest identification is a crucial aspect of pest control in agriculture. However, most farmers are not capable of accurately identifying pests in the field, and there is a limited number of structured data sources available for rapid querying. In this work, we explored using domain-agnostic general pre-trained large language model(LLM) to extract structured data from agricultural documents with minimal or no human intervention. We propose a methodology that involves text retrieval and filtering using embedding-based retrieval, followed by LLM question-answering to automatically extract entities and attributes from the documents, and transform them into structured data. In comparison to existing methods, our approach achieves consistently better accuracy in the benchmark while maintaining efficiency.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReLayout: Integrating Relation Reasoning for Content-aware Layout Generation with Multi-modal Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ReLayout adds relation-based chain-of-thought annotations and a prototype-rebalance sampler to an InternVL-based layout generator, improving structural quality and diversity on PKU and CGL poster datasets.

  2. Do LLMs Understand Why We Write Diaries? A Method for Purpose Extraction and Clustering

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The paper proposes an LLM-based pipeline to extract and cluster diary-writing purposes from Soviet-era diaries, finding GPT-4o and o1-mini most accurate.

  3. Generative AI for CAD Automation: Leveraging Large Language Models for 3D Modelling

    cs.HC 2025-07 conditional novelty 3.0 of 10

    An LLM-powered FreeCAD pipeline with error-driven re-prompting succeeds on simple and moderate 3D shapes but fails on highly constrained parameterized models.

Pith tools