Pith. sign in

REVIEW 1 cited by

Can we Trust Chatbots for now? Accuracy, reproducibility, traceability; a Case Study on Leonardo da Vinci's Contribution to Astronomy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11852 v1 pith:EHOSEMU6 submitted 2023-04-24 cs.CY

classification cs.CY
keywords accuracyastronomycasechatbotscontributionleonardoproblemsreproducibility
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLM) are studied. Applications to chatbots and education are considered. A case study on Leonardo's contribution to astronomy is presented. Major problems with accuracy, reproducibility and traceability of answers are reported for ChatGPT, GPT-4, BLOOM and Google Bard. Possible reasons for problems are discussed and some solutions are proposed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

    cs.CV 2025-07 reject novelty 4.0 of 10

    Med-GRIM, combining the BIND encoder with graph retrieval and small language models, reports 83.33% accuracy on a new 30-question DermaGraph dermatology benchmark, outperforming several larger medical VLMs in a zero-s...

Pith tools