Pith. sign in

REVIEW 1 cited by

Empowering the Deaf and Hard of Hearing Community: Enhancing Video Captions Using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.00342 v2 pith:HWAQOYBX submitted 2024-11-30 cs.AI

classification cs.AI
keywords captionsvideocommunitylanguagellmsmodelsaccuracychallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In today's digital age, video content is prevalent, serving as a primary source of information, education, and entertainment. However, the Deaf and Hard of Hearing (DHH) community often faces significant challenges in accessing video content due to the inadequacy of automatic speech recognition (ASR) systems in providing accurate and reliable captions. This paper addresses the urgent need to improve video caption quality by leveraging Large Language Models (LLMs). We present a comprehensive study that explores the integration of LLMs to enhance the accuracy and context-awareness of captions generated by ASR systems. Our methodology involves a novel pipeline that corrects ASR-generated captions using advanced LLMs. It explicitly focuses on models like GPT-3.5 and Llama2-13B due to their robust performance in language comprehension and generation tasks. We introduce a dataset representative of real-world challenges the DHH community faces to evaluate our proposed pipeline. Our results indicate that LLM-enhanced captions significantly improve accuracy, as evidenced by a notably lower Word Error Rate (WER) achieved by ChatGPT-3.5 (WER: 9.75%) compared to the original ASR captions (WER: 23.07%), ChatGPT-3.5 shows an approximate 57.72% improvement in WER compared to the original ASR captions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code

    cs.SE 2025-07 conditional novelty 5.0 of 10

    AccessGuru combines accessibility testing tools and LLM prompting to correct syntactic, semantic, and layout HTML accessibility violations, reporting up to 84% average violation score decrease on a new benchmark.

Pith tools