Pith. sign in

REVIEW 1 cited by

Efficient Urdu Caption Generation using Attention based LSTM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.01663 v4 pith:N3IBKMGZ submitted 2020-08-02 cs.CL cs.LG

classification cs.CLcs.LG
keywords urdulanguagecaptiondatasetgenerationresearchdeepdone
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in deep learning have created many opportunities to solve real-world problems that remained unsolved for more than a decade. Automatic caption generation is a major research field, and the research community has done a lot of work on it in most common languages like English. Urdu is the national language of Pakistan and also much spoken and understood in the sub-continent region of Pakistan-India, and yet no work has been done for Urdu language caption generation. Our research aims to fill this gap by developing an attention-based deep learning model using techniques of sequence modeling specialized for the Urdu language. We have prepared a dataset in the Urdu language by translating a subset of the "Flickr8k" dataset containing 700 'man' images. We evaluate our proposed technique on this dataset and show that it can achieve a BLEU score of 0.83 in the Urdu language. We improve on the previous state-of-the-art by using better CNN architectures and optimization techniques. Furthermore, we provide a discussion on how the generated captions can be made correct grammar-wise.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A machine-translated, quality-filtered Urdu caption set covering 59,000 MS COCO images with 319,000 captions, presented as the largest public Urdu image-caption dataset.

Pith tools