Pith. sign in

REVIEW 12 cited by

BBC-Oxford British Sign Language Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.03635 v1 pith:3RGXGUCX submitted 2021-11-05 cs.CV

classification cs.CV
keywords datasetsignlanguagebobslbritishavailablebbc-oxforddata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work, we introduce the BBC-Oxford British Sign Language (BOBSL) dataset, a large-scale video collection of British Sign Language (BSL). BOBSL is an extended and publicly released dataset based on the BSL-1K dataset introduced in previous work. We describe the motivation for the dataset, together with statistics and available annotations. We conduct experiments to provide baselines for the tasks of sign recognition, sign language alignment, and sign language translation. Finally, we describe several strengths and limitations of the data from the perspectives of machine learning and linguistics, note sources of bias present in the dataset, and discuss potential applications of BOBSL in the context of sign language technology. The dataset is available at https://www.robots.ox.ac.uk/~vgg/data/bobsl/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Isharah is a new 30,000-clip, multi-scene Saudi Sign Language dataset with gloss and translation annotations, plus signer-independent and unseen-sentence benchmarks for continuous sign language recognition and translation.

  2. Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Hard negatives selected by visual confusability in sign embeddings, not linguistic similarity, substantially raise fine-grained sign-language retrieval accuracy without collapsing coarse performance.

  3. SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning

    cs.CV 2026-03 accept novelty 6.0 of 10

    Sparse keyframe-conditioned Conditional Flow Matching produces fluid, articulate 3D sign language motion across four languages while enabling precise Keyframe-to-Pose editing.

  4. Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dual visual encoder with contrastive visual-text pretraining achieves the best reported BLEU-4 score among gloss-free sign language translation methods on Phoenix-2014T.

  5. iLSU-T: an Open Dataset for Uruguayan Sign Language Translation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    iLSU-T is a 187-hour Uruguayan Sign Language video dataset with Spanish text, 18 interpreters, and first baseline translation results.

  6. Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LLM-generated pseudo glosses, reordered via weak video supervision, enable sign language translation that rivals gloss-supervised models while needing only 30 gloss examples.

  7. 2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset

    cs.CL 2024-12 conditional novelty 6.0 of 10

    2M-BELEBELE is a new multilingual speech and ASL comprehension benchmark built from BELEBELE and FLEURS, with human recordings for 74 spoken languages and ASL video with glosses.

  8. SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A masked cluster-prediction transformer over four sign-language streams sets state-of-the-art results on multiple ASL translation and recognition benchmarks using only public pre-training data.

  9. ODE-Based Transformer Decoders for Iterative Sign Language Translation

    cs.CL 2026-08 conditional novelty 5.0 of 10

    Replacing residual decoder updates with RK-2 and RK-4 numerical integration steps improves BLEU-4 on two sign language benchmarks against a matched iterative refinement baseline without adding decoder parameters.

  10. Sign Spotting Disambiguation using Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LLM-based beam search disambiguation improves dictionary sign spotting WER from 47.2% to 44.4% on an internal BSL dataset.

  11. Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.

  12. Deaf in AI: AI language technologies and the erosion of linguistic rights

    cs.CY 2025-05 conditional novelty 5.0 of 10

    The paper argues that current AI sign language technologies, trained on interpreter-mediated data and framed as substitutes for interpreters, threaten deaf people's linguistic rights and require deaf-led design and go...

Pith tools