{"id":"b87fd0fe-465f-4734-8f32-670308b285b8","arxiv_id":"2411.12865","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AzSLD, a new public Azerbaijani Sign Language dataset with fingerspelling, 100 word classes, 500 sentence videos from two camera views, and a data loader, is introduced.","lead":"This paper introduces AzSLD, a public video dataset for Azerbaijani Sign Language with fingerspelling, word, and sentence level samples, plus a data loader. It is the first sentence-level resource for this low-resource sign language and could support translation systems for Azerbaijani Deaf communities.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset statistics are internally inconsistent (30,000 vs. 30,312 videos; 65 vs. 65.2 hours) and the reported annotation pipeline cannot verify the claimed frame-level annotation accuracy.","rationale":"The reader correctly flagged the annotation-accuracy assumption as the weakest premise. My stress-test agrees but refines it: the more actionable, falsifiable concern is that the paper's own reported statistics do not cohere (30,000 vs. 30,312; 65 vs. 65.2 hours; no component-level breakdown summing to the total), and the annotation-pipeline description (upload to Supervisely, labels added by signing actors) is too thin to support the 'meticulously annotated' and 'frame-number-aligned' claims. The concrete test is decisive because it checks whether the artifact actually contains the advertised videos and whether the annotations are self-consistent with the videos. This does not rise to REJECT because dataset papers can be repaired with a corrected supplement, a full inventory, and a small annotation-consistency audit; the underlying collection effort seems real and ethically documented. I recommend keeping the CONDITIONAL verdict, adding the specific conditions: reconcile the counts, describe the overlap with [54], and publish either inter-annotator agreement or a spot-check protocol.","tokens_in":13959,"tokens_out":1829,"duration_ms":14333,"concrete_test":"Download the Zenodo deposit (10.5281/zenodo.13627300) and programmatically count: (a) number of video files per component (fingerspelling, words, sentences, Cam1/Cam2), (b) total duration, and (c) number of annotation JSON files whose start/end frame indices fall within the corresponding video's frame count. If the counts don't match 30,312 videos / 65.2 hours, or if any annotation file contains out-of-range or empty frame alignments, the dataset-as-claimed does not exist as described.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that AzSLD is a comprehensive, carefully annotated, first-of-its-kind dataset for Azerbaijani Sign Language AI. The load-bearing issue is that the paper does not establish that the dataset is actually what it claims to be: (1) The headline number is 30,000 videos, but Table 1 and Section 3 report 30,312 videos and 65.2 hours; this may indicate the dataset bundle does not match the paper. (2) The claimed frame-level, time-aligned annotations are the core value proposition, but the paper provides no inter-annotator agreement, no annotation consistency check, and no example showing a sign label actually aligns with its stated start/end frame. Section 4 explicitly admits that inter-annotator agreement metrics 'were not exhaustively detailed.' (3) Section 3.1 says the word-level data comes from the top 100 most frequent words in the sentence data, with 30,312 videos total, but the sentence/fingerspelling/word components are never enumerated in a way that sums to the total, so the reader cannot tell which partition contains what. (4) The 'first resource for Azerbaijani Sign Language with AI purposes' claim is weakened by reference [54], which already describes an AzSL dactyl-alphabet dataset and word recognition system; the overlap is never clarified. These are not consensus disagreements but internal inconsistencies and missing verification of the dataset's central artifact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AzSLD, a dataset for Azerbaijani Sign Language (AzSL) that comprises three components: fingerspelling (static images and dynamic videos for the 32-letter Azerbaijani alphabet), isolated word videos for the 100 most frequent words in the sentence corpus, and sentence-level videos recorded from two camera views with frame-level, time-aligned gloss annotations. The dataset is released on Zenodo with a companion data loader on GitHub. The paper describes the collection setup, annotation workflow via the Supervisely tool, ethical consent, and limitations, and positions AzSLD as the first AI-oriented resource for Azerbaijani Sign Language.","tokens_in":14262,"tokens_out":6210,"duration_ms":54714,"significance":"If the dataset is as described, AzSLD fills a clear gap for a low-resource sign language and offers a multi-view, time-aligned, sentence-level resource that could support sign recognition, translation, and linguistic research. The public release under CC-BY, the inclusion of a data loader, and the explicit attention to FATE/ethics are commendable and align with best practices for dataset papers. However, the manuscript currently contains internal inconsistencies in the reported statistics and does not provide sufficient evidence for the claimed frame-level annotation accuracy, so the dataset's actual contents and reliability cannot be fully verified from the paper alone.","major_comments":[{"comment":"The total dataset size is stated as 30,000 videos in the abstract, 30,312 videos in Table 1, and 'approximately 65 hours' in the introduction versus 65.2 hours in Table 1. More critically, the paper never breaks down the total into the three components (fingerspelling videos, word videos, sentence videos). Without explicit per-component counts, the reader cannot verify whether the sum is consistent. Please provide exact counts for each component, including the number of unique sentence videos, the number of camera views, and the number of samples per word class, so the total reconciles.","section":"Section 3.1 and Table 1"},{"comment":"The text reports 10,864 images and 3,587 videos for the fingerspelling component, summing to 14,451, but the per-letter counts in Table 2 sum to 14,453. This small discrepancy should be reconciled. In addition, Table 2 does not indicate which letters are static (image) and which are dynamic (video), so the image/video split cannot be cross-checked against the table. Please add this distinction or clarify the classification of each letter.","section":"Section 3.1 and Table 2"},{"comment":"The paper's core value proposition is the frame-level, time-aligned annotation of the sentence videos, yet Section 4 explicitly acknowledges that inter-annotator agreement metrics and consistency checks 'were not exhaustively detailed.' Since the dataset is the primary contribution, this is load-bearing: without any evidence about annotation reliability, the claim of 'meticulously annotated' and 'accurate sign labels' is not supported. Please provide a concrete annotation-quality check, such as agreement statistics on a double-annotated subset, a description of how disagreements were resolved, or at least a worked example showing actual start/end frame numbers from a JSON annotation file for one sentence.","section":"Section 3.2 and Section 4"},{"comment":"The paper claims AzSLD is 'the first resource for Azerbaijani Sign Language with AI purposes,' but reference [54] already describes an AzSL dactyl-alphabet dataset and a word recognition system. The authors even cite [54] for the fingerspelling component. This undermines the novelty claim unless the paper clarifies exactly what new content AzSLD provides beyond [54]. Please state the relationship between the two resources and what is newly released in AzSLD.","section":"Introduction and Section 3.1"},{"comment":"The number of signers is reported inconsistently: Section 3.2 says the dataset was developed by '42 DHH individuals who are native users of AzSLD, one CODA,' then later mentions '40 participants involved in the project,' and elsewhere states that 'Each sentence was recorded by at least 18 of the 40 participants' while Section 3.1 says the sentence videos were performed by '18 to 25 different signers.' These figures need to be unified so the reader understands the actual signer pool and the per-sentence signer counts.","section":"Section 3.2"}],"minor_comments":[{"comment":"The Zenodo DOI is given as 10.5281/ZENODO.13627301 in reference [24] and as zenodo.org/doi/10.5281/zenodo.13627300 in the Data availability statement; please ensure the correct DOI is used consistently.","section":"Data availability and reference [24]"},{"comment":"The language entry 'Chenese SL' appears to be a typo for 'Chinese SL'.","section":"Table 1"},{"comment":"There are several typos: 'staticly' should be 'statically', 'letteres' should be 'letters', and 'mactivity' appears in the description of the camera views and should likely be 'activity'.","section":"Section 3.1 and Section 3.2"},{"comment":"Figures 2 and 3 are referenced without descriptive captions in the text; please add clear captions and axis labels so the reader knows what is plotted (e.g., histogram of frame lengths, histogram of word counts per sentence).","section":"Figures 2 and 3"},{"comment":"The data loader description states that labels are extracted from directory names, but for the sentence-level component the labels are stored in JSON annotation files; please clarify how the loader handles this difference or whether the sentence-level data require a separate loading path.","section":"Section 3.4"},{"comment":"The AzSLD Words component is described as having 100 classes but no sample counts are given; please provide a link to a file listing the word classes and their counts, or include a summary table in the paper.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a potentially valuable dataset, but the internal inconsistencies in the statistics and the lack of annotation verification are significant enough that the manuscript should not be accepted in its current form. The issues appear fixable within the scope of a revision: the authors can reconcile the counts, provide a clear component breakdown, add a basic annotation-quality check, and clarify the relationship to their prior work in [54]. I would recommend sending back for major revision rather than rejecting, provided the authors can address the load-bearing concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Couple of things to know. AzSLD is not dead on arrival. The sentence-level part—500 signed sentences in Azerbaijani Sign Language with two camera views, per-sign start/end frame annotations, and a 100-class word subset—is a real new resource, and the authors have actually shipped it on Zenodo with a data loader on GitHub. That combination is the first public sentence-level AzSL resource I'm aware of.\n\nWhat the paper does well: the collection story is credible and community-grounded. Having native signers, social service workers, and a CODA compose the 500 sentences is the right way to build a first resource for a low-resource sign language. The FATE section is honest, and the limitations section admits the main weakness up front: no detailed inter-annotator agreement. The data loader appears to be a modest but real practical contribution.\n\nNow the soft spots, in proportion. The numbers don't add up. The abstract says 30,000 videos; Table 1 says 30,312 and 65.2 hours; Section 1 says 65 hours. The paper never enumerates how many videos are in each component, so you can't verify the total. That's sloppy in a dataset paper. Also, the 'first resource for Azerbaijani Sign Language with AI purposes' claim collides with their own [54], which already released a fingerspelling dataset. The new part is the sentence corpus and the word subset; the claim should say that.\n\nThe bigger issue is annotation quality. The frame-level, time-aligned annotations are the main selling point, but there is no inter-annotator agreement, no spot check, no example showing a label's start/end frames against the video. The paper admits this. For a dataset meant to serve as a benchmark, that's a real gap. It's fixable—a small agreement study or a sanity evaluation on a subset would help. I don't think the dataset is fake; the description is too detailed and the limitations text too honest for that. But the annotation accuracy claim is unverified.\n\nOne correction to the stress-test note: the claim that Table 2 lists 31 letters missing 'k' and totals 13,961 is wrong. In the version I read, 'k' is present (464 samples) and the table includes all 32 letters; the sum is about 14,453, close to the reported 10,864 images + 3,587 videos (14,451). So that particular inconsistency isn't there. The real ones are above.\n\nBottom line: this deserves a serious referee. The resource fills a genuine gap, and the issues are addressable. I'd recommend major revision, not desk rejection, with a clean, self-consistent statistics table and at least one annotation reliability check required.","headline":"A genuinely useful low-resource sign language dataset that needs a clean second pass on its own numbers and some annotation validation before benchmarking claims hold.","tokens_in":14750,"tokens_out":6024,"would_cite":true,"duration_ms":48544,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces AzSLD, a 30,000-video Azerbaijani Sign Language dataset with frame-aligned word and sentence annotations, positioned as the first AI-oriented resource for the language.","keywords":["Azerbaijani Sign Language","sign language dataset","sign language recognition","sign language translation","fingerspelling","frame-level annotation","low-resource language","video dataset"],"falsifier":"Take a random sample of about 50 sentence videos, have native Azerbaijani Sign Language signers who were not part of the original annotation process independently re-annotate the labels and frame boundaries, and compare; if label mismatches or boundary disagreements are frequent, the claim of accurate and reliable annotations fails.","tokens_in":13790,"feed_emoji":"🤟","tokens_out":7972,"duration_ms":73700,"temperature":0.7,"pith_summary":"AzSLD is a newly released video dataset for Azerbaijani Sign Language, built from recordings by 43 signers and totaling about 65 hours. The paper's central claim is that this is the first resource for the language made with AI purposes in mind, containing fingerspelling data, word-level videos, and sentence-level videos with time-aligned gloss labels. The sentence component consists of 500 sentences recorded from two camera angles, annotated so each sign is mapped to the frames where it occurs, which is what makes the data usable for both recognition and translation. The authors also provide a data loader that standardizes preprocessing, label encoding, and train/test splitting. A sympathetic reader should take the contribution to be the existence of the resource itself — publicly released, consent-cleared, and documented — rather than a new model result.","feed_headline":"First Azerbaijani sign language video dataset brings 30,000 clips","feed_subtitle":"Frame-aligned word and sentence labels, two camera angles, and an open data loader target recognition and translation.","key_machinery":"The object carrying the argument is AzSLD itself, structured as three components: fingerspelling, words, and sentences. The load-bearing design choices are frame-level alignment, dual-camera capture, and a configurable data loader. In the sentence component, each label is tied to the exact video frames where the sign appears, letting downstream models do sign-boundary detection and learn temporal structure rather than treating each video as one undifferentiated clip. The two camera angles are intended as complementary views, so a sign hidden from the front may be visible from the side. The data loader extracts a fixed number of frames, resizes them to 224x224 pixels, converts categorical labels to one-hot vectors, and splits data 80/20 into training and testing sets, which is the mechanism that turns raw video into a usable benchmark.","core_discovery":"The central claim is that AzSLD is the first dataset purpose-built for Azerbaijani Sign Language AI work, and that its annotation design supports isolated and continuous recognition as well as translation. In the paper's own summary table the dataset contains 30,312 videos: 10,864 images for the 24 statically expressed fingerspelling letters, 3,587 videos for the 8 dynamically expressed letters, 100 word classes drawn from the most frequent labels in the sentence set, and 500 sentences performed by 18 to 25 of the 43 participants. Every sentence video is captured from a frontal and a side camera, and each comes with a JSON annotation file that aligns every gloss label to its start and end frames; 687 unique labels appear across the sentence set. The authors state that all contributors gave informed consent and that the dataset is released publicly with documentation and source code. If the claim is right, researchers now have a documented, ready-to-load benchmark for a sign language that previously had no AI-oriented public resource.","pith_inferences":["I infer that the sentence component's vocabulary, built around social service interactions, makes AzSLD a targeted record of a specific register of Azerbaijani Sign Language rather than a general-language corpus.","I infer that because inter-annotator agreement was not exhaustively documented, downstream users should re-validate a random sample of frame alignments before trusting model outputs; the paper itself lists this as a limitation.","I infer that a useful test the authors did not run is cross-view evaluation — training on one camera and testing on the other — to measure how much the two angles contribute independently.","I infer that if AzSLD becomes a benchmark, the most informative early baselines will be cross-lingual transfer from geographically or typologically related sign language datasets, a comparison the paper does not provide."],"forward_implications":["Isolated sign recognition, fingerspelling recognition, sentence-level translation, and sign-boundary detection can be trained and evaluated on Azerbaijani Sign Language data for the first time.","The frame-aligned sentence annotations let researchers study how individual sign tokens combine into sentence-level translations, and compare temporal placement of the same label across different signers.","The dual-camera capture gives models a second, complementary view that can recover signs occluded in the frontal view.","The standardized data loader lowers the barrier to using the dataset and makes results across teams more directly comparable.","Because recording happened in a controlled laboratory setting, models trained on AzSLD are expected to need fine-tuning before deployment in naturalistic settings, as the paper's limitations section concedes."],"supporting_citations":[{"why":"It is the released dataset object that makes the resource publicly available under a permissive license.","marker":"[24]"},{"why":"It is the accompanying data loader code that standardizes frame extraction, label encoding, and dataset splitting.","marker":"[25]"},{"why":"It is the earlier work that produced the fingerspelling collection and a word recognition demo, which the dataset extends.","marker":"[54]"},{"why":"It provides the collection guidelines that motivate the recording setup and annotation practice.","marker":"[20]"},{"why":"It is the annotation platform used for associating each sign with its start and end frames.","marker":"[56]"},{"why":"It supplies the fairness, accountability, transparency, and ethics considerations the authors say they followed.","marker":"[63]"}],"fun_headline_variants":["AzSLD: first Azerbaijani sign language dataset with 30k clips","Azerbaijani sign language dataset: 30,000 videos for AI","New AzSLD: 30k clips for Azeri sign language recognition","AzSLD dataset includes fingerspelling, words, and sentences","First Azerbaijani sign language video dataset for AI translation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the human annotations — the sign labels and the frame alignments — are accurate and consistent enough to serve as ground truth for training and evaluation, and the paper itself concedes that inter-annotator agreement metrics were not exhaustively documented.","fun_headline_variants_meta":{"raw":{"variants":["AzSLD: first Azerbaijani sign language dataset with 30k clips","Azerbaijani sign language dataset: 30,000 videos for AI","New AzSLD: 30k clips for Azeri sign language recognition","AzSLD dataset includes fingerspelling, words, and sentences","First Azerbaijani sign language video dataset for AI translation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000497,"raw_usage":{"total_tokens":2446,"prompt_tokens":963,"completion_tokens":1483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":1389}},"tokens_in":579,"tokens_out":1483,"duration_ms":11821,"temperature":1.0,"reasoning_tokens":1389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:06:19.797497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of about 50 sentence videos, have native Azerbaijani Sign Language signers who were not part of the original annotation process independently re-annotate the labels and frame boundaries, and compare; if label mismatches or boundary disagreements are frequent, the claim of accurate and reliable annotations fails.","supporting_citations":[{"cited_title":"Alishzade, J","cited_arxiv_id":null,"evidence_quote":"It is the released dataset object that makes the resource publicly available under a permissive license."},{"cited_title":"Hasanov, A data loader code for azsl, https://github.com/ ADA-SITE-JML/azsl_dataloader (2023)","cited_arxiv_id":null,"evidence_quote":"It is the accompanying data loader code that standardizes frame extraction, label encoding, and dataset splitting."},{"cited_title":"Hasanov, N","cited_arxiv_id":null,"evidence_quote":"It is the earlier work that produced the fingerspelling collection and a word recognition demo, which the dataset extends."},{"cited_title":"Forster, D","cited_arxiv_id":null,"evidence_quote":"It provides the collection guidelines that motivate the recording setup and annotation practice."},{"cited_title":"URL https://docs.supervisely.com/","cited_arxiv_id":null,"evidence_quote":"It is the annotation platform used for associating each sign with its start and end frames."}],"review_version":1}