{"id":"cfb9955c-5908-4611-9c5b-2bf5265b4690","arxiv_id":"2608.04186","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors present the first holistic conceptual architecture for an LLM-driven electronic explanatory dictionary of Tajik, but the system is neither built nor evaluated.","lead":"This paper proposes a modular, LLM-based design for an electronic explanatory dictionary of Tajik, combining existing morphological tools, word embeddings, and fine-tuned language models. It is a conceptual blueprint, with no implemented prototype or evaluation yet.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The training and quality-assessment modules assume the 2008 explanatory dictionary is machine-readable, but Section 1 says it exists only in printed or limited electronic form.","rationale":"The manuscript is a coherent conceptual proposal, not an empirical result, and I take the central claim to be the availability of enough existing linguistic infrastructure to make the proposed pipeline feasible. The reader is right that the self-cited resources (morpheme database, Tajik Web Corpus, TajikNLP, PEFT benchmark) are the load-bearing substrate and are not independently verified. I would sharpen that into the paper's own internal evidence: the training and evaluation modules presuppose an electronic version of Shukurov et al. (2008), while Section 1 explicitly states the existing dictionaries are printed or limited electronic. This is not an external quibble; it is a dependency that the manuscript itself identifies as missing, and the limitation paragraph in Section 3.6 confirms the labeled pair scarcity. A pilot with synthetic data could still work, but the quality-assessment loop needs gold references, so the architecture as written has an unresolved data dependency. No new verdict is needed beyond the reader's CONDITIONAL; the condition should explicitly require a digitization or gold-reference plan.","tokens_in":17420,"tokens_out":6940,"duration_ms":62597,"concrete_test":"Attempt to obtain a machine-readable copy of the 2008 Explanatory Dictionary of the Tajik Language and build the validation set of 100-200 entries described in Section 3.3 plus the reference entries required by Section 3.4. If fewer than 100 entries can be extracted without launching a new OCR or digitization effort, then the training and quality-assessment modules depend on a resource the paper itself says is unavailable, and the architecture should be revised to include digitization or an independently constructed gold-standard set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 says the training set will use existing explanatory Tajik dictionaries in electronic form, and Section 3.4 says the same dictionaries supply reference entries for BLEU, ROUGE, METEOR, and BERTScore. Section 1, however, describes these dictionaries as available in printed form or limited electronic versions, not supporting dynamic updating and not integrated with automatic text processing. The architecture therefore presupposes a machine-readable lexicographic substrate that the paper itself identifies as missing, and no digitization or OCR step is specified. Section 3.6 partially concedes the problem when it says labeled word-dictionary-entry pairs are not available in sufficient volume and that synthetic training data may be needed. Synthetic generation does not remove the need for a gold reference set for the quality-assessment module; without such a set, BERTScore and the other metrics cannot be anchored to expert lexicographic judgments. So the central claim that the proposal is grounded in existing resources is only partially supported: the corpus and morphological resources are cited, but the lexicographic training and evaluation component depends on an unplanned conversion of a print dictionary.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conceptual architecture for an electronic explanatory dictionary of Tajik, integrating morphological analysis and lemmatization, semantic clustering, LLM-based dictionary entry generation, and quality assessment. It surveys existing Tajik linguistic and corpus resources, justifies subword tokenization and parameter-efficient fine-tuning, and presents an end-to-end illustrative example of processing the word form китобҳоямро. The stated contribution is a holistic blueprint that can ground the first comprehensive digital lexicographic resource for Tajik; no implementation, prototype, or evaluation of generated entries is reported.","tokens_in":17602,"tokens_out":3825,"duration_ms":36914,"significance":"If the proposed architecture were implemented and validated, it could provide a systematic template for LLM-based lexicography in Tajik and, by analogy, other low-resource agglutinative languages. The paper's strengths are its broad and systematic literature survey, a clear modular decomposition, a sensible choice of subword tokenization and PEFT methods given Tajik's morphology and scarce annotated data, and a candid discussion of limitations. The manuscript is also honest that this is a conceptual first stage rather than a finished system. However, the central claim of grounding in existing resources is only partially supported, and the evaluation design presupposes a machine-readable lexicographic substrate that the paper itself says is missing.","major_comments":[{"comment":"The architecture presupposes machine-readable explanatory dictionaries: Section 3.3 says the training set will use existing explanatory Tajik dictionaries in electronic form, and Section 3.4 says the same dictionaries will supply reference entries for BLEU, ROUGE, METEOR, and BERTScore. Yet Section 1 states that existing Tajik explanatory dictionaries are available in printed form or as limited electronic versions that do not support dynamic updating or integration with automatic text processing. The paper never specifies a digitization, OCR, or expert-revision step to convert these dictionaries into the required electronic format. This is a load-bearing gap: without a concrete plan for obtaining the machine-readable lexicographic substrate, the training and quality-assessment modules cannot be realized as described.","section":"Sections 1, 3.3, 3.4"},{"comment":"The limitations discussion admits that labeled word-dictionary-entry pairs for Tajik are not available in sufficient volume and proposes synthetic generation of training pairs. However, synthetic data cannot anchor the automatic metrics in Section 3.4, which require reference dictionary entries to compute BLEU, ROUGE, METEOR, and BERTScore. The paper does not specify how a gold-standard reference set will be built or expert-validated, or how synthetic pairs will be kept distinct from reference entries in evaluation. Without such a protocol, the quality-assessment module has no ground truth against which generated entries can be measured.","section":"Section 3.6"},{"comment":"The feasibility argument depends heavily on a set of author-maintained or co-authored resources (Tajik Web Corpus, TajikNLP, the PEFT benchmark in [Arabov 2026c], and the Soro models) that are cited as arXiv preprints or HuggingFace datasets but are not independently verified in this manuscript. The architecture's modules rely on the existence, completeness, and accessibility of these resources; in particular, the morpheme database (81 prefixes, 76,539 roots, 128,760 postfixes) is a quantitative load-bearing input to the morphological analysis module. The paper should provide a minimal verification plan: exact dataset identifiers, access conditions, and a reproducibility audit of at least the morpheme counts and corpus size, or an explicit statement that these numbers are taken from the cited sources without independent verification.","section":"Sections 2.3 and 3.5"}],"minor_comments":[{"comment":"The generated dictionary entry for китобҳоямро is presented as if it demonstrates the pipeline's output, but it appears to be an illustrative hand-crafted example. This should be stated explicitly, for instance by labeling it 'illustrative output, not produced by the implemented system', to avoid overstating the current level of feasibility.","section":"Section 3.3"},{"comment":"The claim that BERTScore 'better correlates with expert evaluation' for lexicographic tasks is made without a citation or a planned experiment. Either add a supporting reference or frame this as a hypothesis to be tested during the validation stage.","section":"Section 3.4"},{"comment":"The reported POS-tagging benchmark result (weighted F1 = 0.62) is modest, yet the proposed architecture's morphological module feeds directly into entry generation. The paper should discuss the implications of this accuracy level for the downstream dictionary-entry quality and, if relevant, how morphological disambiguation errors will be handled.","section":"Section 2.4"},{"comment":"Many citations are to 2026 preprints and HuggingFace datasets. Before publication, please verify that all cited arXiv identifiers and dataset URLs resolve and that the associated numbers (corpus size, number of aligned sentences, etc.) match the cited versions, since the paper's feasibility claims rely on these resources.","section":"References and data citations"},{"comment":"The semantic clustering module proposes K-means or hierarchical clustering without specifying the number of clusters or a validation criterion. For a conceptual framework this is acceptable, but adding a sentence on how cluster granularity will be chosen (for example, by intrinsic clustering metrics or expert review) would improve reproducibility.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's core resources are almost exclusively the authors' own preprints and HuggingFace datasets. Editors may wish to confirm that these resources are publicly accessible and that the reported corpus and morpheme statistics are reproducible, since the feasibility argument depends on them. The paper is a conceptual framework rather than an implemented system, so the absence of experimental results is not by itself disqualifying; the load-bearing issue is the missing machine-readable lexicographic substrate for training and evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious survey-and-blueprint paper, not an empirical result. It makes a decent case that Tajik has enough morphology, corpora, and LLM components to support a first electronic explanatory dictionary, and the four-module architecture is sensible. What is genuinely new is the assembly: nobody has connected Dovudov's morpheme database, the TajikNLP toolkit, and PEFT-tuned LLMs into a lexicographic pipeline. The worked example for китобҳоямро is a nice touch—it shows the modules actually interlock.\n\nCredit: the related-work section is the most useful part. It documents a substantial Tajik computational-linguistics school and maps global LLM lexicography efforts to it. The limitations section is honest: it admits labeled word–entry pairs are scarce and that expert control is necessary. The architecture is modular enough to be implementable.\n\nSoft spots, in proportion: the paper's biggest internal tension is the one flagged in the stress test. Sections 3.3 and 3.4 treat the 2008 Shukurov dictionary as if it will supply electronic training and reference entries, while Section 1 says existing dictionaries are printed or limited electronic versions. No OCR or digitization step is specified. Section 3.6 partially concedes this by proposing synthetic pairs, but then the BLEU/ROUGE/METEOR/BERTScore evaluation lacks a gold anchor. That is a genuine gap, not a minor one. A serious revision should either add a digitization/conversion module or limit the evaluation claim to expert-judgment plus synthetic-reference comparison.\n\nOther concerns are less severe. The heavy reliance on self-cited TajikNLP, Tajik Web Corpus, and PEFT benchmark results makes the foundation hard to verify from the paper alone, but those resources are at least named with URLs and arXiv IDs. The novelty claim is honest—it is architectural assembly, not a new algorithm. No implementation means no data, so the empirical sections are proposals; the authors say as much.\n\nWho gets value: computational lexicographers and anyone working on low-resource NLP for Iranian languages, plus Tajik NLP researchers who want a roadmap. It deserves a serious referee on the strength of the survey and architecture, with the request that the dictionary-substrate problem be addressed before acceptance. It should not be desk-rejected, but it should not be published as is without a fix.","headline":"A solid, well-written blueprint that assembles existing Tajik NLP resources for a dictionary no one has built; the main hole is that it assumes an electronic dictionary that Section 1 admits does not exist.","tokens_in":18115,"tokens_out":1971,"would_cite":false,"duration_ms":18973,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes the first holistic conceptual architecture for an electronic explanatory dictionary of Tajik, integrating morphological analysis, semantic clustering, LLM-based entry generation, and quality assessment in one pipeline.","keywords":["large language models","electronic explanatory dictionary","Tajik language","low-resource languages","parameter-efficient fine-tuning","subword tokenization","morphological analysis","semantic clustering"],"falsifier":"Attempt to build the proposed prototype against the cited resources and evaluate it on a validation set of 100 to 200 dictionary entries using BERTScore and expert linguists. The central claim would be falsified if the morpheme database cannot be obtained or reproduced, if the reported corpus and toolkit counts differ materially, or if the generated entries consistently fail the expert check despite acceptable automatic metric scores.","tokens_in":17240,"feed_emoji":"📖","tokens_out":6812,"duration_ms":51546,"temperature":0.7,"pith_summary":"The paper argues that Tajik, while still lacking a comprehensive digital explanatory dictionary, has accumulated enough corpus data and morphological descriptions to support building one with modern large language models. It proposes a multi-stage architecture whose modules perform morphological analysis and lemmatization, semantic clustering, LLM-based entry generation, and quality assessment. The stated intent is to unify classical lexicography, statistical text analysis, and generative capabilities into a single system. The paper is explicitly a first conceptual stage: it does not yet report a running prototype or experimental evaluation of generated entries.","feed_headline":"LLM blueprint for a Tajik explanatory dictionary proposed","feed_subtitle":"Morphology, corpus data, and PEFT-tuned models combine to fill Tajik's missing digital lexicographic resource.","key_machinery":"The central object is a four-module pipeline. The morphological analysis and lemmatization module uses the morpheme database of 81 prefixes, 76,539 roots, and 128,760 postfixes together with word-formation classifications to reduce word forms to lemmas. The semantic clustering module maps lemmas to vectors with pretrained Word2Vec and FastText embeddings from the TajikNLP toolkit and groups them into semantic fields. The LLM-based generation module builds structured prompts from the lemma, grammatical labels, and semantic cluster, and fine-tunes an open-weight model with LoRA or QLoRA. The quality assessment module scores generated entries with BLEU, ROUGE, METEOR, and BERTScore, then routes them through expert validation. The architecture also relies on the Tajik Web Corpus of 168.5 million words and the Tajik National Corpus for context and verification, and it justifies subword tokenization and parameter-efficient fine-tuning by the agglutinative nature and high morphological variability of Tajik.","core_discovery":"The paper's central claim is that a complete electronic explanatory dictionary of Tajik is architecturally feasible today, and it specifies how to build one. The load-bearing design choice is to anchor LLM generation on a formal morphological model, with a database of 81 prefixes, 76,539 roots, and 128,760 postfixes, rather than letting the model process Tajik's agglutinative word forms from raw text. Generation is assigned to an open-weight LLM adapted with LoRA or QLoRA, with Mistral 7B plus QLoRA at rank 16 identified as the current benchmark leader at perplexity 5.03 and the Gemma-based Soro models as the alternative; entries are then checked by automatic metrics and expert linguists in an iterative loop. The paper illustrates the intended flow on the word form китобҳоямро, segmented as китоб + -ҳо + -ям + -ро, and shows how a full entry with definition, examples, synonyms, and thematic group would emerge. The authors state plainly that the architecture awaits prototype implementation and experimental comparison.","pith_inferences":["Because the entire data foundation comes from resources attributed to the authors' own prior work, the blueprint stands or falls on whether those resources are publicly available and exactly as described; a reader should treat the architecture as speculative until a prototype runs.","If the pipeline works, it would give a reusable template for other low-resource agglutinative languages that have strong morphological descriptions but no digital dictionary.","The reported script barrier, in which multilingual LLMs degrade sharply on Tajik Cyrillic, suggests that the morphological analysis and tokenizer components, not model scale, may carry most of the performance in this design.","A direct test of the design would be to compare generated entries against a sample of the printed Explanatory Dictionary of the Tajik Language using the same evaluation metrics, to see whether the LLM output is usable or merely resembles lexicography."],"forward_implications":["If implemented as described, Tajik would gain its first comprehensive electronic explanatory dictionary, built largely from existing corpus and morphological resources.","The dictionary would serve as a foundational resource for machine translation, automatic summarization, sentiment analysis, and question-answering systems in Tajik.","The PEFT strategy should permit the generation module to be trained with limited annotated word-entry pairs, supplemented by synthetic pairs and transfer from Persian parallels.","The choice between the Gemma-based Soro models and Mistral 7B would be settled experimentally on a validation set of 100 to 200 entries, judged by BERTScore and expert evaluation.","Automatic metrics alone would not be trusted: expert linguists remain part of the quality loop, so the resulting entries could be vetted lexicographically before publication."],"supporting_citations":[{"why":"Supplies the morpheme database of 81 prefixes, 76,539 roots, and 128,760 postfixes that anchors morphological analysis and lemmatization.","marker":"[Dovudov, 2018]"},{"why":"Supplies the conceptual model of automatic morphological analysis that the first module implements.","marker":"[Usmanov and Dovudov, 2014]"},{"why":"Supplies the TajikNLP toolkit with pretrained Word2Vec and FastText embeddings used for semantic clustering.","marker":"[Arabov et al., 2026]"},{"why":"Supplies the Tajik Web Corpus of 168.5 million words and 1.11 billion characters used for context extraction, examples, and fine-tuning data.","marker":"[Arabov, 2026a]"},{"why":"Supplies a morphologically annotated corpus of 58.4 million word occurrences with 96 percent parsing coverage for verifying grammatical labels.","marker":"[Tajik National Corpus]"},{"why":"Provides the PEFT benchmark result, Mistral 7B with QLoRA rank 16 at perplexity 5.03, that justifies the generation module's base model choice.","marker":"[Arabov, 2026c]"},{"why":"Supplies the Tajik-specialized Soro family of models as the alternative candidate for dictionary entry generation.","marker":"[Liashkov et al., 2026]"},{"why":"Demonstrates that subword-based models can generate dictionary definitions for a low-resource language, supporting the feasibility of the generation module.","marker":"[Bear and Cook, 2021]"},{"why":"Provides the modular dictionary-creation pipeline with optional human intervention that the quality assessment module adapts.","marker":"[Widmann, 2025]"},{"why":"Documents the limitations of LLM-generated lexicographic content, motivating the combination of automatic metrics and expert validation.","marker":"[Jakubicek and Rundell, 2023]"}],"fun_headline_variants":["Tajik dictionary blueprint: LLMs anchored on morphology","LLMs + 76k roots: blueprint for Tajik e-dictionary","First holistic LLM architecture for Tajik dictionary","PEFT-tuned LLMs to fill Tajik lexicographic gap","Morphology-first LLM design for Tajik dictionary"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The architecture assumes that the cited Tajik language resources, the morpheme database of 81 prefixes, 76,539 roots, and 128,760 postfixes, the Tajik Web Corpus, the TajikNLP toolkit, and the PEFT benchmark results, are real, complete, and accessible as described; if any of them are unavailable or not independently reproducible, the data foundation for the dictionary collapses.","fun_headline_variants_meta":{"raw":{"variants":["Tajik dictionary blueprint: LLMs anchored on morphology","LLMs + 76k roots: blueprint for Tajik e-dictionary","First holistic LLM architecture for Tajik dictionary","PEFT-tuned LLMs to fill Tajik lexicographic gap","Morphology-first LLM design for Tajik dictionary"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1493,"prompt_tokens":1029,"completion_tokens":464,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":382}},"tokens_in":645,"tokens_out":464,"duration_ms":3796,"temperature":1.0,"reasoning_tokens":382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:41:38.963210+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Attempt to build the proposed prototype against the cited resources and evaluate it on a validation set of 100 to 200 dictionary entries using BERTScore and expert linguists. The central claim would be falsified if the morpheme database cannot be obtained or reproduced, if the reported corpus and toolkit counts differ materially, or if the generated entries consistently fail the expert check despite acceptable automatic metric scores.","supporting_citations":[],"review_version":1}