{"id":"a39da58c-b334-440e-96f0-de293a28180e","arxiv_id":"2506.23394","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"LoRA fine-tuning of BgGPT models on a bilingual Bulgarian function-calling dataset yields large gains on a self-built 120-case benchmark while keeping knowledge benchmarks stable.","lead":"The paper fine-tunes Bulgarian BgGPT models so they can call external tools, creating TUCAN models with up to 28.75 percentage points higher function-calling accuracy than the base models. It releases the training dataset and an evaluation framework, presenting the recipe as a template for other languages.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported gains rest on a 120-case self-authored eval set with no demonstrated separation from the 10,035-case training set; until leakage/format memorization is ruled out, the headline accuracy numbers are not secure.","rationale":"The reader's weakest assumption is that Tucan-BG-Eval-v1.0 is an unbiased, representative measure of function-calling skill, with no leakage or format overlap with the training set. That is exactly the load-bearing point in my reading as well. The paper's central quantitative claims consist of three accuracy improvements on this 120-case benchmark, so any failure of the benchmark to measure transferable skill directly undermines the headline results. The absence of any reported deduplication or overlap analysis is particularly consequential because both train and eval sets are generated through the same synthetic pipeline and share the same structured tag format. My independent check of Section 6.4 found that the text's reported deviations and averages do not match Table 6, which adds a second layer of uncertainty to the 'no catastrophic forgetting' half of the claim. I do not see deliberate misrepresentation; the inconsistencies and missing overlap checks are more consistent with an early-stage, self-authored evaluation lacking standard safeguards. The released models, dataset, and framework are a genuine asset and make the proposed checks feasible. Because the reader already issued a conditional verdict and my concern reinforces that condition rather than moving it in a new direction, I recommend the verdict remain unchanged.","tokens_in":11969,"tokens_out":3908,"duration_ms":37498,"concrete_test":"Compute the maximum token n-gram overlap and tool-schema similarity between the 120 Tucan-BG-Eval cases and the 10,035 training conversations, then re-run the same 120 semantic tasks with renamed tools and re-parameterized schemas while keeping the Bulgarian queries unchanged; if the TUCAN-vs-base accuracy gap shrinks materially under renamed tools, the reported gains reflect format memorization rather than transferable tool use. Also recompute every Table 6 difference to verify the claimed 'comparable performance' averages.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LoRA fine-tuning adds reliable tool-use skill in Bulgarian without catastrophic forgetting. The evidence for 'reliable tool-use' is the Tucan-BG-Eval-v1.0 score, but that benchmark is the weakest load-bearing element. Section 3.1 describes a hybrid manual/synthetic generation pipeline using GPT-4.1, Gemini 2.5 Pro, and Claude Sonnet 4; Section 5.1.5 describes a 'carefully curated' 120-case eval set across six scenario types. No overlap analysis, n-gram similarity check, or template-distance check is reported between the 10,035 training conversations and the 120 eval cases, and both share the same XML/tool_call format. With only 120 cases, a few memorized examples can move accuracy substantially: 1 case equals 0.83 percentage points, so the 27B improvement is exactly one test case, and the headline 2.6B gain of 28.75 percentage points corresponds to 34.5 cases. If the eval set shares tool schemas, topics, or generation seeds with training, the gains on this benchmark do not demonstrate generalizable function-calling. A second support for the claim is knowledge retention, but Section 6.4's arithmetic is unreliable: the text claims a maximum deviation of 0.0382 and a WinograndeBG gain of +0.0635 for the 2.6B model, while Table 6 shows +0.0053 and +0.0150 respectively, and the stated averages do not match the table either. Thus the 'no forgetting' evidence is not internally consistent. The appropriate response is to treat the accuracy claims as conditional pending a leakage check and external evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TUCAN, a set of LoRA fine-tuned BgGPT-Gemma-2 models (2.6B, 9B, and 27B) trained on a newly constructed bilingual Bulgarian/English function-calling dataset of 10,035 conversations. The authors report accuracy improvements on their own 120-case Tucan-BG-Eval-v1.0 benchmark, up to +28.75 percentage points for the 2.6B model, and claim that core language understanding is preserved on Bulgarian HellaSwag, Winogrande, and ARC benchmarks. The models, dataset, and evaluation framework are released for replication.","tokens_in":12277,"tokens_out":4505,"duration_ms":50036,"significance":"If the reported results are secure, this is a practically useful contribution: it provides a replicable recipe for adding function-calling behavior to a non-English language model family, with released artifacts and a direct base-versus-fine-tuned comparison. The knowledge-retention check on established Bulgarian benchmarks, with standard errors reported, is a genuine strength, as is the open release of the training data and evaluation harness. However, the central quantitative claim depends on a small, self-authored evaluation set with no demonstrated separation from the training data, and the retention analysis in Section 6.4 contains arithmetic inconsistencies. The contribution's significance is therefore conditional on correcting these issues.","major_comments":[{"comment":"The headline function-calling gains rest entirely on the Tucan-BG-Eval-v1.0 dataset of 120 cases, but the paper does not report any overlap or leakage analysis between these cases and the 10,035 training conversations. No n-gram overlap, template-distance, or schema-overlap check is described, and both sets use the same XML/tool_call format. With 120 cases, one case equals 0.83 percentage points, so the 27B improvement is exactly one test case and the 2.6B improvement corresponds to about 34.5 cases. Under these conditions, the reported improvements could reflect format memorization rather than generalizable tool-use competence. Please provide a quantitative leakage analysis, report confidence intervals or run-level variance for the accuracy scores, and ideally validate on an independently constructed or held-out set.","section":"§5.1.5 and §6.1"},{"comment":"The narrative in Section 6.4 is inconsistent with the numbers in Table 6. The text claims a maximum deviation of 0.0382 on HellaSwagBG for the 2.6B model, but Table 6 gives a difference of only 0.0053 for that cell. It also claims a WinograndeBG gain of +0.0635 for the 2.6B model, while the table shows +0.0150. The stated average improvement of +0.0176 for the 2.6B model does not match the table values, which average +0.0059. These discrepancies directly affect the 'no catastrophic forgetting' conclusion and must be corrected, with all derived statements recomputed accordingly.","section":"§6.4 and Table 6"}],"minor_comments":[{"comment":"Section 3.6 appears twice, once for topic distribution and once for message length, and Section 4.3 appears twice, once for the prompt template and once for hyperparameters; please renumber these sections.","section":"§3 and §4"},{"comment":"Table 3 is used both for the hyperparameter configuration and for the function-calling accuracy results; the accuracy table should be renumbered to avoid ambiguity.","section":"§6.1"},{"comment":"The in-text figure references appear to be off by one: Section 6.1 refers to 'Figure 5' for the overall accuracy results, while the corresponding caption is Figure 4, and Section 6.2 refers to 'Figure 6' for scenario breakdown, while the corresponding caption is Figure 5.","section":"§6.1 and §6.2"},{"comment":"In the prompt template, the token `toll_response` appears to be a typo for `tool_response`; please verify whether this is a typo in the manuscript or in the actual template used for training.","section":"§4.3"},{"comment":"The benchmark name 'Winogrande' is misspelled as 'Winograde' in the Introduction and Section 2.1.","section":"§1 and §2.1"},{"comment":"The discussion already acknowledges that the 120-case evaluation represents controlled conditions; this limitation statement is appropriate and should be retained after the leakage analysis is added.","section":"§6.7"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the self-authored 120-case benchmark lands: the absence of any leakage analysis is a load-bearing gap, and the Section 6.4 arithmetic inconsistencies need correction. The paper's central idea remains defensible, and the open releases are a positive feature, but the evidence as presented does not yet support the headline accuracy claims. A major revision with a proper leakage analysis, confidence intervals, and corrected retention numbers would make the contribution publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is a straightforward LoRA fine-tuning application for Bulgarian function calling, with genuinely useful released artifacts. The central accuracy claims, however, rest on a 120-case self-built evaluation set that has no demonstrated separation from the 10,035-case training set. Until that is checked, treat the headline gains as upper bounds.\n\nWhat's new: this is the first open Bulgarian function-calling dataset, model series, and eval framework. The methodology (synthetic multi-model generation, XML/tool_call formatting, LoRA) is standard, but the artifacts are real and can seed work for other low-resource languages. The knowledge-retention check uses external Bulgarian benchmarks (HellaSwagBG, WinograndeBG, ARC-BG), which is the right instinct and gives some confidence that fine-tuning did not wreck the base models. The paper is also candid: Section 6.7 admits the 120-case eval is a “controlled laboratory” and calls for larger-scale and human assessment.\n\nSoft spots, in order of importance. First, the eval set: 120 cases, six scenarios, 20 each. No overlap, n-gram, or template-distance analysis versus training data. Both share the same tag format, so the risk of format memorization is real. With n=120, one case is 0.83 points—the 27B “improvement” is exactly one test case. The 2.6B jump of 28.75 points is 34.5 cases; that could still be real, but without leakage checks the number is not secure. Second, Section 6.4's arithmetic does not match Table 6: the text says max deviation 0.0382 on HellaSwagBG and +0.0635 on WinograndeBG for 2.6B, but the table shows +0.0053 and +0.0150. The stated averages do not match either. This is sloppy and undercuts the “no forgetting” claim even though the table itself looks fine. Third, no external baseline beyond the base BgGPT models—no prompt-engineered BgGPT, no multilingual model with tool support. That limits the claim of “reliable tool use” to a relative gain.\n\nMy verdict aligns with the reader's conditional. The artifacts and direction are valuable; the reported numbers are not yet well-supported. A serious referee should ask for leakage analysis, confidence intervals, and at least one external benchmark before accepting the accuracy headline. For the paper's own stated scope (a practical blueprint for low-resource tool use), the core idea likely holds—the eval just needs to be made trustworthy.\n\nRecommendation: send to peer review. It is a legitimate new application with open artifacts, and the flaws are fixable. I'd cite the dataset if I were working on multilingual tool use.","headline":"Useful Bulgarian function-calling artifacts, but the headline accuracy numbers rest on a 120-case self-built eval with no leakage check; the core idea likely holds.","tokens_in":12816,"tokens_out":2591,"would_cite":true,"duration_ms":26652,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning on a bilingual dataset of 10,035 conversations lifts Bulgarian models' tool-calling accuracy by up to 28.75 percentage points without harming their language benchmarks.","keywords":["function calling","tool use","LoRA fine-tuning","multilingual language models","Bulgarian","low-resource language adaptation","Model Context Protocol","benchmark evaluation"],"falsifier":"Write a new Bulgarian function-calling test set independently of the 10,035 training conversations — fresh functions, fresh queries, no topic or phrasing overlap — and rerun the six scenario types on TUCAN-2.6B and its base model. If the 28.75-point gap collapses toward the base model's level on this held-out set, the improvement is format memorization; if it survives, fine-tuning installed a generalizable skill. A cheaper first step is a near-duplicate scan (n-gram or embedding similarity) between the 120 evaluation cases and the training corpus.","tokens_in":11717,"feed_emoji":"🔧","tokens_out":9666,"duration_ms":86283,"temperature":0.7,"pith_summary":"This paper tries to establish that a language model can be taught to call external tools reliably in a non-English language by continued fine-tuning on a small bilingual dataset, without erasing its existing language skills. Using Bulgarian as the case study, the author adapts the BgGPT model family at three scales, 2.6B, 9B, and 27B parameters, with Low-Rank Adaptation (LoRA) on 10,035 conversations in which tool schemas are in English while user and assistant turns are in Bulgarian. On a 120-case evaluation spanning six tool-use scenarios, the fine-tuned TUCAN models beat their base models by 28.75, 8.34, and 0.83 percentage points, with the smallest model improving most. The broader point is a recipe: if this holds, the same pipeline can extend tool-augmented AI to languages that today's English-centric function-calling ecosystem largely ignores.","feed_headline":"Tool-calling accuracy jumps 28.75% after Bulgarian fine-tuning","feed_subtitle":"LoRA-tuned BgGPT models emit clean, parseable function calls in Bulgarian without losing language skills.","key_machinery":"The load-bearing mechanism is a parameter-efficient continued-training recipe: Low-Rank Adaptation (LoRA) fine-tunes only 0.79% to 1.2% of each BgGPT model's parameters for three epochs. The dataset is the other half: 10,035 bilingual conversations that pair English function definitions with Bulgarian dialogue, mark tools, calls, and results with a lightweight tag convention, and deliberately include negative cases where the correct behavior is to decline a tool or request clarification. The same structured prompt template is enforced at training and inference, and the released Tucan-Eval harness grades outputs into five error classes (no call when expected, unexpected call, wrong function, wrong parameters, malformed JSON). Together these pieces convert 'whether and when to call' from a prompted instruction into a learned behavior.","core_discovery":"The paper's central claim is that specialized fine-tuning, not prompt engineering, is what gives a language model dependable function-calling behavior in Bulgarian. TUCAN-2.6B reaches 78.75% overall accuracy versus 50.00% for its base model; TUCAN-9B reaches 86.67% versus 78.33%; and TUCAN-27B reaches 87.50% versus 86.67%, averaged over multiple runs of the six-scenario benchmark. The author attributes the gains to training the model on when to call a tool, when to decline one, and when to ask for missing parameters, using a fixed prompt template with XML-style tags (<tools>, <tool_code>, <tool_response>). Knowledge-retention checks on HellaSwagBG, WinograndeBG, ARC-Easy-BG, and ARC-Challenge-BG show deviations within roughly ±0.04 points, which the paper reads as measurement noise rather than catastrophic forgetting. The paper further claims a production benefit: TUCAN emits clean, parseable JSON function calls, whereas base models surround their calls with verbose prose that complicates machine consumption.","pith_inferences":["The sharp inverse scaling of the gains suggests the small model was missing a discrete behavioral rule — call versus do not call — that fine-tuning installs; a natural extension is to test whether even smaller models or other language families show the same jump.","The cleanest external check of the claim is a fresh Bulgarian evaluation set written independently of the 10,035 training conversations; if the large gains persist on such a distribution-shifted set, the competence is general rather than template-specific.","The paper keeps tool schemas in English while localizing dialogue; languages with their own developer ecosystems would likely need a second pass that localizes function names and parameter descriptions, which the released dataset format already supports.","The single-model-family comparison leaves open how much of the gain is specific to fine-tuning versus formatting; benchmarking prompt-engineered versions of the same base models would separate those effects."],"forward_implications":["The full pipeline — dataset, models, LoRA adapters, and evaluation framework — is released openly, so another language needs only its own bilingual function-calling corpus to repeat the experiment.","Smaller models gain the most (28.75 points for 2.6B versus 0.83 points for 27B), so the recipe is cheapest exactly where base tool-use ability is weakest.","TUCAN outputs are directly machine-parseable, removing the post-processing step that verbose base-model responses require in production agents.","Residual failures concentrate in wrong or incomplete parameters rather than malformed JSON, identifying parameter extraction as the next bottleneck for training data.","Because the training format follows MCP-style structured tool definitions, the adapted models integrate with standardized agent protocols rather than proprietary interfaces."],"supporting_citations":[{"why":"Supplies the BgGPT base models that TUCAN fine-tunes and the Branch-and-Merge continued-pretraining precedent for the approach.","marker":"(Alexandrov, et al. 2024)"},{"why":"Provides LoRA, the parameter-efficient fine-tuning method claimed to add function-calling while preserving base knowledge.","marker":"(Hu, et al. 2022)"},{"why":"Defines the Model Context Protocol whose structured tool-definition format the training dataset is designed to support.","marker":"(Hou, et al. 2025)"},{"why":"Source of the HellaSwag benchmark whose Bulgarian translation checks commonsense reasoning retention.","marker":"(Zellers, et al. 2019)"},{"why":"Source of the Winograd-style coreference benchmark used in the knowledge-retention suite.","marker":"(Sakaguchi, et al. 2021)"},{"why":"Source of the ARC easy and challenge sets used to check factual-reasoning retention.","marker":"(Clark, et al. 2018)"},{"why":"Provides the evaluation-harness protocols on which the Bulgarian benchmark knowledge-retention runs are based.","marker":"(Gao, et al. 2024)"}],"fun_headline_variants":["Bulgarian tool calls: fine-tuning lifts accuracy 28.75 points","TUCAN models call tools in Bulgarian with 87.5% accuracy","Teaching BgGPT to call tools: 28.75-point accuracy gain","Bulgarian language model learns tool calling via LoRA tuning","Tool use for Bulgarian: fine-tuned models beat base by 28.75%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 120 test cases in the Bulgarian evaluation set are an unbiased measure of function-calling skill — specifically, that they do not overlap with or share a tell-tale format with the 10,035 training conversations, so the reported accuracy gains are genuine tool-use competence rather than memorization of the training template.","fun_headline_variants_meta":{"raw":{"variants":["Bulgarian tool calls: fine-tuning lifts accuracy 28.75 points","TUCAN models call tools in Bulgarian with 87.5% accuracy","Teaching BgGPT to call tools: 28.75-point accuracy gain","Bulgarian language model learns tool calling via LoRA tuning","Tool use for Bulgarian: fine-tuned models beat base by 28.75%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000691,"raw_usage":{"total_tokens":3159,"prompt_tokens":1008,"completion_tokens":2151,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":2051}},"tokens_in":624,"tokens_out":2151,"duration_ms":14524,"temperature":1.0,"reasoning_tokens":2051,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:44:14.184568+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Write a new Bulgarian function-calling test set independently of the 10,035 training conversations — fresh functions, fresh queries, no topic or phrasing overlap — and rerun the six scenario types on TUCAN-2.6B and its base model. If the 28.75-point gap collapses toward the base model's level on this held-out set, the improvement is format memorization; if it survives, fine-tuning installed a generalizable skill. A cheaper first step is a near-duplicate scan (n-gram or embedding similarity) between the 120 evaluation cases and the training corpus.","supporting_citations":[],"review_version":1}