{"id":"da015849-1335-41a7-8f40-4212d0c45d95","arxiv_id":"1906.09675","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Task-specific decoder parameters outperform fully shared decoder parameters in both supervised and zero-shot multilingual translation performance.","lead":"This paper evaluates methods for sharing decoder parameters in multilingual neural machine translation models on supervised and zero-shot tasks across 110 directions using only WMT 2019 data. A smart generalist might read it to understand practical trade-offs when building systems that translate between many languages without parallel data for every pair.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Zero-shot performance claims rest on repurposed unsupervised MT metrics whose correlation to actual quality on these pairs is unverified.","rationale":"The reader's weakest assumption matches the load-bearing point exactly. The supervised results alone are not described as driving the headline claim; the paper's emphasis on zero-shot diversity makes the proxy-metric validity central. No other internal inconsistency (parameter counts, training data, etc.) is visible from the provided text.","tokens_in":1618,"tokens_out":326,"duration_ms":14339,"concrete_test":"Identify any zero-shot pair in the paper that has an independent held-out parallel test set (even if small); recompute both the paper's repurposed metric and standard BLEU on that gold set for the task-specific and shared-decoder models; if the ranking reverses or the gap shrinks below significance, the concern lands.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim compares task-specific vs. fully shared decoder parameters on translation performance across 110 directions. Supervised directions use standard WMT test sets, but the paper explicitly states it repurposes unsupervised MT evaluation methods (e.g., round-trip or back-translation consistency) for the many zero-shot pairs lacking gold parallels. For the outperformance result to be load-bearing, these proxy metrics must preserve relative model rankings. If the proxies systematically favor models with extra task-specific parameters (more capacity for the proxy task), the headline finding could be an artifact. The abstract positions the zero-shot scale and diversity as a key contribution, so this assumption directly supports the claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript evaluates several methods for full or partial sharing of decoder parameters in multilingual NMT models trained on WMT 2019 parallel datasets. It assesses both supervised performance on standard test sets and zero-shot performance on 110 directions, using repurposed unsupervised MT evaluation methods for pairs without gold-standard parallels. The central finding is that models with task-specific decoder parameters outperform those with fully shared decoder parameters across these settings.","tokens_in":1749,"tokens_out":435,"duration_ms":31586,"significance":"If the results hold, the work provides a large-scale empirical comparison of decoder parameter sharing strategies in multilingual translation, highlighting trade-offs and supporting the use of task-specific parameters. The diversity of evaluated zero-shot pairs is notable, though dependent on the validity of the proxy metrics.","major_comments":[{"comment":"The paper relies on repurposed unsupervised MT metrics (such as round-trip or back-translation consistency) for zero-shot pairs lacking gold data. However, there is no verification that these proxies correlate with actual translation quality or preserve model rankings between different decoder sharing configurations. Since models with task-specific parameters have more capacity, they may perform better on the proxy tasks artifactually, undermining the load-bearing outperformance claim for zero-shot translation.","section":"Zero-shot evaluation methods"},{"comment":"Details on data balancing, language pair selection criteria, and controls for per-direction training data volume are needed to confirm that performance differences are attributable to decoder sharing rather than imbalances in the 110-direction setup.","section":"Experimental setup"}],"minor_comments":[{"comment":"The claim of conducting the 'largest evaluation' would be strengthened by quantitative comparison of training data volume and zero-shot pair count against prior multilingual NMT studies.","section":"Abstract"},{"comment":"Notation for the different decoder sharing configurations (task-specific vs. fully shared) could be clarified with a table or diagram in the methods section.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our evaluation of decoder parameter sharing in multilingual NMT. We address each major comment below and will incorporate revisions where appropriate to strengthen the manuscript.","responses":[{"response":"We agree that the manuscript does not contain an explicit verification (e.g., correlation analysis) of the proxy metrics against gold-standard BLEU or human judgments on directions where both are available. These proxies are drawn from established unsupervised MT evaluation practices, and the same trend of task-specific decoder superiority appears in our supervised results (where gold data exists). Nevertheless, to directly address the concern about capacity bias and ranking preservation, we will add a new subsection that computes proxy-to-BLEU correlations on the supervised test sets and checks whether the relative ordering of models is preserved under the proxies.","revision_made":"yes","referee_comment":"[Zero-shot evaluation methods] The paper relies on repurposed unsupervised MT metrics (such as round-trip or back-translation consistency) for zero-shot pairs lacking gold data. However, there is no verification that these proxies correlate with actual translation quality or preserve model rankings between different decoder sharing configurations. Since models with task-specific parameters have more capacity, they may perform better on the proxy tasks artifactually, undermining the load-bearing outperformance claim for zero-shot translation."},{"response":"The current manuscript summarizes the use of WMT 2019 parallel data but does not provide exhaustive per-direction statistics or explicit balancing procedures. We will expand the experimental setup section with the requested details: language-pair selection criteria from WMT 2019, any data balancing or upsampling applied during training, and tables or text reporting training data volume per direction to allow readers to assess whether differences are due to decoder sharing rather than data imbalance.","revision_made":"yes","referee_comment":"[Experimental setup] Details on data balancing, language pair selection criteria, and controls for per-direction training data volume are needed to confirm that performance differences are attributable to decoder sharing rather than imbalances in the 110-direction setup."}],"tokens_in":1249,"tokens_out":443,"duration_ms":26216,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper runs the biggest reported comparison so far of decoder parameter sharing strategies across 110 translation directions, all trained on WMT 2019 data. Models with task-specific decoder parameters beat fully shared ones on the supervised directions, which use standard test sets. That part of the result looks straightforward and worth noting for anyone building these systems.","headline":"Large-scale decoder-sharing comparison in multilingual NMT shows task-specific parameters win on supervised tests, but zero-shot results hinge on unverified proxy metrics.","tokens_in":2232,"tokens_out":149,"would_cite":false,"duration_ms":14698,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Multilingual NMT decoder-sharing study is orthogonal to RS framework","alignment":"orthogonal","rationale":"The paper's central machinery (task-specific vs. shared decoder parameters in transformer NMT, zero-shot evaluation via pivoting/back-translation proxies on WMT data) operates entirely within empirical NLP/ML. It contains no reference to, or structural parallel with, J-cost functions, ratio symmetry, phi-ladder identities, 8-tick periodicity, or any forcing chain from a single distinction. RS modules such as Cost.FunctionalEquation, Foundation.RealityFromDistinction, and Foundation.DimensionForcing have no bearing on parameter-sharing trade-offs or BLEU-based rankings.","tokens_in":52187,"confidence":"high","tokens_out":158,"duration_ms":4275,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Models with task-specific decoder parameters outperform those with fully shared decoders across supervised and zero-shot multilingual translation tasks.","keywords":["multilingual NMT","decoder parameter sharing","zero-shot translation","supervised translation","WMT shared task"],"falsifier":"Human evaluation or new gold parallel test sets for several zero-shot pairs that directly compare BLEU or other automatic scores against human judgments of translation adequacy.","tokens_in":2518,"feed_emoji":"🌐","tokens_out":520,"duration_ms":22880,"temperature":0.7,"pith_summary":"The paper trains multilingual neural machine translation models on WMT 2019 parallel data and compares different ways of sharing decoder parameters across translation directions. It measures performance in 110 unique directions, including many zero-shot pairs that lack direct training data, by adapting evaluation techniques from unsupervised machine translation. The central result is that allowing some decoder parameters to remain unique to each task produces higher quality output than forcing all decoder parameters to be identical across tasks. This finding addresses a practical design choice in building systems that must handle many languages at once without separate models for each pair.","feed_headline":"Task-specific decoder parameters improve multilingual MT results","feed_subtitle":"Across 110 directions on WMT 2019 data, models with dedicated decoder components beat fully shared decoders in supervised and zero-shot use.","key_machinery":"Methods for full or partial sharing of decoder parameters in multilingual NMT, where task-specific parameters allow separate adaptation per translation direction while shared parameters capture cross-lingual patterns.","core_discovery":"Models which have task-specific decoder parameters outperform models where decoder parameters are fully shared across all tasks.","pith_inferences":["Designs that keep a modest number of decoder parameters private per task could reduce the need for separate models in production multilingual systems.","The same partial-sharing pattern may apply to encoder parameters or other components if similar ablation studies were run."],"forward_implications":["Partial decoder sharing yields better results than full sharing in both supervised and zero-shot settings.","Trade-offs exist between the amount of parameter sharing and translation quality across the 110 directions tested.","The approach scales to large training data volumes while maintaining gains from task-specific components."],"fun_headline_variants":["Task-specific decoder parameters outperform shared in multilingual MT","Task-specific decoders outperform fully shared in 110 directions","Dedicated decoder parameters outperform shared across all tasks","Models with task-specific decoders outperform shared in MT evaluation"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Repurposed evaluation methods from unsupervised machine translation accurately reflect true zero-shot translation quality for language pairs without gold-standard parallel data.","fun_headline_variants_meta":{"raw":{"variants":["Task-specific decoder parameters outperform shared in multilingual MT","Task-specific decoders outperform fully shared in 110 directions","Dedicated decoder parameters outperform shared across all tasks","Models with task-specific decoders outperform shared in MT evaluation"]},"model":"grok-4.3","cost_usd":0.006517,"raw_usage":{"total_tokens":2988,"prompt_tokens":547,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":65174500,"prompt_tokens_details":{"text_tokens":547,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2381,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":547,"tokens_out":60,"duration_ms":19044,"temperature":1.0,"reasoning_tokens":2381,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T18:01:49.532347+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Human evaluation or new gold parallel test sets for several zero-shot pairs that directly compare BLEU or other automatic scores against human judgments of translation adequacy.","supporting_citations":[],"review_version":1}