{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2016:IEDUSXUNSRHHWCLZYQR7MYZ6LQ","short_pith_number":"pith:IEDUSXUN","schema_version":"1.0","canonical_sha256":"4107495e8d944e7b0979c423f6633e5c37fd1fe1cfb8fa27e6cc97c6099f2e03","source":{"kind":"arxiv","id":"1609.08144","version":2},"attestation_state":"computed","paper":{"title":"Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"GNMT, a deep LSTM neural machine translation system with wordpieces and coverage penalties, reduces translation errors by an average of 60% compared to phrase-based systems.","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Alex Rudnick, Apurva Shah, Cliff Young, George Kurian, Greg Corrado, Hideto Kazawa, Jason Riesa, Jason Smith, Jeff Klingner, Jeffrey Dean, Keith Stevens, Klaus Macherey, {\\L}ukasz Kaiser, Macduff Hughes, Maxim Krikun, Melvin Johnson, Mike Schuster, Mohammad Norouzi, Nishant Patil, Oriol Vinyals, Qin Gao, Quoc V. Le, Stephan Gouws, Taku Kudo, Wei Wang, Wolfgang Macherey, Xiaobing Liu, Yonghui Wu, Yoshikiyo Kato, Yuan Cao, Zhifeng Chen","submitted_at":"2016-09-26T19:59:55Z","abstract_excerpt":"Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive both in training and in translation inference. Also, most NMT systems have difficulty with rare words. These issues have hindered NMT's use in practical deployments and services, where both accuracy and speed are essential. In this work, we present GNMT, Google's Neural Machine Translation system, which attempts to address many of"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"1609.08144","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2016-09-26T19:59:55Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"7e1f30af6bb78f3b01e138dd4c95c0225405b345dc64dcb2de9ce5d5e167c230","abstract_canon_sha256":"f8728cf6bc353f04780df8a769aa5b6792264ff8dd0e04d7bfd682bf2a9cfe92"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-04T21:27:34.571619Z","signature_b64":"RFTXQOnaSHojP4PC7kEGtQVuitgLs4KM7YVo1VrJKM24itkhJcq9lDTRDolFfpUUXtE2IJWapssYbpsIJ6j2Cw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4107495e8d944e7b0979c423f6633e5c37fd1fe1cfb8fa27e6cc97c6099f2e03","last_reissued_at":"2026-07-04T21:27:34.571107Z","signature_status":"signed_v1","first_computed_at":"2026-07-04T21:27:34.571107Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"GNMT, a deep LSTM neural machine translation system with wordpieces and coverage penalties, reduces translation errors by an average of 60% compared to phrase-based systems.","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Alex Rudnick, Apurva Shah, Cliff Young, George Kurian, Greg Corrado, Hideto Kazawa, Jason Riesa, Jason Smith, Jeff Klingner, Jeffrey Dean, Keith Stevens, Klaus Macherey, {\\L}ukasz Kaiser, Macduff Hughes, Maxim Krikun, Melvin Johnson, Mike Schuster, Mohammad Norouzi, Nishant Patil, Oriol Vinyals, Qin Gao, Quoc V. Le, Stephan Gouws, Taku Kudo, Wei Wang, Wolfgang Macherey, Xiaobing Liu, Yonghui Wu, Yoshikiyo Kato, Yuan Cao, Zhifeng Chen","submitted_at":"2016-09-26T19:59:55Z","abstract_excerpt":"Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive both in training and in translation inference. Also, most NMT systems have difficulty with rare words. These issues have hindered NMT's use in practical deployments and services, where both accuracy and speed are essential. In this work, we present GNMT, Google's Neural Machine Translation system, which attempts to address many of"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Using a human side-by-side evaluation on a set of isolated simple sentences, it reduces translation errors by an average of 60% compared to Google's phrase-based production system.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the measured gains are attributable to the described architectural choices (attention placement, wordpieces, coverage penalty) rather than differences in training data scale or compute, and that results on simple sentences generalize to complex, domain-specific text.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"GNMT deploys 8-layer LSTMs with attention, wordpieces, low-precision inference, and coverage-penalized beam search to match state-of-the-art on WMT'14 En-Fr and En-De while cutting translation errors by 60% in human evaluations.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"GNMT, a deep LSTM neural machine translation system with wordpieces and coverage penalties, reduces translation errors by an average of 60% compared to phrase-based systems.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"bc17a0358a1506ab42c63d7038ea35f80b9ae4eef5a72712acaa7fe7cafe4360"},"source":{"id":"1609.08144","kind":"arxiv","version":2},"verdict":{"id":"4882b53b-1857-42cf-9d4a-e0d4055e262b","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-12T15:16:34.439988Z","strongest_claim":"Using a human side-by-side evaluation on a set of isolated simple sentences, it reduces translation errors by an average of 60% compared to Google's phrase-based production system.","one_line_summary":"GNMT deploys 8-layer LSTMs with attention, wordpieces, low-precision inference, and coverage-penalized beam search to match state-of-the-art on WMT'14 En-Fr and En-De while cutting translation errors by 60% in human evaluations.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the measured gains are attributable to the described architectural choices (attention placement, wordpieces, coverage penalty) rather than differences in training data scale or compute, and that results on simple sentences generalize to complex, domain-specific text.","pith_extraction_headline":"GNMT, a deep LSTM neural machine translation system with wordpieces and coverage penalties, reduces translation errors by an average of 60% compared to phrase-based systems."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/1609.08144/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":44,"sample":[{"doi":"","year":2016,"title":"G., Steiner, B., Tucker, P., V asudevan, V., W arden, P., Wicke, M., Yu, Y., and Zheng, X","work_id":"4a906750-f28c-420f-ab94-4be981a5cb73","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2015,"title":"Neural machine translation by jointly learning to align and translate","work_id":"cdbe1816-fb6e-4e5b-a76b-bf2ab8eff431","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":1988,"title":"Brown, P., Cocke, J., Pietra, S. D., Pietra, V. D., Jelinek, F., Mercer, R., and Roossin, P. A statistical approach to language translation. InProceedings of the 12th Conference on Computational Lingu","work_id":"c9f3c1a1-e17d-4115-9006-a26acfd498d1","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":1990,"title":"F., Cocke, J., Pietra, S","work_id":"8cb8dd79-9ad3-4933-9a22-18b29574d8a4","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":1993,"title":"Brown, P. F., Pietra, V. J. D., Pietra, S. A. D., and Mercer, R. L. The mathematics of statistical machine translation: Parameter estimation.Comput. Linguist. 19, 2 (June 1993), 263–311","work_id":"dd4f9230-a684-4065-9c21-b5bc9dd315bd","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":44,"snapshot_sha256":"aff6e61d831b44bf7b6732cc387e18d561d3f875eb3d6d2b94780ab6da9504d9","internal_anchors":4},"formal_canon":{"evidence_count":2,"snapshot_sha256":"6705fa0ad9d0d2ef0b9e33476f38147a33a2d7958e51f254f62e2b15b2c8788b"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"1609.08144","created_at":"2026-07-04T21:27:34.571169+00:00"},{"alias_kind":"arxiv_version","alias_value":"1609.08144v2","created_at":"2026-07-04T21:27:34.571169+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.1609.08144","created_at":"2026-07-04T21:27:34.571169+00:00"},{"alias_kind":"pith_short_12","alias_value":"IEDUSXUNSRHH","created_at":"2026-07-04T21:27:34.571169+00:00"},{"alias_kind":"pith_short_16","alias_value":"IEDUSXUNSRHHWCLZ","created_at":"2026-07-04T21:27:34.571169+00:00"},{"alias_kind":"pith_short_8","alias_value":"IEDUSXUN","created_at":"2026-07-04T21:27:34.571169+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":87,"internal_anchor_count":87,"sample":[{"citing_arxiv_id":"2607.06818","citing_title":"Ad Headline Generation using Self-Critical Masked Language Model","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25432","citing_title":"Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation","ref_index":57,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17358","citing_title":"OTRO: Oblivious Tokenization Path with Square-Root ORAM","ref_index":54,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12348","citing_title":"MATLAB-Based Layerwise Self-Adaptive Physics-Informed Neural Network in Applications to Multidimensional Coupled Burgers' Equations with High Reynolds Numbers","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.08604","citing_title":"Gryphon: A Unified Architecture for Semantic-ID Generation and Item-Level Scoring in Industrial Recommendations","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2606.08728","citing_title":"Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery","ref_index":108,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25432","citing_title":"Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation","ref_index":56,"is_internal_anchor":true},{"citing_arxiv_id":"2605.24842","citing_title":"Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01172","citing_title":"Revisiting Neural Processes via Fourier Transform and Volterra Series","ref_index":114,"is_internal_anchor":true},{"citing_arxiv_id":"1906.08996","citing_title":"Incremental Adaptation of NMT for Professional Post-editors: A User Study","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09000","citing_title":"Demonstration of a Neural Machine Translation System with Online Learning for Translators","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"1906.11018","citing_title":"Integration of TensorFlow based Acoustic Model with Kaldi WFST Decoder","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09302","citing_title":"Neural Machine Translating from Natural Language to SPARQL","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09444","citing_title":"Retrieving Sequential Information for Non-Autoregressive Neural Machine Translation","ref_index":37,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09675","citing_title":"Evaluating the Supervised and Zero-shot Performance of Multi-lingual Translation Models","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"1906.09795","citing_title":"Conversational Response Re-ranking Based on Event Causality and Role Factored Tensor Event Embedding","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"1906.11024","citing_title":"Sharing Attention Weights for Fast Transformer","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"1906.11751","citing_title":"The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation","ref_index":30,"is_internal_anchor":true},{"citing_arxiv_id":"1907.01686","citing_title":"Machine Reading Comprehension: a Literature Review","ref_index":68,"is_internal_anchor":true},{"citing_arxiv_id":"1907.00570","citing_title":"Do Transformer Attention Heads Provide Transparency in Abstractive Summarization?","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"1907.00874","citing_title":"System Misuse Detection via Informed Behavior Clustering and Modeling","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"1907.01300","citing_title":"Learning to Reformulate the Queries on the WEB","ref_index":42,"is_internal_anchor":true},{"citing_arxiv_id":"1907.03040","citing_title":"BERT-DST: Scalable End-to-End Dialogue State Tracking with Bidirectional Encoder Representations from Transformer","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"1907.04648","citing_title":"EPNAS: Efficient Progressive Neural Architecture Search","ref_index":46,"is_internal_anchor":true},{"citing_arxiv_id":"2104.05565","citing_title":"Survey on reinforcement learning for language processing","ref_index":139,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ","json":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ.json","graph_json":"https://pith.science/api/pith-number/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/graph.json","events_json":"https://pith.science/api/pith-number/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/events.json","paper":"https://pith.science/paper/IEDUSXUN"},"agent_actions":{"view_html":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ","download_json":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ.json","view_paper":"https://pith.science/paper/IEDUSXUN","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=1609.08144&json=true","fetch_graph":"https://pith.science/api/pith-number/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/graph.json","fetch_events":"https://pith.science/api/pith-number/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/action/storage_attestation","attest_author":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/action/author_attestation","sign_citation":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/action/citation_signature","submit_replication":"https://pith.science/pith/IEDUSXUNSRHHWCLZYQR7MYZ6LQ/action/replication_record"}},"created_at":"2026-07-04T21:27:34.571169+00:00","updated_at":"2026-07-04T21:27:34.571169+00:00"}