Pith. sign in
Pith Number

pith:TQTJPPBT

pith:2015:TQTJPPBT3U5DVCSV7KLEZWAR6I
not attested not anchored not stored refs resolved

Neural Machine Translation of Rare Words with Subword Units

Alexandra Birch, Barry Haddow, Rico Sennrich

Encoding rare words as subword sequences enables open-vocabulary neural machine translation.

arxiv:1508.07909 v5 · 2015-08-31 · cs.CL

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{TQTJPPBT3U5DVCSV7KLEZWAR6I}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

subword models improve over a back-off dictionary baseline for the WMT 15 translation tasks English-German and English-Russian by 1.1 and 1.3 BLEU, respectively.

C2weakest assumption

That segmenting words into subword units via byte pair encoding or character n-grams preserves enough linguistic structure for the neural model to learn accurate translations of rare words without introducing excessive ambiguity or length.

C3one line summary

Subword segmentation via byte pair encoding enables open-vocabulary neural machine translation and improves BLEU scores by 1.1 on English-German and 1.3 on English-Russian WMT 2015 tasks over dictionary back-off baselines.

References

36 extracted · 36 resolved · 0 Pith anchors

[1] Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate . In Proceedings of the International Conference on Learning Representat 2015
[2] Issam Bazzi and James R. Glass. 2000. Modeling out-of-vocabulary words for robust speech recognition . In Sixth International Conference on Spoken Language Processing, ICSLP 2000 / INTERSPEECH 2000 , 2000
[3] Botha and Phil Blunsom 2014
[4] Rohan Chitnis and John DeNero. 2015. Variable-Length Word Encodings for Neural Translation Models . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP) 2015
[5] Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder--Decoder for Statisti 2014

Cited by

58 papers in Pith

Receipt and verification
First computed 2026-07-04T21:04:06.804399Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

9c2697bc33dd3a3a8a55fa964cd811f236b952e599ed006333bb29fc4f96efeb

Aliases

arxiv: 1508.07909 · arxiv_version: 1508.07909v5 · doi: 10.48550/arxiv.1508.07909 · pith_short_12: TQTJPPBT3U5D · pith_short_16: TQTJPPBT3U5DVCSV · pith_short_8: TQTJPPBT
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/TQTJPPBT3U5DVCSV7KLEZWAR6I \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 9c2697bc33dd3a3a8a55fa964cd811f236b952e599ed006333bb29fc4f96efeb
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "eb4c8c60b6f1e36b8bdf8179098349e27db1aeb18eb99d3ec08aa6ffe2fb8153",
    "cross_cats_sorted": [],
    "license": "http://creativecommons.org/licenses/by/4.0/",
    "primary_cat": "cs.CL",
    "submitted_at": "2015-08-31T16:37:31Z",
    "title_canon_sha256": "6722be1c762778c2577b341007623b5971b87983909ff213c2a3d9e1558fd668"
  },
  "schema_version": "1.0",
  "source": {
    "id": "1508.07909",
    "kind": "arxiv",
    "version": 5
  }
}