Pith. sign in

REVIEW 1 cited by

Character-based NMT with Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.04997 v1 pith:JB2LBVHI submitted 2019-11-12 cs.CL

classification cs.CL
keywords character-basedtransformermodelstexttranslatingwhenadvantagesappealing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Character-based translation has several appealing advantages, but its performance is in general worse than a carefully tuned BPE baseline. In this paper we study the impact of character-based input and output with the Transformer architecture. In particular, our experiments on EN-DE show that character-based Transformer models are more robust than their BPE counterpart, both when translating noisy text, and when translating text from a different domain. To obtain comparable BLEU scores in clean, in-domain data and close the gap with BPE-based models we use known techniques to train deeper Transformer models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HarMoEny: Efficient Multi-GPU Inference of MoE Models

    cs.DC 2025-06 conditional novelty 5.0 of 10

    HarMoEny uses dynamic token redistribution and asynchronous expert prefetching to achieve near-perfect GPU load balance in multi-GPU MoE inference.

Pith tools