Pith. sign in

REVIEW 1 cited by

Knowledge of Pretrained Language Models on Surface Information of Tokens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09808 v2 pith:KW7RDYWG submitted 2024-02-15 cs.CL

classification cs.CL
keywords modelsknowledgelanguagepretrainedtokeninformationregardingsurface
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Do pretrained language models have knowledge regarding the surface information of tokens? We examined the surface information stored in word or subword embeddings acquired by pretrained language models from the perspectives of token length, substrings, and token constitution. Additionally, we evaluated the ability of models to generate knowledge regarding token surfaces. We focused on 12 pretrained language models that were mainly trained on English and Japanese corpora. Experimental results demonstrate that pretrained language models have knowledge regarding token length and substrings but not token constitution. Additionally, the results imply that there is a bottleneck on the decoder side in terms of effectively utilizing acquired knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs reconstruct token spelling in intermediate and later transformer layers rather than reading it directly from embeddings.

Pith tools