{"id":"b83c3c5f-9606-42d0-acac-bfb7a5eb377b","arxiv_id":"2606.17299","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Word2Vec on Toki Pona shows distributional patterns suffice for semantic structure even at extreme vocabulary reduction, and incidental non-core tokens tighten rather than disrupt clusters.","lead":"The paper trains Word2Vec on 1.4 million Toki Pona sentences to test semantic embeddings in a language with roughly 130 words. A smart generalist might read it to see how little vocabulary is needed before distributional models stop working.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Evaluation metrics may measure embedding consistency rather than independent semantic capture","rationale":"The reader's weakest assumption directly identifies the same evaluation-validity issue as the load-bearing concern. Because the abstract alone was available to the reader, the UNVERDICTED verdict is appropriate; the concrete test above would allow a more definitive assessment without requiring changes to the reported experiments.","tokens_in":1689,"tokens_out":350,"duration_ms":32654,"concrete_test":"Obtain pairwise semantic similarity ratings for 50–100 Toki Pona word pairs from at least three fluent speakers or linguists familiar with the language; compute Spearman correlation between these ratings and the cosine similarities from both the noisy and clean embeddings. If the correlation is near zero or statistically indistinguishable from a frequency-matched random baseline, the claim that semantics are captured is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that Word2Vec captures semantic relationships via distributional patterns even at ~130-word lexicon size—requires the three quantitative methods to reflect semantics rather than corpus artifacts. Proximity to category centroids presupposes independently defined categories whose validity is not cross-checked against external criteria. Silhouette scores from agglomerative clustering on the embeddings themselves quantify clusterability of the learned space, which any sufficiently consistent co-occurrence model would produce. Representational similarity matrices versus English compare two embedding spaces without establishing that either aligns with ground-truth semantics. The noisy-vs-clean ablation shows only that incidental tokens do not disrupt relative structure; it does not test whether that structure encodes meaning beyond frequency patterns. With a constructed language whose words have deliberately broad, overlapping senses, high intra-category proximity and cross-lingual RSM similarity can arise from usage regularities alone.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript trains Word2Vec on a corpus of 1.4 million Toki Pona sentences (7.95 million tokens) drawn from community sources to test whether semantic embeddings can be learned from an extremely small ~130-word lexicon. Two models are compared (one retaining incidental non-Toki Pona tokens, one filtering them), and evaluation uses three methods: proximity of words to semantic category centroids, silhouette scores from agglomerative clustering, and representational similarity matrices versus English embeddings. The central claim is that Word2Vec captures semantic relationships via distributional patterns even at this extreme lower bound, that incidental tokens do not disrupt relative structure (and may draw similar words closer), and that lexicon size is less important than distributional statistics.","tokens_in":1854,"tokens_out":523,"duration_ms":44529,"significance":"If the quantitative evaluations are shown to measure semantics rather than corpus artifacts and the numerical results support the claims, the work would provide evidence that Word2Vec remains effective under extreme vocabulary reduction, using Toki Pona's deliberately broad senses as a strong test case. This could clarify the minimal conditions for distributional semantics and the robustness of embedding methods to lexicon size.","major_comments":[{"comment":"Abstract: the abstract states results from training and three evaluation approaches but provides no numerical values, error bars, statistical tests, or details on how category centroids were defined; the central claim therefore rests on unshown quantitative support.","section":"Abstract"},{"comment":"Evaluation methods: proximity to semantic category centroids presupposes independently defined categories whose validity is not cross-checked against external criteria; with Toki Pona's broad, overlapping senses, this risks capturing usage regularities alone rather than independent semantic capture.","section":"Evaluation methods"},{"comment":"Evaluation methods: silhouette scores from agglomerative clustering on the embeddings themselves quantify clusterability of the learned space, which any sufficiently consistent co-occurrence model would produce, without establishing alignment with ground-truth semantics.","section":"Evaluation methods"}],"minor_comments":[{"comment":"The noisy-vs-clean ablation shows only that incidental tokens do not disrupt relative structure; it does not test whether that structure encodes meaning beyond frequency patterns.","section":"Abstract"},{"comment":"The representational similarity matrices versus English compare two embedding spaces without establishing that either aligns with ground-truth semantics.","section":"Evaluation methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight important aspects of our evaluation approach. We address each major comment below and indicate where revisions will be made to strengthen the manuscript.","responses":[{"response":"We agree that the abstract would be strengthened by including key numerical results. In the revised version, we will add specific values such as mean distances to category centroids, silhouette scores with standard deviations, and details on the definition of the 10 semantic categories (drawn from the official Toki Pona dictionary). Where appropriate, we will report statistical tests comparing the two models.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the abstract states results from training and three evaluation approaches but provides no numerical values, error bars, statistical tests, or details on how category centroids were defined; the central claim therefore rests on unshown quantitative support."},{"response":"The categories were constructed from the canonical Toki Pona word list and documented community usage to reflect the language's deliberately broad senses. To address the concern about external validation, we will add a mapping of these categories to English semantic equivalents and report alignment with resources such as WordNet synsets. This provides an independent check while preserving the focus on Toki Pona's reduced lexicon.","revision_made":"partial","referee_comment":"[Evaluation methods] Evaluation methods: proximity to semantic category centroids presupposes independently defined categories whose validity is not cross-checked against external criteria; with Toki Pona's broad, overlapping senses, this risks capturing usage regularities alone rather than independent semantic capture."},{"response":"We acknowledge that silhouette scores primarily assess internal structure. In the manuscript, these scores are interpreted in conjunction with the predefined semantic categories and the representational similarity analysis against English embeddings, which serves as an external semantic reference. We will revise the methods and discussion sections to explicitly state how the three evaluation approaches are combined to link clusterability to semantic alignment rather than generic co-occurrence patterns.","revision_made":"partial","referee_comment":"[Evaluation methods] Evaluation methods: silhouette scores from agglomerative clustering on the embeddings themselves quantify clusterability of the learned space, which any sufficiently consistent co-occurrence model would produce, without establishing alignment with ground-truth semantics."}],"tokens_in":1413,"tokens_out":492,"duration_ms":29854,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"Toki Pona gives a clean lower-bound test for static embeddings because its vocabulary is fixed at around 130 words by design. The paper trains Word2Vec on 1.4 million sentences and compares a clean model against one that keeps the 23 percent noisy tokens. The main result is that the noise does not break relative structure and may even tighten some clusters.\n\nThe controlled noise ablation is the clearest addition. Most Word2Vec studies skip this isolation of incidental tokens, and running it on an extreme small-vocab constructed language is a focused extension. The three checks—centroid proximity, silhouette scores from clustering, and representational similarity against English—also give more angles than the usual single metric.\n\nThe soft spots are in how those checks are interpreted. All three stay inside the learned space or compare two learned spaces. Category centroids assume the categories are valid on their own; silhouette scores simply measure how well the vectors group, which follows from any consistent co-occurrence model; the RSM against English shows similarity between two embedding spaces without tying either to external meaning. Toki Pona's broad, overlapping senses make it especially plausible that usage patterns alone produce these results. The abstract supplies no numbers, error bars, or statistical tests, so effect sizes remain unknown.\n\nThis is mainly for people working on low-resource or minimal-vocabulary embeddings. A reader testing whether distributional methods bottom out at very small lexicons will find the ablation useful. The experiment is straightforward and the edge case is real, so the paper deserves a serious referee even though the metric claims will need tightening.","headline":"Word2Vec still produces clusterable structure on Toki Pona's 130-word lexicon through co-occurrence, but the three evaluation methods stay internal to the embedding space and do not independently confirm semantic capture.","tokens_in":2298,"tokens_out":406,"would_cite":false,"duration_ms":36887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Word2Vec captures semantic structure in Toki Pona's 130-word lexicon when trained on large text volumes.","keywords":["Word2Vec","Toki Pona","word embeddings","distributional semantics","low-resource languages","vocabulary size","corpus noise","semantic clustering"],"falsifier":"Finding that the similarity matrices or cluster structures for the Toki Pona embeddings differ markedly from English patterns in a manner attributable to vocabulary size alone, or that removing incidental tokens produces substantially worse clustering scores.","tokens_in":2595,"feed_emoji":"","tokens_out":732,"duration_ms":43480,"temperature":0.7,"pith_summary":"The paper tests Word2Vec on Toki Pona, a constructed language limited to roughly 130 words, to determine if embeddings can still reflect semantic relationships. Researchers gathered 1.4 million sentences totaling 7.95 million tokens and trained two versions of the model, one keeping incidental non-Toki Pona tokens and one removing them. They measured performance through word proximity to category centroids, silhouette scores from agglomerative clustering, and representational similarity matrices compared to English embeddings. Results show the model succeeds by relying on distributional patterns in usage rather than requiring a large vocabulary, and that sparse extra tokens tighten clusters without altering overall structure. This setup directly probes the lower limits of what word embeddings need to function.","feed_headline":"Word2Vec embeds 130-word Toki Pona effectively","feed_subtitle":"Large corpus shows distributional patterns suffice for semantics even at extreme vocabulary reduction.","key_machinery":"Two parallel Word2Vec models on the same Toki Pona corpus, one retaining and one filtering non-core tokens, evaluated via semantic category centroid proximity, agglomerative clustering silhouette scores, and English representational similarity matrices.","core_discovery":"Word2Vec successfully generates embeddings that capture semantic relationships in Toki Pona despite its extreme vocabulary reduction to approximately 130 words. Training on 1.4 million sentences reveals that effectiveness stems primarily from distributional patterns in the corpus rather than lexicon size. Retaining incidental tokens such as named entities and loanwords draws similar words closer together in vector space while leaving the relative embedding structure intact, as confirmed by centroid proximity measures, agglomerative clustering silhouette scores, and similarity matrices aligned with English.","pith_inferences":["The findings point toward testing the same approach on other constructed languages or pidgins to see if corpus size consistently overrides lexicon limits.","One could examine whether the observed benefit from incidental tokens holds when the extra items are systematically varied in type or frequency.","This setup invites direct comparison with other embedding methods to check if the pattern is specific to Word2Vec or general to distributional models.","Results suggest that data collection efforts for small languages should prioritize volume of usage examples over vocabulary expansion."],"forward_implications":["Embeddings remain stable in relative structure when sparse non-core tokens are retained in training data.","Incidental tokens improve the tightness of similar-word clusters without harming overall organization.","Semantic relationships in embeddings arise chiefly from co-occurrence statistics rather than total unique word count.","Word2Vec can be applied to other minimal-vocabulary constructed or low-resource languages given adequate text volume."],"fun_headline_variants":["Word2Vec succeeds with Toki Pona's 130 words","Toki Pona tests Word2Vec at vocabulary extreme","Distributional patterns power Word2Vec in Toki Pona","Extra tokens tighten embeddings in Toki Pona Word2Vec","Word2Vec captures semantics despite Toki Pona's size"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The chosen quantitative metrics of centroid proximity, silhouette scores, and similarity matrices to English measure genuine semantic capture instead of artifacts from the small vocabulary or mixed corpus composition.","fun_headline_variants_meta":{"raw":{"variants":["Word2Vec succeeds with Toki Pona's 130 words","Toki Pona tests Word2Vec at vocabulary extreme","Distributional patterns power Word2Vec in Toki Pona","Extra tokens tighten embeddings in Toki Pona Word2Vec","Word2Vec captures semantics despite Toki Pona's size"]},"model":"grok-4.3","cost_usd":0.004109,"raw_usage":{"total_tokens":2091,"prompt_tokens":680,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":41087000,"prompt_tokens_details":{"text_tokens":680,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1327,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":680,"tokens_out":84,"duration_ms":19407,"temperature":1.0,"reasoning_tokens":1327,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T02:58:58.476697+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding that the similarity matrices or cluster structures for the Toki Pona embeddings differ markedly from English patterns in a manner attributable to vocabulary size alone, or that removing incidental tokens produces substantially worse clustering scores.","supporting_citations":[],"review_version":1}