REVIEW 3 major objections 4 minor 101 references
RALS: Resources and Baselines for Romanian Automatic Lexical Simplification
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper claims that a transparent, dictionary-based pipeline outperforms large language models on Romanian lexical simplification, and releases the first joint complexity-prediction and simplification dataset for the language.
desk verdict First Romanian LCP/LS dataset plus a hybrid system that looks competitive, but the system comparison needs a declared split and uncertainty bounds before the 'consistently outperforms' claim holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
DexFlex is the carrying mechanism: it uses part-of-speech and morphological analysis of the target word, retrieves synonym candidates from a Romanian dictionary by matching the sentence's BERT contextual embedding to cached embeddings of dictionary example sentences via approximate nearest-neighbour search, then inflects the chosen synonyms to match the sentence context in number, gender, and person. The ranking gold standard comes from a Bradley–Terry model applied to human pairwise simplicity judgments, replacing the usual frequency-based candidate ranking.
What would settle it
A direct test would take a set of polysemous Romanian words in context, run DexFlex's retrieval, and compare the selected synonyms against human gold senses; if retrieval frequently selects synonyms from the wrong sense, the MAP/Potential advantage would likely collapse on unrestricted text. Additionally, expanding the RoLS candidate collection to the RoLCP original-text samples and re-evaluating DexFlex would show whether the reported scores transfer beyond the human-translated genre.
Extended reading notes
Core claim
The paper's central claim is that DexFlex, a quasi-rule-based Lexical Simplification system, achieves the best ranking and coverage scores for Romanian when compared against open-weight and closed-source LLMs and a fine-tuned BERT model. Its novelty lies in combining a morphological tagger, a dictionary of synonyms with cached example-sentence embeddings, and approximate nearest-neighbour retrieval to select context-appropriate synonyms, followed by rule-based inflection and a handcrafted-feature complexity re-ranker. The authors further show that the three LCP datasets they release — human-translated, word-translated, and original Romanian text — reach Pearson correlations of 0.56, 0.68, an
Load-bearing premise
The claim rests on the assumption that matching a sentence's contextual BERT embedding to the nearest cached example embedding in the dictionary reliably identifies the intended word sense and yields suitable synonyms; the authors acknowledge they did not run exhaustive word-sense-disambiguation evaluation.
Editorial extensions
If this is right
- If the claim holds, Romanian text simplification can proceed without relying on closed LLMs, supporting privacy-preserving applications in health, education, and government.
- The released dataset enables training and evaluation of LCP models for Romanian, with published baseline correlations for comparison.
- The pairwise ranking methodology can be reused to build simplification gold standards for other under-resourced languages.
- DexFlex's competitive performance suggests that dictionary- and rule-based systems should be considered strong baselines in multilingual simplification shared tasks.
- The finding that cross-lingual complexity transfer distorts annotations cautions against using translated data as a proxy for original-language complexity.
Reading between the lines
- If the DexFlex advantage generalizes beyond the human-translated subset to unrestricted Romanian texts, it implies that a curated dictionary plus a small amount of human pairwise judgment may deliver simplification quality comparable to expensive LLM APIs in other under-resourced languages — a testable claim for each new language.
- The paper's data show LLMs still edge out DexFlex on accuracy at the top gold candidate, suggesting a hybrid that uses LLM suggestions but re-ranks them with a dictionary-grounded complexity model could outperform either alone; the paper does not test this combination.
- Because simplification candidates were only collected for the human-translated subset, the LS benchmark covers a narrow genre; expanding RoLS to original-text RoLCP sentences would reveal whether the ranking methodology survives varied text types and whether DexFlex retains its edge.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RALS, a Romanian resource combining lexical complexity prediction (LCP) annotations for 3,921 word-in-context samples (HT, WT, and RoLCP subsets) with a lexical substitution dataset (RoLS) for the HT subset. The authors propose a Bradley–Terry ranking procedure to order substitution candidates from pairwise simplicity judgments, evaluate a Ridge-regressor LCP baseline under grouped cross-validation and cross-dataset transfer, and compare several simplification systems — Apertus-8B, RoLlama-8B, GPT-4o, a fine-tuned Romanian BERT, and the dictionary-based DexFlex — on MAP, Potential, and accuracy metrics. The central claim is that DexFlex consistently outperforms all other approaches on MAP and Potential metrics, and that this constitutes a strong non-LLM baseline for Romanian lexical simplification.
Significance. If the evaluation is sound, the resource contribution is valuable: Romanian currently lacks LCP/LS datasets, and the construction methodology — frequency sampling, sense-clustering for sentence selection, grouped cross-validation, and human-ranked substitution gold data — is careful and well contextualized against shared-task results. The paper also ships open code and data, and its Limitations section honestly discloses the moderate inter-annotator agreement, the HT-only scope of RoLS, and the unevaluated WSD component. The claim that a transparent dictionary-based pipeline can beat prompt-based LLMs on ranking and coverage metrics is interesting and practically relevant for low-resource settings. However, the headline comparison is not yet fully supported because the evaluation protocol for DexFlex's LCP-based ranker is underspecified and no uncertainty quantification is provided.
major comments (3)
- [Section 4, Table 3] The DexFlex row uses "the LCP pipeline ... to rank the candidate synonyms," but the manuscript does not state how that LCP ranker was trained for the Lexical Simplification evaluation. The LCP experiments in Table 2 use grouped 10-fold cross-validation; no analogous split is described for the candidate-ranking protocol. If the Ridge regressor was fitted on HT LCP labels (the same sentences that define the RoLS test set), the DexFlex comparisons are trained-on-test. Please specify the training data for this ranker, whether any HT or RoLCP labels were seen, and the exact split/grouping used in computing Table 3.
- [Section 4, Table 3] No confidence intervals, standard deviations, or significance tests are reported for any MAP/Potential/ACC entry. The lead over RoLlama is only 0.01 at MAP@1 (0.41 vs 0.40); without uncertainty estimates the phrase "consistently outperforms" is not supported. Please report bootstrap or per-sample paired statistics, clustering by the 190 HT sentences.
- [Appendix A] The WSD component is central to DexFlex, yet the appendix states "we did not run exhaustive word-sense-disambiguation evaluation." The reliability of the nearest-neighbor sense matching is asserted only through "summary evaluations." Because the MAP/Potential superiority of DexFlex depends on retrieving the correct sense from dexonline, this is not a peripheral limitation. Please provide a quantitative WSD evaluation on a representative sample (including polysemous/rare words), report coverage of dexonline entries, and discuss failure cases.
minor comments (4)
- [References] Reference formatting is inconsistent: "Codrut, et al., 2024" should be normalized, "Horacio Saggio" in the Štajner et al. (2024) reference should be "Horacio Saggion," and "DellFlex" in the Limitations section should be "DexFlex."
- [Section 4, Evaluation] The definitions of MAP@N, Potential@N, and Accuracy@N@top_gold_1 are imprecise. MAP@N is not simply "precision" but mean average precision over candidate lists; Potential@N and the "top_gold_1" notation should be defined with formulas or a precise reference.
- [Abstract / Conclusions] The phrase "first text simplification system for Romanian" overstates the contribution: the implemented and evaluated component is lexical simplification (candidate substitution), not full sentence-level text simplification. Consider qualifying the claim.
- [Section 3.2] The sentence "The final list is the union of both sets of suggestions, verified by a third annotator" would benefit from a statement of how disagreements with the third annotator were resolved and how the union was pruned.
Circularity Check
No significant circularity: the central claims rest on new annotated data and external baselines; the only mild self-citation is not load-bearing, and the DexFlex evaluation-split gap is a correctness concern, not a circular reduction.
full rationale
The paper does not exhibit a circular derivation chain. The LCP models are evaluated with grouped 10-fold cross-validation (Table 2), and the RoLS gold rankings are produced from independent pairwise human simplicity judgments using an external Bradley-Terry methodology (Jerdee and Newman, 2024). DexFlex's use of the LCP pipeline to rank candidate synonyms is a system component rather than a fitted quantity presented as a prediction, and the paper does not state that this ranker was trained on the HT subset that constitutes the RoLS test set; the missing train/test split is an evaluation-protocol weakness, not a definitional circularity. The self-citations to Cristea and Nisioi (2024) are used as a negative baseline and as a caution about LLM prompting, so they are not load-bearing support for the paper's main claims. Appendix A's admission that exhaustive WSD evaluation was not run is a limitation, not a circular step. The paper is largely self-contained against external benchmarks and its central contribution, the RALS dataset and the DexFlex comparison, does not reduce by construction to its own inputs.
Assumptions & free parameters
assumptions (4)
- standard math Logistic Bradley-Terry model with Gaussian prior variance 1/2 converts pairwise simplicity judgments into a ranking.
- domain assumption Human translation of the MultiLS English samples preserves enough lexical-complexity comparability to anchor Romanian annotations, despite documented distribution shifts.
- domain assumption Dexonline together with Romanian BERT contextual embeddings provides reliable synonyms and word-sense disambiguation.
- domain assumption University-student annotators (ages 20-33) provide complexity and simplicity judgments that can serve as a benchmark for broader populations.
Cite this review
Pith. "Pith review of RALS: Resources and Baselines for Romanian Automatic Lexical Simplification." pith.science (2026). https://pith.science/paper/55QPG5LP
@misc{pith2026260720078,
author = {Pith},
title = {Pith review of: RALS: Resources and Baselines for Romanian Automatic Lexical Simplification},
year = {2026},
howpublished = {\url{https://pith.science/paper/55QPG5LP}},
note = {Machine review of arXiv:2607.20078}
}
read the original abstract
We introduce the first dataset that jointly covers both lexical complexity prediction (LCP) annotations and lexical simplification (LS) for Romanian, along with a comparison of lexical simplification approaches. We propose a methodology for ordering simplification suggestions using a pairwise ranking approximation method, arranging candidates from simple to complex based on a separate set of human judgments. In addition, we provide human lexical complexity annotations for 3,921 word samples in context. Finally, we explore several novel pipelines for complexity prediction and simplification and present the first text simplification system for Romanian.
Figures
Reference graph
Works this paper leans on
-
[1]
Semeval-2012 task 1: English lexical simplification , author=. * SEM 2012: The First Joint Conference on Lexical and Computational Semantics--Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012) , pages=
2012
-
[2]
S em E val 2016 Task 11: Complex Word Identification
Paetzold, Gustavo and Specia, Lucia. S em E val 2016 Task 11: Complex Word Identification. Proceedings of the 10th International Workshop on Semantic Evaluation ( S em E val-2016). 2016. doi:10.18653/v1/S16-1085
-
[3]
A Report on the Complex Word Identification Shared Task 2018
Yimam, Seid Muhie and Biemann, Chris and Malmasi, Shervin and Paetzold, Gustavo and Specia, Lucia and. A Report on the Complex Word Identification Shared Task 2018. Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications. 2018. doi:10.18653/v1/W18-0507
-
[4]
Findings of the TSAR -2022 Shared Task on Multilingual Lexical Simplification
Saggion, Horacio and S tajner, Sanja and Ferr \'e s, Daniel and Sheang, Kim Cheng and Shardlow, Matthew and North, Kai and Zampieri, Marcos. Findings of the TSAR -2022 Shared Task on Multilingual Lexical Simplification. Proceedings of the Workshop on Text Simplification, Accessibility, and Readability (TSAR-2022). 2022. doi:10.18653/v1/2022.tsar-1.31
-
[5]
Proceedings of the Third Workshop on Text Simplification, Accessibility and Readability. 2024
2024
-
[6]
S em E val-2021 Task 1: Lexical Complexity Prediction
Shardlow, Matthew and Evans, Richard and Paetzold, Gustavo Henrique and Zampieri, Marcos. S em E val-2021 Task 1: Lexical Complexity Prediction. Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021). 2021. doi:10.18653/v1/2021.semeval-1.1
-
[7]
Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA) , year=
Shardlow, Matthew and Alva-Manchego, Fernando and Batista-Navarro, Riza and Bott, Stefan and Calderon Ramirez, Saul and Cardon, Rémi and François, Thomas and Hayakawa, Akio and Horbach, Andrea and Huelsing, Anna and Ide, Yusuke and Imperial, Joseph Marvin and Nohejl, Adam and North, Kai and Occhipinti, Laura and Peréz Rojas, Nelson and Raihan, Nishat and ...
-
[8]
An Extensible Massively Multilingual Lexical Simplification Pipeline Dataset using the M ulti LS Framework
Shardlow, Matthew and Alva-Manchego, Fernando and Batista-Navarro, Riza and Bott, Stefan and Calderon Ramirez, Saul and Cardon, R. An Extensible Massively Multilingual Lexical Simplification Pipeline Dataset using the M ulti LS Framework. Proceedings of the 3rd Workshop on Tools and Resources for People with REAding DIfficulties (READI) @ LREC-COLING 2024. 2024
2024
Show all 101 references
-
[9]
Computational Processing of the Portuguese Language: 9th International Conference, PROPOR 2010, Porto Alegre, RS, Brazil, April 27-30, 2010
Translating from complex to simplified sentences , author=. Computational Processing of the Portuguese Language: 9th International Conference, PROPOR 2010, Porto Alegre, RS, Brazil, April 27-30, 2010. Proceedings 9 , pages=. 2010 , organization=
2010
-
[10]
Exploring Neural Text Simplification Models
Nisioi, Sergiu and S tajner, Sanja and Ponzetto, Simone Paolo and Dinu, Liviu P. Exploring Neural Text Simplification Models. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2017. doi:10.18653/v1/P17-2014
2017 doi
-
[11]
L e SS : A Computationally-Light Lexical Simplifier for S panish
Stajner, Sanja and Ibanez, Daniel and Saggion, Horacio. L e SS : A Computationally-Light Lexical Simplifier for S panish. Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing. 2023
2023
-
[12]
ArXiv , year=
Albert Qiaochu Jiang and Alexandre Sablayrolles and Arthur Mensch and Chris Bamford and Devendra Singh Chaplot and Diego de Las Casas and Florian Bressand and Gianna Lengyel and Guillaume Lample and Lucile Saulnier and L'elio Renard Lavaud and Marie-Anne Lachaux and Pierre Sto...
-
[13]
2022 , version =
Chase, Harrison , title =. 2022 , version =
2022
-
[14]
Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining , pages=
Xgboost: A scalable tree boosting system , author=. Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining , pages=
-
[15]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 5: Tutorial Abstracts) , pages=
Automatic and Human-AI Interactive Text Generation (with a focus on Text Simplification and Revision) , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 5: Tutorial Abstracts) , pages=
-
[16]
On the Ethical Considerations of Text Simplification
Gooding, Sian. On the Ethical Considerations of Text Simplification. Ninth Workshop on Speech and Language Processing for Assistive Technologies (SLPAT-2022). 2022. doi:10.18653/v1/2022.slpat-1.7
2022 doi
-
[17]
Proceedings of the AAAI-98 workshop on integrating artificial intelligence and assistive technology , pages=
Practical simplification of English newspaper text to assist aphasic readers , author=. Proceedings of the AAAI-98 workshop on integrating artificial intelligence and assistive technology , pages=. 1998 , organization=
1998
-
[18]
Proceedings of the SIGIR workshop on accessible search systems , pages=
Text simplification for children , author=. Proceedings of the SIGIR workshop on accessible search systems , pages=
-
[19]
Simplifying Lexical Simplification: Do We Need Simplified Corpora?
Glava s , Goran and S tajner, Sanja. Simplifying Lexical Simplification: Do We Need Simplified Corpora?. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2:...
2015 doi
-
[20]
Personalizing Lexical Simplification
Lee, John and Yeung, Chak Yan. Personalizing Lexical Simplification. Proceedings of the 27th International Conference on Computational Linguistics. 2018
2018
-
[21]
One Size Does Not Fit All: The Case for Personalised Word Complexity Models
Gooding, Sian and Tragut, Manuel. One Size Does Not Fit All: The Case for Personalised Word Complexity Models. Findings of the Association for Computational Linguistics: NAACL 2022. 2022. doi:10.18653/v1/2022.findings-naacl.27
2022 doi
-
[22]
Controllable Lexical Simplification for E nglish
Sheang, Kim Cheng and Ferr \'e s, Daniel and Saggion, Horacio. Controllable Lexical Simplification for E nglish. Proceedings of the Workshop on Text Simplification, Accessibility, and Readability (TSAR-2022). 2022. doi:10.18653/v1/2022.tsar-1.19
2022 doi
-
[23]
Language Resources and Evaluation , volume=
Predicting lexical complexity in English texts: the Complex 2.0 dataset , author=. Language Resources and Evaluation , volume=. 2022 , publisher=
2022
-
[24]
ACM Computing Surveys , volume=
Lexical complexity prediction: An overview , author=. ACM Computing Surveys , volume=. 2023 , publisher=
2023
-
[25]
Journal of Intelligent Information Systems , pages=
Deep learning approaches to lexical simplification: A survey , author=. Journal of Intelligent Information Systems , pages=. 2024 , publisher=
2024
-
[26]
Combining Expert Knowledge with Frequency Information to Infer CEFR Levels for Words
Pintard, Alice and Fran c ois, Thomas. Combining Expert Knowledge with Frequency Information to Infer CEFR Levels for Words. Proceedings of the 1st Workshop on Tools and Resources to Empower People with REAding DIfficulties (READI). 2020
2020
-
[27]
Frontiers in Communication , volume=
Automatic text simplification for German , author=. Frontiers in Communication , volume=. 2022 , publisher=
2022
-
[28]
Controlled and Balanced Dataset for J apanese Lexical Simplification
Kodaira, Tomonori and Kajiwara, Tomoyuki and Komachi, Mamoru. Controlled and Balanced Dataset for J apanese Lexical Simplification. Proceedings of the ACL 2016 Student Research Workshop. 2016
2016
-
[29]
Japanese Lexical Complexity for Non-Native Readers: A New Dataset
Ide, Yusuke and Mita, Masato and Nohejl, Adam and Ouchi, Hiroki and Watanabe, Taro. Japanese Lexical Complexity for Non-Native Readers: A New Dataset. Proceedings of the Eighteenth Workshop on Innovative Use of NLP for Building Educational Applications
-
[30]
Hartmann, Nathan Siegle and Alu. Adapta. Linguam
-
[31]
IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=
Chinese lexical simplification , author=. IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=. 2021 , publisher=
2021
-
[32]
ALEXSIS : A Dataset for Lexical Simplification in S panish
Ferr \'e s, Daniel and Saggion, Horacio. ALEXSIS : A Dataset for Lexical Simplification in S panish. Proceedings of the Thirteenth Language Resources and Evaluation Conference. 2022
2022
-
[33]
Plos one , volume=
EASIER corpus: A lexical simplification resource for people with cognitive impairments , author=. Plos one , volume=. 2023 , publisher=
2023
-
[34]
27th International Conference on Computational Linguistics (COLING 2018) , year=
ReSyf: a French lexicon with ranked synonyms , author=. 27th International Conference on Computational Linguistics (COLING 2018) , year=
2018
-
[35]
Proceedings of the ACL-IJCNLP 2015 Student Research Workshop , pages=
Evaluation dataset and system for Japanese lexical simplification , author=. Proceedings of the ACL-IJCNLP 2015 Student Research Workshop , pages=
2015
-
[36]
Language learning , volume=
Universals of lexical simplification , author=. Language learning , volume=. 1978 , publisher=
1978
-
[37]
Researching translation in the age of technology and global conflict , pages=
Corpus linguistics and translation studies*: Implications and applications , author=. Researching translation in the age of technology and global conflict , pages=. 2019 , publisher=
2019
-
[38]
Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies , pages=
Translationese and its dialects , author=. Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies , pages=
-
[39]
Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
On the Similarities Between Native, Non-native and Translated Texts , author=. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[40]
High-Confidence Computing , pages=
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly , author=. High-Confidence Computing , pages=. 2024 , publisher=
2024
-
[41]
Findings of the Association for Computational Linguistics: NAACL 2024 , pages=
RoDia: A New Dataset for Romanian Dialect Identification from Speech , author=. Findings of the Association for Computational Linguistics: NAACL 2024 , pages=
2024
-
[42]
A Novel Cartography-Based Curriculum Learning Method Applied on R o NLI : The First R omanian Natural Language Inference Corpus
Poesina, Eduard and Caragea, Cornelia and Ionescu, Radu. A Novel Cartography-Based Curriculum Learning Method Applied on R o NLI : The First R omanian Natural Language Inference Corpus. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
2024 doi
-
[43]
RETUYT - INCO at MLSP 2024: Experiments on Language Simplification using Embeddings, Classifiers and Large Language Models
Sastre, Ignacio and Alfonso, Leandro and Fleitas, Facundo and Gil, Federico and Lucas, Andr \'e s and Spoturno, Tom \'a s and G \'o ngora, Santiago and Ros \'a , Aiala and Chiruzzo, Luis. RETUYT - INCO at MLSP 2024: Experiments on Language Simplification using Embeddings, Clas...
2024
-
[44]
TMU - HIT at MLSP 2024: How Well Can GPT -4 Tackle Multilingual Lexical Simplification?
Enomoto, Taisei and Kim, Hwichan and Hirasawa, Tosho and Nagai, Yoshinari and Sato, Ayako and Nakajima, Kyotaro and Komachi, Mamoru. TMU - HIT at MLSP 2024: How Well Can GPT -4 Tackle Multilingual Lexical Simplification?. Proceedings of the 19th Workshop on Innovative Use of N...
2024
-
[45]
Archaeology at MLSP 2024: Machine Translation for Lexical Complexity Prediction and Lexical Simplification
Cristea, Petru and Nisioi, Sergiu. Archaeology at MLSP 2024: Machine Translation for Lexical Complexity Prediction and Lexical Simplification. Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024). 2024
2024
-
[46]
A Lexical Simplification Tool for Promoting Health Literacy
Zilio, Leonardo and Braga Paraguassu, Liana and Leiva Hercules, Luis Antonio and Ponomarenko, Gabriel and Berwanger, Laura and Bocorny Finatto, Maria Jos \'e. A Lexical Simplification Tool for Promoting Health Literacy. Proceedings of the 1st Workshop on Tools and Resources to...
2020
-
[47]
Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=
Automatic text simplification for social good: Progress and challenges , author=. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=
2021
-
[48]
Shardlow, Matthew and Alva-Manchego, Fernando and Batista-Navarro, Riza and Bott, Stefan and Calderon Ramirez, Saul and Cardon, Rémi and François, Thomas and Hayakawa, Akio and Horbach, Andrea and Huelsing, Anna and Ide, Yusuke and Imperial, Joseph Marvin and Nohejl, Adam and ...
-
[49]
arXiv preprint arXiv:2402.14972 , year=
MultiLS: A Multi-task Lexical Simplification Framework , author=. arXiv preprint arXiv:2402.14972 , year=
-
[50]
Behavior research methods, instruments, & computers , volume=
MRC psycholinguistic database: Machine-usable dictionary, version 2.00 , author=. Behavior research methods, instruments, & computers , volume=. 1988 , publisher=
1988
-
[51]
, author=
Handbook of semantic word norms. , author=. 1978 , publisher=
1978
-
[52]
, author=
The teacher's word book of 30,000 words. , author=. 1944 , publisher=
1944
-
[53]
Behavior Research Methods, Instruments, & Computers , volume=
A frequency count of 190,000 words in the London-Lund Corpus of English Conversation , author=. Behavior Research Methods, Instruments, & Computers , volume=. 1984 , publisher=
1984
-
[54]
, author=
Concreteness, imagery, and meaningfulness values for 925 nouns. , author=. Journal of experimental psychology , volume=. 1968 , publisher=
1968
-
[55]
W ord N et: A Lexical Database for E nglish
Miller, George A. W ord N et: A Lexical Database for E nglish. H uman L anguage T echnology: Proceedings of a Workshop held at P lainsboro, N ew J ersey, M arch 8-11, 1994. 1994
1994
-
[56]
doi:10.5281/zenodo.10009823 , url =
Ines Montani and Matthew Honnibal and Matthew Honnibal and Adriane Boyd and Sofie Van Landeghem and Henning Peters , title =. doi:10.5281/zenodo.10009823 , url =
-
[57]
Bird, Steven and Klein, Ewan and Loper, Edward , year=
-
[58]
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16) , pages=
A corpus of native, non-native and translated texts , author=. Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16) , pages=
- [59]
-
[60]
(No Title) , year=
Computational analysis of present-day American English , author=. (No Title) , year=
-
[61]
Educational research bulletin , pages=
A formula for predicting readability: Instructions , author=. Educational research bulletin , pages=. 1948 , publisher=
1948
-
[62]
Linguistic databases , year=
The use of a psycholinguistic database in the simplification of text for aphasic readers , author=. Linguistic databases , year=
-
[63]
Behavior research methods , volume=
The Glasgow Norms: Ratings of 5,500 words on nine scales , author=. Behavior research methods , volume=. 2019 , publisher=
2019
-
[64]
, author=
The meaning-familiarity relationship. , author=. Psychological Review , volume=. 1953 , publisher=
1953
-
[65]
Journal of Verbal Learning and Verbal Behavior , volume=
Age-of-acquisition norms for 220 picturable nouns , author=. Journal of Verbal Learning and Verbal Behavior , volume=. 1973 , publisher=
1973
-
[66]
Journal of Verbal Learning and Verbal Behavior , volume=
Parameters of abstraction, meaningfulness, and pronunciability for 329 nouns , author=. Journal of Verbal Learning and Verbal Behavior , volume=. 1966 , publisher=
1966
-
[67]
Behavior research methods & instrumentation , volume=
Age-of-acquisition, imagery, concreteness, familiarity, and ambiguity measures for 1,944 words , author=. Behavior research methods & instrumentation , volume=. 1980 , publisher=
1980
-
[68]
Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications , pages=
CAMB at CWI shared task 2018: Complex word identification with ensemble-based voting , author=. Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications , pages=
2018
-
[69]
and Varoquaux, G
Pedregosa, F. and Varoquaux, G. and Gramfort, A. and Michel, V. and Thirion, B. and Grisel, O. and Blondel, M. and Prettenhofer, P. and Weiss, R. and Dubourg, V. and Vanderplas, J. and Passos, A. and Cournapeau, D. and Brucher, M. and Perrot, M. and Duchesnay, E. , journal=
-
[70]
C o C o: A Tool for Automatically Assessing Conceptual Complexity of Texts
Stajner, Sanja and Nisioi, Sergiu and Hulpu s , Ioana. C o C o: A Tool for Automatically Assessing Conceptual Complexity of Texts. Proceedings of the Twelfth Language Resources and Evaluation Conference. 2020
2020
-
[71]
R o B o C o P : A Comprehensive RO mance BO rrowing CO gnate Package and Benchmark for Multilingual Cognate Identification
Dinu, Liviu and Uban, Ana and Cristea, Alina and Dinu, Anca and Iordache, Ioan-Bogdan and Georgescu, Simona and Zoicas, Laurentiu. R o B o C o P : A Comprehensive RO mance BO rrowing CO gnate Package and Benchmark for Multilingual Cognate Identification. Proceedings of the 202...
2023 doi
-
[72]
Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence , pages =
Speer, Robyn and Chin, Joshua and Havasi, Catherine , title =. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence , pages =. 2017 , publisher =
2017
-
[73]
“Vorbești Rom
Masala, Mihai and Ilie-Ablachim, Denis and Dima, Alexandru and Corlatescu, Dragos Georgian and Zavelca, Miruna-Andreea and Olaru, Ovio and Terian, Simina-Maria and Terian, Andrei and Leordeanu, Marius and Velicu, Horia and others , booktitle=. “Vorbești Rom
-
[74]
Unsupervised Lexical Simplification with Context Augmentation
Wada, Takashi and Baldwin, Timothy and Lau, Jey. Unsupervised Lexical Simplification with Context Augmentation. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.627
2023 doi
-
[75]
`` Geen makkie '' : Interpretable Classification and Simplification of D utch Text Complexity
Hobo, Eliza and Pouw, Charlotte and Beinborn, Lisa. `` Geen makkie '' : Interpretable Classification and Simplification of D utch Text Complexity. Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023). 2023. doi:10.18653/v1/...
2023 doi
-
[76]
Thirty-Fourth AAAI Conference on Artificial Intelligence , pages=
Lexical Simplification with Pretrained Encoders , author =. Thirty-Fourth AAAI Conference on Artificial Intelligence , pages=
-
[77]
Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=
The birth of Romanian BERT , author=. Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=
2020
-
[78]
EMNLP 2017 , pages=
An Adaptable Lexical Simplification Architecture for Major Ibero-Romance Languages , author=. EMNLP 2017 , pages=
2017
-
[79]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Lexical simplification with pretrained encoders , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[80]
IEEE/ACM transactions on audio, speech, and language processing , volume=
Lsbert: Lexical simplification based on bert , author=. IEEE/ACM transactions on audio, speech, and language processing , volume=. 2021 , publisher=
2021
-
[81]
LSL lama: Fine-Tuned LL a MA for Lexical Simplification
Baez, Anthony and Saggion, Horacio. LSL lama: Fine-Tuned LL a MA for Lexical Simplification. Proceedings of the Second Workshop on Text Simplification, Accessibility and Readability. 2023
2023
-
[82]
Procesamiento del Lenguaje Natural , volume=
Multilingual Controllable Transformer-Based Lexical Simplification , author=. Procesamiento del Lenguaje Natural , volume=. 2023 , publisher=
2023
-
[83]
LEX enstein: A Framework for Lexical Simplification
Paetzold, Gustavo and Specia, Lucia. LEX enstein: A Framework for Lexical Simplification. Proceedings of ACL - IJCNLP 2015 System Demonstrations. 2015. doi:10.3115/v1/P15-4015
2015 doi
-
[84]
Proceedings-International Conference on Computational Linguistics, COLING , volume=
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification , author=. Proceedings-International Conference on Computational Linguistics, COLING , volume=
-
[85]
U ni HD at TSAR -2022 Shared Task: Is Compute All We Need for Lexical Simplification?
Aumiller, Dennis and Gertz, Michael. U ni HD at TSAR -2022 Shared Task: Is Compute All We Need for Lexical Simplification?. Proceedings of the Workshop on Text Simplification, Accessibility, and Readability (TSAR-2022). 2022. doi:10.18653/v1/2022.tsar-1.28
2022 doi
-
[86]
and Newman, M
Jerdee, M. and Newman, M. E. J. , title =. Science Advances , year =. doi:10.1126/sciadv.adn2654 , url =
-
[87]
Hiroki Nakayama and Takahiro Kubo and Junya Kamura and Yasufumi Taniguchi and Xu Liang , year=
-
[88]
Hollenstein, Nora and Kasperė, Ramunė and Jäger, Lena A. , year=. Opportunities and Challenges in the MultiplEYE Data Collection , url=. doi:10.23668/psycharchives.15479 , publisher=
-
[89]
Resources in Underrepresented Languages: Building a Representative R omanian Corpus
Midrigan - Ciochina, Ludmila and Boyd, Victoria and Sanchez-Ortega, Lucila and Malancea \_ Malac, Diana and Midrigan, Doina and Corina, David P. Resources in Underrepresented Languages: Building a Representative R omanian Corpus. Proceedings of the Twelfth Language Resources a...
2020
-
[90]
Computerstrategien f
Validity in content analysis , author=. Computerstrategien f
-
[91]
K-Alpha Calculator–Krippendorff's Alpha Calculator: A user-friendly tool for computing Krippendorff's Alpha inter-rater reliability coefficient , journal =
Giacomo Marzi and Marco Balzano and Davide Marchiori , keywords =. K-Alpha Calculator–Krippendorff's Alpha Calculator: A user-friendly tool for computing Krippendorff's Alpha inter-rater reliability coefficient , journal =. 2024 , issn =. doi:https://doi.org/10.1016/j.mex.2023...
2024
-
[92]
2024 , eprint=
Gemini: A Family of Highly Capable Multimodal Models , author=. 2024 , eprint=
2024
-
[93]
COLM , year=
FineWeb2: One Pipeline to Scale Them All--Adapting Pre-Training Data Processing to Every Language , author=. COLM , year=
-
[94]
2025 , eprint=
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments , author=. 2025 , eprint=
2025
-
[95]
2013 , publisher =
OECD skills outlook 2013: first results from the survey of adult skills , author =. 2013 , publisher =
2013
-
[96]
Frontiers in Artificial Intelligence , volume =
Lexical simplification benchmarks for English, Portuguese, and Spanish , author =. Frontiers in Artificial Intelligence , volume =
-
[97]
2020 , doi =
Improving educational equity in Romania , author =. 2020 , doi =
2020
-
[98]
Research Involvement and Engagement , volume =
Preparing accessible and understandable clinical research participant information leaflets and consent forms: a set of guidelines from an expert consensus conference , author =. Research Involvement and Engagement , volume =
-
[99]
International Journal of Environmental Research and Public Health , volume =
Health Literacy Level and Comprehension of Prescription and Nonprescription Drug Information , author =. International Journal of Environmental Research and Public Health , volume =. 2022 , doi =
2022
-
[100]
Patient Education and Counseling , volume =
Antenatal group consultations: Facilitating patient-patient education , author =. Patient Education and Counseling , volume =
-
[101]
2022 , publisher =
Automatic Text Simplification , author =. 2022 , publisher =
2022
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.