{"id":"0c4cd1c3-d6d0-4aa0-b76a-9aa907dea3b8","arxiv_id":"2501.16112","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 23-respondent expert survey finds performance is the most important criterion for choosing NER tools, with cloud and local tools posing different challenges.","lead":"This paper surveys 23 machine learning experts about how they choose named entity recognition (NER) tools and the challenges they face. It finds that performance is the top selection criterion, while cloud-based and locally installed tools raise different concerns, such as cost and ease of use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's own inclusion criteria (Section 3) are not enforced: 50% of respondents have <1 year NER experience (Figure 5), so the claim to describe 'ML experts' rests on a self-selected, non-expert majority.","rationale":"The paper is a clearly written, small-scale exploratory survey. Its central contribution is a set of descriptive findings about the criteria that ML experts use when selecting NER tools, and the challenges they face. The most load-bearing condition for that claim is that the respondents actually are ML experts with relevant NER experience. The paper's own Section 3 defines the expert population as people with more than three years of experience in ML, NLP, and NER or a PhD in those areas. The demographics in Section 5.1 show that this condition is not met: half the sample has under one year of NER experience and over a third rate their NER experience as poor. The authors attempt to wave this away by pointing to the high educational level of the group, but that is a different criterion and does not establish expertise in NER. This is not merely a representativeness or sampling-size issue, which the authors openly acknowledge; it is an internal inconsistency between the study's own inclusion criteria and its actual sample. As a result, the aggregate figures and cloud/local comparisons in Section 5.2 cannot be interpreted as measures of ML experts' preferences unless the analysis is repeated on the expert subgroup. The paper deserves credit for following a named survey methodology, reporting its sampling plan transparently, making the questionnaire available online, and including a threats-to-validity section. However, the threats-to-validity section understates this issue by calling the risk 'minor' or 'low.' A concrete subgroup analysis would settle whether the non-expert respondents drive the results. If they do, the paper's conclusions must be reframed as descriptive of the 23 respondents, not of ML experts; if they do not, the conclusions are more robust. Given the paper's exploratory nature and its stated aim to inform future system design rather than to prove a general law, conditional acceptance remains appropriate. The reader's verdict is therefore unchanged.","tokens_in":15836,"tokens_out":4290,"duration_ms":41979,"concrete_test":"Split the 23 respondents into those meeting the Section 3 expert criterion (>=3 years NER experience, or a PhD in ML/NLP/NER) versus those who do not, and recompute Figures 11-12 and Tables 2-3 for each subgroup separately. If the ranking of criteria or the cloud/local deltas (e.g., User interface and ease of use delta 1.09; Cost delta 1.79) change materially or reverse between the two subgroups, the aggregate conclusions are an artifact of including non-experts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is about criteria used by ML experts to evaluate NER tools (RQ1/RQ2). Section 3 defines eligible experts as people with more than three years of experience in ML, NLP, and NER, or a PhD in these areas. Yet Section 5.1, Figure 5 shows 50% of the 23 respondents have less than one year of NER experience, and Figure 6 shows that 35.71% self-rate their NER experience as 'Poor' (2/5). The paper reclassifies this group as 'experienced' because they are highly educated, but a general Computer Science degree is not the same as demonstrated NER expertise, and the stated criterion was specifically >3 years or a PhD in ML/NLP/NER. The sampling plan explicitly waived representativeness because the findings were not intended to be generalized, yet the conclusions section generalizes to 'ML experts' and Section 5.4 assesses the residual risk as low. This internal inconsistency is the load-bearing weakness: the aggregate ratings (e.g., performance average 4.57 in Figure 11; cloud/local deltas in Tables 2 and 3) mix responses from people who do not meet the paper's own expert definition with those who do. If the two groups differ systematically, the paper's descriptive findings cannot be attributed to ML experts. The caveats in Section 5.4 do not resolve this because they label the resulting risk as 'minor' and 'low' rather than testing whether the non-expert responses drive the results.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a survey of 23 self-selected participants, recruited by email and through a research institute mailing list, about how they compare and select Named Entity Recognition (NER) tools and frameworks. The study is organized around two research questions: which criteria ML experts use to evaluate NER tools (RQ1) and what challenges they face when selecting a tool (RQ2). Using Kasunic's survey methodology, the authors designed a questionnaire, piloted it with three experts, and analyzed responses with descriptive statistics and figures. The main reported findings are that performance is the most important selection criterion, that all listed criteria are considered important by at least some respondents, that cloud-based services are particularly affected by cost and user-friendliness, and that locally installed tools are particularly affected by the time and effort needed to learn the system. The paper concludes with design implications for a future system to support non-experts in choosing NER tools.","tokens_in":16092,"tokens_out":4173,"duration_ms":40220,"significance":"If the findings held up, this would fill a genuine gap: most prior NER tool comparisons are benchmark studies, and the authors are right that user-centered evidence on how practitioners choose NER tools is scarce. The paper has real strengths: it follows a recognized survey methodology, reports a pilot test, makes the questionnaire publicly available, provides primary data, and explicitly discusses several threats to validity. The authors are also appropriately cautious in places, noting that the computer-science-heavy sample limits generalization to domain experts. However, the contribution is currently tentative because the respondents do not consistently match the paper's own definition of an ML/NER expert, the sample is small (N=23), several per-tool and per-category analyses rest on one to nine responses, and no inferential statistics or uncertainty measures are provided. The descriptive patterns are plausible and useful as hypothesis-generating evidence, but they are not yet sufficient to support the paper's stronger conclusions about what 'ML experts' need.","major_comments":[{"comment":"The paper's stated inclusion criterion for experts is more than three years of experience in ML, NLP, and NER, or a PhD in one of these areas. The reported sample does not satisfy this criterion: Figure 5 shows that 50% of the 23 respondents have less than one year of NER experience, and Figure 6 shows that 35.71% self-rate their NER experience as 'Poor' (2/5). In Section 5.1 the authors reclassify this group as experienced because of their high education level, but a general Computer Science degree is not a substitute for the stated NER expertise requirement. Since RQ1 and RQ2 are explicitly about ML experts, the aggregate results in Figure 11 and Tables 2 and 3 mix respondents who meet the stated expert definition with those who do not. The authors should either restrict the analysis to the criterion-defined experts or report a subgroup comparison demonstrating that the two groups do not differ systematically.","section":"Section 3 (Target Audience) and Section 5.1 (Figures 5 and 6)"},{"comment":"There is an internal inconsistency between the sampling plan and the conclusions. Section 3 states that no representative cross section is required because the findings are not intended to be generalized, yet Section 6 concludes that the research objectives have been 'successfully achieved' and presents general statements about what 'ML experts' find important, and Section 5.4 assesses the remaining risk as 'low' despite the acknowledged low sample size. Calling the low-N threat a 'minor risk' is not justified by the evidence in the paper. The conclusions should be reframed as descriptive, hypothesis-generating findings from a convenience sample, or the authors should provide a concrete argument for why the low response count and self-selected sample do not materially affect the reported patterns.","section":"Section 5.4 and Section 6"},{"comment":"The cloud-versus-local comparisons and the per-tool averages are presented as meaningful differences without any measure of uncertainty. Several cells are based on a single response (Microsoft Azure Cognitive Services and Flair in Figure 10), and even the larger per-tool cells contain only five to nine responses. For example, Table 3 reports a Cost delta of 1.79 and Table 2 reports a User interface and ease of use delta of 1.09, but with these sample sizes the deltas could easily be driven by a single respondent. The authors should report the number of responses per cell and either provide confidence intervals or a nonparametric test, or explicitly label all deltas and averages as purely descriptive with no claim of statistical reliability.","section":"Tables 2 and 3; Figures 10-12"},{"comment":"The claim that 'each challenge was mentioned at least once as hindering or very hindering' supports the conclusion that requirements are project specific is too strong. In a small convenience sample, a single endorsement of each item is almost guaranteed and does not demonstrate meaningful variability. Similarly, the statement that 'reducing the time and effort required to learn new frameworks is essential' is extrapolated from an average of 2.84 on a 1-5 scale, which is between 'slightly hindering' and 'moderately hindering'. The interpretation should be calibrated to the magnitude of the observed averages and to the small number of responses per category.","section":"Section 5.2, Figure 12 and Table 3"}],"minor_comments":[{"comment":"The text reports that 14.20% of participants have one to two years of NER experience, while Figure 5 shows 14.29%; the numbers should be reconciled.","section":"Section 5.1, Figure 5"},{"comment":"The narrative lists only Software Developer, Data Scientist, Machine Learning Engineer, Domain Expert, and Project Manager, but the figure also includes Data Engineer (8%), Researcher (8%), and Student (4%); the full set of roles should be described.","section":"Section 5.1, Figure 7"},{"comment":"The text says Microsoft Azure Cognitive Services has an average of 3.1 from one response, while Figure 10 reports an average of 3 from one response; the discrepancy should be corrected.","section":"Section 5.2, Figure 10"},{"comment":"The statement that OpenAI GPT-4, 'designed primarily for text generation, is highly adaptable for NER tasks' is presented without a supporting citation or evidence from the survey; either provide a reference or soften the claim.","section":"Section 5.2"},{"comment":"The heading 'Average Priority per Selection Criteria' is imprecise; the tables report average importance ratings, not a priority ranking, and Table 3 reports hindrance ratings rather than priorities.","section":"Tables 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The core idea is sound and the survey is honestly reported, but the sampling problem is load-bearing: with half the respondents having less than one year of NER experience, the paper's central descriptive claims about 'ML experts' cannot be taken at face value. The fix is feasible within the manuscript's scope: re-analyze the data for respondents meeting the stated expert criterion, report subgroup comparisons, add uncertainty measures or explicitly restrict all conclusions to descriptive observations, and align the conclusion with the non-representative sampling design. If the authors are unwilling to narrow the claims, the paper would be better positioned as a pilot-study report rather than a full empirical article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful pilot survey, not a definitive statement about ML experts. The new thing is primary data on how people choose NER tools—which criteria they weight, and where cloud and local tools differ. The survey is well-structured (Kasunic's method followed, questionnaire pretested, response rate reported), and the results are presented cleanly. I believe the descriptive findings are real: performance dominates, cost and ease-of-use matter for cloud, learning curve matters for local tools. That is not surprising, but it is now documented from actual users rather than inferred from benchmarks.\n\nThe soft spot is the one the stress-test note flags, and it is central. Section 3 defines eligible experts as people with more than three years in ML/NLP/NER or a PhD in those areas. Section 5.1 then shows 50% of the 23 respondents have under one year of NER experience, and 35.7% self-rate as 'poor'. The paper reclassifies the group as experienced because they are highly educated, but that is not the stated criterion. The aggregate ratings therefore mix people who meet the paper's own expert definition with people who do not, and there is no subgroup analysis. The sampling plan explicitly says representativeness is not needed, which is fine for an exploratory study, but the conclusions still generalize to 'ML experts' and the threats-to-validity section ends by calling the residual risk 'low.' That is an overstatement.\n\nThe rest of the weaknesses are minor in comparison: N=23, per-tool cells as small as one response, no confidence intervals or significance tests, and no raw data posted for independent checking. None of these would be fatal if the paper were framed as a pilot, but they do not support the current generality.\n\nWho is this for: anyone building NER tool recommender systems or studying how practitioners pick NLP libraries. It is a reasonable starting point, not a settled result. I would send it to peer review with a clear request to reframe the claims, release the data, and either enforce the expert definition or analyze the experience subgroups. If the authors do that, the survey could be a useful reference.","headline":"Small, clearly reported expert survey on NER tool choice; new primary data but the sample undercuts the 'ML expert' label.","tokens_in":16650,"tokens_out":2042,"would_cite":false,"duration_ms":20604,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of ML experts finds performance is the top criterion for choosing NER tools.","keywords":["Expert Survey","Named Entity Recognition","Machine Learning","Tool Selection","Evaluation Criteria","Cloud Computing","Natural Language Processing"],"falsifier":"A replication survey using the same per-tool criteria questions with a larger sample of ML experts who each have more than three years of NER experience would settle the priority claim if performance no longer ranked first overall, or if cloud users did not rate cost and ease of use above local-tool concerns.","tokens_in":15571,"feed_emoji":"📊","tokens_out":5633,"duration_ms":48361,"temperature":0.7,"pith_summary":"This paper reports a survey of 23 people who work with machine learning, asking how they compare and choose named-entity-recognition (NER) tools. Its aim is to establish which selection criteria matter most and which obstacles are the most hindering, so that future software can help non-experts make the same choices. The results say performance is the dominant criterion, but every listed criterion was rated important or very important by at least one respondent, meaning the right tool depends on the project. The paper also distinguishes cloud-based services from locally installed tools: cost and user-friendliness drive cloud choices, while documentation and the effort of learning a new system matter most for local tools. A sympathetic reader would take the study's contribution to be a criterion list plus deployment-specific priorities, not a general ranking of tools.","feed_headline":"Performance tops what ML experts want from NER tools","feed_subtitle":"23-expert survey finds selection criteria vary by project, with cost and usability leading for cloud tools.","key_machinery":"The load-bearing mechanism is the per-tool questionnaire: for each NER tool a respondent had used, the survey asked them to rate nine selection criteria and six potential hindrances on five-point scales, then aggregated the ratings across respondents. This design lets the authors split the same criteria by cloud-based versus locally installable tools and compute average-priority differences, which is what produces the deployment-specific conclusions. The per-tool structure was itself a product of pilot testing, after an earlier global question proved unable to link a criterion like performance or privacy to a specific tool choice.","core_discovery":"On the paper's own terms, the central discovery is that ML experts evaluate NER tools primarily by performance, with an average importance rating of 4.57 on a 5-point scale, while no criterion can be dismissed as universally unimportant. When the responses are split by deployment type, cloud-based services are judged especially on user interface and ease of use (average 4.43, a 1.09-point gap over local tools) and on licensing and cost, whereas locally installed systems are judged especially on documentation and support (a 1.14-point gap favoring local tools) and on minimizing the time and effort needed to learn the framework. The survey also finds that the most frequently hindering challenge is the time and effort to learn a new framework, and that locally operated open-source large language models are already used for NER and should be included in future tool comparisons.","pith_inferences":["Editorial inference: the project-specific spread of criteria suggests a configurable comparison matrix rather than a single recommended tool; a decision aid would need to elicit the user's deployment context first.","Editorial inference: the high priority given to performance may partly reflect the respondents' computer-science background; a domain-expert sample could shift priorities toward privacy and knowledge-domain fit, and the paper's own data leave this open.","Editorial inference: the finding that open-source local LLMs are used for NER, combined with the cost barrier for cloud services, points toward a two-tier selection space where local open models and paid cloud APIs serve different users; this distinction could be tested by asking respondents directly which tool class they ended up choosing.","Editorial inference: a testable extension would ask respondents to rank tools for a fixed hypothetical task, which would disentangle project-specificity from personal preference."],"forward_implications":["Any support system for choosing NER tools must let users weigh multiple criteria, because all nine surveyed criteria were rated important or very important at least once; performance alone does not decide.","For cloud-based NER offerings, controlling cost and simplifying the user interface are concrete levers that would address the two highest-priority cloud concerns.","For locally installed NER tools, improving documentation and reducing the effort of learning the framework would remove the most hindering obstacle.","Locally operated open-source large language models should be treated as a legitimate NER option in comparisons and selection aids.","Future surveys should recruit domain experts outside computer science, since the paper itself cautions that its results cannot be directly generalized to them."],"supporting_citations":[{"why":"Supplies the seven-stage survey-design process the paper follows and cites as the methodological basis.","marker":"[48]"},{"why":"A replicable comparison of NER software that the paper uses to motivate why comparing tools is difficult and why documentation matters.","marker":"[25]"},{"why":"A systematic review of NLP applied to radiology reports, cited as evidence of the growing number of ML-based NER tools and difficulty of comparison.","marker":"[18]"},{"why":"An enterprise-app benchmarking study that underlines tool selection as a critical step in developing NLP applications.","marker":"[26]"},{"why":"The cloud-based information extraction project that motivates the cloud questions and supports the claim that managing cloud resources is hard for newbies and ML experts.","marker":"[23]"},{"why":"A study on friendly cloud user interfaces, used to interpret the cloud-specific importance of user interface and ease of use.","marker":"[54]"},{"why":"A comparison of NLP toolkits on formal and social media text, cited to show tool choice depends on text kind and source.","marker":"[42]"},{"why":"A survey of software engineers on AI/ML that supplies the related finding that automation and training support matter, contextualizing the paper's survey.","marker":"[46]"}],"fun_headline_variants":["Performance tops ML experts' NER tool wish list","Cloud NER tools judged on ease, local on docs","NER tool survey: Learning curve is top hurdle","ML experts rate NER tools: Performance first","Survey: Performance wins for NER tool choice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that 23 volunteers who answered an emailed survey, half of whom reported less than one year of NER experience, can speak for the challenges of experienced ML experts; the paper's own sampling plan says no representative cross-section was required.","fun_headline_variants_meta":{"raw":{"variants":["Performance tops ML experts' NER tool wish list","Cloud NER tools judged on ease, local on docs","NER tool survey: Learning curve is top hurdle","ML experts rate NER tools: Performance first","Survey: Performance wins for NER tool choice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":1116,"prompt_tokens":858,"completion_tokens":258,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":184}},"tokens_in":474,"tokens_out":258,"duration_ms":2835,"temperature":1.0,"reasoning_tokens":184,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:42:35.340781+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication survey using the same per-tool criteria questions with a larger sample of ML experts who each have more than three years of NER experience would settle the priority claim if performance no longer ranked first overall, or if cloud users did not rate cost and ease of use above local-tool concerns.","supporting_citations":[{"cited_title":"Designing an Effective Survey","cited_arxiv_id":null,"evidence_quote":"Supplies the seven-stage survey-design process the paper follows and cites as the methodological basis."},{"cited_title":"A Replicable Comparison Study of NER Software: StanfordNLP, NLTK, OpenNLP, SpaCy, Gate","cited_arxiv_id":null,"evidence_quote":"A replicable comparison of NER software that the paper uses to motivate why comparing tools is difficult and why documentation matters."},{"cited_title":"A systematic review of natural language process- ing applied to radiology reports","cited_arxiv_id":null,"evidence_quote":"A systematic review of NLP applied to radiology reports, cited as evidence of the growing number of ML-based NER tools and difficulty of comparison."},{"cited_title":"Bench- marking NLP Toolkits for Enterprise Application","cited_arxiv_id":null,"evidence_quote":"An enterprise-app benchmarking study that underlines tool selection as a critical step in developing NLP applications."},{"cited_title":"Cie: A cloud-based information extraction system for named entity recognition in aws, azure, and medical domain","cited_arxiv_id":null,"evidence_quote":"The cloud-based information extraction project that motivates the cloud questions and supports the claim that managing cloud resources is hard for newbies and ML experts."},{"cited_title":"Kurdi, Safwat Hamad, and Amal Khalifa","cited_arxiv_id":null,"evidence_quote":"A study on friendly cloud user interfaces, used to interpret the cloud-specific importance of user interface and ease of use."},{"cited_title":"Comparing the per- formance of different NLP toolkits in formal and social media text","cited_arxiv_id":null,"evidence_quote":"A comparison of NLP toolkits on formal and social media text, cited to show tool choice depends on text kind and source."},{"cited_title":"Soft- ware Engineering for Machine Learning: A Case Study","cited_arxiv_id":null,"evidence_quote":"A survey of software engineers on AI/ML that supplies the related finding that automation and training support matter, contextualizing the paper's survey."}],"review_version":1}