REVIEW 3 major objections 2 minor 110 cited by
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
T0 review · 3 major / 2 minor · reviewed 2026-05-15 · grok-4.3
Pith's one-line read Multiple stochastic samples from a black-box LLM reveal which generated facts are hallucinations by checking their consistency.
desk verdict SelfCheckGPT uses sampling consistency to flag hallucinations in black-box LLMs and beats grey-box baselines on WikiBio, but the signal may mix factual errors with normal output variation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Consistency check across multiple independently sampled responses to the same prompt
What would settle it
A collection of prompts where the model repeats the identical incorrect fact in every sample, causing the consistency metric to score the output as factual despite it being wrong.
Extended reading notes
Core claim
SelfCheckGPT is a sampling-based approach for zero-resource hallucination detection in black-box LLMs. It rests on the premise that when an LLM possesses knowledge of a concept, stochastically sampled responses tend to be similar and factually consistent, whereas hallucinated facts produce divergent and contradictory samples. Applied to GPT-3 generations on the WikiBio dataset with human-annotated factuality labels, the method detects non-factual sentences and ranks passages by factuality, delivering higher AUC-PR scores at the sentence level and stronger correlation scores at the passage level than grey-box alternatives.
Load-bearing premise
Divergence among sampled responses primarily signals hallucinated facts rather than stylistic differences or partial but consistent knowledge.
Editorial extensions
If this is right
- Allows factuality assessment for closed models such as ChatGPT that expose only generated text.
- Achieves higher precision-recall in sentence-level hallucination detection than methods using token probabilities.
- Enables ranking of generated passages by overall factuality without any external knowledge source.
- Requires only repeated sampling from the target model, making it applicable to any LLM supporting stochastic generation.
Reading between the lines
- The method could be layered with other lightweight signals to catch cases where the model consistently repeats the same error.
- It suggests a general principle that internal model uncertainty may be approximated through output variation alone.
- Extensions to longer-form or multi-turn outputs would require adapting the consistency metric to handle accumulating context.
- In deployment, the approach could lower reliance on curated fact-checking databases for routine verification tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SelfCheckGPT, a zero-resource black-box method for hallucination detection in generative LLMs. It samples multiple responses from the model (e.g., GPT-3 on WikiBio prompts) and measures consistency via metrics such as BERTScore, QA-based overlap, and n-gram overlap; low consistency is taken to indicate hallucinated facts. On manually annotated generations, the method reports higher sentence-level AUC-PR for hallucination detection and higher passage-level correlation with human factuality judgments than grey-box baselines.
Significance. If the consistency signal can be shown to isolate factual errors rather than stylistic or partial-knowledge variation, the approach would offer a practical, external-database-free tool for fact-checking black-box LLMs, addressing a key deployment barrier. The zero-resource design and direct comparison to grey-box methods are clear strengths.
major comments (3)
- [§4] §4 (Experiments): the central AUC-PR and correlation claims rest on human factuality annotations, yet no inter-annotator agreement statistics, annotation guidelines, or controls for annotator bias are reported; this directly affects the reliability of the ground-truth labels used to compute all performance numbers.
- [§3] §3 (Method): the premise that sample divergence signals hallucination is not isolated from other sources of variation (alternative valid phrasings, differing levels of detail, or stylistic choices). No ablation or control experiment is described that holds factual content fixed while varying only style or completeness, leaving open whether the reported gains are inflated by conflating multiple variation types.
- [§4.2] §4.2 (Evaluation metrics): the exact aggregation formulas for the consistency scores (e.g., how BERTScore or QA overlap is averaged or thresholded across the 5–20 samples) are not fully specified, nor are statistical significance tests for the AUC-PR improvements over baselines provided.
minor comments (2)
- [§3] The number of samples, sampling temperature, and prompt templates used for generation should be stated explicitly in §3 and §4.1 for reproducibility.
- [§4] Figure 2 or the corresponding table should include error bars or confidence intervals on the AUC-PR and correlation values.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. We address each major comment below, clarifying our approach where possible and committing to revisions that strengthen the presentation of the experimental details and method assumptions.
read point-by-point responses
-
Referee: [§4] §4 (Experiments): the central AUC-PR and correlation claims rest on human factuality annotations, yet no inter-annotator agreement statistics, annotation guidelines, or controls for annotator bias are reported; this directly affects the reliability of the ground-truth labels used to compute all performance numbers.
Authors: We agree that inter-annotator agreement and annotation protocol details are necessary to establish label reliability. Although omitted from the initial submission for brevity, annotations were performed by three independent annotators following explicit guidelines that defined hallucination as any non-supported factual claim. In the revised manuscript we will add a new subsection in §4 reporting the inter-annotator agreement statistics, reproducing the full annotation guidelines in the appendix, and describing bias-mitigation steps such as randomized presentation order and independent adjudication of disagreements. revision: yes
-
Referee: [§3] §3 (Method): the premise that sample divergence signals hallucination is not isolated from other sources of variation (alternative valid phrasings, differing levels of detail, or stylistic choices). No ablation or control experiment is described that holds factual content fixed while varying only style or completeness, leaving open whether the reported gains are inflated by conflating multiple variation types.
Authors: We acknowledge the concern that consistency metrics could be influenced by non-factual sources of variation. Our QA-based and entity-overlap metrics are deliberately chosen to emphasize factual content rather than surface form; however, we concede that a controlled ablation isolating style while fixing facts would provide stronger evidence. We will expand the discussion in §3 to articulate why we expect stylistic variation to have limited impact on the reported metrics for the WikiBio task, and we will explicitly list the absence of such an ablation as a limitation. Performing the ablation would require substantial new annotation and is left for future work. revision: partial
-
Referee: [§4.2] §4.2 (Evaluation metrics): the exact aggregation formulas for the consistency scores (e.g., how BERTScore or QA overlap is averaged or thresholded across the 5–20 samples) are not fully specified, nor are statistical significance tests for the AUC-PR improvements over baselines provided.
Authors: We apologize for the incomplete specification. The sentence-level consistency score is the mean of the pairwise similarity values (BERTScore, QA overlap, or n-gram overlap) between the candidate sentence and each of the N sampled responses; no additional thresholding is applied. In the revised §4.2 we will state these aggregation formulas explicitly and will report statistical significance of the AUC-PR gains over baselines via paired bootstrap resampling with 10,000 iterations. revision: yes
Circularity Check
No significant circularity in SelfCheckGPT sampling consistency heuristic
full rationale
The paper defines SelfCheckGPT as a direct sampling procedure: generate multiple stochastic responses from a black-box LLM and measure consistency (via BERTScore, QA, n-gram overlap) to flag hallucinations. This is tested against independent human factuality labels on WikiBio passages. No equations reduce a claimed prediction to a fitted input by construction, no self-citation chains justify the core premise, and the consistency assumption is presented as a testable heuristic rather than a self-definition. The method is self-contained against external benchmarks.
Assumptions & free parameters
assumptions (1)
- domain assumption If an LLM has knowledge of a concept, sampled responses are likely to be similar and contain consistent facts; for hallucinated facts, responses are likely to diverge.
Cite this review
Pith. "Pith review of SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models." pith.science (2026). https://pith.science/paper/3ELOZYL4
@misc{pith2026230308896,
author = {Pith},
title = {Pith review of: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/3ELOZYL4}},
note = {Machine review of arXiv:2303.08896}
}
read the original abstract
Generative Large Language Models (LLMs) such as GPT-3 are capable of generating highly fluent responses to a wide variety of user prompts. However, LLMs are known to hallucinate facts and make non-factual statements which can undermine trust in their output. Existing fact-checking approaches either require access to the output probability distribution (which may not be available for systems such as ChatGPT) or external databases that are interfaced via separate, often complex, modules. In this work, we propose "SelfCheckGPT", a simple sampling-based approach that can be used to fact-check the responses of black-box models in a zero-resource fashion, i.e. without an external database. SelfCheckGPT leverages the simple idea that if an LLM has knowledge of a given concept, sampled responses are likely to be similar and contain consistent facts. However, for hallucinated facts, stochastically sampled responses are likely to diverge and contradict one another. We investigate this approach by using GPT-3 to generate passages about individuals from the WikiBio dataset, and manually annotate the factuality of the generated passages. We demonstrate that SelfCheckGPT can: i) detect non-factual and factual sentences; and ii) rank passages in terms of factuality. We compare our approach to several baselines and show that our approach has considerably higher AUC-PR scores in sentence-level hallucination detection and higher correlation scores in passage-level factuality assessment compared to grey-box methods.
Forward citations
Showing 60 of 110 Pith papers that cite this
-
The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale
Across 0.8B-12B models plus a frontier arm, apparent self-correction effects are dominated by format-recovery and format-loss artifacts, with near-zero content-level change at capable scale.
-
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
VLMs recover common ERD elements at F1>0.74 but drop to 0.07–0.28 on N-ary relationships, multivalued attributes, and weak entities; reasoning models gain 15–25% yet stay prior- and complexity-sensitive.
-
Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement
Models delayed verification in multi-agent LLMs as graph consensus, derives stability thresholds (inverse golden ratio for delay two) via grounded Laplacian, and gives a supermodular greedy rule for corrector placemen...
-
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs
Operadic consistency is a new per-question signal that correlates strongly with accuracy (r 0.86-0.94) across four multi-hop QA datasets and improves selective prediction over CoT-SC baselines.
-
Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees
Introduces object-level semantic uncertainty for VLM memory, the UQ-DAAAM refinement system, and probabilistic guarantees that selected high-quality views reduce uncertainty more effectively.
-
Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation
Introduces CHARM framework that detects cascading hallucinations in agentic RAG at 89.4% rate with 5.3% false positives and reduces error propagation by 82.1% on multi-hop QA benchmarks.
-
Before and After Temperature: A Distributional View of Creative LLM Generation
A per-token feature from temperature-induced changes in LLM token distributions predicts within-prompt creativity rank at Spearman rho 0.918 vs LLM judges and 0.870 vs humans, outperforming perplexity, entropy, top-1 ...
-
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
EEG study reveals distinct ERP patterns for AI hallucinations, with misjudged ones failing to trigger standard neurocognitive verification pathways.
-
Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
Compositional selective specificity (CSS) improves overcommitment-aware utility from 0.846 to 0.913 on LongFact while retaining 0.938 specificity by calibrating claim-level backoffs in agentic AI responses.
-
LLM4Log: A Systematic Review of Large Language Model-based Log Analysis
Systematic review of 145 papers on LLM-based log analysis, providing a unified taxonomy, common design patterns, evaluation practices, and challenges for deployment under drift and limited labels.
-
ScrapeGraphAI-100k: Dataset for Schema-Constrained LLM Generation
ScrapeGraphAI-100k releases 93,695 real telemetry examples pairing web page content with prompts, schemas, and LLM responses to support training and benchmarking of schema-constrained generation.
-
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
LatentRefusal predicts answerability of text-to-SQL queries from LLM hidden states using a Tri-Residual Gated Encoder, reaching 88.5% average F1 across four benchmarks with about 2ms overhead.
-
Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
TARG uses uncertainty scores from a short no-context draft to gate retrieval in RAG, matching Always-RAG accuracy while cutting retrievals by 70-90% on QA benchmarks.
-
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
Clotho ranks LLM test inputs by failure likelihood using pre-generation hidden states and GMMs, achieving 0.716 ROC-AUC after labeling 5.4% of inputs on average across eight tasks and three models, with transfer to pr...
-
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
Directly editing a large reasoning model's chain-of-thought and feeding back a distilled version of the edit improves correction success by over 25% and cuts token usage by roughly 40% versus dialogue-based correction.
-
Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering
Prompt-point activations carry a graded, steerable entity-familiarity signal that is robust to Polish/English stem changes and is stronger in Polish-adapted models than in base models.
-
Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States
Among 8/8 self-consistent answers on FinQA, residual-stream probes detect wrong answers at 0.68–0.77 AUROC versus 0.55–0.63 for the best cheap output baselines across three 8–9B models.
-
Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment
Distilling an 8B reasoning teacher into a 0.6B student recovers most summary quality at ~50× speed, but teacher type—not scale alone—determines which capabilities transfer.
-
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
Self-consistency is a weak, regime-dependent proxy for correctness: positive but small correlations (rho 0.20–0.59), with the most self-consistent frontier model over-confident and wrong 48% of the time at high agreement.
-
Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale
Unsupervised MLP activation dispersion separates known from fabricated entities at AUROC 0.95–1.00 across Bielik scales, while factual reliability scales separately and refusals stay near zero.
-
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
A large-scale multilingual evaluation of LLM uncertainty estimation methods across 22 languages and 9 models finds that English reasoning closes the UE gap for low-resource languages and that optimal UE method choice ...
-
When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs
Global calibration metrics like ECE are confounded by accuracy; the proposed ACE framework with three accuracy-controlled views shows many prior calibration advantages weaken or reverse.
-
Grad Detect: Gradient-Based Hallucination Detection in LLMs
Grad Detect uses internal gradient patterns from one inference pass to predict LLM hallucinations and abstention, outperforming confidence and sampling baselines on Q&A benchmarks with most signal in the final five layers.
-
Constrained Paraphrase Consistency for LLM Hallucination Detection
CCHD formulates hallucination detector training as constrained optimization with paraphrase-consistency and label-preservation rules solved via gradient descent-ascent, outperforming baselines on factuality benchmarks.
-
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation
DeepSurvey introduces an agentic system for automated survey generation that improves depth through full-text keynotes, cross-paper clustering, and code analysis, while boosting citation reliability via graph expansio...
-
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
Introduces functional equivalence methods and functional entropy to predict functional correctness of LLM-generated code via uncertainty quantification, outperforming NLI-based baselines in most tested settings.
-
Evaluating the False Trust Engendered by LLM Explanations
LLM reasoning traces and post-hoc explanations increase false trust in incorrect predictions, whereas contrastive dual explanations enhance users' ability to distinguish correct from incorrect AI outputs.
-
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
Confidence estimators that score well on calibration and discrimination still fail to stay stable under rephrasing and to react to changes in answer meaning.
-
Geometry-Aware Hallucination Detection in Large Language Models
A manifold-based prototype sampling method for choosing in-context examples improves hallucination-detection accuracy over several ICL baselines in the majority of tested settings.
-
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
RUDDER creates a persistent visual anchor by extracting CARD from prefill residuals and modulating its injection via an adaptive Beta Gate, cutting CHAIR_S by 24.4% and CHAIR_i by 23.6% on average across LLaVA, Idefic...
-
False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
Confidently wrong LLM answers behave like locally stable fixed points: no fragility gap vs correct answers, and abstention-style self-critique trades coverage for confidence.
-
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
ReFACT benchmark reveals LLMs show a persistent salient distractor failure mode where 61% of incorrect error span predictions are semantically unrelated to actual errors, persisting across model sizes, and comparative...
-
Neural Message-Passing on Attention Graphs for Hallucination Detection
CHARM trains graph neural networks on token-attention graphs built from LLM computational traces and outperforms prior hallucination detectors on five benchmarks at token and response level.
-
What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale
An instruction-tuned LLaMA 3.1 8B model, trained on LLM-generated pseudo-labels, is claimed to identify IoT device vendors from passive network metadata with 98.25% top-1 accuracy across 2,015 vendors.
-
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
IRIS detects LLM hallucinations by training a lightweight probe on hidden states elicited during the model's own step-by-step verification, using the model's verbalized confidence as soft pseudolabels.
-
Cross-Layer Attention Probing for Fine-Grained Hallucination Detection
CLAP, a cross-layer attention probe over all LLM layer activations, improves hallucination detection and enables a detect-then-mitigate decoding strategy.
-
Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
A decoding pipeline reuses cached tokens and anneals sampling temperature to accelerate self-consistency hallucination detection by up to 3x without meaningful AUROC loss.
-
Principled Detection of Hallucinations in Large Language Models via Multiple Testing
The method aggregates multiple hallucination evaluation scores via conformal p-values to enable calibrated detection with controlled false alarm rates across LLMs and datasets.
-
Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
A two-stage ideation-plus-projection evaluation narrows measured performance gaps between LLMs relative to standard QA benchmarks, implying standard QA overstates model differences.
-
Extension Decisions in Open Source Software Ecosystem
A GitHub Actions graph study reports that most new CI tools duplicate existing functionality and that a handful of early tools become the templates for later copies, although the supporting calculation is missing from...
-
Hallucination Detection and Mitigation with Diffusion in Multi-Variate Time-Series Foundation Models
Pre-trained multivariate time-series imputation models frequently return values that violate known relations between variables, and a diffusion-based score can detect and filter these errors.
-
The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems
A framework for agentic recommender systems plus a pilot study showing multi-agent pipelines beat a single-shot LLM only on high-diversity user histories.
-
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
AggTruth uses four ways of aggregating attention scores over the retrieved passage to train a logistic regression detector, achieving stable cross-task AUROC across four LLMs.
-
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
Reasoning language models are systematically overconfident, deeper reasoning makes them more overconfident, and a two-stage introspective prompting method improves calibration for some models.
-
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
A dual-level attention intervention that boosts salient visual-token attention and suppresses text/system attention during decoding reduces hallucination rates in LLaVA, MiniGPT-4, and mPLUG-Owl2 on POPE and CHAIR.
-
Brevity is the soul of sustainability: Characterizing LLM response lengths
LLMs produce longer-than-needed answers to factual questions, and simple prompt instructions such as 'provide only the minimal answer' cut response length and inference energy by about 25-60% without hurting automated...
-
Your Agent Can Defend Itself against Backdoor Attacks
A two-level consistency defense detects backdoored LLM agents by matching thoughts to actions and reconstructed instructions to the user's instruction, reducing attack success rates on tested tasks.
-
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
The paper proposes a query-centric evaluation of LLM "intent hallucination" via constraint decomposition, but the headline metric comparison is undermined by a self-referential human evaluation design.
-
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.
-
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
A clinician-validated taxonomy and 35-benchmark suite show that large language models vary widely across medical tasks, with reasoning models leading overall.
-
DoctorRAG: Medical RAG Fusing Knowledge with Patient Analogy through Textual Gradients
Combining knowledge retrieval, analogous patient case retrieval, and iterative textual-gradient refinement improves medical RAG accuracy across Chinese, English, and French benchmarks.
-
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
Adapts multi-layer token-level Mahalanobis distance with supervised linear regression to yield improved uncertainty scores for LLM truthfulness tasks.
-
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
A three-method auditing framework detects with roughly 87 to 97 percent accuracy whether classifiers, generators, and t-SNE plots were trained on or derived from LLM-generated synthetic data.
-
Ragas: Automated Evaluation of Retrieval Augmented Generation
Ragas supplies reference-free metrics for measuring context relevance, faithfulness to retrieved passages, and answer quality in RAG pipelines.
-
Chain-of-Verification Reduces Hallucination in Large Language Models
Chain-of-Verification reduces hallucinations in large language models by drafting responses, planning independent verification questions, answering them separately, and generating a final verified output.
-
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
DoLa reduces hallucinations in LLMs by contrasting logits from later versus earlier layers during decoding, improving truthfulness on TruthfulQA by 12-17 absolute points without fine-tuning or retrieval.
-
Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI
HALO is a layered oversight architecture that grounds, constrains, verifies, abstains, traces, and monitors LLM outputs to make hallucinations containable rather than eliminated.
-
LLM-as-a-Verifier: A General-Purpose Verification Framework
Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.
-
Strategic Decision Support for AI Agents
The paper introduces an optimization framework for AI agents to strategically seek support, proving a threshold policy on support value and providing an online algorithm to control missed-support error without distrib...
-
Cross Paraphrastic Invariance Learning for Hallucination Detection
CPIL is a contrastive two-stage method that enforces paraphrase invariance on limited labeled data to outperform baselines in hallucination detection across 11 tasks.
Reference graph
Works this paper leans on
-
[3]
Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. https://doi.org/10.18653/v1/2022.bigscience-1.9 GPT - N eo X -20 B : An open-source autoreg...
-
[4]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
work page 2020
-
[6]
Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20:37 -- 46
work page 1960
-
[8]
Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. https://doi.org/10.1162/tacl_a_00454 A survey on automated fact-checking . Transactions of the Association for Computational Linguistics, 10:178--206
-
[9]
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023. https://openreview.net/forum?id=sE7-XhLxHA De BERT av3: Improving de BERT a using ELECTRA -style pre-training with gradient-disentangled embedding sharing . In The Eleventh International Conference on Learning Representations
work page 2023
-
[11]
Ganesh Jawahar, Beno \^ t Sagot, and Djam \'e Seddah. 2019. https://doi.org/10.18653/v1/P19-1356 What does BERT learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3651--3657, Florence, Italy. Association for Computational Linguistics
-
[14]
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.750 Evaluating the factual consistency of abstractive text summarization . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332--9346, Online. Association for Computational Linguistics
-
[15]
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. https://openreview.net/forum?id=VD-AYtP0dve Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation . In The Eleventh International Conference on Learning Representations
work page 2023
Show all 288 references
-
[16]
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. https://doi.org/10.18653/v1/D17-1082 RACE : Large-scale R e A ding comprehension dataset from examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages...
2017 doi
-
[17]
R \' e mi Lebret, David Grangier, and Michael Auli. 2016. http://arxiv.org/abs/1603.07771 Generating text from structured data with application to the biography domain . CoRR, abs/1603.07771
2016 arXiv
-
[18]
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2022. https://doi.org/10.18653/v1/2022.acl-long.464 A token-level reference-free hallucination detection benchmark for free-form text generation . In Proceedings of the 60th Annual Meeti...
2022 doi
-
[20]
Adian Liusie, Vatsal Raina, and Mark Gales. 2023. https://aclanthology.org/2023.fever-1.5 `` world knowledge '' in multiple choice reading comprehension . In Proceedings of the Sixth Fact Extraction and VERification Workshop (FEVER), pages 49--57, Dubrovnik, Croatia. Associati...
2023
-
[22]
Andrey Malinin and Mark Gales. 2021. https://openreview.net/forum?id=jN5y-zb5Q7m Uncertainty estimation in autoregressive structured prediction . In International Conference on Learning Representations
2021
-
[23]
Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. MQAG : Multiple-choice question answering and generation for assessing information consistency in summarization. arXiv preprint arXiv:2301.12307
2023
-
[24]
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. https://doi.org/10.18653/v1/2020.acl-main.173 On faithfulness and factuality in abstractive summarization . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1...
2020 doi
-
[26]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485--5551
2020
-
[27]
Vatsal Raina and Mark Gales. 2022. https://doi.org/10.18653/v1/2022.findings-acl.82 Answer uncertainty and unanswerability in multiple-choice machine reading comprehension . In Findings of the Association for Computational Linguistics: ACL 2022, pages 1020--1034, Dublin, Irela...
2022 doi
-
[28]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. https://doi.org/10.18653/v1/D16-1264 SQ u AD : 100,000+ questions for machine comprehension of text . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383-...
2016 doi
-
[29]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.320 Retrieval augmentation reduces hallucination in conversation . In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3784--...
2021 doi
-
[30]
James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. 2018. The Fact Extraction and VERification (FEVER) shared task. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER)
2018
-
[32]
Anthony J Viera, Joanne M Garrett, et al. 2005. Understanding interobserver agreement: the kappa statistic. Fam med, 37(5):360--363
2005
-
[33]
Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model . https://github.com/kingoflolz/mesh-transformer-jax
2021
-
[34]
Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. https://openreview.net/forum?id=1PL1NIMMrw Self-consistency improves chain of thought reasoning in language models . In The Eleventh International Conferenc...
2023
-
[35]
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computation...
2018 doi
-
[36]
Yijun Xiao and William Yang Wang. 2021. https://doi.org/10.18653/v1/2021.eacl-main.236 On hallucination and predictive uncertainty in conditional language generation . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistic...
2021 doi
-
[37]
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems, 34:27263--27277
2021
-
[39]
Zhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li, Shuailiang Zhang, Xi Zhou, and Xiang Zhou. 2020. Semantics-aware bert for language understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9628--9635
2020
-
[40]
Wanjun Zhong, Jingjing Xu, Duyu Tang, Zenan Xu, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin. 2020. https://doi.org/10.18653/v1/2020.acl-main.549 Reasoning over semantic-level graph for fact checking . In Proceedings of the 58th Annual Meeting of the Association for Computat...
2020 doi
-
[41]
Manakul, Potsawee and Liusie, Adian and Gales, Mark JF , journal=
-
[42]
Advances in neural information processing systems , volume=
Language models are few-shot learners , author=. Advances in neural information processing systems , volume=
-
[43]
Transactions on Machine Learning Research , year=
Emergent Abilities of Large Language Models , author=. Transactions on Machine Learning Research , year=
-
[44]
arXiv preprint arXiv:2204.02311 , year=
Palm: Scaling language modeling with pathways , author=. arXiv preprint arXiv:2204.02311 , year=
-
[45]
The Journal of Machine Learning Research , volume=
Exploring the limits of transfer learning with a unified text-to-text transformer , author=. The Journal of Machine Learning Research , volume=. 2020 , publisher=
2020
-
[46]
`` World Knowledge '' in Multiple Choice Reading Comprehension
Liusie, Adian and Raina, Vatsal and Gales, Mark. `` World Knowledge '' in Multiple Choice Reading Comprehension. Proceedings of the Sixth Fact Extraction and VERification Workshop (FEVER). 2023
2023
-
[47]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Semantics-aware BERT for language understanding , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[48]
and Vitrià, Jordi , journal=
Brando, Axel and Torres, Damià and Rodríguez-Serrano, Jose A. and Vitrià, Jordi , journal=. Building Uncertainty Models on Top of Black-Box Predictive APIs , year=
-
[49]
Generating Text from Structured Data with Application to the Biography Domain , journal =
R. Generating Text from Structured Data with Application to the Biography Domain , journal =. 2016 , url =
2016
-
[50]
Advances in Neural Information Processing Systems , volume=
Bartscore: Evaluating generated text as text generation , author=. Advances in Neural Information Processing Systems , volume=
-
[51]
arXiv preprint arXiv:2205.01068 , year=
Opt: Open pre-trained transformer language models , author=. arXiv preprint arXiv:2205.01068 , year=
-
[52]
Wang, Ben and Komatsuzaki, Aran , title =
-
[53]
GPTScore: Evaluate as You Desire , publisher =
Fu, Jinlan and Ng, See-Kiong and Jiang, Zhengbao and Liu, Pengfei , keywords =. GPTScore: Evaluate as You Desire , publisher =. 2023 , copyright =. doi:10.48550/ARXIV.2302.04166 , url =
2023 doi
-
[54]
Proceedings of the First Workshop on
Thorne, James and Vlachos, Andreas and Cocarascu, Oana and Christodoulopoulos, Christos and Mittal, Arpit , title =. Proceedings of the First Workshop on
- [55]
-
[56]
and Vinyals, Oriol and Sifre, Laurent , keywords =
Hoffmann, Jordan and Borgeaud, Sebastian and Mensch, Arthur and Buchatskaya, Elena and Cai, Trevor and Rutherford, Eliza and Casas, Diego de Las and Hendricks, Lisa Anne and Welbl, Johannes and Clark, Aidan and Hennigan, Tom and Noland, Eric and Millican, Katie and Driessche, ...
-
[57]
Language Models are Unsupervised Multitask Learners , author=
-
[58]
ACM Comput
Ji, Ziwei and Lee, Nayeon and Frieske, Rita and Yu, Tiezheng and Su, Dan and Xu, Yan and Ishii, Etsuko and Bang, Ye Jin and Madotto, Andrea and Fung, Pascale , title =. ACM Comput. Surv. , month =. 2023 , issue_date =. doi:10.1145/3571730 , abstract =
2023 doi
-
[59]
The Factual Inconsistency Problem in Abstractive Text Summarization: A Survey , publisher =
Huang, Yichong and Feng, Xiachong and Feng, Xiaocheng and Qin, Bing , keywords =. The Factual Inconsistency Problem in Abstractive Text Summarization: A Survey , publisher =. 2021 , copyright =. doi:10.48550/ARXIV.2104.14839 , url =
2021 doi
-
[60]
International Conference on Learning Representations , year=
Uncertainty Estimation in Autoregressive Structured Prediction , author=. International Conference on Learning Representations , year=
- [61]
-
[62]
Educational and Psychological Measurement , year=
A Coefficient of Agreement for Nominal Scales , author=. Educational and Psychological Measurement , year=
-
[63]
Fam med , volume=
Understanding interobserver agreement: the kappa statistic , author=. Fam med , volume=
-
[64]
arXiv preprint arXiv:2304.13734 , year=
The Internal State of an LLM Knows When its Lying , author=. arXiv preprint arXiv:2304.13734 , year=
-
[65]
The Eleventh International Conference on Learning Representations , year=
Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , author=. The Eleventh International Conference on Learning Representations , year=
-
[66]
arXiv preprint arXiv:2207.05221 , year=
Language models (mostly) know what they know , author=. arXiv preprint arXiv:2207.05221 , year=
-
[67]
arXiv preprint arXiv:2302.13971 , year=
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
-
[68]
arXiv preprint arXiv:2303.15621 , year=
Chatgpt as a factual inconsistency evaluator for abstractive text summarization , author=. arXiv preprint arXiv:2303.15621 , year=
-
[69]
arXiv preprint arXiv:1907.11692 , year=
Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=
1907 arXiv
-
[70]
arXiv preprint arXiv:2305.14251 , year=
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation , author=. arXiv preprint arXiv:2305.14251 , year=
-
[71]
Pengcheng He and Jianfeng Gao and Weizhu Chen , booktitle=. De. 2023 , url=
2023
-
[72]
The Eleventh International Conference on Learning Representations , year=
Self-Consistency Improves Chain of Thought Reasoning in Language Models , author=. The Eleventh International Conference on Learning Representations , year=
-
[73]
Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[74]
Neural Approaches for Data Driven Dependency Parsing in S anskrit
Krishna, Amrith and Gupta, Ashim and Garasangi, Deepak and Sandhan, Jeevnesh and Satuluri, Pavankumar and Goyal, Pawan. Neural Approaches for Data Driven Dependency Parsing in S anskrit. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented...
2023
-
[75]
Evaluating Neural Word Embeddings for S anskrit
Sandhan, Jivnesh and Paranjay, Om Adideva and Digumarthi, Komal and Behra, Laxmidhar and Goyal, Pawan. Evaluating Neural Word Embeddings for S anskrit. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Confer...
2023
-
[76]
Validation and Normalization of DCS corpus and Development of the S anskrit Heritage Engine ' s Segmenter
Sriram, Krishnan and Kulkarni, Amba and Huet, G \'e rard. Validation and Normalization of DCS corpus and Development of the S anskrit Heritage Engine ' s Segmenter. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S a...
2023
-
[77]
Pre-annotation Based Approach for Development of a S anskrit Named Entity Recognition Dataset
Sujoy, Sarkar and Krishna, Amrith and Goyal, Pawan. Pre-annotation Based Approach for Development of a S anskrit Named Entity Recognition Dataset. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[78]
Disambiguation of Instrumental, Dative and Ablative Case suffixes in S anskrit
Maity, Malay and Panchal, Sanjeev and Kulkarni, Amba. Disambiguation of Instrumental, Dative and Ablative Case suffixes in S anskrit. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[79]
Creation of a Digital Rig V edic Index (Anukramani) for Computational Linguistic Tasks
Mahesh, A V S D S and Bhattacharya, Arnab. Creation of a Digital Rig V edic Index (Anukramani) for Computational Linguistic Tasks. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[80]
Skrutable: Another Step Toward Effective S anskrit Meter Identification
Neill, Tyler. Skrutable: Another Step Toward Effective S anskrit Meter Identification. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[81]
Chandojnanam: A S anskrit Meter Identification and Utilization System
Terdalkar, Hrishikesh and Bhattacharya, Arnab. Chandojnanam: A S anskrit Meter Identification and Utilization System. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[82]
Ajotikar, Tanuja P and Scharf, Peter M. Development of a TEI standard for digital S anskrit texts containing commentaries: A pilot study of Bhaṭṭti ' s R \=a vaṇavadha with Mallin \=a tha ' s commentary on the first canto. Proceedings of the Computational S anskrit & Digital H...
2023
-
[83]
R \=a mop \=a khy \=a na: A Web-based reader and index
Scharf, Peter M and Chauhan, Dhruv. R \=a mop \=a khy \=a na: A Web-based reader and index. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[84]
Semantic Annotation and Querying Framework based on Semi-structured Ayurvedic Text
Terdalkar, Hrishikesh and Bhattacharya, Arnab and Dubey, Madhulika and Ramamurthy, S and Singh, Bhavna Naneria. Semantic Annotation and Querying Framework based on Semi-structured Ayurvedic Text. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers ...
2023
-
[85]
Shaastra Maps: Enabling Conceptual Exploration of I ndic Shaastra Texts
Susarla, Sai and Jammalamadaka, Suryanarayana and Nishankar, Vaishnavi and Panuganti, Siva and Ryali, Anupama and Sushrutha, S. Shaastra Maps: Enabling Conceptual Exploration of I ndic Shaastra Texts. Proceedings of the Computational S anskrit & Digital Humanities: Selected pa...
2023
-
[86]
The V edic corpus as a graph
Hellwig, Oliver and Sellmer, Sven and Amano, Kyoko. The V edic corpus as a graph. An updated version of Bloomfields V edic Concordance. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[87]
The transmission of the Buddha ' s teachings in the digital age
Harnsukworapanich, Sumachaya and Supphipat, Phatchareporn. The transmission of the Buddha ' s teachings in the digital age. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[88]
Distinguishing Commentary from Canon: Experiments in P \=a li Computational Linguistics
Zigmond, Dan. Distinguishing Commentary from Canon: Experiments in P \=a li Computational Linguistics. Proceedings of the Computational S anskrit & Digital Humanities: Selected papers presented at the 18th World S anskrit Conference. 2023
2023
-
[89]
Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[90]
Analyzing Zero-Shot transfer Scenarios across S panish variants for Hate Speech Detection
Castillo-l \'o pez, Galo and Riabi, Arij and Seddah, Djam \'e. Analyzing Zero-Shot transfer Scenarios across S panish variants for Hate Speech Detection. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[91]
Optimizing the Size of Subword Vocabularies in Dialect Classification
Kanjirangat, Vani and Samard z i \'c , Tanja and Dolamic, Ljiljana and Rinaldi, Fabio. Optimizing the Size of Subword Vocabularies in Dialect Classification. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[92]
Murreviikko - A Dialectologically Annotated and Normalized Dataset of F innish Tweets
Kuparinen, Olli. Murreviikko - A Dialectologically Annotated and Normalized Dataset of F innish Tweets. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[93]
Does Manipulating Tokenization Aid Cross-Lingual Transfer? A Study on POS Tagging for Non-Standardized Languages
Blaschke, Verena and Sch. Does Manipulating Tokenization Aid Cross-Lingual Transfer? A Study on POS Tagging for Non-Standardized Languages. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[94]
Temporal Domain Adaptation for Historical I rish
Dereza, Oksana and Fransen, Theodorus and Mccrae, John P. Temporal Domain Adaptation for Historical I rish. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[95]
Variation and Instability in Dialect-Based Embedding Spaces
Dunn, Jonathan. Variation and Instability in Dialect-Based Embedding Spaces. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[96]
PALI : A Language Identification Benchmark for P erso- A rabic Scripts
Ahmadi, Sina and Agarwal, Milind and Anastasopoulos, Antonios. PALI : A Language Identification Benchmark for P erso- A rabic Scripts. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[97]
Get to Know Your Parallel Data: Performing E nglish Variety and Genre Classification over M a C o C u Corpora
Kuzman, Taja and Rupnik, Peter and Ljube s i \'c , Nikola. Get to Know Your Parallel Data: Performing E nglish Variety and Genre Classification over M a C o C u Corpora. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[98]
Reconstructing Language History by Using a Phonological Ontology
Fischer, Hanna and Engsterhold, Robert. Reconstructing Language History by Using a Phonological Ontology. An Analysis of G erman Surnames. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[99]
BENCH i \'c -lang: A Benchmark for Discriminating between B osnian, C roatian, M ontenegrin and S erbian
Rupnik, Peter and Kuzman, Taja and Ljube s i \'c , Nikola. BENCH i \'c -lang: A Benchmark for Discriminating between B osnian, C roatian, M ontenegrin and S erbian. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[100]
Comparing and Predicting Eye-tracking Data of M andarin and C antonese
Li, Junlin and Peng, Bo and Hsu, Yu-yin and Chersoni, Emmanuele. Comparing and Predicting Eye-tracking Data of M andarin and C antonese. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[101]
A Measure for Linguistic Coherence in Spatial Language Variation
Lameli, Alfred and Sch. A Measure for Linguistic Coherence in Spatial Language Variation. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[102]
Dialect and Variant Identification as a Multi-Label Classification Task: A Proposal Based on Near-Duplicate Analysis
Bernier-colborne, Gabriel and Goutte, Cyril and Leger, Serge. Dialect and Variant Identification as a Multi-Label Classification Task: A Proposal Based on Near-Duplicate Analysis. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[103]
Fine-Tuning BERT with Character-Level Noise for Zero-Shot Transfer to Dialects and Closely-Related Languages
Srivastava, Aarohi and Chiang, David. Fine-Tuning BERT with Character-Level Noise for Zero-Shot Transfer to Dialects and Closely-Related Languages. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[104]
Lemmatization Experiments on Two Low-Resourced Languages: L ow S axon and O ccitan
Mileti \'c , Aleksandra and Siewert, Janine. Lemmatization Experiments on Two Low-Resourced Languages: L ow S axon and O ccitan. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[105]
The Use of Khislavichi Lect Morphological Tagging to Determine its Position in the E ast S lavic Group
Afanasev, Ilia. The Use of Khislavichi Lect Morphological Tagging to Determine its Position in the E ast S lavic Group. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[106]
D iatop I t: A Corpus of Social Media Posts for the Study of Diatopic Language Variation in I taly
Ramponi, Alan and Casula, Camilla. D iatop I t: A Corpus of Social Media Posts for the Study of Diatopic Language Variation in I taly. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[107]
Dialect Representation Learning with Neural Dialect-to-Standard Normalization
Kuparinen, Olli and Scherrer, Yves. Dialect Representation Learning with Neural Dialect-to-Standard Normalization. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[108]
V ar D ial in the Wild: Industrial Applications of LID Systems for Closely-Related Language Varieties
Hohl, Fritz and Shim, Soh-eun. V ar D ial in the Wild: Industrial Applications of LID Systems for Closely-Related Language Varieties. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[109]
Two-stage Pipeline for Multilingual Dialect Detection
Vaidya, Ankit and Kane, Aditya. Two-stage Pipeline for Multilingual Dialect Detection. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[110]
Using Ensemble Learning in Language Variety Identification
Gaman, Mihaela. Using Ensemble Learning in Language Variety Identification. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[111]
SIDLR : Slot and Intent Detection Models for Low-Resource Language Varieties
Kwon, Sang Yun and Bhatia, Gagan and Nagoudi, Elmoatez Billah and Alcoba Inciarte, Alcides and Abdul-mageed, Muhammad. SIDLR : Slot and Intent Detection Models for Low-Resource Language Varieties. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[112]
Findings of the V ar D ial Evaluation Campaign 2023
Aepli, No. Findings of the V ar D ial Evaluation Campaign 2023. Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023). 2023
2023
-
[113]
Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[114]
Introducing U ber T ext 2.0: A Corpus of M odern U krainian at Scale
Chaplynskyi, Dmytro. Introducing U ber T ext 2.0: A Corpus of M odern U krainian at Scale. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[115]
Contextual Embeddings for U krainian: A Large Language Model Approach to Word Sense Disambiguation
Laba, Yurii and Mudryi, Volodymyr and Chaplynskyi, Dmytro and Romanyshyn, Mariana and Dobosevych, Oles. Contextual Embeddings for U krainian: A Large Language Model Approach to Word Sense Disambiguation. Proceedings of the Second Ukrainian Natural Language Processing Workshop ...
2023
-
[116]
Learning Word Embeddings for U krainian: A Comparative Study of F ast T ext Hyperparameters
Romanyshyn, Nataliia and Chaplynskyi, Dmytro and Zakharov, Kyrylo. Learning Word Embeddings for U krainian: A Comparative Study of F ast T ext Hyperparameters. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[117]
GPT -2 Metadata Pretraining Towards Instruction Finetuning for U krainian
Kyrylov, Volodymyr and Chaplynskyi, Dmytro. GPT -2 Metadata Pretraining Towards Instruction Finetuning for U krainian. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[118]
The Evolution of Pro-Kremlin Propaganda From a Machine Learning and Linguistics Perspective
Solopova, Veronika and Benzm. The Evolution of Pro-Kremlin Propaganda From a Machine Learning and Linguistics Perspective. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[119]
Abstractive Summarization for the U krainian Language: Multi-Task Learning with Hromadske.ua News Dataset
Galeshchuk, Svitlana. Abstractive Summarization for the U krainian Language: Multi-Task Learning with Hromadske.ua News Dataset. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[120]
Extension M ulti30 K : Multimodal Dataset for Integrated Vision and Language Research in U krainian
Saichyshyna, Nataliia and Maksymenko, Daniil and Turuta, Oleksii and Yerokhin, Andriy and Babii, Andrii and Turuta, Olena. Extension M ulti30 K : Multimodal Dataset for Integrated Vision and Language Research in U krainian. Proceedings of the Second Ukrainian Natural Language ...
2023
-
[121]
Silver Data for Coreference Resolution in U krainian: Translation, Alignment, and Projection
Kuchmiichuk, Pavlo. Silver Data for Coreference Resolution in U krainian: Translation, Alignment, and Projection. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[122]
Exploring Word Sense Distribution in U krainian with a Semantic Vector Space Model
Cheilytko, Nataliia and von Waldenfels, Ruprecht. Exploring Word Sense Distribution in U krainian with a Semantic Vector Space Model. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[123]
The Parliamentary Code-Switching Corpus: Bilingualism in the U krainian Parliament in the 1990s-2020s
Kanishcheva, Olha and Kovalova, Tetiana and Shvedova, Maria and von Waldenfels, Ruprecht. The Parliamentary Code-Switching Corpus: Bilingualism in the U krainian Parliament in the 1990s-2020s. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[124]
Creating a POS Gold Standard Corpus of M odern U krainian
Starko, Vasyl and Rysin, Andriy. Creating a POS Gold Standard Corpus of M odern U krainian. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[125]
UA - GEC : Grammatical Error Correction and Fluency Corpus for the U krainian Language
Syvokon, Oleksiy and Nahorna, Olena and Kuchmiichuk, Pavlo and Osidach, Nastasiia. UA - GEC : Grammatical Error Correction and Fluency Corpus for the U krainian Language. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[126]
Comparative Study of Models Trained on Synthetic Data for U krainian Grammatical Error Correction
Bondarenko, Maksym and Yushko, Artem and Shportko, Andrii and Fedorych, Andrii. Comparative Study of Models Trained on Synthetic Data for U krainian Grammatical Error Correction. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[127]
A Low-Resource Approach to the Grammatical Error Correction of U krainian
Gomez, Frank and Rozovskaya, Alla and Roth, Dan. A Low-Resource Approach to the Grammatical Error Correction of U krainian. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[128]
R ed P en N et for Grammatical Error Correction: Outputs to Tokens, Attentions to Spans
Didenko, Bohdan and Sameliuk, Andrii. R ed P en N et for Grammatical Error Correction: Outputs to Tokens, Attentions to Spans. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[129]
The UNLP 2023 Shared Task on Grammatical Error Correction for U krainian
Syvokon, Oleksiy and Romanyshyn, Mariana. The UNLP 2023 Shared Task on Grammatical Error Correction for U krainian. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023
2023
-
[130]
Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[131]
Building a U niversal D ependencies Treebank for a Polysynthetic Language: the Case of A baza
Koshevoy, Alexey and Panova, Anastasia and Makarchuk, Ilya. Building a U niversal D ependencies Treebank for a Polysynthetic Language: the Case of A baza. Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[132]
Universalising L atin U niversal D ependencies: a harmonisation of L atin treebanks in UD
Gamba, Federica and Zeman, Daniel. Universalising L atin U niversal D ependencies: a harmonisation of L atin treebanks in UD. Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[133]
S inhala Dependency Treebank ( STB )
Liyanage, Chamila and Sarveswaran, Kengatharaiyer and Nadungodage, Thilini and Pushpananda, Randil. S inhala Dependency Treebank ( STB ). Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[134]
Constantinides, Nicolaos and Stamou, Vivian and Arampatzakis, Vasileios and G
Markantonatou, Stella and Th. Constantinides, Nicolaos and Stamou, Vivian and Arampatzakis, Vasileios and G. Krimpas, Panagiotis and Pavlidis, George. Methodological issues regarding the semi-automatic UD treebank creation of under-resourced languages: the case of Pomak. Proce...
2023
-
[135]
Analysis of Corpus-based Word-Order Typological Methods
Alves, Diego and Bekavac, Bo z o and Zeman, Daniel and Tadi \'c , Marko. Analysis of Corpus-based Word-Order Typological Methods. Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[136]
Findlay, Jamie and Salimifar, Saeedeh and Y ld r m, Ahmet and T
Y. Findlay, Jamie and Salimifar, Saeedeh and Y ld r m, Ahmet and T. T. Haug, Dag. Rule-based semantic interpretation for U niversal D ependencies. Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[137]
Are UD Treebanks Getting More Consistent? A Report Card for E nglish UD
Zeldes, Amir and Schneider, Nathan. Are UD Treebanks Getting More Consistent? A Report Card for E nglish UD. Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[138]
Introducing Morphology in U niversal D ependencies J apanese
Taguchi, Chihiro and Chiang, David. Introducing Morphology in U niversal D ependencies J apanese. Proceedings of the Sixth Workshop on Universal Dependencies (UDW, GURT/SyntaxFest 2023). 2023
2023
-
[139]
Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[140]
Corpus-Based Multilingual Event-type Ontology: Annotation Tools and Principles
Fu c \' kov \'a , Eva and Haji c , Jan and Ure s ov \'a , Zde n ka. Corpus-Based Multilingual Event-type Ontology: Annotation Tools and Principles. Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[141]
S panish Verbal Synonyms in the S yn S em C lass Ontology
Fern \'a ndez-Alcaina, Cristina and Fu c \' kov \'a , Eva and Haji c , Jan and Ure s ov \'a , Zde n ka. S panish Verbal Synonyms in the S yn S em C lass Ontology. Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[142]
Hedging in diachrony: the case of V edic S anskrit iva
Biagetti, Erica and Hellwig, Oliver and Sellmer, Sven. Hedging in diachrony: the case of V edic S anskrit iva. Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[143]
Is J apanese CCGB ank empirically correct? A case study of passive and causative constructions
Bekki, Daisuke and Yanaka, Hitomi. Is J apanese CCGB ank empirically correct? A case study of passive and causative constructions. Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[144]
ICON : Building a Large-Scale Benchmark Constituency Treebank for the I ndonesian Language
Suan Lim, Ee and Qi Leong, Wei and Thanh Nguyen, Ngan and Adhista, Dea and Ming Kng, Wei and Chandra Tjh, William and Purwarianti, Ayu. ICON : Building a Large-Scale Benchmark Constituency Treebank for the I ndonesian Language. Proceedings of the 21st International Workshop on...
2023
-
[145]
Parsing Early N ew H igh G erman: Benefits and limitations of cross-dialectal training
Sapp, Christopher and Dakota, Daniel and Evans, Elliott. Parsing Early N ew H igh G erman: Benefits and limitations of cross-dialectal training. Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[146]
Manning, Christopher
Bauer, John and Kiddon, Chlo \'e and Yeh, Eric and Shan, Alex and D. Manning, Christopher. Semgrex and Ssurgeon, Searching and Manipulating Dependency Graphs. Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023). 2023
2023
-
[147]
Bonn, Julia and Myers, Skatje and E. L. Van Gysel, Jens and Denk, Lukas and Vigus, Meagan and Zhao, Jin and Cowell, Andrew and Croft, William and Haji c , Jan and H. Martin, James and Palmer, Alexis and Palmer, Martha and Pustejovsky, James and Ure s ov \'a , Zdenka and Vallej...
2023
-
[148]
Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[149]
You Can Have Your Data and Balance It Too: Towards Balanced and Efficient Multilingual Models
Limisiewicz, Tomasz and Malkin, Dan and Stanovsky, Gabriel. You Can Have Your Data and Balance It Too: Towards Balanced and Efficient Multilingual Models. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[150]
Multilingual End-to-end Dependency Parsing with Linguistic Typology knowledge
Choudhary, Chinmay and O ' riordan, Colm. Multilingual End-to-end Dependency Parsing with Linguistic Typology knowledge. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[151]
Identifying the Correlation Between Language Distance and Cross-Lingual Transfer in a Multilingual Representation Space
Philippy, Fred and Guo, Siwen and Haddadan, Shohreh. Identifying the Correlation Between Language Distance and Cross-Lingual Transfer in a Multilingual Representation Space. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[152]
Using Modern Languages to Parse Ancient Ones: a Test on O ld E nglish
Brigada Villa, Luca and Giarda, Martina. Using Modern Languages to Parse Ancient Ones: a Test on O ld E nglish. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[153]
The Denglisch Corpus of G erman- E nglish Code-Switching
Osmelak, Doreen and Wintner, Shuly. The Denglisch Corpus of G erman- E nglish Code-Switching. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[154]
Trimming Phonetic Alignments Improves the Inference of Sound Correspondence Patterns from Multilingual Wordlists
Blum, Frederic and List, Johann-Mattis. Trimming Phonetic Alignments Improves the Inference of Sound Correspondence Patterns from Multilingual Wordlists. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[155]
Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP
A Crosslinguistic Database for Combinatorial and Semantic Properties of Attitude Predicates. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[156]
Corpus-based Syntactic Typological Methods for Dependency Parsing Improvement
Alves, Diego and Bekavac, Bo z o and Zeman, Daniel and Tadi \'c , Marko. Corpus-based Syntactic Typological Methods for Dependency Parsing Improvement. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[157]
Cross-lingual Transfer Learning with P ersian
Mollanorozy, Sepideh and Tanti, Marc and Nissim, Malvina. Cross-lingual Transfer Learning with P ersian. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[158]
and Klakow, Dietrich
Steuer, Julius and List, Johann-Mattis and Abdullah, Badr M. and Klakow, Dietrich. Information-Theoretic Characterization of Vowel Harmony: A Cross-Linguistic Study on Word Lists. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual...
2023
-
[159]
Revisiting Dependency Length and Intervener Complexity Minimisation on a Parallel Corpus in 35 Languages
Dyer, Andrew Thomas. Revisiting Dependency Length and Intervener Complexity Minimisation on a Parallel Corpus in 35 Languages. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[160]
Does Topological Ordering of Morphological Segments Reduce Morphological Modeling Complexity? A Preliminary Study on 13 Languages
Shcherbakov, Andreas and Vylomova, Ekaterina. Does Topological Ordering of Morphological Segments Reduce Morphological Modeling Complexity? A Preliminary Study on 13 Languages. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[161]
Findings of the SIGTYP 2023 Shared task on Cognate and Derivative Detection For Low-Resourced Languages
Rani, Priya and Goswami, Koustava and Doyle, Adrian and Fransen, Theodorus and Stearns, Bernardo and McCrae, John P. Findings of the SIGTYP 2023 Shared task on Cognate and Derivative Detection For Low-Resourced Languages. Proceedings of the 5th Workshop on Research in Computat...
2023
-
[162]
\'U FAL Submission for SIGTYP Supervised Cognate Detection Task
Limisiewicz, Tomasz. \'U FAL Submission for SIGTYP Supervised Cognate Detection Task. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[163]
and Iordache, Ioan-Bogdan and Uban, Ana Sabina
Dinu, Liviu P. and Iordache, Ioan-Bogdan and Uban, Ana Sabina. C o T o H i L i at SIGTYP 2023: Ensemble Models for Cognate and Derivative Words Detection. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[164]
Multilingual BERT has an Accent: Evaluating E nglish Influences on Fluency in Multilingual Models
Papadimitriou, Isabel and Lopez, Kezia and Jurafsky, Dan. Multilingual BERT has an Accent: Evaluating E nglish Influences on Fluency in Multilingual Models. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[165]
and Blasi, Dami \'a n and Skirg rd, Hedvig and Greenhill, Simon J
Haynie, Hannah J. and Blasi, Dami \'a n and Skirg rd, Hedvig and Greenhill, Simon J. and Atkinson, Quentin D. and Gray, Russell D. Grambank ' s Typological Advances Support Computational Research on Diverse Languages. Proceedings of the 5th Workshop on Research in Computationa...
2023
-
[166]
and Goldwater, Sharon
Haley, Coleman and Ponti, Edoardo M. and Goldwater, Sharon. Language-Agnostic Measures Discriminate Inflection and Derivation. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[167]
Gradual Language Model Adaptation Using Fine-Grained Typology
Fekete, Marcell Richard and Bjerva, Johannes. Gradual Language Model Adaptation Using Fine-Grained Typology. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[168]
and Shaik, Mohammed Maqsood and Klakow, Dietrich
Abdullah, Badr M. and Shaik, Mohammed Maqsood and Klakow, Dietrich. On the Nature of Discrete Speech Representations in Multilingual Self-supervised Models. Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP. 2023
2023
-
[169]
Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[170]
Ableist Language Teching over Sign Language Research
B. Ableist Language Teching over Sign Language Research. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[171]
The DA - ELEXIS Corpus - a Sense-Annotated Corpus for D anish with Parallel Annotations for Nine E uropean Languages
Pedersen, Bolette and Nimb, Sanni and Olsen, Sussi and Troelsg. The DA - ELEXIS Corpus - a Sense-Annotated Corpus for D anish with Parallel Annotations for Nine E uropean Languages. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languag...
2023
-
[172]
Sentiment Analysis Using Aligned Word Embeddings for U ralic Languages
Alnajjar, Khalid and H. Sentiment Analysis Using Aligned Word Embeddings for U ralic Languages. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[173]
What Causes Unemployment? Unsupervised Causality Mining from S wedish Governmental Reports
D. What Causes Unemployment? Unsupervised Causality Mining from S wedish Governmental Reports. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[174]
Are There Any Limits to E nglish- S wedish Language Transfer? A Fine-grained Analysis Using Natural Language Inference
Morger, Felix. Are There Any Limits to E nglish- S wedish Language Transfer? A Fine-grained Analysis Using Natural Language Inference. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[175]
Word Substitution with Masked Language Models as Data Augmentation for Sentiment Analysis
Kolesnichenko, Larisa and Velldal, Erik and vrelid, Lilja. Word Substitution with Masked Language Models as Data Augmentation for Sentiment Analysis. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[176]
A Large N orwegian Dataset for Weak Supervision ASR
Solberg, Per Erik and Beauguitte, Pierre and Kummervold, Per Egil and Wetjen, Freddy. A Large N orwegian Dataset for Weak Supervision ASR. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[177]
Lexical Semantics with Vector Symbolic Architectures
Roussel, Adam. Lexical Semantics with Vector Symbolic Architectures. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[178]
Linked Open Data compliant Representation of the Interlinking of N ordic Wordnets and Sign Language Data
Declerck, Thierry and Olsen, Sussi. Linked Open Data compliant Representation of the Interlinking of N ordic Wordnets and Sign Language Data. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[179]
Part-of-Speech tagging S panish S ign L anguage data and its applications in Sign Language machine translation
McGill, Euan and Chiruzzo, Luis and Egea G \'o mez, Santiago and Saggion, Horacio. Part-of-Speech tagging S panish S ign L anguage data and its applications in Sign Language machine translation. Proceedings of the Second Workshop on Resources and Representations for Under-Reso...
2023
-
[180]
A Diagnostic Dataset for Sentiment and Negation Modeling for N orwegian
M hlum, Petter and Velldal, Erik and vrelid, Lilja. A Diagnostic Dataset for Sentiment and Negation Modeling for N orwegian. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[181]
Building O kinawan Lexicon Resource for Language Reclamation/Revitalization and Natural Language Processing Tasks such as U niversal D ependencies Treebanking
Miyagawa, So and Kato, Kanji and Zlazli, Miho and Carlino, Salvatore and Machida, Seira. Building O kinawan Lexicon Resource for Language Reclamation/Revitalization and Natural Language Processing Tasks such as U niversal D ependencies Treebanking. Proceedings of the Second Wo...
2023
-
[182]
Bridging the Resource Gap: Exploring the Efficacy of E nglish and Multilingual LLM s for S wedish
Holmstr. Bridging the Resource Gap: Exploring the Efficacy of E nglish and Multilingual LLM s for S wedish. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[183]
Phonotactics as an Aid in Low Resource Loan Word Detection and Morphological Analysis in Sakha
M hlum, Petter and Ivanova, Sardana. Phonotactics as an Aid in Low Resource Loan Word Detection and Morphological Analysis in Sakha. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[184]
Vector Norms as an Approximation of Syntactic Complexity
Ek, Adam and Ilinykh, Nikolai. Vector Norms as an Approximation of Syntactic Complexity. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023). 2023
2023
-
[185]
Low-Resource Techniques for Analysing the Rhetorical Structure of S wedish Historical Petitions
Lindqvist, Ellinor and Pettersson, Eva and Nivre, Joakim. Low-Resource Techniques for Analysing the Rhetorical Structure of S wedish Historical Petitions. Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL...
2023
-
[186]
Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[187]
Automated Claim Detection for Fact-checking: A Case Study using N orwegian Pre-trained Language Models
Sheikhi, Ghazaal and Touileb, Samia and Khan, Sohail. Automated Claim Detection for Fact-checking: A Case Study using N orwegian Pre-trained Language Models. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[188]
Evaluating the Impact of Text De-Identification on Downstream NLP Tasks
Lothritz, Cedric and Lebichot, Bertrand and Allix, Kevin and Ezzini, Saad and Bissyand \'e , Tegawend \'e and Klein, Jacques and Boytsov, Andrey and Lefebvre, Cl \'e ment and Goujon, Anne. Evaluating the Impact of Text De-Identification on Downstream NLP Tasks. Proceedings of ...
2023
-
[189]
Abstractive Text Summarization for I celandic
Sverrisson, \'o r and Einarsson, Hafsteinn. Abstractive Text Summarization for I celandic. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[190]
ASR Language Resources for F aroese
Hern \'a ndez Mena, Carlos and Simonsen, Annika and Gudnason, Jon. ASR Language Resources for F aroese. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[191]
Good Reads and Easy Novels: Readability and Literary Quality in a Corpus of US -published Fiction
Bizzoni, Yuri and Moreira, Pascale and Dwenger, Nicole and Lassen, Ida and Thomsen, Mads and Nielbo, Kristoffer. Good Reads and Easy Novels: Readability and Literary Quality in a Corpus of US -published Fiction. Proceedings of the 24th Nordic Conference on Computational Lingui...
2023
-
[192]
Detection and attribution of quotes in F innish news media: BERT vs
Janicki, Maciej and Kanner, Antti and M. Detection and attribution of quotes in F innish news media: BERT vs. rule-based approach. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[193]
Dyslexia Prediction from Natural Reading of D anish Texts
Bj. Dyslexia Prediction from Natural Reading of D anish Texts. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[194]
Is Part-of-Speech Tagging a Solved Problem for I celandic?
K. Is Part-of-Speech Tagging a Solved Problem for I celandic?. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[195]
Multi- C ross RE A Multi-Lingual Multi-Domain Dataset for Relation Extraction
Bassignana, Elisa and Ginter, Filip and Pyysalo, Sampo and Goot, Rob and Plank, Barbara. Multi- C ross RE A Multi-Lingual Multi-Domain Dataset for Relation Extraction. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[196]
Microservices at Your Service: Bridging the Gap between NLP Research and Industry
Lindh-Knuutila, Tiina and Loftsson, Hrafn and Alonso Doval, Pedro and Andersson, Sebastian and Barkarson, Bjarni and Cerezo-Costas, H. Microservices at Your Service: Bridging the Gap between NLP Research and Industry. Proceedings of the 24th Nordic Conference on Computational ...
2023
-
[197]
Slaapte or Sliep? Extending Neural-Network Simulations of E nglish Past Tense Learning to D utch and G erman
Yang, Xiulin and Chen, Jingyan and van Eerden, Arjan and Samin, Ahnaf and Bisazza, Arianna. Slaapte or Sliep? Extending Neural-Network Simulations of E nglish Past Tense Learning to D utch and G erman. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoD...
2023
-
[198]
Class Explanations: the Role of Domain-Specific Content and Stop Words
Saynova, Denitsa and Bruinsma, Bastiaan and Johansson, Moa and Johansson, Richard. Class Explanations: the Role of Domain-Specific Content and Stop Words. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[199]
Constructing Pseudo-parallel S wedish Sentence Corpora for Automatic Text Simplification
Holmer, Daniel and Rennes, Evelina. Constructing Pseudo-parallel S wedish Sentence Corpora for Automatic Text Simplification. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[200]
Who said what? Speaker Identification from Anonymous Minutes of Meetings
Holmer, Daniel and Ahrenberg, Lars and Monsen, Julius and J. Who said what? Speaker Identification from Anonymous Minutes of Meetings. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[201]
On the Concept of Resource-Efficiency in NLP
D. On the Concept of Resource-Efficiency in NLP. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[202]
Identifying Token-Level Dialectal Features in Social Media
Barnes, Jeremy and Touileb, Samia and M hlum, Petter and Lison, Pierre. Identifying Token-Level Dialectal Features in Social Media. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[203]
N or Q u AD : N orwegian Question Answering Dataset
Ivanova, Sardana and Andreassen, Fredrik and Jentoft, Matias and Wold, Sondre and vrelid, Lilja. N or Q u AD : N orwegian Question Answering Dataset. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[204]
Extracting Sign Language Articulation from Videos with M edia P ipe
B. Extracting Sign Language Articulation from Videos with M edia P ipe. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[205]
Named Entity layer in E stonian UD treebanks
Muischnek, Kadri and M. Named Entity layer in E stonian UD treebanks. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[206]
S cand E val: A Benchmark for S candinavian Natural Language Processing
Nielsen, Dan. S cand E val: A Benchmark for S candinavian Natural Language Processing. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[207]
BRENT : Bidirectional Retrieval Enhanced N orwegian Transformer
Charpentier, Lucas and Wold, Sondre and Samuel, David and R nningstad, Egil. BRENT : Bidirectional Retrieval Enhanced N orwegian Transformer. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[208]
Machine vs
Shaitarova, Anastassia and G. Machine vs. Human: Exploring Syntax and Lexicon in G erman Translations, with a Spotlight on Anglicisms. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[209]
Training and Evaluating N orwegian Sentence Embedding Models
N dland, Bernt Ivar Utst l. Training and Evaluating N orwegian Sentence Embedding Models. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[210]
Dozens of Translation Directions or Millions of Shared Parameters? Comparing Two Types of Multilinguality in Modular Machine Translation
Boggia, Michele and Gr. Dozens of Translation Directions or Millions of Shared Parameters? Comparing Two Types of Multilinguality in Modular Machine Translation. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[211]
D an S um T 5: Automatic Abstractive Summarization for D anish
Kolding, Sara and Nymann, Katrine and Hansen, Ida and Enevoldsen, Kenneth and Kristensen-McLachlan, Ross. D an S um T 5: Automatic Abstractive Summarization for D anish. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[212]
C aptain A - A mobile app for practising F innish pronunciation
Phan, Nhan and Gr \'o sz, Tam \'a s and Kurimo, Mikko. C aptain A - A mobile app for practising F innish pronunciation. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[213]
D an T ok: Domain Beats Language for D anish Social Media POS Tagging
Kirstein Hansen, Kia and Barrett, Maria and M. D an T ok: Domain Beats Language for D anish Social Media POS Tagging. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[214]
Comparison of Current Approaches to Lemmatization: A Case Study in E stonian
Dorkin, Aleksei and Sirts, Kairit. Comparison of Current Approaches to Lemmatization: A Case Study in E stonian. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[215]
Generating Errors: OCR Post-Processing for I celandic
Jasonarson, Atli and Steingr \' msson, Stein \'o r and Sigur sson, Einar and Magn \'u sson, \'A rni and Ingimundarson, Finnur. Generating Errors: OCR Post-Processing for I celandic. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[216]
Generation of Replacement Options in Text Sanitization
Olstad, Annika Willoch and Papadopoulou, Anthi and Lison, Pierre. Generation of Replacement Options in Text Sanitization. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[217]
M e D a- BERT : A medical D anish pretrained transformer model
Pedersen, Jannik and Laursen, Martin and Vinholt, Pernille and Savarimuthu, Thiusius Rajeeth. M e D a- BERT : A medical D anish pretrained transformer model. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[218]
Standardising Pronunciation for a Grapheme-to-Phoneme Converter for F aroese
Lamhauge, Sandra and Debess, Iben and Hern \'a ndez Mena, Carlos and Simonsen, Annika and Gudnason, Jon. Standardising Pronunciation for a Grapheme-to-Phoneme Converter for F aroese. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[219]
Using Membership Inference Attacks to Evaluate Privacy-Preserving Language Modeling Fails for Pseudonymizing Data
Vakili, Thomas and Dalianis, Hercules. Using Membership Inference Attacks to Evaluate Privacy-Preserving Language Modeling Fails for Pseudonymizing Data. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[220]
Sentiment Classification of Historical D anish and N orwegian Literary Texts
Allaith, Ali and Degn, Kirstine and Conroy, Alexander and Pedersen, Bolette and Bjerring-Hansen, Jens and Hershcovich, Daniel. Sentiment Classification of Historical D anish and N orwegian Literary Texts. Proceedings of the 24th Nordic Conference on Computational Linguistics (...
2023
-
[221]
Parser Evaluation for Analyzing S wedish 19th-20th Century Literature
Stymne, Sara and. Parser Evaluation for Analyzing S wedish 19th-20th Century Literature. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[222]
An Empirical Study of Multitask Learning to Improve Open Domain Dialogue Systems
Farahani, Mehrdad and Johansson, Richard. An Empirical Study of Multitask Learning to Improve Open Domain Dialogue Systems. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[223]
Uncertainty-Aware Natural Language Inference with Stochastic Weight Averaging
Talman, Aarne and Celikkanat, Hande and Virpioja, Sami and Heinonen, Markus and Tiedemann, J. Uncertainty-Aware Natural Language Inference with Stochastic Weight Averaging. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[224]
Alignment of W ikidata lexemes and Det Centrale Ordregister
Nielsen, Finn. Alignment of W ikidata lexemes and Det Centrale Ordregister. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[225]
Low-resource Bilingual Dialect Lexicon Induction with Large Language Models
Artemova, Katya and Plank, Barbara. Low-resource Bilingual Dialect Lexicon Induction with Large Language Models. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[226]
Constructing a Knowledge Graph from Textual Descriptions of Software Vulnerabilities in the National Vulnerability Database
H st, Anders and Lison, Pierre and Moonen, Leon. Constructing a Knowledge Graph from Textual Descriptions of Software Vulnerabilities in the National Vulnerability Database. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[227]
A Survey of Corpora for G ermanic Low-Resource Languages and Dialects
Blaschke, Verena and Schuetze, Hinrich and Plank, Barbara. A Survey of Corpora for G ermanic Low-Resource Languages and Dialects. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[228]
You say tomato, I say the same: A large-scale study of linguistic accommodation in online communities
Berdicevskis, Aleksandrs and Erbro, Viktor. You say tomato, I say the same: A large-scale study of linguistic accommodation in online communities. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[229]
Integrating rules and neural nets for morphological tagging of N orwegian - Results and challenges
Haug, Dag and Yildirim, Ahmet and Hagen, Kristin and N klestad, Anders. Integrating rules and neural nets for morphological tagging of N orwegian - Results and challenges. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[230]
Comparing Methods for Segmenting Elementary Discourse Units in a F rench Conversational Corpus
Prevot, Laurent and Hunter, Julie and Muller, Philippe. Comparing Methods for Segmenting Elementary Discourse Units in a F rench Conversational Corpus. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[231]
Multi-way Variational NMT for UGC : Improving Robustness in Zero-shot Scenarios via Mixture Density Networks
Rosales N \'u \ n ez, Jos \'e and Seddah, Djam \'e and Wisniewski, Guillaume. Multi-way Variational NMT for UGC : Improving Robustness in Zero-shot Scenarios via Mixture Density Networks. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[232]
Multilingual Automatic Speech Recognition for S candinavian Languages
Cerniavski, Rafal and Stymne, Sara. Multilingual Automatic Speech Recognition for S candinavian Languages. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[233]
A character-based analysis of impacts of dialects on end-to-end N orwegian ASR
Parsons, Phoebe and Kvale, Knut and Svendsen, Torbj rn and Salvi, Giampiero. A character-based analysis of impacts of dialects on end-to-end N orwegian ASR. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[234]
Quasi: a synthetic Question-Answering dataset in S wedish using GPT -3 and zero-shot learning
Kalpakchi, Dmytro and Boye, Johan. Quasi: a synthetic Question-Answering dataset in S wedish using GPT -3 and zero-shot learning. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[235]
Automatic Closed Captioning for E stonian Live Broadcasts
Alum. Automatic Closed Captioning for E stonian Live Broadcasts. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[236]
The Effect of Data Encoding on Relation Triplet Identification
Fri riksd \'o ttir, Steinunn and Einarsson, Hafsteinn. The Effect of Data Encoding on Relation Triplet Identification. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[237]
Improving Generalization of N orwegian ASR with Limited Linguistic Resources
Solberg, Per Erik and Ortiz, Pablo and Parsons, Phoebe and Svendsen, Torbj rn and Salvi, Giampiero. Improving Generalization of N orwegian ASR with Limited Linguistic Resources. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[238]
The Finer They Get: Combining Fine-Tuned Models For Better Semantic Change Detection
Zhou, Wei and Tahmasebi, Nina and Dubossarsky, Haim. The Finer They Get: Combining Fine-Tuned Models For Better Semantic Change Detection. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[239]
Question Answering and Question Generation for F innish
Kylli. Question Answering and Question Generation for F innish. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[240]
Probing structural constraints of negation in Pretrained Language Models
Kletz, David and Candito, Marie and Amsili, Pascal. Probing structural constraints of negation in Pretrained Language Models. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[241]
Boosting N orwegian Automatic Speech Recognition
De La Rosa, Javier and Braaten, Rolv-Arild and Kummervold, Per and Wetjen, Freddy. Boosting N orwegian Automatic Speech Recognition. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[242]
Length Dependence of Vocabulary Richness
Zechner, Niklas. Length Dependence of Vocabulary Richness. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[243]
A query engine for L 1- L 2 parallel dependency treebanks
Masciolini, Arianna. A query engine for L 1- L 2 parallel dependency treebanks. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[244]
Filtering Matters: Experiments in Filtering Training Sets for Machine Translation
Steingr \' msson, Stein \'o r and Loftsson, Hrafn and Way, Andy. Filtering Matters: Experiments in Filtering Training Sets for Machine Translation. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[245]
Gamli - I celandic Oral History Corpus: Design, Collection and Evaluation
O ' Brien, Luke and Ingimundarson, Finnur and Gu nasson, J \'o n and Steingr \' msson, Stein \'o r. Gamli - I celandic Oral History Corpus: Design, Collection and Evaluation. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[246]
N o C o LA : The N orwegian Corpus of Linguistic Acceptability
Jentoft, Matias and Samuel, David. N o C o LA : The N orwegian Corpus of Linguistic Acceptability. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[247]
N or B ench -- A Benchmark for N orwegian Language Models
Samuel, David and Kutuzov, Andrey and Touileb, Samia and Velldal, Erik and vrelid, Lilja and R nningstad, Egil and Sigdel, Elina and Palatkina, Anna. N or B ench -- A Benchmark for N orwegian Language Models. Proceedings of the 24th Nordic Conference on Computational Linguisti...
2023
-
[248]
Making Instruction Finetuning Accessible to Non- E nglish Languages: A Case Study on S wedish Models
Holmstr. Making Instruction Finetuning Accessible to Non- E nglish Languages: A Case Study on S wedish Models. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[249]
G iella LT --- a stable infrastructure for N ordic minority languages and beyond
Pirinen, Flammie and Moshagen, Sjur and Hiovain-Asikainen, Katri. G iella LT --- a stable infrastructure for N ordic minority languages and beyond. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[250]
Adapting an I celandic morphological database to F aroese
R \'u narsson, Kristj \'a n and Bjarnadottir, Kristin. Adapting an I celandic morphological database to F aroese. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[251]
D anish Clinical Named Entity Recognition and Relation Extraction
Laursen, Martin and Pedersen, Jannik and Hansen, Rasmus and Savarimuthu, Thiusius Rajeeth and Vinholt, Pernille. D anish Clinical Named Entity Recognition and Relation Extraction. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[252]
Scaling-up the Resources for a Freely Available S wedish VADER (sv VADER )
Kokkinakis, Dimitrios and Mu \ n oz S \'a nchez, Ricardo and Hammarlin, Mia-Marie. Scaling-up the Resources for a Freely Available S wedish VADER (sv VADER ). Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[253]
C olex2 L ang: Language Embeddings from Semantic Typology
Chen, Yiyi and Biswas, Russa and Bjerva, Johannes. C olex2 L ang: Language Embeddings from Semantic Typology. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[254]
Toxicity Detection in F innish Using Machine Translation
Eskelinen, Anni and Silvala, Laura and Ginter, Filip and Pyysalo, Sampo and Laippala, Veronika. Toxicity Detection in F innish Using Machine Translation. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[255]
Evaluating a U niversal D ependencies Conversion Pipeline for I celandic
Arnard \'o ttir, \'o runn and Hafsteinsson, Hinrik and Jasonarson, Atli and Ingaon, Anton and Steingr \' msson, Stein \'o r. Evaluating a U niversal D ependencies Conversion Pipeline for I celandic. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLi...
2023
-
[256]
Automatic Transcription for E stonian Children ' s Speech
Luhtaru, Agnes and Jaaska, Rauno and Kruusam. Automatic Transcription for E stonian Children ' s Speech. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[257]
Translated Benchmarks Can Be Misleading: the Case of E stonian Question Answering
Kuulmets, Hele-Andra and Fishel, Mark. Translated Benchmarks Can Be Misleading: the Case of E stonian Question Answering. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[258]
Predicting the presence of inline citations in academic text using binary classification
Vajdecka, Peter and Callegari, Elena and Xhura, Desara and \'A smundsson, Atli. Predicting the presence of inline citations in academic text using binary classification. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[259]
Neural Text-to-Speech Synthesis for V \ o ro
R. Neural Text-to-Speech Synthesis for V \ o ro. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[260]
Transfer to a Low-Resource Language via Close Relatives: The Case Study on F aroese
Sn bjarnarson, V \'e steinn and Simonsen, Annika and Glava s , Goran and Vuli \'c , Ivan. Transfer to a Low-Resource Language via Close Relatives: The Case Study on F aroese. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[261]
Evaluating Morphological Generalisation in Machine Translation by Distribution-Based Compositionality Assessment
Moisio, Anssi and Creutz, Mathias and Kurimo, Mikko. Evaluating Morphological Generalisation in Machine Translation by Distribution-Based Compositionality Assessment. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[262]
E stonian Named Entity Recognition: New Datasets and Models
Sirts, Kairit. E stonian Named Entity Recognition: New Datasets and Models. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[263]
Machine Translation for Low-resource F inno- U gric Languages
Yankovskaya, Lisa and Tars, Maali and T. Machine Translation for Low-resource F inno- U gric Languages. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[264]
Distilling E stonian Text Domains for Production-Oriented Machine Translation
Korotkova, Elizaveta and Fishel, Mark. Distilling E stonian Text Domains for Production-Oriented Machine Translation. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[265]
Spelling Correction for E stonian Learner Language
Allkivi-Metsoja, Kais and Kippar, Jaagup. Spelling Correction for E stonian Learner Language. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa). 2023
2023
-
[266]
Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[267]
M ulti GED -2023 shared task at NLP 4 CALL : Multilingual Grammatical Error Detection
Volodina, Elena and Bryant, Christopher and Caines, Andrew and De Clercq, Orph \'e e and Frey, Jennifer-Carmen and Ershova, Elizaveta and Rosen, Alexandr and Vinogradova, Olga. M ulti GED -2023 shared task at NLP 4 CALL : Multilingual Grammatical Error Detection. Proceedings o...
2023
-
[268]
NTNU - TRH system at the M ulti GED -2023 Shared on Multilingual Grammatical Error Detection
Bungum, Lars and Gamb. NTNU - TRH system at the M ulti GED -2023 Shared on Multilingual Grammatical Error Detection. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[269]
E li C o D e at M ulti GED 2023: fine-tuning XLM - R o BERT a for multilingual grammatical error detection
Colla, Davide and Delsanto, Matteo and Di Nuovo, Elisa. E li C o D e at M ulti GED 2023: fine-tuning XLM - R o BERT a for multilingual grammatical error detection. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[270]
A distantly supervised Grammatical Error Detection/Correction system for S wedish
Kurfal. A distantly supervised Grammatical Error Detection/Correction system for S wedish. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[271]
Two Neural Models for Multilingual Grammatical Error Detection
Le-Hong, Phuong and Ngo, The Quyen and Nguyen, Thi Minh Huyen. Two Neural Models for Multilingual Grammatical Error Detection. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[272]
Experiments on Automatic Error Detection and Correction for Uruguayan Learners of E nglish
Brown, Romina and Paez, Santiago and Herrera, Gonzalo and Chiruzzo, Luis and Ros \'a , Aiala. Experiments on Automatic Error Detection and Correction for Uruguayan Learners of E nglish. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[273]
Sequence Tagging in EFL Email Texts as Feedback for Language Learners
Ding, Yuning and Tr. Sequence Tagging in EFL Email Texts as Feedback for Language Learners. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[274]
Speech Technology to Support Phonics Learning for Kindergarten Children at Risk of Dyslexia
Fuglsang Engmose, Stine and Henrichsen, Peter Juel. Speech Technology to Support Phonics Learning for Kindergarten Children at Risk of Dyslexia. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[275]
On the relevance and learner dependence of co-text complexity for exercise difficulty
Heck, Tanja and Meurers, Detmar. On the relevance and learner dependence of co-text complexity for exercise difficulty. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[276]
Manual and Automatic Identification of Similar Arguments in EFL Learner Essays
Mousa, Ahmed and Laarmann-Quante, Ronja and Horbach, Andrea. Manual and Automatic Identification of Similar Arguments in EFL Learner Essays. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[277]
D a LAJ - GED - a dataset for Grammatical Error Detection tasks on S wedish
Volodina, Elena and Mohammed, Yousuf Ali and Berdicevskis, Aleksandrs and Bouma, Gerlof and. D a LAJ - GED - a dataset for Grammatical Error Detection tasks on S wedish. Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning. 2023
2023
-
[278]
Automated Assessment of Task Completion in Spontaneous Speech for F innish and F inland S wedish Language Learners
Voskoboinik, Ekaterina and Getman, Yaroslav and Al-Ghezi, Ragheb and Kurimo, Mikko and Grosz, Tamas. Automated Assessment of Task Completion in Spontaneous Speech for F innish and F inland S wedish Language Learners. Proceedings of the 12th Workshop on NLP for Computer Assiste...
2023
-
[279]
Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[280]
Token-level Identification of Multiword Expressions using Pre-trained Multilingual Language Models
Swaminathan, Raghuraman and Cook, Paul. Token-level Identification of Multiword Expressions using Pre-trained Multilingual Language Models. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[281]
R omanian Multiword Expression Detection Using Multilingual Adversarial Training and Lateral Inhibition
Avram, Andrei and Barbu Mititelu, Verginica and Cercel, Dumitru-Clementin. R omanian Multiword Expression Detection Using Multilingual Adversarial Training and Lateral Inhibition. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[282]
Predicting Compositionality of Verbal Multiword Expressions in P ersian
Sarlak, Mahtab and Yarandi, Yalda and Shamsfard, Mehrnoush. Predicting Compositionality of Verbal Multiword Expressions in P ersian. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[283]
PARSEME corpus release 1.3
Savary, Agata and Ben Khelil, Cherifa and Ramisch, Carlos and Giouli, Voula and Barbu Mititelu, Verginica and Hadj Mohamed, Najet and Krstev, Cvetana and Liebeskind, Chaya and Xu, Hongzhi and Stymne, Sara and G. PARSEME corpus release 1.3. Proceedings of the 19th Workshop on M...
2023
-
[284]
Investigating the Effects of MWE Identification in Structural Topic Modelling
Kokkinakis, Dimitrios and S \'a nchez, Ricardo and Bruinsma, Sebastianus and Hammarlin, Mia-Marie. Investigating the Effects of MWE Identification in Structural Topic Modelling. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[285]
Idioms, Probing and Dangerous Things: Towards Structural Probing for Idiomaticity in Vector Space
Klubi c ka, Filip and Nedumpozhimana, Vasudevan and Kelleher, John. Idioms, Probing and Dangerous Things: Towards Structural Probing for Idiomaticity in Vector Space. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[286]
Graph-based multi-layer querying in Parseme Corpora
Guillaume, Bruno. Graph-based multi-layer querying in Parseme Corpora. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[287]
Enriching Multiword Terms in W iktionary with Pronunciation Information
Bajcetic, Lenka and Declerck, Thierry and S \'e rasset, Gilles. Enriching Multiword Terms in W iktionary with Pronunciation Information. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[288]
Detecting Idiomatic Multiword Expressions in Clinical Terminology using Definition-Based Representation Learning
Remy, Fran c ois and Khabibullina, Alfiya and Demeester, Thomas. Detecting Idiomatic Multiword Expressions in Clinical Terminology using Definition-Based Representation Learning. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[289]
Automatic Generation of Vocabulary Lists with Multiword Expressions
Lee, John and Uvaliyev, Adilet. Automatic Generation of Vocabulary Lists with Multiword Expressions. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[290]
Rambelli, Giulia and Chersoni, Emmanuele and Senaldi, Marco S. G. and Blache, Philippe and Lenci, Alessandro. Are Frequent Phrases Directly Retrieved like Idioms? An Investigation with Self-Paced Reading and Language Models. Proceedings of the 19th Workshop on Multiword Expres...
2023
-
[291]
Annotation of lexical bundles with discourse functions in a S panish academic corpus
Guzzi, Eleonora and Alonso-Ramos, Margarita and Garcia, Marcos and Garc \' a Salido, Marcos. Annotation of lexical bundles with discourse functions in a S panish academic corpus. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[292]
A Survey of MWE Identification Experiments: The Devil is in the Details
Ramisch, Carlos and Walsh, Abigail and Blanchard, Thomas and Taslimipoor, Shiva. A Survey of MWE Identification Experiments: The Devil is in the Details. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[293]
A MWE lexicon formalism optimised for observational adequacy
Lion-Bouton, Adam and Savary, Agata and Antoine, Jean-Yves. A MWE lexicon formalism optimised for observational adequacy. Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). 2023
2023
-
[294]
Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2023). 2023
2023
-
[295]
Train Global, Tailor Local: Minimalist Multilingual Translation into Endangered Languages
Zhou, Zhong and Niehues, Jan and Waibel, Alexander. Train Global, Tailor Local: Minimalist Multilingual Translation into Endangered Languages. Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2023). 2023
2023
-
[296]
Multilingual Bidirectional Unsupervised Translation through Multilingual Finetuning and Back-Translation
Li, Bryan and Rasooli, Mohammad Sadegh and Patel, Ajay and Callison-burch, Chris. Multilingual Bidirectional Unsupervised Translation through Multilingual Finetuning and Back-Translation. Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Reso...
2023
-
[297]
PEACH : Pre-Training Sequence-to-Sequence Multilingual Models for Translation with Semi-Supervised Pseudo-Parallel Document Generation
Salemi, Alireza and Abaskohi, Amirhossein and Tavakoli, Sara and Shakery, Azadeh and Yaghoobzadeh, Yadollah. PEACH : Pre-Training Sequence-to-Sequence Multilingual Models for Translation with Semi-Supervised Pseudo-Parallel Document Generation. Proceedings of the The Sixth Wor...
2023
-
[298]
and Allemann, Alexis and Dolamic, Ljiljana and Popescu-Belis, Andrei
Atrio, \`A lex R. and Allemann, Alexis and Dolamic, Ljiljana and Popescu-Belis, Andrei. A Simplified Training Pipeline for Low-Resource and Unsupervised Machine Translation. Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages...
2023
-
[299]
Language-Family Adapters for Low-Resource Multilingual Neural Machine Translation
Chronopoulou, Alexandra and Stojanovski, Dario and Fraser, Alexander. Language-Family Adapters for Low-Resource Multilingual Neural Machine Translation. Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2023). 2023
2023
-
[300]
Improving Neural Machine Translation of Indigenous Languages with Multilingual Transfer Learning
Chen, Wei-rui and Abdul-mageed, Muhammad. Improving Neural Machine Translation of Indigenous Languages with Multilingual Transfer Learning. Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2023). 2023
2023
Reviewed May 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.