Pith. sign in

REVIEW 2 major objections 2 minor 71 cited by

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

T0 review · 2 major / 2 minor · reviewed 2026-05-16 · grok-4.3

Pith's one-line read AGIEval benchmark shows GPT-4 surpassing average humans on SAT math at 95 percent and LSAT.

desk verdict AGIEval assembles a useful benchmark from real standardized exams and releases the data, but GPT-4's reported scores need checks for training-data overlap before the generalization claims hold. read the letter →

arxiv 2304.06364 v2 pith:FRJ3HHYI submitted 2023-04-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords AGIEvalfoundationmodelsbenchmarkGPT-4standardizedexamsSATLSATreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces AGIEval, a benchmark that draws questions directly from standardized human exams such as college entrance tests, law school admissions, math competitions, and lawyer qualification exams to measure foundation model abilities on human-level tasks. Evaluations of GPT-4, ChatGPT, and Text-Davinci-003 find that GPT-4 exceeds average human scores on SAT, LSAT, and math competitions, reaching 95 percent accuracy on SAT math and 92.5 percent on the English section of the Chinese national college entrance exam. The models show weaker results on problems that require complex reasoning chains or narrow domain knowledge. The authors break performance into categories of understanding, knowledge, reasoning, and calculation to expose specific strengths and gaps. The approach replaces artificial datasets with tasks tied to real human cognition and decision-making.

What carries the argument

The AGIEval benchmark, assembled from standardized human exams to test foundation models on understanding, knowledge, reasoning, and calculation in human-relevant contexts.

What would settle it

A controlled comparison in which models achieve high AGIEval scores yet fail on equivalent non-exam problems that test the same underlying skills in open-ended or novel settings.

Watch

Extended reading notes

Core claim

AGIEval evaluates foundation models on collections of real standardized exams including SAT, LSAT, math competitions, and lawyer qualification tests. GPT-4 surpasses average human performance on SAT, LSAT, and math competitions, attaining 95 percent accuracy on the SAT Math test and 92.5 percent accuracy on the English test of the Chinese national college entrance exam, while remaining less proficient on tasks that demand complex reasoning or specific domain knowledge.

Load-bearing premise

Standardized human exams serve as valid and unbiased proxies for general cognitive capabilities without favoring current model training methods or test formats.

Editorial extensions

If this is right

  • Foundation models can now solve many exam-style questions at or above average human levels across multiple subjects.
  • Performance gaps appear most clearly in complex reasoning and domain-specific knowledge, guiding targeted improvements.
  • Capability breakdowns by category supply concrete directions for strengthening general abilities.
  • Human-exam benchmarks connect model results more directly to real-world cognitive demands than synthetic tests do.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Sustained high scores could support deployment of models as automated tutors or graders for these exact exams.
  • Gaps in complex reasoning may require architectural additions rather than further scaling alone.
  • Extending the benchmark with harder or culturally varied exam variants could track whether gains generalize.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper introduces AGIEval, a benchmark assembled from publicly available human standardized exams (SAT, LSAT, math competitions, Gaokao, lawyer qualification tests). It evaluates GPT-4, ChatGPT, and Text-Davinci-003, reporting that GPT-4 exceeds average human performance on several tests (95% on SAT Math, 92.5% on Gaokao English) while showing weaker results on complex reasoning and domain-knowledge tasks. Capability breakdowns (understanding/knowledge/reasoning/calculation) and full data/code/output release are provided.

Significance. If the headline numbers survive decontamination checks, the work supplies a more ecologically valid signal of foundation-model progress than synthetic benchmarks and supplies concrete capability diagnostics plus reproducible artifacts. The public release of all model outputs strengthens the contribution.

major comments (2)
  1. [Abstract] Abstract and evaluation section: the 95% SAT-Math and 92.5% Gaokao-English figures are presented without the number of items per test, sampling protocol, exact prompt templates, or any statistical testing; these omissions leave the central claim that GPT-4 surpasses humans only moderately supported.
  2. [Evaluation] Evaluation methodology: no membership-inference, decontamination, or paraphrased-variant experiments are reported for the publicly circulated exam questions, even though the central claim (surpassing humans via reasoning) requires that performance not be explained by training-data overlap.
minor comments (2)
  1. [Figures] Figure captions and axis labels could more explicitly state the human baseline source and sample size for each exam.
  2. [Appendix] A short table summarizing prompt templates per task type would improve reproducibility.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback on our AGIEval benchmark paper. The comments highlight valuable opportunities to strengthen the presentation of results and the evaluation methodology. We have revised the manuscript to incorporate additional details and experiments where feasible, and we respond to each major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract and evaluation section: the 95% SAT-Math and 92.5% Gaokao-English figures are presented without the number of items per test, sampling protocol, exact prompt templates, or any statistical testing; these omissions leave the central claim that GPT-4 surpasses humans only moderately supported.

    Authors: We agree that these supporting details are necessary to substantiate the central claims. In the revised manuscript, we have expanded the evaluation section to report the exact number of items per test (SAT Math: 58 questions; Gaokao English: 40 questions), clarified that evaluations used the full publicly available test sets with no subsampling, included the precise prompt templates in a new appendix, and added statistical testing via binomial proportion tests to confirm that GPT-4's accuracies significantly exceed the reported human averages. These changes provide stronger empirical grounding for the headline figures. revision: yes

  2. Referee: [Evaluation] Evaluation methodology: no membership-inference, decontamination, or paraphrased-variant experiments are reported for the publicly circulated exam questions, even though the central claim (surpassing humans via reasoning) requires that performance not be explained by training-data overlap.

    Authors: We acknowledge the importance of ruling out data contamination to support interpretations of reasoning ability. While full membership-inference or decontamination experiments are not feasible without access to the proprietary training data of the evaluated models, we have added paraphrased-variant experiments on subsets of the SAT and Gaokao questions in the revision; these maintain high performance, indicating robustness beyond exact memorization. We have also expanded the limitations and discussion sections to address contamination risks explicitly, noting the public nature of the exams and known training cutoffs, and we release all model outputs to support community-led analyses. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: benchmark is direct measurement on newly assembled external exam items

full rationale

The paper constructs AGIEval by collecting questions from public standardized exams (SAT, LSAT, Gaokao, math contests) and reports model accuracies as direct empirical measurements against published human averages. No equations, fitted parameters, or predictions are derived; the central claims (e.g., GPT-4 at 95% SAT Math) are simple accuracy counts on the collected items. No self-citations, uniqueness theorems, or ansatzes are invoked to justify results. The derivation chain is therefore self-contained as straightforward benchmarking.

Assumptions & free parameters 0 free parameters · 0 assumptions · 1 invented entities

The paper introduces a new evaluation benchmark without introducing fitted parameters, unstated mathematical axioms, or new physical entities; the only added construct is the benchmark collection itself.

invented entities (1)
  • AGIEval benchmark
    purpose: To provide human-centric standardized exam questions for evaluating foundation models
    The benchmark is assembled and released in this work.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models." pith.science (2026). https://pith.science/paper/FRJ3HHYI

@misc{pith2026230406364,
  author       = {Pith},
  title        = {Pith review of: AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FRJ3HHYI}},
  note         = {Machine review of arXiv:2304.06364}
}
read the original abstract

Evaluating the general abilities of foundation models to tackle human-level tasks is a vital aspect of their development and application in the pursuit of Artificial General Intelligence (AGI). Traditional benchmarks, which rely on artificial datasets, may not accurately represent human-level capabilities. In this paper, we introduce AGIEval, a novel benchmark specifically designed to assess foundation model in the context of human-centric standardized exams, such as college entrance exams, law school admission tests, math competitions, and lawyer qualification tests. We evaluate several state-of-the-art foundation models, including GPT-4, ChatGPT, and Text-Davinci-003, using this benchmark. Impressively, GPT-4 surpasses average human performance on SAT, LSAT, and math competitions, attaining a 95% accuracy rate on the SAT Math test and a 92.5% accuracy on the English test of the Chinese national college entrance exam. This demonstrates the extraordinary performance of contemporary foundation models. In contrast, we also find that GPT-4 is less proficient in tasks that require complex reasoning or specific domain knowledge. Our comprehensive analyses of model capabilities (understanding, knowledge, reasoning, and calculation) reveal these models' strengths and limitations, providing valuable insights into future directions for enhancing their general capabilities. By concentrating on tasks pertinent to human cognition and decision-making, our benchmark delivers a more meaningful and robust evaluation of foundation models' performance in real-world scenarios. The data, code, and all model outputs are released in https://github.com/ruixiangcui/AGIEval.

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Forward citations

Showing 60 of 71 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 71 Pith citations

  1. TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

    cs.CL 2026-08 reject novelty 7.0 of 10

    The paper introduces TCS-Bench, a 300-task proof-generation benchmark from top TCS papers, and reports frontier LLM accuracies from 30% to 68% using an automated verifier.

  2. BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    BLUEX v2 supplies 919 graded subquestions from UNICAMP and USP exams (2022-2025) and reports LLM-as-a-judge scores for 21 models showing a 4.92-point spread, with math reasoning and image understanding as the weakest areas.

  3. VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models

    cs.CL 2025-12 conditional novelty 7.0 of 10

    VLegal-Bench supplies 10,450 expert-validated samples for evaluating LLMs on Vietnamese legal questions, retrieval, multi-step reasoning, and scenario solving.

  4. Towards Efficient and Effective Alignment of Large Language Models

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.

  5. Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese

    cs.CL 2025-05 accept novelty 7.0 of 10

    A new benchmark shows LLMs are more accurate in Simplified Chinese for regional terms but favor Taiwanese names in simulated hiring, revealing task-dependent bias between Chinese script variants.

  6. PRIMETIME : Limits of LLMs in Temporal Primitives

    cs.NE 2025-04 unverdicted novelty 7.0 of 10

    PRIMETIME generator reveals that LLM datetime parsing and arithmetic primitives are individually unreliable but fully learnable via fine-tuning, enabling frontier-level accuracy on event planning with small LoRA models.

  7. L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

    cs.CL 2025-03 unverdicted novelty 7.0 of 10

    LCPO trains L1 reasoning models to adhere to prompt-specified CoT lengths, supporting accuracy-compute trade-offs and yielding short reasoning models that outperform larger baselines at matched lengths.

  8. Scaling Native Multimodal Pre-Training From Scratch

    cs.CL 2026-07 conditional novelty 6.0 of 10

    In models trained from scratch on text plus images, the text-objective scaling law is data-mix-invariant while the image-conditioned objective shifts toward many more tokens relative to parameters as the multimodal sh...

  9. Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Zero-RL multi-stage constructive safety alignment with SERL and long-context training lets a 14B model match much larger models on safety without collapsing helpfulness or style.

  10. Do Value Vectors in Deep Layers Need Context from the Residual Stream?

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Deep transformer layers can replace context-dependent value vectors with per-token lookup tables (Bank of Values), improving validation loss and the 21-benchmark average at 135M–780M while cutting FLOPs and the value cache.

  11. The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    LLM routers across 21 methods on 5 benchmarks converge to similar accuracy below oracle due to learning global performance trends rather than fine-grained query signals.

  12. Confidence Calibration in Large Language Models

    cs.AI 2026-04 conditional novelty 6.0 of 10

    LLMs show average overconfidence moderated by a strong hard-easy effect, quantified across tasks and via the new LifeEval actuarial-probability benchmark.

  13. SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

    cs.AI 2026-02 reject novelty 6.0 of 10

    A new benchmark of 2,703 automatically generated multimodal questions for scanning probe microscopy, plus a modified F1 metric that penalizes over-selection and labels model 'personalities'.

  14. Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study

    cs.CY 2026-02 reject novelty 6.0 of 10

    A 10-question Chinese/English benchmark of 17 LLMs on Taiwan sovereignty finds only GPT-4o Mini passes, with Chinese models uniformly failing and most flagged language bias unconfirmed by the paper's own statistics.

  15. Dr.LLM: Dynamic Layer Routing in LLMs

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    Dr. LLM retrofits frozen LLMs with MCTS-supervised per-layer routers for skip/execute/repeat decisions, delivering up to +3.4% accuracy and 5-layer savings on reasoning tasks with strong out-of-domain generalization.

  16. Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners

    cs.HC 2025-09 unverdicted novelty 6.0 of 10

    A human-centered design workshop with journalism practitioners yields an evaluation cookbook and design requirements for contextualized, value-aligned generative AI benchmarks.

  17. Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models

    cs.CL 2025-09 unverdicted novelty 6.0 of 10

    PAS automates activation steering for LLMs using labeled data to improve behavior control on tasks like bias and alignment, with gains over ICL and SFT but limited effect on intelligence tasks.

  18. UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

    cs.CL 2025-08 conditional novelty 6.0 of 10

    UI-Bench is the first large-scale benchmark that ranks AI text-to-app tools by blinded expert pairwise preference, using TrueSkill to produce a leaderboard with confidence intervals.

  19. Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models

    cs.AI 2025-08 reject novelty 6.0 of 10

    A new MCP benchmark across six LLMs finds that proactive tool use is rare on first prompts, instructed tool use mainly improves in two-turn dialogues, MCP context degrades accuracy by about 9.5%, and input-token overh...

  20. League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models

    cs.AI 2025-07 unverdicted novelty 6.0 of 10

    League of LLMs organizes LLMs into a self-governed mutual evaluation league using dynamic, transparent, objective, and professional criteria to distinguish model capabilities with 70.7% top-k ranking stability.

  21. Language Models Improve When Pretraining Data Matches Target Tasks

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.

  22. Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Math reasoning gains in LLMs rarely transfer to general domains; RL tuning generalizes while SFT causes forgetting and representation drift.

  23. FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An adaptive, per-language data filtering and deduplication pipeline produces multilingual LLM pre-training corpora that beat prior public datasets on 11 of 14 evaluated languages, and a 20TB, 1,868 language-script dat...

  24. Think Clearly: Improving Reasoning via Redundant Token Pruning

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A training-free test-time method prunes low-attention reasoning tokens from the KV cache, guided by an injected end-of-thinking token, and reports accuracy gains on math competition benchmarks.

  25. Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MoE models with activation rates in an optimal region outperform dense LLMs of identical total parameter count, training compute, and data budget, with the optimal region consistent across scales.

  26. Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The monotonicity of token probabilities during initial decoding predicts chain-of-thought gains, enabling dynamic selection between CoT and direct answers.

  27. MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MaXIFE is a 23-language, 1,667-task benchmark for multilingual and cross-lingual instruction-following evaluation, with baseline scores for five commercial LLMs.

  28. PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

    cs.LG 2025-05 reject novelty 6.0 of 10

    A new benchmark of 380 principle-based physics problems shows that state-of-the-art LLMs struggle to apply symmetry, conservation, and dimensional-analysis shortcuts, achieving under 50 percent average accuracy with h...

  29. Scalable Complexity Control Facilitates Reasoning Ability of LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Controlling model complexity through smaller initialization rates and stronger weight decay improved LLM benchmark scores and made loss-versus-scale curves descend faster.

  30. SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis

    cs.SE 2025-05 conditional novelty 6.0 of 10

    Large language models perform poorly on a new C-code vulnerability benchmark, indicating they rely on pattern matching rather than genuine reasoning.

  31. STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    STAR-R1 uses single-stage reinforcement learning with fine-grained rewards to improve spatial transformation reasoning in multimodal LLMs, outperforming supervised fine-tuning on cross-view TVR tasks.

  32. RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    RoSTE couples quantization-aware supervised fine-tuning with per-layer Hadamard rotation selection, reducing quantization outliers and improving 4-bit quantized LLM accuracy over SFT-then-PTQ baselines.

  33. Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A large synthetic instruction corpus with guidelines, preference rules, and format variants improves LLM performance on five NLU benchmarks by an average of 3.1%.

  34. Minerva: A Programmable Memory Test Benchmark for Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Minerva is a programmable memory-test benchmark showing that LLMs at 4k tokens perform well on search but drop sharply on editing, counting, state tracking, and composite tasks, revealing that retrieval ability does n...

  35. Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A fine-tuned 14B LLM judge, trained with scenario-based prompts and controlled instruction generation, approaches GPT-4's human-agreement performance, and the paper documents why scaling distillation data can fail.

  36. Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

    cs.CR 2025-02 conditional novelty 6.0 of 10

    Model tampering attacks, especially few-shot fine-tuning, reliably re-elicit unlearned capabilities in Llama-3-8B and can bound the success of held-out input-space attacks.

  37. UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    UGPhysics is a new bilingual benchmark of 5,520 undergraduate physics problems; the strongest tested LLM, OpenAI o1-mini, reaches only 49.8% accuracy.

  38. SedarEval: Automated Evaluation using Self-Adaptive Rubrics

    cs.CV 2025-01 reject novelty 6.0 of 10

    A benchmark and judge model that uses per-question custom rubrics to score LLM outputs, claiming better alignment with human grading than GPT-4.

  39. A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks

    cs.AI 2025-01 conditional novelty 6.0 of 10

    A survey that unifies LLM search-based inference frameworks under an MDP-based taxonomy and modular search procedures.

  40. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  41. Ultra-Sparse Memory Network

    cs.LG 2024-11 conditional novelty 6.0 of 10

    UltraMem, a sparse memory layer with Tucker-decomposed retrieval and virtual memory expansion, outperforms Mixture-of-Experts at equal compute, with up to 6x lower inference latency.

  42. MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MEMO-Bench scores 12 text-to-image models and 16 multimodal LLMs on emotion generation and recognition, finding stronger performance on positive emotions and weak fine-grained intensity estimation.

  43. DataComp-LM: In search of the next generation of training sets for language models

    cs.LG 2024-06 unverdicted novelty 6.0 of 10

    DCLM-Baseline dataset lets a 7B model reach 64% 5-shot MMLU accuracy after 2.6T tokens, beating prior open-data models by 6.6 points on MMLU with 40% less compute.

  44. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    cs.CL 2024-02 unverdicted novelty 6.0 of 10

    DeepSeekMath 7B reaches 51.7% on MATH via continued pretraining on curated web math data and Group Relative Policy Optimization.

  45. MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

    cs.CL 2023-09 conditional novelty 6.0 of 10

    MAmmoTH models trained via hybrid CoT-PoT instruction tuning on MathInstruct outperform prior open-source LLMs by 16-32% average accuracy on nine math datasets, reaching 33% and 44% on MATH for 7B and 34B scales.

  46. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.5 of 10

    A monolithic multimodal LLM that cuts pre-training data by 58% and first-token latency by up to 69% while matching or beating its predecessor on 15 benchmarks.

  47. SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    SingGuard presents a policy-adaptive multimodal LLM guardrail family with hybrid reasoning regimes and a new benchmark of 56,340 examples, claiming SOTA F1 across 35 datasets and improved policy adherence under runtim...

  48. Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    An integrated survey organizing AI mathematical reasoning into informal, formal, discovery, and technique axes while cataloging benchmarks and assessing failure modes.

  49. Position: AI Evaluations Should be Grounded on a Theory of Capability

    cs.AI 2025-09 conditional novelty 5.0 of 10

    AI evaluations should be reframed as inference tasks grounded in an explicit theory of capability, with an empirical demonstration that results depend on modeling assumptions and a proposed Evaluation Card for transparency.

  50. Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    Grove MoE uses unequal-size adjugate experts with complexity-based activation to run 33B-parameter models at roughly 3.1 to 3.3B active parameters while matching larger open models in benchmarks.

  51. Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SFT on curated data is a lower bound on a sparse-reward RL objective, and an importance-weighted variant, iw-SFT, tightens the bound and beats plain SFT on AIME 2024 and GPQA.

  52. HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong

    cs.CL 2025-07 reject novelty 5.0 of 10

    A DeepSeek-based model fine-tuned for Hong Kong outperforms general models on Hong Kong benchmarks, but most of those benchmarks are self-authored and unreleased.

  53. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  54. SciDA: Scientific Dynamic Assessor of LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    SciDA is a dynamically initialized, multi-discipline olympiad benchmark that shows LLMs perform substantially worse when problem variables are randomized, which the authors attribute to memorization of fixed numerical...

  55. dots.llm1 Technical Report

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A 14B-active MoE model roughly matches Qwen2.5-72B on a broad benchmark suite while reporting about a 4x reduction in training GPU-hours.

  56. PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    PC-MoE shards the expert layers of an MoE LLM across parties and routes only sparse top-k activations between them, achieving near-centralized accuracy with about 70% memory savings and resistance to one partial-gradi...

  57. Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

    cs.CL 2025-02 conditional novelty 5.0 of 10

    TELLME edits an LLM's hidden representations so similar behaviors cluster and different behaviors separate, improving safety monitoring and detoxification while preserving general ability.

  58. AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity

    cs.CV 2024-11 reject novelty 5.0 of 10

    A sales-weighted feature influence score, a Random Forest popularity predictor, and diffusion-based image edits are combined to rank fashion features, with a human survey that only weakly confirms the ranking.

  59. VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A data composition method that aligns SFT data proportions with a model's detected domain knowledge distribution and dynamically reweights domains by learnable potential improves multi-domain performance versus unifor...

  60. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

    cs.CV 2024-08 unverdicted novelty 5.0 of 10

    mPLUG-Owl3 introduces hyper attention blocks to integrate vision and language for long image-sequence understanding and reports SOTA results on single-image, multi-image, and video benchmarks.

See all 71 Pith citations

Reference graph

Works this paper leans on

287 extracted references · 287 canonical work pages · cited by 71 Pith papers (see all)

  1. [1]

    Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,

    Reasoning over Hybrid Chain for Table-and-Text Open Domain Question Answering , author =. Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,. 2022 , month =. doi:10.24963/ijcai.2022/629 , url =

  2. [2]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

    Reasoning Over Semantic-Level Graph for Fact Checking , author=. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

  3. [3]

    2023 , publisher =

    Beeching, Edward and Han, Sheon and Lambert, Nathan and Rajani, Nazneen and Sanseviero, Omar and Tunstall, Lewis and Wolf, Thomas , title =. 2023 , publisher =

  4. [4]

    Communications of the ACM , volume=

    Datasheets for datasets , author=. Communications of the ACM , volume=. 2021 , publisher=

  5. [5]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    GLM: General Language Model Pretraining with Autoregressive Blank Infilling , author=. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  6. [6]

    Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

    Bold: Dataset and metrics for measuring biases in open-ended language generation , author=. Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

  7. [9]

    and Stoica, Ion and Xing, Eric P

    Chiang, Wei-Lin and Li, Zhuohan and Lin, Zi and Sheng, Ying and Wu, Zhanghao and Zhang, Hao and Zheng, Lianmin and Zhuang, Siyuan and Zhuang, Yonghao and Gonzalez, Joseph E. and Stoica, Ion and Xing, Eric P. , month =. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90\ url =

  8. [10]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

    LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module Network , author=. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Show all 287 references
  1. [11]

    Syntax-Enhanced Pre-trained Model , author=. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , pages=

  2. [12]

    Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

    Neural Deepfake Detection with Factual Structure of Text , author=. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

  3. [13]

    Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

    ProQA: Structural Prompt-based Pre-training for Unified Question Answering , author=. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

  4. [14]

    Findings of the Association for Computational Linguistics: NAACL 2022 , pages=

    Analytical Reasoning of Text , author=. Findings of the Association for Computational Linguistics: NAACL 2022 , pages=

  5. [15]

    Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=

    UserAdapter: Few-shot user learning in sentiment analysis , author=. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=

  6. [16]

    Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=

    Improving Question Answering by Commonsense-Based Pre-training , author=. Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=

  7. [17]

    IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

    From lsat: The progress and challenges of complex reasoning , author=. IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=. 2022 , publisher=

  8. [18]

    Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=

    LogiQA: a challenge dataset for machine reading comprehension with logical reasoning , author=. Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=

  9. [19]

    Sort , volume=

    Measuring Mathematical Problem Solving With the MATH Dataset , author=. Sort , volume=

  10. [20]

    Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems , author=. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  11. [21]

    Proceedings of AAAI , year=

    JEC-QA: A Legal-Domain Question Answering Dataset , author=. Proceedings of AAAI , year=

  12. [24]

    2023 , eprint=

    GPT-4 Technical Report , author=. 2023 , eprint=

  13. [25]

    2019 , publisher=

    Rebooting AI: Building artificial intelligence we can trust , author=. 2019 , publisher=

  14. [26]

    Advances in neural information processing systems , volume=

    Language models are few-shot learners , author=. Advances in neural information processing systems , volume=

  15. [29]

    Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=

    SQuAD: 100,000+ Questions for Machine Comprehension of Text , author=. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=

  16. [30]

    Proceedings of the 44

    Ehrmann, Maud and Romanello, Matteo and Clematide, Simon and Doucet, Antoine , year =. Proceedings of the 44

  17. [31]

    Advances in neural information processing systems , volume=

    Superglue: A stickier benchmark for general-purpose language understanding systems , author=. Advances in neural information processing systems , volume=

  18. [36]

    Advances in Neural Information Processing Systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in Neural Information Processing Systems , volume=

  19. [37]

    Conference on Empirical Methods in Natural Language Processing , year=

    Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering , author=. Conference on Empirical Methods in Natural Language Processing , year=

  20. [39]

    Available at SSRN , year=

    Chatgpt goes to law school , author=. Available at SSRN , year=

  21. [40]

    2022 , eprint=

    Solving Quantitative Reasoning Problems with Language Models , author=. 2022 , eprint=

  22. [43]

    Open llm leaderboard

    Edward Beeching, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. Open llm leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023

  23. [44]

    Bowman, Gabor Angeli, Christopher Potts, and Christopher D

    Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp.\ 632--642, Lisbon, Portugal, Septembe...

  24. [45]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020

  25. [46]

    Sparks of artificial general intelligence: Early experiments with gpt-4

    S \'e bastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023

  26. [47]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90\ quality, March 2023. URL https://lmsys.org/blog/2023...

  27. [48]

    Chatgpt goes to law school

    Jonathan H Choi, Kristin E Hickman, Amy Monahan, and Daniel Schwarcz. Chatgpt goes to law school. Available at SSRN, 2023

  28. [49]

    Scaling instruction-finetuned language models

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022

  29. [50]

    Boolq: Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019

  30. [51]

    S ent E val: An evaluation toolkit for universal sentence representations

    Alexis Conneau and Douwe Kiela. S ent E val: An evaluation toolkit for universal sentence representations. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 2018) , Miyazaki, Japan, May 2018. European Language Resources Associa...

  31. [52]

    BERT : Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Lan...

  32. [53]

    Bold: Dataset and metrics for measuring biases in open-ended language generation

    Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. Bold: Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparen...

  33. [54]

    Glm: General language model pretraining with autoregressive blank infilling

    Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...

  34. [55]

    Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents

    Maud Ehrmann, Matteo Romanello, Simon Clematide, and Antoine Doucet. Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents . In Proceedings of the 44 d European Conference on IR Research ( ECIR 2022) , Stavanger, Norway...

  35. [56]

    Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection

    Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509, 2022

  36. [57]

    Measuring mathematical problem solving with the math dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. Sort, 2 0 (4): 0 0--6

  37. [58]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020

  38. [59]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  39. [60]

    Solving quantitative reasoning problems with language models, 2022

    Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models, 2022

  40. [61]

    Holistic evaluation of language models

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022

  41. [62]

    Program induction by rationale generation: Learning to solve and explain algebraic word problems

    Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 15...

  42. [63]

    Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

    Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. Logiqa: a challenge dataset for machine reading comprehension with logical reasoning. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelli...

  43. [64]

    Rebooting AI: Building artificial intelligence we can trust

    Gary Marcus and Ernest Davis. Rebooting AI: Building artificial intelligence we can trust. Vintage, 2019

  44. [65]

    The natural language decathlon: Multitask learning as question answering

    Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730, 2018

  45. [66]

    Can a suit of armor conduct electricity? a new dataset for open book question answering

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Conference on Empirical Methods in Natural Language Processing, 2018

  46. [67]

    Gpt-4 technical report, 2023

    OpenAI. Gpt-4 technical report, 2023

  47. [68]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 0 2...

  48. [69]

    The lambada dataset: Word prediction requiring a broad discourse context

    Denis Paperno, Germ \'a n Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern \'a ndez. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016

  49. [70]

    Squad: 100,000+ questions for machine comprehension of text

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp.\ 2383--2392, 2016

  50. [71]

    Internlm, 2023

    SenseTime . Internlm, 2023. https://github.com/InternLM/InternLM-techreport/

  51. [72]

    FEVER : a large-scale dataset for fact extraction and VER ification

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. FEVER : a large-scale dataset for fact extraction and VER ification. In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Langu...

  52. [73]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  53. [74]

    GLUE : A multi-task benchmark and analysis platform for natural language understanding

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. GLUE : A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop B lackbox NLP : Analyzing and Interpreting Neural Networks fo...

  54. [75]

    Superglue: A stickier benchmark for general-purpose language understanding systems

    Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32, 2019

  55. [76]

    From lsat: The progress and challenges of complex reasoning

    Siyuan Wang, Zhongkun Liu, Wanjun Zhong, Ming Zhou, Zhongyu Wei, Zhumin Chen, and Nan Duan. From lsat: The progress and challenges of complex reasoning. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 0 2201--2216, 2022

  56. [77]

    Chain of thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022

  57. [78]

    Opt: Open pre-trained transformer language models

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022 a

  58. [79]

    Automatic chain of thought prompting in large language models

    Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493, 2022 b

  59. [80]

    Jec-qa: A legal-domain question answering dataset

    Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. Jec-qa: A legal-domain question answering dataset. In Proceedings of AAAI, 2020

  60. [81]

    Analytical reasoning of text

    Wanjun Zhong, Siyuan Wang, Duyu Tang, Zenan Xu, Daya Guo, Yining Chen, Jiahai Wang, Jian Yin, Ming Zhou, and Nan Duan. Analytical reasoning of text. In Findings of the Association for Computational Linguistics: NAACL 2022, pp.\ 2306--2319, Seattle, United States, July 2022. As...

  61. [82]

    Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021

  62. [83]

    Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers

    Singh, Sumer and Li, Sheng. Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.1

  63. [84]

    Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces

    Hahn, Vanessa and Ruiter, Dana and Kleinbauer, Thomas and Klakow, Dietrich. Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.2

  64. [85]

    H ate BERT : Retraining BERT for Abusive Language Detection in E nglish

    Caselli, Tommaso and Basile, Valerio and Mitrovi \'c , Jelena and Granitzer, Michael. H ate BERT : Retraining BERT for Abusive Language Detection in E nglish. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.3

  65. [86]

    Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset

    Kirk, Hannah and Jun, Yennie and Rauba, Paulius and Wachtel, Gal and Li, Ruining and Bai, Xingjian and Broestl, Noah and Doff-Sotta, Martin and Shtedritski, Aleksandar and Asano, Yuki M. Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset. ...

  66. [87]

    Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation

    Kivlichan, Ian and Lin, Zi and Liu, Jeremiah and Vasserman, Lucy. Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.5

  67. [88]

    DALC : the D utch Abusive Language Corpus

    Caselli, Tommaso and Schelhaas, Arjan and Weultjes, Marieke and Leistra, Folkert and van der Veen, Hylke and Timmerman, Gerben and Nissim, Malvina. DALC : the D utch Abusive Language Corpus. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18...

  68. [89]

    and Dulal, Saurab and Koirala, Diwa

    Niraula, Nobal B. and Dulal, Saurab and Koirala, Diwa. Offensive Language Detection in N epali Social Media. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.7

  69. [90]

    MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms

    Fortuna, Paula and Cortez, Vanessa and Sozinho Ramalho, Miguel and P \'e rez-Mayos, Laura. MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.8

  70. [91]

    Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist

    Manerba, Marta Marchiori and Tonelli, Sara. Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.9

  71. [92]

    Improving Counterfactual Generation for Fair Hate Speech Detection

    Mostafazadeh Davani, Aida and Omrani, Ali and Kennedy, Brendan and Atari, Mohammad and Ren, Xiang and Dehghani, Morteza. Improving Counterfactual Generation for Fair Hate Speech Detection. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.1865...

  72. [93]

    Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon

    Zad, Samira and Jimenez, Joshuan and Finlayson, Mark. Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.11

  73. [94]

    Mitigating Biases in Toxic Language Detection through Invariant Rationalization

    Chuang, Yung-Sung and Gao, Mingye and Luo, Hongyin and Glass, James and Lee, Hung-yi and Chen, Yun-Nung and Li, Shang-Wen. Mitigating Biases in Toxic Language Detection through Invariant Rationalization. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 20...

  74. [95]

    Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments

    Aksenov, Dmitrii and Bourgonje, Peter and Zaczynska, Karolina and Ostendorff, Malte and Moreno-Schneider, Julian and Rehm, Georg. Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments. Proceedings of the 5th Workshop on Online Abuse ...

  75. [96]

    Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse

    Sodhi, Ravsimar and Pant, Kartikey and Mamidi, Radhika. Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.14

  76. [97]

    Context Sensitivity Estimation in Toxicity Detection

    Xenos, Alexandros and Pavlopoulos, John and Androutsopoulos, Ion. Context Sensitivity Estimation in Toxicity Detection. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.15

  77. [98]

    A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection

    Salawu, Semiu and Lumsden, Jo and He, Yulan. A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.16

  78. [99]

    Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format

    Risch, Julian and Schmidt, Philipp and Krestel, Ralf. Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.17

  79. [100]

    and H \'e bert-Dufresne, Laurent and Roth, Allison M

    Trujillo, Milo and Rosenblatt, Sam and de Anda J \'a uregui, Guillermo and Moog, Emily and Samson, Briane Paul V. and H \'e bert-Dufresne, Laurent and Roth, Allison M. When the Echo Chamber Shatters: Examining the Use of Community-Specific Language Post-Subreddit Ban. Proceedi...

  80. [101]

    Targets and Aspects in Social Media Hate Speech

    Shvets, Alexander and Fortuna, Paula and Soler, Juan and Wanner, Leo. Targets and Aspects in Social Media Hate Speech. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.19

  81. [102]

    Abusive Language on Social Media Through the Legal Looking Glass

    Bertaglia, Thales and Grigoriu, Andreea and Dumontier, Michel and van Dijck, Gijs. Abusive Language on Social Media Through the Legal Looking Glass. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.20

  82. [103]

    Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection

    Mathias, Lambert and Nie, Shaoliang and Mostafazadeh Davani, Aida and Kiela, Douwe and Prabhakaran, Vinodkumar and Vidgen, Bertie and Waseem, Zeerak. Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection. Proceedings of the 5th Workshop on Online Abuse and...

  83. [104]

    VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes

    Aggarwal, Piush and Liman, Michelle Espranita and Gold, Darina and Zesch, Torsten. VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.22

  84. [105]

    Racist or Sexist Meme? Classifying Memes beyond Hateful

    Zia, Haris Bin and Castro, Ignacio and Tyson, Gareth. Racist or Sexist Meme? Classifying Memes beyond Hateful. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.23

  85. [106]

    Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes

    Kougia, Vasiliki and Pavlopoulos, John. Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes. Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021). 2021. doi:10.18653/v1/2021.woah-1.24

  86. [107]

    Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021

  87. [108]

    Text Simplification for Comprehension-based Question-Answering

    Dadu, Tanvi and Pant, Kartikey and Nagar, Seema and Barbhuiya, Ferdous and Dey, Kuntal. Text Simplification for Comprehension-based Question-Answering. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.1

  88. [109]

    Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets

    Olsen, Benjamin and Plank, Barbara. Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.2

  89. [110]

    Detecting Depression in T hai Blog Posts: a Dataset and a Baseline

    H. Detecting Depression in T hai Blog Posts: a Dataset and a Baseline. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.3

  90. [111]

    Keyphrase Extraction with Incomplete Annotated Training Data

    Lei, Yanfei and Hu, Chunming and Ma, Guanghui and Zhang, Richong. Keyphrase Extraction with Incomplete Annotated Training Data. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.4

  91. [112]

    Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks

    Tran Phu, Minh and Nguyen, Minh Van and Nguyen, Thien Huu. Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.5

  92. [113]

    Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks

    Le, Duong and Nguyen, Thien Huu. Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.6

  93. [114]

    G oogle-trickers, Yaminjeongeum, and Leetspeak: An Empirical Taxonomy for Intentionally Noisy User-Generated Text

    Cho, Won Ik and Kim, Soomin. G oogle-trickers, Yaminjeongeum, and Leetspeak: An Empirical Taxonomy for Intentionally Noisy User-Generated Text. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.7

  94. [115]

    Description-based Label Attention Classifier for Explainable ICD -9 Classification

    Feucht, Malte and Wu, Zhiliang and Althammer, Sophia and Tresp, Volker. Description-based Label Attention Classifier for Explainable ICD -9 Classification. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.8

  95. [116]

    A Text Editing Approach to Joint J apanese Word Segmentation, POS Tagging, and Lexical Normalization

    Higashiyama, Shohei and Utiyama, Masao and Watanabe, Taro and Sumita, Eiichiro. A Text Editing Approach to Joint J apanese Word Segmentation, POS Tagging, and Lexical Normalization. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.186...

  96. [117]

    Intrinsic evaluation of language models for code-switching

    Cheong, Sik Feng and Chieu, Hai Leong and Lim, Jing. Intrinsic evaluation of language models for code-switching. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.10

  97. [118]

    Can images help recognize entities? A study of the role of images for Multimodal NER

    Chen, Shuguang and Aguilar, Gustavo and Neves, Leonardo and Solorio, Thamar. Can images help recognize entities? A study of the role of images for Multimodal NER. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.11

  98. [119]

    Perceived and Intended Sarcasm Detection with Graph Attention Networks

    Plepi, Joan and Flek, Lucie. Perceived and Intended Sarcasm Detection with Graph Attention Networks. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.12

  99. [120]

    Hierarchical Character Tagger for Short Text Spelling Error Correction

    Gao, Mengyi and Xu, Canran and Shi, Peng. Hierarchical Character Tagger for Short Text Spelling Error Correction. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.13

  100. [121]

    Common Sense Bias in Semantic Role Labeling

    Lent, Heather and S gaard, Anders. Common Sense Bias in Semantic Role Labeling. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.14

  101. [122]

    P oli WAM : An Exploration of a Large Scale Corpus of Political Discussions on W hats A pp Messenger

    Srivastava, Vivek and Singh, Mayank. P oli WAM : An Exploration of a Large Scale Corpus of Political Discussions on W hats A pp Messenger. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.15

  102. [123]

    P ars T wi NER : A Corpus for Named Entity Recognition at Informal P ersian

    Aghajani, MohammadMahdi and Badri, AliAkbar and Beigy, Hamid. P ars T wi NER : A Corpus for Named Entity Recognition at Informal P ersian. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.16

  103. [124]

    D ream D rug - A crowdsourced NER dataset for detecting drugs in darknet markets

    Bogensperger, Johannes and Schlarb, Sven and Hanbury, Allan and Recski, G \'a bor. D ream D rug - A crowdsourced NER dataset for detecting drugs in darknet markets. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.17

  104. [125]

    Comparing Grammatical Theories of Code-Mixing

    Pratapa, Adithya and Choudhury, Monojit. Comparing Grammatical Theories of Code-Mixing. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.18

  105. [126]

    Improving Punctuation Restoration for Speech Transcripts via External Data

    Fu, Xue-Yong and Chen, Cheng and Laskar, Md Tahmid Rahman and Bhushan, Shashi and Corston-Oliver, Simon. Improving Punctuation Restoration for Speech Transcripts via External Data. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.1865...

  106. [127]

    Learning to Rank Question Answer Pairs with Bilateral Contrastive Data Augmentation

    Deng, Yang and Zhang, Wenxuan and Lam, Wai. Learning to Rank Question Answer Pairs with Bilateral Contrastive Data Augmentation. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.20

  107. [128]

    Mitigation of Diachronic Bias in Fake News Detection Dataset

    Murayama, Taichi and Wakamiya, Shoko and Aramaki, Eiji. Mitigation of Diachronic Bias in Fake News Detection Dataset. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.21

  108. [129]

    Understanding the Impact of UGC Specificities on Translation Quality

    Rosales N \'u \ n ez, Jos \'e Carlos and Seddah, Djam \'e and Wisniewski, Guillaume. Understanding the Impact of UGC Specificities on Translation Quality. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.22

  109. [130]

    Noisy UGC Translation at the Character Level: Revisiting Open-Vocabulary Capabilities and Robustness of Char-Based Models

    Rosales N \'u \ n ez, Jos \'e Carlos and Wisniewski, Guillaume and Seddah, Djam \'e. Noisy UGC Translation at the Character Level: Revisiting Open-Vocabulary Capabilities and Robustness of Char-Based Models. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-N...

  110. [131]

    Changes in T witter geolocations: Insights and suggestions for future usage

    Kruspe, Anna and H. Changes in T witter geolocations: Insights and suggestions for future usage. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.24

  111. [132]

    C on Q uest: Contextual Question Paraphrasing through Answer-Aware Synthetic Question Generation

    Mirshekari, Mostafa and Gu, Jing and Sisto, Aaron. C on Q uest: Contextual Question Paraphrasing through Answer-Aware Synthetic Question Generation. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.25

  112. [133]

    NADE : A Benchmark for Robust Adverse Drug Events Extraction in Face of Negations

    Scaboro, Simone and Portelli, Beatrice and Chersoni, Emmanuele and Santus, Enrico and Serra, Giuseppe. NADE : A Benchmark for Robust Adverse Drug Events Extraction in Face of Negations. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10...

  113. [134]

    S pan A lign: Efficient Sequence Tagging Annotation Projection into Translated Data applied to Cross-Lingual Opinion Mining

    Jacqmin, L \'e o and Marzinotto, Gabriel and Gromada, Justyna and Szczekocka, Ewelina and Ko ody \'n ski, Robert and Damnati, G \'e raldine. S pan A lign: Efficient Sequence Tagging Annotation Projection into Translated Data applied to Cross-Lingual Opinion Mining. Proceedings...

  114. [135]

    A Novel Framework for Detecting Important Subevents from Crisis Events via Dynamic Semantic Graphs

    Spiliopoulou, Evangelia and Saha, Tanay Kumar and Tetreault, Joel and Jaimes, Alejandro. A Novel Framework for Detecting Important Subevents from Crisis Events via Dynamic Semantic Graphs. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi...

  115. [136]

    Synthetic Data Generation and Multi-Task Learning for Extracting Temporal Information from Health-Related Narrative Text

    Shim, Heereen and Lowet, Dietwig and Luca, Stijn and Vanrumste, Bart. Synthetic Data Generation and Multi-Task Learning for Extracting Temporal Information from Health-Related Narrative Text. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. ...

  116. [137]

    Neural-based RST Parsing And Analysis In Persuasive Discourse

    Li, Jinfen and Xiao, Lu. Neural-based RST Parsing And Analysis In Persuasive Discourse. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.30

  117. [138]

    BART for Post-Correction of OCR Newspaper Text

    Soper, Elizabeth and Fujimoto, Stanley and Yu, Yen-Yun. BART for Post-Correction of OCR Newspaper Text. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.31

  118. [139]

    Coping with Noisy Training Data Labels in Paraphrase Detection

    Vahtola, Teemu and Creutz, Mathias and Sj. Coping with Noisy Training Data Labels in Paraphrase Detection. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.32

  119. [140]

    Knowledge Distillation with Noisy Labels for Natural Language Understanding

    Bhardwaj, Shivendra and Ghaddar, Abbas and Rashid, Ahmad and Bibi, Khalil and Li, Chengyang and Ghodsi, Ali and Langlais, Phillippe and Rezagholizadeh, Mehdi. Knowledge Distillation with Noisy Labels for Natural Language Understanding. Proceedings of the Seventh Workshop on No...

  120. [141]

    Integrating Transformers and Knowledge Graphs for T witter Stance Detection

    Clark, Thomas and Conforti, Costanza and Liu, Fangyu and Meng, Zaiqiao and Shareghi, Ehsan and Collier, Nigel. Integrating Transformers and Knowledge Graphs for T witter Stance Detection. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:...

  121. [142]

    Detecting Cross-Geographic Biases in Toxicity Modeling on Social Media

    Ghosh, Sayan and Baker, Dylan and Jurgens, David and Prabhakaran, Vinodkumar. Detecting Cross-Geographic Biases in Toxicity Modeling on Social Media. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.35

  122. [143]

    Detection of Puffery on the E nglish W ikipedia

    Bertsch, Amanda and Bethard, Steven. Detection of Puffery on the E nglish W ikipedia. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.36

  123. [144]

    Robustness and Sensitivity of BERT Models Predicting A lzheimer ' s Disease from Text

    Novikova, Jekaterina. Robustness and Sensitivity of BERT Models Predicting A lzheimer ' s Disease from Text. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.37

  124. [145]

    Understanding Model Robustness to User-generated Noisy Texts

    N \'a plava, Jakub and Popel, Martin and Straka, Milan and Strakov \'a , Jana. Understanding Model Robustness to User-generated Noisy Texts. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.38

  125. [146]

    CIDE r- R : Robust Consensus-based Image Description Evaluation

    Oliveira dos Santos, Gabriel and Colombini, Esther Luna and Avila, Sandra. CIDE r- R : Robust Consensus-based Image Description Evaluation. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.39

  126. [147]

    Improved Named Entity Recognition for Noisy Call Center Transcripts

    Davidson, Sam and Hosier, Jordan and Zhou, Yu and Gurbani, Vijay. Improved Named Entity Recognition for Noisy Call Center Transcripts. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.40

  127. [148]

    Contrapositive Local Class Inference

    Kashefi, Omid and Hwa, Rebecca. Contrapositive Local Class Inference. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.41

  128. [149]

    Improved Multilingual Language Model Pretraining for Social Media Text via Translation Pair Prediction

    Mishra, Shubhanshu and Haghighi, Aria. Improved Multilingual Language Model Pretraining for Social Media Text via Translation Pair Prediction. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.42

  129. [150]

    Co-training for Commit Classification

    Lee, Jian Yi David and Chieu, Hai Leong. Co-training for Commit Classification. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.43

  130. [151]

    Study of Manifestation of Civil Unrest on T witter

    Chinta, Abhinav and Zhang, Jingyu and DeLucia, Alexandra and Dredze, Mark and Buczak, Anna L. Study of Manifestation of Civil Unrest on T witter. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.44

  131. [152]

    The K orean Morphologically Tight-Fitting Tokenizer for Noisy User-Generated Texts

    Lee, Sangah and Shin, Hyopil. The K orean Morphologically Tight-Fitting Tokenizer for Noisy User-Generated Texts. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.45

  132. [153]

    Character Transformations for Non-Autoregressive GEC Tagging

    Straka, Milan and N \'a plava, Jakub and Strakov \'a , Jana. Character Transformations for Non-Autoregressive GEC Tagging. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.46

  133. [154]

    Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?

    Riabi, Arij and Sagot, Beno \^ t and Seddah, Djam \'e. Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2...

  134. [155]

    `` Something Something Hota Hai! '' An Explainable Approach towards Sentiment Analysis on I ndian Code-Mixed Data

    Priyanshu, Aman and Vardhan, Aleti and Sivakumar, Sudarshan and Vijay, Supriti and Chhabra, Nipuna. `` Something Something Hota Hai! '' An Explainable Approach towards Sentiment Analysis on I ndian Code-Mixed Data. Proceedings of the Seventh Workshop on Noisy User-generated Te...

  135. [156]

    BERT weet FR : Domain Adaptation of Pre-Trained Language Models for F rench Tweets

    Guo, Yanzhu and Rennard, Virgile and Xypolopoulos, Christos and Vazirgiannis, Michalis. BERT weet FR : Domain Adaptation of Pre-Trained Language Models for F rench Tweets. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021...

  136. [157]

    To What Extent Does Lexical Normalization Help E nglish-as-a-Second Language Learners to Read Noisy E nglish Texts?

    Ehara, Yo. To What Extent Does Lexical Normalization Help E nglish-as-a-Second Language Learners to Read Noisy E nglish Texts?. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.50

  137. [158]

    Multilingual Sequence Labeling Approach to solve Lexical Normalization

    Kubal, Divesh and Nagvenkar, Apurva. Multilingual Sequence Labeling Approach to solve Lexical Normalization. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.51

  138. [159]

    Sesame Street to Mount Sinai: BERT -constrained character-level M oses models for multilingual lexical normalization

    Scherrer, Yves and Ljube s i \'c , Nikola. Sesame Street to Mount Sinai: BERT -constrained character-level M oses models for multilingual lexical normalization. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.52

  139. [160]

    Sequence-to-Sequence Lexical Normalization with Multilingual Transformers

    Bucur, Ana-Maria and Cosma, Adrian and Dinu, Liviu P. Sequence-to-Sequence Lexical Normalization with Multilingual Transformers. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.53

  140. [161]

    \'U FAL at M ulti L ex N orm 2021: Improving Multilingual Lexical Normalization by Fine-tuning B y T 5

    Samuel, David and Straka, Milan. \'U FAL at M ulti L ex N orm 2021: Improving Multilingual Lexical Normalization by Fine-tuning B y T 5. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.54

  141. [162]

    M ulti L ex N orm: A Shared Task on Multilingual Lexical Normalization

    van der Goot, Rob and Ramponi, Alan and Zubiaga, Arkaitz and Plank, Barbara and Muller, Benjamin and San Vicente Roncal, I. M ulti L ex N orm: A Shared Task on Multilingual Lexical Normalization. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 20...

  142. [163]

    CL - M o N oise: Cross-lingual Lexical Normalization

    van der Goot, Rob. CL - M o N oise: Cross-lingual Lexical Normalization. Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021. doi:10.18653/v1/2021.wnut-1.56

  143. [164]

    Proceedings of the Fifth Workshop on Widening Natural Language Processing. 2021

  144. [165]

    Dossou, Bonaventure F. P. and Emezue, Chris Chinenye. O kwu G b \'e : End-to-End Speech Recognition for F on and I gbo. Proceedings of the Fifth Workshop on Widening Natural Language Processing. 2021

  145. [166]

    TEET ! T unisian Dataset for Toxic Speech Detection

    Gharbi, Slim and Haddad, Hatem and Kchaou, Mayssa and Arfaoui, Heger. TEET ! T unisian Dataset for Toxic Speech Detection. Proceedings of the Fifth Workshop on Widening Natural Language Processing. 2021

  146. [167]

    Developing Keyboards for the Endangered L ivonian Language

    H. Developing Keyboards for the Endangered L ivonian Language. Proceedings of the Fifth Workshop on Widening Natural Language Processing. 2021

  147. [168]

    and Bucur, Ana-Maria and Todea, Diana and Fodor, Liviu and Luca, Andreea and Dinu, Liviu P

    Podin a , Ioana R. and Bucur, Ana-Maria and Todea, Diana and Fodor, Liviu and Luca, Andreea and Dinu, Liviu P. and Boian, Rare s. Natural language processing as a tool to identify the R eddit particularities of cancer survivors around the time of diagnosis and remission: A pil...

  148. [169]

    The Development of Pre-processing Tools and Pre-trained Embedding Models for A mharic

    Destaw, Tadesse and Ayele, Abinew and Yimam, Seid Muhie. The Development of Pre-processing Tools and Pre-trained Embedding Models for A mharic. Proceedings of the Fifth Workshop on Widening Natural Language Processing. 2021

  149. [170]

    Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021

  150. [171]

    Overview of the 8th Workshop on A sian Translation

    Nakazawa, Toshiaki and Nakayama, Hideki and Ding, Chenchen and Dabre, Raj and Higashiyama, Shohei and Mino, Hideya and Goto, Isao and Pa Pa, Win and Kunchukuttan, Anoop and Parida, Shantipriya and Bojar, Ond r ej and Chu, Chenhui and Eriguchi, Akiko and Abe, Kaori and Oda, Yus...

  151. [172]

    NHK ' s Lexically-Constrained Neural Machine Translation at WAT 2021

    Mino, Hideya and Kinugawa, Kazutaka and Ito, Hitoshi and Goto, Isao and Yamada, Ichiro and Tokunaga, Takenobu. NHK ' s Lexically-Constrained Neural Machine Translation at WAT 2021. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.2

  152. [173]

    Input Augmentation Improves Constrained Beam Search for Neural Machine Translation: NTT at WAT 2021

    Chousa, Katsuki and Morishita, Makoto. Input Augmentation Improves Constrained Beam Search for Neural Machine Translation: NTT at WAT 2021. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.3

  153. [174]

    NICT ' s Neural Machine Translation Systems for the WAT 21 Restricted Translation Task

    Li, Zuchao and Utiyama, Masao and Sumita, Eiichiro and Zhao, Hai. NICT ' s Neural Machine Translation Systems for the WAT 21 Restricted Translation Task. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.4

  154. [175]

    Machine Translation with Pre-specified Target-side Words Using a Semi-autoregressive Model

    Kondo, Seiichiro and Koyama, Aomi and Kiyuna, Tomoshige and Hirasawa, Tosho and Komachi, Mamoru. Machine Translation with Pre-specified Target-side Words Using a Semi-autoregressive Model. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/20...

  155. [176]

    NECTEC ' s Participation in WAT -2021

    Hlaing, Zar Zar and Thu, Ye Kyaw and Myint Oo, Thazin and Ei San, Mya and Usanavasin, Sasiporn and Netisopakul, Ponrudee and Supnithi, Thepchai. NECTEC ' s Participation in WAT -2021. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.6

  156. [177]

    Hybrid Statistical Machine Translation for E nglish- M yanmar: UTYCC Submission to WAT -2021

    Thu, Ye Kyaw and Oo, Thazin Myint and Nwe, Hlaing Myat and Mon, Khaing Zar and Kyaw, Nang Aeindray and Phyo, Naing Linn and Khun, Nann Hwan and Thant, Hnin Aye. Hybrid Statistical Machine Translation for E nglish- M yanmar: UTYCC Submission to WAT -2021. Proceedings of the 8th...

  157. [178]

    NICT -2 Translation System at WAT -2021: Applying a Pretrained Multilingual Encoder-Decoder Model to Low-resource Language Pairs

    Imamura, Kenji and Sumita, Eiichiro. NICT -2 Translation System at WAT -2021: Applying a Pretrained Multilingual Encoder-Decoder Model to Low-resource Language Pairs. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.8

  158. [179]

    Rakuten ' s Participation in WAT 2021: Examining the Effectiveness of Pre-trained Models for Multilingual and Multimodal Machine Translation

    Susanto, Raymond Hendy and Wang, Dongzhe and Yadav, Sunil and Jain, Mausam and Htun, Ohnmar. Rakuten ' s Participation in WAT 2021: Examining the Effectiveness of Pre-trained Models for Multilingual and Multimodal Machine Translation. Proceedings of the 8th Workshop on Asian T...

  159. [180]

    BTS : Back T ran S cription for Speech-to-Text Post-Processor using Text-to-Speech-to-Text

    Park, Chanjun and Seo, Jaehyung and Lee, Seolhwa and Lee, Chanhee and Moon, Hyeonseok and Eo, Sugyeong and Lim, Heuiseok. BTS : Back T ran S cription for Speech-to-Text Post-Processor using Text-to-Speech-to-Text. Proceedings of the 8th Workshop on Asian Translation (WAT2021)....

  160. [181]

    Zero-pronoun Data Augmentation for J apanese-to- E nglish Translation

    Ri, Ryokan and Nakazawa, Toshiaki and Tsuruoka, Yoshimasa. Zero-pronoun Data Augmentation for J apanese-to- E nglish Translation. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.11

  161. [182]

    Evaluation Scheme of Focal Translation for J apanese Partially Amended Statutes

    Yamakoshi, Takahiro and Komamizu, Takahiro and Ogawa, Yasuhiro and Toyama, Katsuhiko. Evaluation Scheme of Focal Translation for J apanese Partially Amended Statutes. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.12

  162. [183]

    TMU NMT System with J apanese BART for the Patent task of WAT 2021

    Kim, Hwichan and Komachi, Mamoru. TMU NMT System with J apanese BART for the Patent task of WAT 2021. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.13

  163. [184]

    System Description for Transperfect

    Stribi \.z ew, Wiktor and Bane, Fred and Concei c \ a o, Jos \'e and Zaretskaya, Anna. System Description for Transperfect. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.14

  164. [185]

    Bering Lab ' s Submissions on WAT 2021 Shared Task

    Park, Heesoo and Lee, Dongjun. Bering Lab ' s Submissions on WAT 2021 Shared Task. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.15

  165. [186]

    NLPH ut ' s Participation at WAT 2021

    Parida, Shantipriya and Panda, Subhadarshi and Kotwal, Ketan and Dash, Amulya Ratna and Dash, Satya Ranjan and Sharma, Yashvardhan and Motlicek, Petr and Bojar, Ond r ej. NLPH ut ' s Participation at WAT 2021. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 202...

  166. [187]

    Improved E nglish to H indi Multimodal Neural Machine Translation

    Laskar, Sahinur Rahman and Khilji, Abdullah Faiz Ur Rahman and Kaushik, Darsh and Pakray, Partha and Bandyopadhyay, Sivaji. Improved E nglish to H indi Multimodal Neural Machine Translation. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/...

  167. [188]

    IITP at WAT 2021: System description for E nglish- H indi Multimodal Translation Task

    Gain, Baban and Bandyopadhyay, Dibyanayan and Ekbal, Asif. IITP at WAT 2021: System description for E nglish- H indi Multimodal Translation Task. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.18

  168. [189]

    V i TA : Visual-Linguistic Translation by Aligning Object Tags

    Gupta, Kshitij and Gautam, Devansh and Mamidi, Radhika. V i TA : Visual-Linguistic Translation by Aligning Object Tags. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.19

  169. [190]

    TMEKU System for the WAT 2021 Multimodal Translation Task

    Zhao, Yuting and Komachi, Mamoru and Kajiwara, Tomoyuki and Chu, Chenhui. TMEKU System for the WAT 2021 Multimodal Translation Task. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.20

  170. [191]

    Optimal Word Segmentation for Neural Machine Translation into D ravidian Languages

    Dhar, Prajit and Bisazza, Arianna and van Noord, Gertjan. Optimal Word Segmentation for Neural Machine Translation into D ravidian Languages. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.21

  171. [192]

    Itihasa: A large-scale corpus for S anskrit to E nglish translation

    Aralikatte, Rahul and de Lhoneux, Miryam and Kunchukuttan, Anoop and S gaard, Anders. Itihasa: A large-scale corpus for S anskrit to E nglish translation. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.22

  172. [193]

    NICT -5 ' s Submission To WAT 2021: MBART Pre-training And In-Domain Fine Tuning For Indic Languages

    Dabre, Raj and Chakrabarty, Abhisek. NICT -5 ' s Submission To WAT 2021: MBART Pre-training And In-Domain Fine Tuning For Indic Languages. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.23

  173. [194]

    How far can we get with one GPU in 100 hours? C o AS ta L at M ulti I ndic MT Shared Task

    Aralikatte, Rahul and Murrieta Bello, H \'e ctor Ricardo and de Lhoneux, Miryam and Hershcovich, Daniel and Bollmann, Marcel and S gaard, Anders. How far can we get with one GPU in 100 hours? C o AS ta L at M ulti I ndic MT Shared Task. Proceedings of the 8th Workshop on Asian...

  174. [195]

    IIIT Hyderabad Submission To WAT 2021: Efficient Multilingual NMT systems for I ndian languages

    Kumar, Sourav and Aggarwal, Salil and Sharma, Dipti. IIIT Hyderabad Submission To WAT 2021: Efficient Multilingual NMT systems for I ndian languages. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.25

  175. [196]

    Language Relatedness and Lexical Closeness can help Improve Multilingual NMT : IITB ombay@ M ulti I ndic NMT WAT 2021

    Khatri, Jyotsana and Saini, Nikhil and Bhattacharyya, Pushpak. Language Relatedness and Lexical Closeness can help Improve Multilingual NMT : IITB ombay@ M ulti I ndic NMT WAT 2021. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.26

  176. [197]

    S amsung R & D Institute P oland submission to WAT 2021 Indic Language Multilingual Task

    Dobrowolski, Adam and Szyma \'n ski, Marcin and Chochowski, Marcin and Przybysz, Pawe. S amsung R & D Institute P oland submission to WAT 2021 Indic Language Multilingual Task. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.27

  177. [198]

    Multilingual Machine Translation Systems at WAT 2021: One-to-Many and Many-to-One Transformer based NMT

    Mhaskar, Shivam and Jain, Aditya and Banerjee, Aakash and Bhattacharyya, Pushpak. Multilingual Machine Translation Systems at WAT 2021: One-to-Many and Many-to-One Transformer based NMT. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021...

  178. [199]

    IITP - MT at WAT 2021: Indic- E nglish Multilingual Neural Machine Translation using R omanized Vocabulary

    Appicharla, Ramakrishna and Gupta, Kamal Kumar and Ekbal, Asif and Bhattacharyya, Pushpak. IITP - MT at WAT 2021: Indic- E nglish Multilingual Neural Machine Translation using R omanized Vocabulary. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.1...

  179. [200]

    ANVITA Machine Translation System for WAT 2021 M ulti I ndic MT Shared Task

    Vegi, Pavanpankaj and J, Sivabhavani and Paul, Biswajit and Viswanathan, Chitra and K R, Prasanna Kumar. ANVITA Machine Translation System for WAT 2021 M ulti I ndic MT Shared Task. Proceedings of the 8th Workshop on Asian Translation (WAT2021). 2021. doi:10.18653/v1/2021.wat-1.30

  180. [201]

    Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  181. [202]

    T ox CCI n: Toxic Content Classification with Interpretability

    Xiang, Tong and MacAvaney, Sean and Yang, Eugene and Goharian, Nazli. T ox CCI n: Toxic Content Classification with Interpretability. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  182. [203]

    Language that Captivates the Audience: Predicting Affective Ratings of TED Talks in a Multi-Label Classification Task

    Kerz, Elma and Qiao, Yu and Wiechmann, Daniel. Language that Captivates the Audience: Predicting Affective Ratings of TED Talks in a Multi-Label Classification Task. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media An...

  183. [204]

    Partisanship and Fear are Associated with Resistance to COVID -19 Directives

    Lindow, Mike and DeFranza, David and Mishra, Arul and Mishra, Himanshu. Partisanship and Fear are Associated with Resistance to COVID -19 Directives. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  184. [205]

    Explainable Detection of Sarcasm in Social Media

    Akula, Ramya and Garibay, Ivan. Explainable Detection of Sarcasm in Social Media. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  185. [206]

    Emotion Ratings: How Intensity, Annotation Confidence and Agreements are Entangled

    Troiano, Enrica and Pad \'o , Sebastian and Klinger, Roman. Emotion Ratings: How Intensity, Annotation Confidence and Agreements are Entangled. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  186. [207]

    Disentangling Document Topic and Author Gender in Multiple Languages: Lessons for Adversarial Debiasing

    Dayanik, Erenay and Pad \'o , Sebastian. Disentangling Document Topic and Author Gender in Multiple Languages: Lessons for Adversarial Debiasing. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  187. [208]

    Universal Joy A Data Set and Results for Classifying Emotions Across Languages

    Lamprinidis, Sotiris and Bianchi, Federico and Hardt, Daniel and Hovy, Dirk. Universal Joy A Data Set and Results for Classifying Emotions Across Languages. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  188. [209]

    FEEL - IT : Emotion and Sentiment Classification for the I talian Language

    Bianchi, Federico and Nozza, Debora and Hovy, Dirk. FEEL - IT : Emotion and Sentiment Classification for the I talian Language. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  189. [210]

    An End-to-End Network for Emotion-Cause Pair Extraction

    Singh, Aaditya and Hingane, Shreeshail and Wani, Saim and Modi, Ashutosh. An End-to-End Network for Emotion-Cause Pair Extraction. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  190. [211]

    WASSA 2021 Shared Task: Predicting Empathy and Emotion in Reaction to News Stories

    Tafreshi, Shabnam and De Clercq, Orphee and Barriere, Valentin and Buechel, Sven and Sedoc, Jo \ a o and Balahur, Alexandra. WASSA 2021 Shared Task: Predicting Empathy and Emotion in Reaction to News Stories. Proceedings of the Eleventh Workshop on Computational Approaches to ...

  191. [212]

    PVG at WASSA 2021: A Multi-Input, Multi-Task, Transformer-Based Architecture for Empathy and Distress Prediction

    Kulkarni, Atharva and Somwase, Sunanda and Rajput, Shivam and Marathe, Manisha. PVG at WASSA 2021: A Multi-Input, Multi-Task, Transformer-Based Architecture for Empathy and Distress Prediction. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, S...

  192. [213]

    WASSA @ IITK at WASSA 2021: Multi-task Learning and Transformer Finetuning for Emotion Classification and Empathy Prediction

    Mundra, Jay and Gupta, Rohan and Mukherjee, Sagnik. WASSA @ IITK at WASSA 2021: Multi-task Learning and Transformer Finetuning for Emotion Classification and Empathy Prediction. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Soc...

  193. [214]

    Analyzing Curriculum Learning for Sentiment Analysis along Task Difficulty, Pacing and Visualization Axes

    Rao Vijjini, Anvesh and Anuranjana, Kaveri and Mamidi, Radhika. Analyzing Curriculum Learning for Sentiment Analysis along Task Difficulty, Pacing and Visualization Axes. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Med...

  194. [215]

    Lightweight Models for Multimodal Sequential Data

    Sourav, Soumya and Ouyang, Jessica. Lightweight Models for Multimodal Sequential Data. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  195. [216]

    Exploring Implicit Sentiment Evoked by Fine-grained News Events

    Van Hee, Cynthia and De Clercq, Orphee and Hoste, Veronique. Exploring Implicit Sentiment Evoked by Fine-grained News Events. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  196. [217]

    Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection

    Markov, Ilia and Ljube s i \'c , Nikola and Fi s er, Darja and Daelemans, Walter. Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment a...

  197. [218]

    Emotion-Aware, Emotion-Agnostic, or Automatic: Corpus Creation Strategies to Obtain Cognitive Event Appraisal Annotations

    Hofmann, Jan and Troiano, Enrica and Klinger, Roman. Emotion-Aware, Emotion-Agnostic, or Automatic: Corpus Creation Strategies to Obtain Cognitive Event Appraisal Annotations. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Socia...

  198. [219]

    Hate Towards the Political Opponent: A T witter Corpus Study of the 2020 US Elections on the Basis of Offensive Speech and Stance Detection

    Grimminger, Lara and Klinger, Roman. Hate Towards the Political Opponent: A T witter Corpus Study of the 2020 US Elections on the Basis of Offensive Speech and Stance Detection. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Soc...

  199. [220]

    Synthetic Examples Improve Cross-Target Generalization: A Study on Stance Detection on a T witter corpus

    Conforti, Costanza and Berndt, Jakob and Pilehvar, Mohammad Taher and Giannitsarou, Chryssi and Toxvaerd, Flavio and Collier, Nigel. Synthetic Examples Improve Cross-Target Generalization: A Study on Stance Detection on a T witter corpus. Proceedings of the Eleventh Workshop o...

  200. [221]

    Creating and Evaluating Resources for Sentiment Analysis in the Low-resource Language: S indhi

    Ali, Wazir and Ali, Naveed and Dai, Yong and Kumar, Jay and Tumrani, Saifullah and Xu, Zenglin. Creating and Evaluating Resources for Sentiment Analysis in the Low-resource Language: S indhi. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sen...

  201. [222]

    Towards Emotion Recognition in H indi- E nglish Code-Mixed Data: A Transformer Based Approach

    Wadhawan, Anshul and Aggarwal, Akshita. Towards Emotion Recognition in H indi- E nglish Code-Mixed Data: A Transformer Based Approach. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  202. [223]

    Nearest neighbour approaches for Emotion Detection in Tweets

    Kaminska, Olha and Cornelis, Chris and Hoste, Veronique. Nearest neighbour approaches for Emotion Detection in Tweets. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  203. [224]

    L 3 C ube M aha S ent: A M arathi Tweet-based Sentiment Analysis Dataset

    Kulkarni, Atharva and Mandhane, Meet and Likhitkar, Manali and Kshirsagar, Gayatri and Joshi, Raviraj. L 3 C ube M aha S ent: A M arathi Tweet-based Sentiment Analysis Dataset. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Soci...

  204. [225]

    Multi-Emotion Classification for Song Lyrics

    Edmonds, Darren and Sedoc, Jo \ a o. Multi-Emotion Classification for Song Lyrics. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  205. [226]

    ONE : Toward ONE model, ONE algorithm, ONE corpus dedicated to sentiment analysis of A rabic/ A rabizi and its dialects

    Guellil, Imane and Azouaou, Faical and Benali, Fodil and Ala-Eddine, Hachani. ONE : Toward ONE model, ONE algorithm, ONE corpus dedicated to sentiment analysis of A rabic/ A rabizi and its dialects. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivi...

  206. [227]

    Me, myself, and ire: Effects of automatic transcription quality on emotion, sarcasm, and personality detection

    Culnan, John and Park, Seongjin and Krishnaswamy, Meghavarshini and Sharp, Rebecca. Me, myself, and ire: Effects of automatic transcription quality on emotion, sarcasm, and personality detection. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity,...

  207. [228]

    Emotional R ob BERT and Insensitive BERT je: Combining Transformers and Affect Lexica for D utch Emotion Detection

    De Bruyne, Luna and De Clercq, Orphee and Hoste, Veronique. Emotional R ob BERT and Insensitive BERT je: Combining Transformers and Affect Lexica for D utch Emotion Detection. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Socia...

  208. [229]

    E mp N a at WASSA 2021: A Lightweight Model for the Prediction of Empathy, Distress and Emotions from Reactions to News Stories

    Vettigli, Giuseppe and Sorgente, Antonio. E mp N a at WASSA 2021: A Lightweight Model for the Prediction of Empathy, Distress and Emotions from Reactions to News Stories. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Med...

  209. [230]

    M ila NLP @ WASSA : Does BERT Feel Sad When You Cry?

    Fornaciari, Tommaso and Bianchi, Federico and Nozza, Debora and Hovy, Dirk. M ila NLP @ WASSA : Does BERT Feel Sad When You Cry?. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. 2021

  210. [231]

    Team Phoenix at WASSA 2021: Emotion Analysis on News Stories with Pre-Trained Language Models

    Butala, Yash and Singh, Kanishk and Kumar, Adarsh and Shrivastava, Shrey. Team Phoenix at WASSA 2021: Emotion Analysis on News Stories with Pre-Trained Language Models. Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media...

  211. [232]

    Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  212. [233]

    QADI : A rabic Dialect Identification in the Wild

    Abdelali, Ahmed and Mubarak, Hamdy and Samih, Younes and Hassan, Sabit and Darwish, Kareem. QADI : A rabic Dialect Identification in the Wild. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  213. [234]

    D ia L ex: A Benchmark for Evaluating Multidialectal A rabic Word Embeddings

    Abdul-Mageed, Muhammad and Elbassuoni, Shady and Doughman, Jad and Elmadany, AbdelRahim and Nagoudi, El Moatez Billah and Zoughby, Yorgo and Shaher, Ahmad and Gaba, Iskander and Helal, Ahmed and El-Razzaz, Mohammed. D ia L ex: A Benchmark for Evaluating Multidialectal A rabic ...

  214. [235]

    Benchmarking Transformer-based Language Models for A rabic Sentiment and Sarcasm Detection

    Abu Farha, Ibrahim and Magdy, Walid. Benchmarking Transformer-based Language Models for A rabic Sentiment and Sarcasm Detection. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  215. [236]

    What does BERT Learn from A rabic Machine Reading Comprehension Datasets?

    Albilali, Eman and Altwairesh, Nora and Hosny, Manar. What does BERT Learn from A rabic Machine Reading Comprehension Datasets?. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  216. [237]

    Kawarith: an A rabic T witter Corpus for Crisis Events

    Alharbi, Alaa and Lee, Mark. Kawarith: an A rabic T witter Corpus for Crisis Events. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  217. [238]

    A rabic Compact Language Modelling for Resource Limited Devices

    Alyafeai, Zaid and Ahmad, Irfan. A rabic Compact Language Modelling for Resource Limited Devices. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  218. [239]

    and Hendley, Robert and Smith, Phillip

    Hakami, Shatha Ali A. and Hendley, Robert and Smith, Phillip. A rabic Emoji Sentiment Lexicon ( A rab- ESL ): A Comparison between A rabic and E uropean Emoji Sentiment Lexicons. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  219. [240]

    A r COV 19-Rumors: A rabic COVID -19 T witter Dataset for Misinformation Detection

    Haouari, Fatima and Hasanain, Maram and Suwaileh, Reem and Elsayed, Tamer. A r COV 19-Rumors: A rabic COVID -19 T witter Dataset for Misinformation Detection. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  220. [241]

    A r COV -19: The First A rabic COVID -19 T witter Dataset with Propagation Networks

    Haouari, Fatima and Hasanain, Maram and Suwaileh, Reem and Elsayed, Tamer. A r COV -19: The First A rabic COVID -19 T witter Dataset with Propagation Networks. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  221. [242]

    The Interplay of Variant, Size, and Task Type in A rabic Pre-trained Language Models

    Inoue, Go and Alhafni, Bashar and Baimukan, Nurpeiis and Bouamor, Houda and Habash, Nizar. The Interplay of Variant, Size, and Task Type in A rabic Pre-trained Language Models. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  222. [243]

    Automatic Difficulty Classification of A rabic Sentences

    Khallaf, Nouran and Sharoff, Serge. Automatic Difficulty Classification of A rabic Sentences. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  223. [244]

    Dynamic Ensembles in Named Entity Recognition for Historical A rabic Texts

    Majadly, Muhammad and Sagi, Tomer. Dynamic Ensembles in Named Entity Recognition for Historical A rabic Texts. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  224. [245]

    A rabic Offensive Language on T witter: Analysis and Experiments

    Mubarak, Hamdy and Rashed, Ammar and Darwish, Kareem and Samih, Younes and Abdelali, Ahmed. A rabic Offensive Language on T witter: Analysis and Experiments. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  225. [246]

    Adult Content Detection on A rabic T witter: Analysis and Experiments

    Mubarak, Hamdy and Hassan, Sabit and Abdelali, Ahmed. Adult Content Detection on A rabic T witter: Analysis and Experiments. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  226. [247]

    UL 2 C : Mapping User Locations to Countries on A rabic T witter

    Mubarak, Hamdy and Hassan, Sabit. UL 2 C : Mapping User Locations to Countries on A rabic T witter. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  227. [248]

    Let-Mi: An A rabic L evantine T witter Dataset for Misogynistic Language

    Mulki, Hala and Ghanem, Bilal. Let-Mi: An A rabic L evantine T witter Dataset for Misogynistic Language. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  228. [249]

    Empathetic BERT 2 BERT Conversational Model: Learning A rabic Language Generation with Little Data

    Naous, Tarek and Antoun, Wissam and Mahmoud, Reem and Hajj, Hazem. Empathetic BERT 2 BERT Conversational Model: Learning A rabic Language Generation with Little Data. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  229. [250]

    ALUE : A rabic Language Understanding Evaluation

    Seelawi, Haitham and Tuffaha, Ibraheem and Gzawi, Mahmoud and Farhan, Wael and Talafha, Bashar and Badawi, Riham and Sober, Zyad and Al-Dweik, Oday and Freihat, Abed Alhakim and Al-Natsheh, Hussein. ALUE : A rabic Language Understanding Evaluation. Proceedings of the Sixth Ara...

  230. [251]

    Quranic Verses Semantic Relatedness Using A ra BERT

    Alsaleh, Abdullah and Atwell, Eric and Altahhan, Abdulrahman. Quranic Verses Semantic Relatedness Using A ra BERT. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  231. [252]

    A ra ELECTRA : Pre-Training Text Discriminators for A rabic Language Understanding

    Antoun, Wissam and Baly, Fady and Hajj, Hazem. A ra ELECTRA : Pre-Training Text Discriminators for A rabic Language Understanding. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  232. [253]

    A ra GPT 2: Pre-Trained Transformer for A rabic Language Generation

    Antoun, Wissam and Baly, Fady and Hajj, Hazem. A ra GPT 2: Pre-Trained Transformer for A rabic Language Generation. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  233. [254]

    Q uran T ree.jl: A Julia Package for Quranic A rabic Corpus

    Asaad, Al-Ahmadgaid. Q uran T ree.jl: A Julia Package for Quranic A rabic Corpus. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  234. [255]

    Automatic R omanization of A rabic Bibliographic Records

    Eryani, Fadhl and Habash, Nizar. Automatic R omanization of A rabic Bibliographic Records. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  235. [256]

    SERAG : Semantic Entity Retrieval from A rabic Knowledge Graphs

    Esmeir, Saher. SERAG : Semantic Entity Retrieval from A rabic Knowledge Graphs. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  236. [257]

    Introducing A large T unisian A rabizi Dialectal Dataset for Sentiment Analysis

    Fourati, Chayma and Haddad, Hatem and Messaoudi, Abir and BenHajhmida, Moez and Ben Elhaj Mabrouk, Aymen and Naski, Malek. Introducing A large T unisian A rabizi Dialectal Dataset for Sentiment Analysis. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  237. [258]

    A ra F acts: The First Large A rabic Dataset of Naturally Occurring Claims

    Sheikh Ali, Zien and Mansour, Watheq and Elsayed, Tamer and Al‐Ali, Abdulaziz. A ra F acts: The First Large A rabic Dataset of Naturally Occurring Claims. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  238. [259]

    Improving Cross-Lingual Transfer for Event Argument Extraction with Language-Universal Sentence Structures

    Nguyen, Minh Van and Nguyen, Thien Huu. Improving Cross-Lingual Transfer for Event Argument Extraction with Language-Universal Sentence Structures. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  239. [260]

    NADI 2021: The Second Nuanced A rabic Dialect Identification Shared Task

    Abdul-Mageed, Muhammad and Zhang, Chiyu and Elmadany, AbdelRahim and Bouamor, Houda and Habash, Nizar. NADI 2021: The Second Nuanced A rabic Dialect Identification Shared Task. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  240. [261]

    Adapting MARBERT for Improved A rabic Dialect Identification: Submission to the NADI 2021 Shared Task

    AlKhamissi, Badr and Gabr, Mohamed and ElNokrashy, Muhammad and Essam, Khaled. Adapting MARBERT for Improved A rabic Dialect Identification: Submission to the NADI 2021 Shared Task. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  241. [262]

    Country-level A rabic Dialect Identification Using Small Datasets with Integrated Machine Learning Techniques and Deep Learning Models

    Althobaiti, Maha J. Country-level A rabic Dialect Identification Using Small Datasets with Integrated Machine Learning Techniques and Deep Learning Models. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  242. [263]

    BERT -based Multi-Task Model for Country and Province Level MSA and Dialectal A rabic Identification

    El Mekki, Abdellah and El Mahdaouy, Abdelkader and Essefar, Kabil and El Mamoun, Nabil and Berrada, Ismail and Khoumsi, Ahmed. BERT -based Multi-Task Model for Country and Province Level MSA and Dialectal A rabic Identification. Proceedings of the Sixth Arabic Natural Language...

  243. [264]

    Country-level A rabic Dialect Identification using RNN s with and without Linguistic Features

    Issa, Elsayed and AlShakhori1, Mohammed and Al-Bahrani, Reda and Hahn-Powell, Gus. Country-level A rabic Dialect Identification using RNN s with and without Linguistic Features. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  244. [265]

    A rabic Dialect Identification based on a Weighted Concatenation of TF - IDF Features

    Lichouri, Mohamed and Abbas, Mourad and Lounnas, Khaled and Benaziz, Besma and Zitouni, Aicha. A rabic Dialect Identification based on a Weighted Concatenation of TF - IDF Features. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  245. [266]

    Machine Learning-Based Approach for A rabic Dialect Identification

    Nayel, Hamada and Hassan, Ahmed and Sobhi, Mahmoud and El-Sawy, Ahmed. Machine Learning-Based Approach for A rabic Dialect Identification. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  246. [267]

    Dialect Identification in Nuanced A rabic Tweets Using Farasa Segmentation and A ra BERT

    Wadhawan, Anshul. Dialect Identification in Nuanced A rabic Tweets Using Farasa Segmentation and A ra BERT. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  247. [268]

    Overview of the WANLP 2021 Shared Task on Sarcasm and Sentiment Detection in A rabic

    Abu Farha, Ibrahim and Zaghouani, Wajdi and Magdy, Walid. Overview of the WANLP 2021 Shared Task on Sarcasm and Sentiment Detection in A rabic. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  248. [269]

    WANLP 2021 Shared-Task: Towards Irony and Sentiment Detection in A rabic Tweets using Multi-headed- LSTM - CNN - GRU and M a RBERT

    Abdel-Salam, Reem. WANLP 2021 Shared-Task: Towards Irony and Sentiment Detection in A rabic Tweets using Multi-headed- LSTM - CNN - GRU and M a RBERT. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  249. [270]

    Sarcasm and Sentiment Detection In A rabic Tweets Using BERT -based Models and Data Augmentation

    Abuzayed, Abeer and Al-Khalifa, Hend. Sarcasm and Sentiment Detection In A rabic Tweets Using BERT -based Models and Data Augmentation. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  250. [271]

    and Lee, Mark

    Alharbi, Abdullah I. and Lee, Mark. Multi-task Learning Using a Combination of Contextualised and Static Word Embeddings for A rabic Sarcasm Detection and Sentiment Analysis. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  251. [272]

    A r S arcasm Shared Task: An Ensemble BERT Model for S arcasm D etection in A rabic Tweets

    Bashmal, Laila and AlZeer, Daliyah. A r S arcasm Shared Task: An Ensemble BERT Model for S arcasm D etection in A rabic Tweets. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  252. [273]

    Sarcasm and Sentiment Detection in A rabic: investigating the interest of character-level features

    Ghoul, Dhaou and Lejeune, Ga. Sarcasm and Sentiment Detection in A rabic: investigating the interest of character-level features. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  253. [274]

    Deep Multi-Task Model for Sarcasm Detection and Sentiment Analysis in A rabic Language

    El Mahdaouy, Abdelkader and El Mekki, Abdellah and Essefar, Kabil and El Mamoun, Nabil and Berrada, Ismail and Khoumsi, Ahmed. Deep Multi-Task Model for Sarcasm Detection and Sentiment Analysis in A rabic Language. Proceedings of the Sixth Arabic Natural Language Processing Wo...

  254. [275]

    A Contextual Word Embedding for A rabic Sarcasm Detection with Random Forests

    Elgabry, Hazem and Attia, Shimaa and Abdel-Rahman, Ahmed and Abdel-Ate, Ahmed and Girgis, Sandra. A Contextual Word Embedding for A rabic Sarcasm Detection with Random Forests. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  255. [276]

    S arcasm D et at Sarcasm Detection Task 2021 in A rabic using A ra BERT Pretrained Model

    Faraj, Dalya and Faraj, Dalya and Abdullah, Malak. S arcasm D et at Sarcasm Detection Task 2021 in A rabic using A ra BERT Pretrained Model. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  256. [277]

    Sarcasm and Sentiment Detection in A rabic language A Hybrid Approach Combining Embeddings and Rule-based Features

    Gaanoun, Kamel and Benelallam, Imade. Sarcasm and Sentiment Detection in A rabic language A Hybrid Approach Combining Embeddings and Rule-based Features. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  257. [278]

    Combining Context-Free and Contextualized Representations for A rabic Sarcasm Detection and Sentiment Identification

    Hengle, Amey and Kshirsagar, Atharva and Desai, Shaily and Marathe, Manisha. Combining Context-Free and Contextualized Representations for A rabic Sarcasm Detection and Sentiment Identification. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  258. [279]

    Leveraging Offensive Language for Sarcasm and Sentiment Detection in A rabic

    Husain, Fatemah and Uzuner, Ozlem. Leveraging Offensive Language for Sarcasm and Sentiment Detection in A rabic. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  259. [280]

    The IDC System for Sentiment Classification and Sarcasm Detection in A rabic

    Israeli, Abraham and Nahum, Yotam and Fine, Shai and Bar, Kfir. The IDC System for Sentiment Classification and Sarcasm Detection in A rabic. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  260. [281]

    Preprocessing Solutions for Detection of Sarcasm and Sentiment for A rabic

    Lichouri, Mohamed and Abbas, Mourad and Benaziz, Besma and Zitouni, Aicha and Lounnas, Khaled. Preprocessing Solutions for Detection of Sarcasm and Sentiment for A rabic. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  261. [282]

    i C ompass at Shared Task on Sarcasm and Sentiment Detection in A rabic

    Naski, Malek and Messaoudi, Abir and Haddad, Hatem and BenHajhmida, Moez and Fourati, Chayma and Ben Elhaj Mabrouk, Aymen. i C ompass at Shared Task on Sarcasm and Sentiment Detection in A rabic. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  262. [283]

    Machine Learning-Based Model for Sentiment and Sarcasm Detection

    Nayel, Hamada and Amer, Eslam and Allam, Aya and Abdallah, Hanya. Machine Learning-Based Model for Sentiment and Sarcasm Detection. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  263. [284]

    D eep B lue AI at WANLP - EACL 2021 task 2: A Deep Ensemble-based Method for Sarcasm and Sentiment Detection in A rabic

    Song, Bingyan and Pan, Chunguang and Wang, Shengguang and Luo, Zhipeng. D eep B lue AI at WANLP - EACL 2021 task 2: A Deep Ensemble-based Method for Sarcasm and Sentiment Detection in A rabic. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  264. [285]

    A ra BERT and Farasa Segmentation Based Approach For Sarcasm and Sentiment Detection in A rabic Tweets

    Wadhawan, Anshul. A ra BERT and Farasa Segmentation Based Approach For Sarcasm and Sentiment Detection in A rabic Tweets. Proceedings of the Sixth Arabic Natural Language Processing Workshop. 2021

  265. [286]

    Proceedings of the Fourth Workshop on Visually Grounded Interaction and Language. 2021

  266. [287]

    Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  267. [288]

    Findings of the V ar D ial Evaluation Campaign 2021

    Chakravarthi, Bharathi Raja and Mihaela, Gaman and Ionescu, Radu Tudor and Jauhiainen, Heidi and Jauhiainen, Tommi and Lind \'e n, Krister and Ljube s i \'c , Nikola and Partanen, Niko and Priyadharshini, Ruba and Purschke, Christoph and Rajagopal, Eswari and Scherrer, Yves an...

  268. [289]

    Hierarchical Transformer for Multilingual Machine Translation

    Khusainova, Albina and Khan, Adil and Rivera, Ad \' n Ram \' rez and Romanov, Vitaly. Hierarchical Transformer for Multilingual Machine Translation. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  269. [290]

    Regression Analysis of Lexical and Morpho-Syntactic Properties of Kiezdeutsch

    Frassinelli, Diego and Lapesa, Gabriella and Alatrash, Reem and Schlechtweg, Dominik and Schulte im Walde, Sabine. Regression Analysis of Lexical and Morpho-Syntactic Properties of Kiezdeutsch. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dial...

  270. [291]

    Representations of Language Varieties Are Reliable Given Corpus Similarity Measures

    Dunn, Jonathan. Representations of Language Varieties Are Reliable Given Corpus Similarity Measures. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  271. [292]

    Whit ' s the Richt Pairt o Speech: P o S tagging for S cots

    Lameris, Harm and Stymne, Sara. Whit ' s the Richt Pairt o Speech: P o S tagging for S cots. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  272. [293]

    Efficient Unsupervised NMT for Related Languages with Cross-Lingual Language Models and Fidelity Objectives

    Aly, Rami and Caines, Andrew and Buttery, Paula. Efficient Unsupervised NMT for Related Languages with Cross-Lingual Language Models and Fidelity Objectives. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  273. [294]

    Fine-tuning Distributional Semantic Models for Closely-Related Languages

    Bhatia, Kushagra and Aggarwal, Divyanshu and Vaidya, Ashwini. Fine-tuning Distributional Semantic Models for Closely-Related Languages. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  274. [295]

    Discriminating Between Similar Nordic Languages

    Haas, Ren \'e and Derczynski, Leon. Discriminating Between Similar Nordic Languages. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  275. [296]

    Naive B ayes-based Experiments in R omanian Dialect Identification

    Jauhiainen, Tommi and Jauhiainen, Heidi and Lind \'e n, Krister. Naive B ayes-based Experiments in R omanian Dialect Identification. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  276. [297]

    U nibuc K ernel: Geolocating S wiss G erman Jodels Using Ensemble Learning

    Mihaela, Gaman and Cojocariu, Sebastian and Ionescu, Radu Tudor. U nibuc K ernel: Geolocating S wiss G erman Jodels Using Ensemble Learning. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  277. [298]

    Optimizing a Supervised Classifier for a Difficult Language Identification Problem

    Bestgen, Yves. Optimizing a Supervised Classifier for a Difficult Language Identification Problem. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  278. [299]

    Comparing the Performance of CNN s and Shallow Models for Language Identification

    Ceolin, Andrea. Comparing the Performance of CNN s and Shallow Models for Language Identification. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects. 2021

  279. [300]

    Dialect Identification through Adversarial Learning and Knowledge Distillation on R omanian BERT

    Zaharia, George-Eduard and Avram, Andrei-Marius and Cercel, Dumitru-Clementin and Rebedea, Traian. Dialect Identification through Adversarial Learning and Knowledge Distillation on R omanian BERT. Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and D...

Pith tools

Reviewed May 16, 2026 · model on record in the stance chip above.