Authors extend an existing Arabic QA dataset into the first parallel open-ended benchmark across dialects and MSA, then benchmark LLMs showing underperformance on dialects and open-ended questions.
Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, and Muhammad Abdul-Mageed
2 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
2
Pith papers citing it
2
external citations · external index
fields
cs.CL 2verdicts
UNVERDICTED 2representative citing papers
Introduces ShopTrajQA long-context benchmark and an RLVR-trained tool-augmented agent that bypasses LLM context limits by external file storage and code-based retrieval for shopping trajectories.
citing papers explorer
-
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
Authors extend an existing Arabic QA dataset into the first parallel open-ended benchmark across dialects and MSA, then benchmark LLMs showing underperformance on dialects and open-ended questions.
-
Customer-Agent: Overcoming Context Limitations in Ultra-Long Shopping Trajectories via Tool-Augmented Agents and RLVR
Introduces ShopTrajQA long-context benchmark and an RLVR-trained tool-augmented agent that bypasses LLM context limits by external file storage and code-based retrieval for shopping trajectories.