Chinese-SkillSpan is the first ESCO-aligned span-level dataset for extracting competencies from over 20,000 Chinese job advertisements.
arXiv preprint arXiv:2311.08526 (2023)
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5verdicts
UNVERDICTED 5representative citing papers
Introduces a 200-document benchmark and character-level R-Score for contextual PII redaction, with model evaluations and human agreement data showing the task remains unsolved.
The authors propose a multitask GLiNER framework with synthetic data and LLM revalidation for scalable monitoring and classification of dataset usage in academic literature.
Opir introduces efficient multi-task encoder models trained on a 996-category safety taxonomy that match or exceed larger baselines on most safety benchmarks while using under 100M parameters for edge variants.
A Telegram data collection pipeline with Parakeet transcription and custom transformer NER models that detects sensitive entities at high F1 scores and provides anonymization metrics preserving structural coherence.
citing papers explorer
-
Chinese-SkillSpan: A Span-Level Dataset for ESCO-Aligned Competency Extraction from Chinese Job Ads
Chinese-SkillSpan is the first ESCO-aligned span-level dataset for extracting competencies from over 20,000 Chinese job advertisements.
-
RedactionBench
Introduces a 200-document benchmark and character-level R-Score for contextual PII redaction, with model evaluations and human agreement data showing the task remains unsolved.
-
AI for Monitoring and Classifying Data Used in Research Literature
The authors propose a multitask GLiNER framework with synthetic data and LLM revalidation for scalable monitoring and classification of dataset usage in academic literature.
-
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
Opir introduces efficient multi-task encoder models trained on a 996-category safety taxonomy that match or exceed larger baselines on most safety benchmarks while using under 100M parameters for edge variants.
-
Identification and Anonymization of Named Entities in Unstructured Information Sources for Use in Social Engineering Detection
A Telegram data collection pipeline with Parakeet transcription and custom transformer NER models that detects sensitive entities at high F1 scores and provides anonymization metrics preserving structural coherence.