Fine-tuning ReaderLM-v2 on rule-based Fundus extractions yields a small model that outperforms zero-shot LLMs on HTML-to-plaintext and HTML-to-JSON overlap metrics, but the evaluation is biased toward the Fundus rules used as both training target and test reference.
W eb IE : Faithful and Robust Information Extraction on the Web
1 Pith paper cite this work, alongside 3 external citations. Polarity classification is still indexing.
1
Pith paper citing it
3
external citations · OpenAlex
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling
Fine-tuning ReaderLM-v2 on rule-based Fundus extractions yields a small model that outperforms zero-shot LLMs on HTML-to-plaintext and HTML-to-JSON overlap metrics, but the evaluation is biased toward the Fundus rules used as both training target and test reference.