Pith. sign in

Paper Citation Record · LEDGER

Multilingual Attribute Extraction from News Web Pages

As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2502.02167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02167 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:08:48.902025Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:30:16.403507Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:30:17.233809Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03e96e32-03aa-4e15-baf1-34334b20e9ef · outbound

This paper cites Markuplm: Pre-training of text and markup language for visually rich document understanding,.

Multilingual Attribute Extraction from News Web Pages Markuplm: Pre-training of text and markup language for visually rich document understanding,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.928019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.579317Z digest=sha256:137ad3461ba90840db3cc6bd594a65901f756483b2ae3465bed29ed267720f8e

Observation e0373d75-3f81-44f2-9f74-82877e0e448e · outbound

This paper cites DOM-LM: Learning Generalizable Representations for HTML Documents.

Multilingual Attribute Extraction from News Web Pages DOM-LM: Learning Generalizable Representations for HTML Documents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T13:08:48.619437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:08:48.619437Z digest=sha256:c3b4d2abff973758c5888ecf35bb849d7e49c6751d79b847e5eca4cf158fdc85

Observation 2dd98728-9837-4149-8187-09ba01206563 · outbound

This paper cites From one tree to a forest: a unified solution for structured web data extraction,.

Multilingual Attribute Extraction from News Web Pages From one tree to a forest: a unified solution for structured web data extraction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.917186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.624731Z digest=sha256:0f3fe3ff2089dc99fc7a7dd6e8ddc9577c5d60000ecadb0eed3e21e976211373

Observation cf338c18-cf25-4114-b2fa-4a1c96946438 · outbound

This paper cites The Klarna Product Page Dataset: Web Element Nomination with Graph Neural Networks and Large Language Models.

Multilingual Attribute Extraction from News Web Pages The Klarna Product Page Dataset: Web Element Nomination with Graph Neural Networks and Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:08:49.177818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.629398Z digest=sha256:605c0066574c4011629113db51902d978878f0212de95cf137c3c75e9b7d98cc

Observation da61feb0-b491-473f-afbf-224dbb069da9 · outbound

This paper cites CoVA: Context-aware Visual Attention for Webpage Information Extraction.

Multilingual Attribute Extraction from News Web Pages CoVA: Context-aware Visual Attention for Webpage Information Extraction

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:08:49.031823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.635283Z digest=sha256:2555386709786dd2f47f6299c5355d39eecaff7e42ffd16153afcb2ca157ed3b

Observation 77d64342-3736-4a6b-8fcc-de5408d31986 · outbound

This paper cites A dataset for information extraction from news web pages,.

Multilingual Attribute Extraction from News Web Pages A dataset for information extraction from news web pages,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.823689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.639951Z digest=sha256:0a1d94177147c052da4147f794ca3766252a14c485dec2af78e0ee19acd61461

Observation 680cef1e-e914-4eba-963e-9dea390b3c45 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Multilingual Attribute Extraction from News Web Pages RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:08:48.644940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:08:48.644940Z digest=sha256:16bf84d5c8c19434327bf51f50988a94fe88cc81befc31f384debc43d0587f7b

Observation 60552130-13a0-4004-976c-e339059d136e · outbound

This paper cites Learning structural co-occurrences for structured web data extraction in low- resource settings,.

Multilingual Attribute Extraction from News Web Pages Learning structural co-occurrences for structured web data extraction in low- resource settings,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.694676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.650530Z digest=sha256:19fbdbbc8922dd75ab5afccd21b6f231a46384f6e5d0edf535e629fa15a3236e

Observation 3ec5dff3-725c-41c9-b65e-d70c6de5f57d · outbound

This paper cites Hierarchi- cal multimodal pre-training for visually rich webpage understanding,.

Multilingual Attribute Extraction from News Web Pages Hierarchi- cal multimodal pre-training for visually rich webpage understanding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.636022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.655372Z digest=sha256:b163942018923501bb4aefddc355b2174ed6f6c97a6cb687c43591a3a46e6fc9

Observation 206ed75b-9da7-4a5c-b60d-ad4450579a4a · outbound

This paper cites Label Studio: Data labeling software,.

Multilingual Attribute Extraction from News Web Pages Label Studio: Data labeling software,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.427473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.678412Z digest=sha256:648caed41731c38d76afa0ef9b008de3139bce9bae9eb3aa7f1d7e865fb16f2e

Observation 81a6b8d9-f36a-4db8-b01d-33df71bc5dfe · outbound

This paper cites Un- supervised cross-lingual representation learning at scale,.

Multilingual Attribute Extraction from News Web Pages Un- supervised cross-lingual representation learning at scale,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.415588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.763569Z digest=sha256:00dd77e4e33dc1efed7849b9600cbbcb79c9a2c153192e9f96449812f4e6bcb0

Observation 5b63ebbe-8f1c-437a-a651-efb203ed131e · outbound

This paper cites Neural machine translation with byte- level subwords,.

Multilingual Attribute Extraction from News Web Pages Neural machine translation with byte- level subwords,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.403821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.894276Z digest=sha256:f2d5844819b20102d5e30b847a98de516c089fa8a0ce06e81862c9662e98c4a1

Observation d7c63974-01b8-48ce-a83d-c3a74a98216b · outbound

This paper cites Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction,.

Multilingual Attribute Extraction from News Web Pages Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.253528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.898364Z digest=sha256:bcd4d097fbebe02f892ad76a79a7549ffab77cc5a190f95875e8745184db89e1

Observation 8bc3801a-862f-49ba-9fcc-48db1740c9c1 · outbound

This paper cites news-please: A generic news crawler and extractor,.

Multilingual Attribute Extraction from News Web Pages news-please: A generic news crawler and extractor,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:08:49.199622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T13:08:48.902025Z digest=sha256:0919f7f4b54c9968cf690bddd2b7a8136166d173d682528757545b3ba536c3a7

Pith citing papers

Observation ddd914dd-db45-4b85-a81d-33c9dba67433 · inbound

WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents cites this paper.

WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents Multilingual Attribute Extraction from News Web Pages

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:30:17.238147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T12:30:16.403507Z digest=sha256:b8d4cebc7a1481d8c9243eb2f5abfe126b90f8ba09fe20a7cb0b59e92f570bec