Pith. sign in

Paper Citation Record · LEDGER

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2412.04026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04026 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:55.976780Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8593755e-5d6b-433e-baff-aa8db76bb587 · outbound

This paper cites DiffusionNER: Boundary diffusion for named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction DiffusionNER: Boundary diffusion for named entity recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.098637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.645493Z digest=sha256:6721e3ed4a3448f7d0e04aaffb911ceac0717f5feca0e043e612f6121c0b38ee

Observation 46c8358c-55d4-491d-87e7-4dafacc56f73 · outbound

This paper cites Dual cache for long document neural coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dual cache for long document neural coreference resolution,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.082147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.651675Z digest=sha256:11caf74aff910d8a1a2ee2662055c8d8eb5dda39da204bb4785278c219305051

Observation 4a8f2460-1eaa-4815-850b-1764e609ff6b · outbound

This paper cites An autoregressive text-to-graph framework for joint entity and relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An autoregressive text-to-graph framework for joint entity and relation extraction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.064499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.656845Z digest=sha256:3e1ad0a458ba8a2eca27cbbe4f82bc6c10398f950f3cc85a662fa1e2a6e8befe

Observation 35ec8c6f-d5fe-4053-ad91-9bf35c1e0844 · outbound

This paper cites Event extraction as question generation and answering,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Event extraction as question generation and answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.662619Z digest=sha256:ec32a77e5fad22cf6eaf2d50ae5a7bc085fba56a2150ba5796442a133bab4202

Observation daa359ee-d090-4e94-80c5-e864f9b48fd6 · outbound

This paper cites Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.029864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.667965Z digest=sha256:867c67985844f0029f89d5c9964160b97c4e00172243e2504093a277f0f9edeb

Observation 0b73483d-7035-4c59-aa5b-37e6ccfba00c · outbound

This paper cites A span-based model for joint overlapped and discontinuous named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based model for joint overlapped and discontinuous named entity recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.010404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.674857Z digest=sha256:22b6d971ea44f9607b8553bfdcb66ba07b0c389689d6c73974361ad3d9b95eda

Observation fac3f90b-679e-4034-aa33-1d6960dc460e · outbound

This paper cites Unified named entity recognition as word-word relation classification,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Unified named entity recognition as word-word relation classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.991325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.681483Z digest=sha256:f92c226f626859599167408b6f6d0a63fd8f8b60161aba5cb0b3e81317238b4c

Observation 085f8686-478d-456b-9cb0-116e5f2da0b2 · outbound

This paper cites Knowledge enhanced coreference resolution via gated attention,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Knowledge enhanced coreference resolution via gated attention,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.972872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.686416Z digest=sha256:f3d6773a728a5bd7acce09786ea41de39d55cdaaa8555aec52b8b8c3444fe3fa

Observation 7dd0163b-1bad-41fc-835e-eaa662995bae · outbound

This paper cites Double graph based reasoning for document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Double graph based reasoning for document-level relation extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.953381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.691524Z digest=sha256:f4f9efc02e9331c2faa512d1284f43487c9fd73150333d4f53433559ef2b6a03

Observation 22b73191-308e-44e7-8b50-839af3353d3a · outbound

This paper cites Coreference resolution without span representations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Coreference resolution without span representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.933021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.696807Z digest=sha256:8c8a8933157dbc743e3bfeb184e534353985dbb4a94ade19c4492480cf02be9f

Observation f7f6b36d-0c67-497c-abdf-afe1d63d9955 · outbound

This paper cites A sequence-to-sequence approach for document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A sequence-to-sequence approach for document-level relation extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.915822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.701825Z digest=sha256:68ca95d50602fc44fda929f52d0165b7e002750653a7f79c281b4d739837272f

Observation 20c8838f-bcaa-4d4a-a2b0-2e48930fb651 · outbound

This paper cites Visual attention model for name tagging in multimodal social media,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual attention model for name tagging in multimodal social media,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.899191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.707133Z digest=sha256:13e320f8983d506fd88be5099adedc6866ccce0a56cc7751f0974193b8fa9fc7

Observation 6d315406-ac50-46b1-aaea-0ba94cd620c1 · outbound

This paper cites Adaptive co-attention network for named entity recognition in tweets,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Adaptive co-attention network for named entity recognition in tweets,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.882580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.712237Z digest=sha256:6fa61147fa9f7be2b77ac9b64e3dabe4c51e64cee51736a6451bf13b8e8acdac

Observation 127096dd-19e9-442e-b007-70ea11d91e36 · outbound

This paper cites A large-scale chinese multimodal ner dataset with speech clues,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A large-scale chinese multimodal ner dataset with speech clues,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.866161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.717006Z digest=sha256:6ceaae4a5dc2e13462fdb3e950c0cbc00a5cf6ec22cac1380fe327c0dd04c9b0

Observation 21f13c5b-8483-4977-b277-ce850b99b0a7 · outbound

This paper cites Who are you referring to? coreference resolution in image narrations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Who are you referring to? coreference resolution in image narrations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.831366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.726731Z digest=sha256:16a3cd653b379112bdddc193bc2fb656d9424bbbbb332cdb505ef1a4bcd3ad6b

Observation d366b976-7698-4089-944c-07597330d6f4 · outbound

This paper cites Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.811261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.731553Z digest=sha256:9a274def86adc0e048f7cb88a66d791812338cf60d1e83926af19b2592e30fae

Observation ee792f39-f53b-48f1-8a22-1e680a79554c · outbound

This paper cites A hierarchical network for multimodal document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A hierarchical network for multimodal document-level relation extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.794259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.736909Z digest=sha256:7e53c34439e925404f495217c48852eed04dd3b4088ff2caa04b689d443e35b6

Observation 64464a7c-2ce7-4c67-b625-b49e06889e5a · outbound

This paper cites Grounded multimodal named entity recognition on social media,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Grounded multimodal named entity recognition on social media,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.849101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.741883Z digest=sha256:7d47cecee17a432dc4884e9049eec254ffa8f7812433ec259f81840ade578359

Observation 3b718cd1-8fc2-412e-be63-439595997b9c · outbound

This paper cites Semi-supervised multimodal coreference resolution in image narrations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Semi-supervised multimodal coreference resolution in image narrations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.775857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.748153Z digest=sha256:4daea9860c77da77eed933a0354ed7665d2db8bb14cfac1d76770750011a87f9

Observation f55bcf04-1a0f-4fb6-869e-cfbbf9ab1d60 · outbound

This paper cites Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.759051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.753323Z digest=sha256:fbf8167de04c0f8c8c84298c1d6520d6cc2275b915f6e0ac08f6596239b7f745

Observation 3bff049b-ab28-422b-b1af-03f0c210678b · outbound

This paper cites Multimodal relation extraction with efficient graph alignment,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal relation extraction with efficient graph alignment,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.740778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.759044Z digest=sha256:3cab4ff7327eb1fecbfbc3101944dde8c1238fb141124ed874d74a7baffed0f5

Observation 1700b069-affd-4949-a99c-3afae2238976 · outbound

This paper cites Docred: A large-scale document-level relation extraction dataset,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Docred: A large-scale document-level relation extraction dataset,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.724088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.764038Z digest=sha256:e56bb05adffdce6b9ce6102183eb5434537e57a1b80d27903819e89b18b48414

Observation ce4a9515-a6dd-414c-a1f1-7c77d4244199 · outbound

This paper cites Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.706299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.768999Z digest=sha256:749819f00974cceda515d691e50c1ee57c3d898ea23ba563e9f376bbecdb6d37

Observation dcea31c6-3b17-4c7d-8673-8791ac23cdef · outbound

This paper cites A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.688592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.774151Z digest=sha256:3d061e1150c8d788e4be825c0a8c4f7502edd9b2ae634cfc48f9bfbcc44a2e93

Observation a925a4fa-ce05-4b19-b1dd-99d624be8349 · outbound

This paper cites Entity- level interaction via heterogeneous graph for multimodal named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Entity- level interaction via heterogeneous graph for multimodal named entity recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.670543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.779488Z digest=sha256:acf3f9be94a16287276fe7eedb2a1792cfa02b09bca34271a4e7df46ed2de51e

Observation 0b949fab-1e68-48aa-b1b8-73d0f3058adc · outbound

This paper cites Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.650947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.785111Z digest=sha256:e50d27009081f264266a102c02d3aac8cb793fc393945db3d40383e09b62ed2a

Observation 6b13dd77-f9ea-4701-b6d7-a924b15827e6 · outbound

This paper cites Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.632958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.790468Z digest=sha256:e8f39e03ccbcb34c132830e58be528d0ae19d3777fe4bc3ec7281ce39df9e25b

Observation 6a183391-336c-4d56-8247-c8f194b86873 · outbound

This paper cites Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.615185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.795316Z digest=sha256:f0b83bc0793040ca6fe1b7d705278e13df115fa91b7015614834e285a2f2c086

Observation 289e4c81-903e-449d-9171-41c4d900c578 · outbound

This paper cites Rethinking multimodal entity and relation extraction from a translation point of view,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking multimodal entity and relation extraction from a translation point of view,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.597643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.802304Z digest=sha256:0ae449c91259575fe5a77ec78e1360205e854d3c3ce3ac5496cfa18700f46d48

Observation aba6abb6-d9f3-43dd-b520-2b7b4884d727 · outbound

This paper cites Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.579968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.807645Z digest=sha256:228a6e6e78911cf9b22ec39393d7a4ffc34dbb1d00633ca651ca1c9f7d5ed9cd

Observation 59b3563e-06c6-444f-be96-14b66d40ae54 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Transvg: End-to-end visual grounding with transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.812633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.812633Z digest=sha256:5ae519b5f4e2e033082733f54065b1e5de4e8f4490bc5fdfd605c69ebcd443eb

Observation 10d8a44e-499c-4f42-9bad-5b21ba5e3538 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.549750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.818110Z digest=sha256:094ecf509d1b82bdc7c8fc5931e3cebb41b7bc4925b58baa343639137e7763c8

Observation b6a867e3-aab2-4351-b594-4ae5589fcb64 · outbound

This paper cites You only look once: Unified, real-time object detection,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction You only look once: Unified, real-time object detection,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.823763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.823763Z digest=sha256:262c826413a6463771c9a3f9e90a28f0a2eb36387c0808d4d885c08929d7398f

Observation e60e0871-744e-4ab5-a7c4-3cc522724eb6 · outbound

This paper cites Ssd: Single shot multibox detector,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ssd: Single shot multibox detector,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.519786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.829059Z digest=sha256:5aa44900ef63b2c482fbbbe288f970e3f3246febb0f2029e870a7f4c22ef9efd

Observation 70f282fc-3a62-43f4-821c-dbd7404185a1 · outbound

This paper cites Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.500810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.834799Z digest=sha256:a9c5c5e96445ef8f018671f081efc9932ec0dff0eb3662f543acf769ee551791

Observation c63856b6-a7dc-4429-b84e-660cddeb2d03 · outbound

This paper cites Missing modalities imputation via cascaded residual autoencoder,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modalities imputation via cascaded residual autoencoder,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.481316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.839867Z digest=sha256:76ffd1f108ea8eceae263f86d4a7c565cf189e2c804094b650146d0e37555a4a

Observation 29585045-df22-4ae7-8611-bcf375eb6047 · outbound

This paper cites Lrmm: Learning to recommend with missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Lrmm: Learning to recommend with missing modalities,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.463516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.844766Z digest=sha256:50f7e7c8b391dfa8b8bd7768124fa7aea6931594efc0acef5df93255fe276441

Observation 3d5846a8-68ac-426e-9596-808564db30f7 · outbound

This paper cites Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.445404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.850013Z digest=sha256:89398a4555d24b388b4d8b6f900cf1e56adb6c203498ba6378cb95fcae2bb06a

Observation 93c3b3cb-0854-4ddc-807a-2ed9696d3f8a · outbound

This paper cites A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.427616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.855160Z digest=sha256:3deb165ae1c0d76b370157926ae87ea272b49b181cc9ab8873baf631084b1901

Observation 398e2ac7-d9c4-4c6b-8fef-6f8e1c49bfbf · outbound

This paper cites Missing modality imagination network for emotion recognition with uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modality imagination network for emotion recognition with uncertain missing modalities,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.860148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.860148Z digest=sha256:b0022a5330bc5f4fe5c022566db4c199854cbc1d3de570d75304cde247c07fc3

Observation 3d58a6d9-b30a-42cb-a99e-9020de48ed42 · outbound

This paper cites Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.865429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.865429Z digest=sha256:8b668b5ebf7ddb5ce15147cc51e36088169ede51915526a303936fb259f63903

Observation 39c5209d-1341-4f87-a005-90bf62d45817 · outbound

This paper cites Multimodal prompting with missing modalities for visual recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompting with missing modalities for visual recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.386234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.870407Z digest=sha256:7d116c930eba4ec24788f9786cfe9bf45b1f6405767ad764e58fe58f647ab578

Observation 1f90a499-dbcd-4942-8037-355dd9c53f14 · outbound

This paper cites Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.368424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.875900Z digest=sha256:33c6c63c2dfedbf462913464dfb5eba7bb0269199a69ee28fef83edb3bab7a08

Observation 29abe825-40e1-4009-826e-5e293f1e5201 · outbound

This paper cites Longformer: The Long-Document Transformer.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Longformer: The Long-Document Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.881703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.881703Z digest=sha256:7b11f20155f09723f4f95b73fccc02e858ec033365238575613a8e42fa9856f9

Observation 2fd49c31-be75-4af4-84ec-3e9a8d5c17d2 · outbound

This paper cites Visual Transformers: Token-based Image Representation and Processing for Computer Vision.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.887844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.887844Z digest=sha256:9257140cf950e2619211dda3c5cd1e2f9c4f8a4db988283ffd5a240bf10816f1

Observation e36b2296-2cc4-4840-883f-2255bf452579 · outbound

This paper cites A primer in bertology: What we know about how bert works,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A primer in bertology: What we know about how bert works,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.349562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.894087Z digest=sha256:14fbb4e1def3cb3b3b32f1c262dea869377910fe730840ff25183338fe4d636d

Observation 548d5047-ec8e-48f8-9686-4e80da3b10bc · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.328333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.899056Z digest=sha256:0f3938510b60047e6c56f9f97fde5fa9b2a6561b0ae8ebc768de55555bfaddfd

Observation 9696bcfe-f93b-4c3c-aff7-4a3c73dba5d9 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.309219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.904139Z digest=sha256:dc037c75d23757372c0ae474cc152d52ba32b3dd5832eba111d316654eb6cb2c

Observation 3609f568-0b54-49d9-94c9-774cda0cecc7 · outbound

This paper cites Auto-encoding variational bayes,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Auto-encoding variational bayes,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.289980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.910780Z digest=sha256:93c0a9d27d0088c5b486d6f48ac27557ea28554eeec921d7a7c061aa26a6d0f8

Observation c8b2cd50-66c0-47ab-a728-8d61aeeca6a8 · outbound

This paper cites Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.267587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.915496Z digest=sha256:54c3b32cd9a95a35a5c1dfc5b6b2ac29db526fa0f23e4669c7336ce53c2840e7

Observation 53c937e7-2558-431c-a72b-ec0c375c9853 · outbound

This paper cites Attention is all you need,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Attention is all you need,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.247428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.920736Z digest=sha256:eca3a83f256471b26a9c3881a4271453600f5ae2d7ff53fef66c8b845b7bdee9

Observation f589483d-954d-4cfd-a5eb-b4ef49dc5c17 · outbound

This paper cites Deep residual learning for image recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Deep residual learning for image recognition,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.925634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.925634Z digest=sha256:93c8c962ac1896361d58f923b184e868ec5d4ce4e52d7746fe7183f1de9226e9

Observation c679d900-ee18-442a-b05f-a707877b0d08 · outbound

This paper cites Conditional random fields: Probabilistic models for segmenting and labeling sequence data,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Conditional random fields: Probabilistic models for segmenting and labeling sequence data,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.215489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.931019Z digest=sha256:e815426a933736b52d2b8f3374dbcdd9a8171b488ab3e97a6a2bdb59f295644b

Observation 1a07ebe0-1223-4048-bec1-804043acef90 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.935726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.935726Z digest=sha256:a2119957cc4ac103f7eecf94877e88882335632fd008a88f19a7478ab39d9bad

Observation 327576fc-5332-447d-a2c5-cdbbc9b28a7f · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.940778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.940778Z digest=sha256:47b21364a1a469e79802a06a6f49770915c7c5aacf2c74959f35690be50e08ac

Observation dc33cd84-8a54-47fb-be2e-8c7e25261b2f · outbound

This paper cites Vivit: A video vision transformer,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vivit: A video vision transformer,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.945512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.945512Z digest=sha256:6e57ce39bfd41cad954edf4257ab0cc8036a5132aa37250cb6977f4ec0f65d53

Observation f4c4426e-de7d-41ac-966a-d60a7f21d47e · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.950362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.950362Z digest=sha256:8ba4d680722a15986f85889d479e0748925f326fa5ef199f4aa073ab58acd7ff

Observation e0ea64b8-955c-4ba8-972f-e48df9251217 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.955213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.955213Z digest=sha256:dbdd47994f0716d6c0130c810a9d5295551889381312ff04a7c15d0207101679

Observation dd061f07-0ff1-49b2-981b-b3c6017764f1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.132655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.960826Z digest=sha256:0a9de003cc496871be0eec74aea5bfd77f83e882eafc3fbd22321102afeb7adf

Observation 060b76a4-32cf-41e9-a904-9eddfe4745bb · outbound

This paper cites Parallel data helps neural entity coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Parallel data helps neural entity coreference resolution,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.115435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.965982Z digest=sha256:38b12fc09b165a4c0669ce4ba3ce1bb40820f5fd7c81560abd8f40af6337cb1d

Observation a6c55131-0d59-431e-ba0e-6846f88eba82 · outbound

This paper cites Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.095439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.971683Z digest=sha256:2b63e409f7a460900c54584fb126f9dfbdca4c347ae56418bd2a98536835932a

Observation acb1842e-afd2-49d1-9f2f-3db8a1d6502e · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A fast and accurate one-stage approach to visual grounding,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.976780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.976780Z digest=sha256:0b6f62977a7db303b2956c1f243fe65233330a651f02257eabb386927cc177dd

Pith citing papers

No inbound Pith citation observations are available.