Pith. sign in

Paper Citation Record · LEDGER

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2606.17950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17950 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:31:57.305121Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6aec185-b25b-4f91-bcc8-e04bc2677e3f · outbound

This paper cites A brief survey on recent advances in coreference resolution,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model A brief survey on recent advances in coreference resolution,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:a7575be25ef0113ee1930217292786355703ab6f4869a68030446ffabae5b048

Observation 85df2914-b136-4958-9ced-fa899663cd65 · outbound

This paper cites Spanbert: Improving pre-training by representing and predicting spans,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Spanbert: Improving pre-training by representing and predicting spans,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:d01304580bfa9a13286c18d0f6f79c8815197cf681509b99eb17bcb12897f4ec

Observation a31a84b2-688e-47ea-b4ea-662f16acc808 · outbound

This paper cites Coreference resolution without span representations,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Coreference resolution without span representations,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:e47381aa7670ff5353f2417095266970637c77386d96523e8776f79005971519

Observation 850387c3-6d29-4d33-bfb8-7b8ab21b885b · outbound

This paper cites Image-based storytelling using deep learning,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Image-based storytelling using deep learning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:7b9f138a1316a120f48d5d82c100b4a698c905ff5fdf3df2e7fdb4a7890e0bff

Observation d4396c91-40ec-4091-adef-3d11a698491f · outbound

This paper cites Pixels to Prose: Understanding the art of Image Captioning.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Pixels to Prose: Understanding the art of Image Captioning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.183301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:482170476ab1737da76ee539ae3bf1e31c9c5454cdca437638834bdd854e9f86

Observation eff252fb-d3a3-4f73-8611-6927cd49c3a8 · outbound

This paper cites Multi-modal self- perception enhanced large language model for 3d region-of-interest captioning with limited data,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Multi-modal self- perception enhanced large language model for 3d region-of-interest captioning with limited data,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:31ac6b893b57a2bde6386e0e89f452b83af208d88c5f01c61816824231044562

Observation 946e652e-70b1-46cb-8d1a-ab71707412e1 · outbound

This paper cites Video storytelling: Textual summaries for events,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Video storytelling: Textual summaries for events,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:622035925c6ae68d6c8b3359b5c2968db877600cd685cf34e860b3757e45d17e

Observation 555f3015-43d7-4709-b37a-37d812054670 · outbound

This paper cites What are you talking about? text-to-image coreference,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model What are you talking about? text-to-image coreference,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:7d886f84b2f022d14175fad06d115f7a3242fd2311913fa505078b9ae7e99500

Observation 2807f6fa-1704-426f-a6e0-ea071d08f41c · outbound

This paper cites Who’s waldo? linking people across text and images,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Who’s waldo? linking people across text and images,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:e649746ddca715a83a51db12f053956d5fc352f16258766002a6b5c8f4afacd4

Observation 17a15f9a-450f-4656-801c-27edd7bf10da · outbound

This paper cites Phrase decoupling cross-modal hierarchical matching and progressive position correction for visual grounding,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Phrase decoupling cross-modal hierarchical matching and progressive position correction for visual grounding,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:c68b0bbe26c10a18f2f0d7fc5dbb985b8e352159d8153589c6da5f93c1d98de5

Observation a9bbb425-810e-4f35-b1be-6a430b9c7c24 · outbound

This paper cites Gravl-bert: Graphical visual-linguistic representations for multimodal coreference resolution,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Gravl-bert: Graphical visual-linguistic representations for multimodal coreference resolution,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:d6201642d25e93de220a1e524adb08a3675f44f60bfc6c2f7188aa7fc165cab4

Observation 8e98a0a0-0712-4d37-ad1b-f0f5c0cc5040 · outbound

This paper cites Reclip: A strong zero-shot baseline for referring ex- pression comprehension,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Reclip: A strong zero-shot baseline for referring ex- pression comprehension,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:4318012818dceaf6210dfb4833003737e5096da5af609bae4339b62bca6f3f16

Observation c5c131b2-2ba4-4571-99fc-b33879a2f4ad · outbound

This paper cites A dual reinforcement learning framework for weakly supervised phrase grounding,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model A dual reinforcement learning framework for weakly supervised phrase grounding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:e9dfd26feb206b019cf23a9ae314faa634b7ce1274d7bcf1b0f0efde112e7cc7

Observation a27f7227-b67a-48be-b9df-8caccb5b8f4c · outbound

This paper cites SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.188183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:1c96f36e04d083327ee853b18184bb35132808bb670e4f12fe53398d8f1559b6

Observation 1e16611c-1532-4682-8064-eb4fcd494130 · outbound

This paper cites Who are you referring to? coreference resolution in image narrations,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Who are you referring to? coreference resolution in image narrations,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:f076787c44d5f59d8e3184184bc6fba375341327b56aa9fbcf90045667c9fe50

Observation 09b744ce-c4bc-45c9-b468-ec84a428527c · outbound

This paper cites Semi-supervised multimodal coreference resolution in image narrations,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Semi-supervised multimodal coreference resolution in image narrations,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:4da6c63d3fc1471a7e774cef047874ed91409738f5bafbbd50501ed6974813a0

Observation 509bbd06-83e0-482d-ad73-89b6aeb0f725 · outbound

This paper cites Self-adaptive fine-grained multi-modal data augmentation for semi- supervised multi-modal coreference resolution,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Self-adaptive fine-grained multi-modal data augmentation for semi- supervised multi-modal coreference resolution,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:14a0e3e622127bfc0ab28486e58527353e3e8a0a2fe27a24d2ac37aa73d72162

Observation 6b7ff081-24b2-49b3-a784-00923daa2c51 · outbound

This paper cites Connecting vision and language with localized narratives,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Connecting vision and language with localized narratives,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:be53a84a5fb0e42025f9bc94f3344fcde44c3c719f46d7c35da293735ee6cb01

Observation 1f4e2554-510b-40e0-8ae4-400add9bc2d8 · outbound

This paper cites Revisiting multi-modal llm evaluation,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Revisiting multi-modal llm evaluation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:8ecd1963ee86ace02f9172b2d445afcd3297dabac8e5fe57752fd02d17e8ed28

Observation 867e7710-eb73-4689-8348-a626bab1c6c9 · outbound

This paper cites Knowledge en- hanced vision and language model for multi-modal fake news detection,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Knowledge en- hanced vision and language model for multi-modal fake news detection,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:68aad8bb02143619c66898bb911fdf5519d4bd35367130418216b14d7a8c14c3

Observation 56c5aacc-e456-4b35-b0ee-8e77c2f4ad4f · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Learning transferable visual models from natural language supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:56e00639c58e2021c3e3a2f4d8c9afaeea3e4abe2bd918e52f8078b866e570e1

Observation 278aa20a-41bc-4d50-9e80-329fdbd20c73 · outbound

This paper cites Combination of evidence in dempster-shafer theory,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Combination of evidence in dempster-shafer theory,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:f06249ed233b84b2520b153724eede1fc3938ae920437d3d4f463084083e85a4

Observation 5d3329ff-afd0-4993-82f1-b270371fdb29 · outbound

This paper cites Jøsang,Subjective logic.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Jøsang,Subjective logic

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:01be8fae7312220e08ee69c327a64104a583ce10a66e5841f75e2dafcff1b14b

Observation 667bb99f-0cda-462e-b6c2-d0afa9091719 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.173941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:cd785274970abc9655b403e3dae0a3b4394568a4ea5112bb96521a61f4de1a1d

Observation 0546ca64-1e77-49a3-8f75-ba137caf22c9 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Uniter: Universal image-text representation learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:6e6080bfca6b5fc55fdb0a63593556b8cc730e69d7ffa624a3bc1623a12c3f36

Observation 38262897-86ce-46b6-9885-bcdadf8caa1d · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Vinvl: Revisiting visual representations in vision-language models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:475d3ea0fd6b7edca5460f978c6266046a5f2a3a19650c052cf741e75538ae05

Observation 04150a8a-f231-4ead-8e04-b590f669c5be · outbound

This paper cites Zero-shot referring expression comprehension via structural similarity between images and captions,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Zero-shot referring expression comprehension via structural similarity between images and captions,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:835adf2256d05dacb20cc80405cf83548cc887fe3386ae496af61f8f55a80795

Observation f513a35c-8d74-4331-b43f-b26c316a00cf · outbound

This paper cites Models overview - anthropic,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Models overview - anthropic,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:bbc14038fe5ff16cd931e551d146b8da4b3fb91d59a723597ec7806e119d7268

Observation 623531fa-1347-47e2-96a1-04176796b4ae · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.170496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:7637399d19e3a41d3ebb7eaa255dd4e20b504ca0a35f5c9b83c58909c5bb00af

Observation ace29cde-5c94-4311-a78c-25e43e5dc1ac · outbound

This paper cites Qwen2.5-VL Technical Report.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Qwen2.5-VL Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.178490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:b277696bb364ef6c2f677c1b9912ece065cb8963993a9f768dad8f3a314cc3a5

Observation 629cdee8-bbec-4051-9148-4e44353cf548 · outbound

This paper cites Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.192887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:bc9ef5d812af4f5213c1f0fbade95bf029899cc3b41122366b7ae3b49b1b0f24

Observation 6b19fbf8-9d5d-418d-b852-c6dd7605ee67 · outbound

This paper cites Gpt-4 in a cancer center—institute-wide deployment challenges and lessons learned,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Gpt-4 in a cancer center—institute-wide deployment challenges and lessons learned,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:bb3465cd653780ec1bf91f0e4ab354c2f40a4d228711971ea9d3ea5a4ddc74aa

Observation adb664dc-0067-47d6-9dfc-1fc49d468276 · outbound

This paper cites 3ur-llm: An end- to-end multimodal large language model for 3d scene understanding,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model 3ur-llm: An end- to-end multimodal large language model for 3d scene understanding,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:4479dffb0a6553f3a9e416cc1276e9ee41c81ad9ee12e2256e674288beda01da

Observation 8ffef104-def1-4c78-8eaf-baa8c44082ed · outbound

This paper cites Hico: A benchmark for recognizing human-object interactions in images,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Hico: A benchmark for recognizing human-object interactions in images,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:f694cc9b399fc697f4588c0b1e11dd699511b7d2821d56f5088b64f591362b71

Observation 0c1bef65-9301-4302-b3c6-0905ea8db195 · outbound

This paper cites Grounded situation recognition,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Grounded situation recognition,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:54efd63e9e393b83a306efba690cb181df2010313c30408c696e1cab631f7218

Observation e12077cd-102e-4813-b29f-e1f7e7684bd6 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:33dd9a40f17ef2ac64558b669009ed0693f8d6fdd21bfe36cf68146816e90804

Observation 359709fc-ac16-4a5a-99bd-7bb25953baf5 · outbound

This paper cites an unresolved cited work.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:5b83c4f48c67525a9ed5d810ae5ea7e44c0029e15277ece4513bc8a0ff04835d

Observation 761881ae-d35e-4327-86b5-74ecc1ad6af2 · outbound

This paper cites Classification-then- grounding: Reformulating video scene graphs as temporal bipartite graphs,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Classification-then- grounding: Reformulating video scene graphs as temporal bipartite graphs,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:25cebe545d1c1ca9df3647fca0f1f7e6c104e8dbd5e816abc3db3cd05f89e149

Observation bf512b9b-a2be-4660-b2b7-2fd03e307c04 · outbound

This paper cites Llm meets scene graph: Can large language models understand and generate scene graphs? a benchmark and empirical study,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Llm meets scene graph: Can large language models understand and generate scene graphs? a benchmark and empirical study,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:d61608f3c4f5a80672dae49e016ed1db83ea0e03e1e52a69136e02650c1b17e5

Observation ddfa099e-bf33-4c02-ad0a-b1a195d30443 · outbound

This paper cites Information extraction,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Information extraction,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:20da0cb6a9473dbac9981ef4bae7130b431552ae998852a55b4671450c06e61c

Observation 5ad7fd5a-5d45-4ac9-9ca9-b470c766a780 · outbound

This paper cites Open Information Extraction: A Review of Baseline Techniques, Approaches, and Applications.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Open Information Extraction: A Review of Baseline Techniques, Approaches, and Applications

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.200198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:f2d5a35602623b0260a2f50d188926917eaad92820f05f1f2d1a07e079efab2a

Observation d02f59e3-f637-4adf-bccf-9b2b47ff3c94 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Faster r-cnn: Towards real-time object detection with region proposal networks,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:6337152e946db2849762e18d6ab69fe9f36d037dcde050eaa66409518d628737

Observation b4ef7bb9-c672-4f0e-9c8b-610710db22f3 · outbound

This paper cites Generalization of dempster–shafer theory: A complex mass function,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Generalization of dempster–shafer theory: A complex mass function,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:bee1b8bcf361f729bc73d6e71f8b573ca2fc05839a6f1ab1027f984782cea07c

Observation 294ef122-8948-4746-bb81-e5131748bd7b · outbound

This paper cites Trusted multi-view classi- fication with dynamic evidential fusion,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Trusted multi-view classi- fication with dynamic evidential fusion,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:e4972fd1b0ca1239d24ce2cdea1a45a036a48adb128b46c3ca9ca6aa1109bcbc

Observation 05a2ecf6-785e-4c8d-93df-da99bba02c73 · outbound

This paper cites Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.196361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:e9b6252147bb6583519a84e102e1e0617c128f41fa22acc28a231cc478058628

Observation e05d12df-0c03-4eb4-b27f-40ccdf1d9543 · outbound

This paper cites Bridge the modality and capability gaps in vision-language model selection,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Bridge the modality and capability gaps in vision-language model selection,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:e36d914fb910be563f1bee3b8e6309601729c95eeb449c0c88492f8b49de163c

Observation cf64065c-2d07-49e6-a035-3f61b046533e · outbound

This paper cites Towards under- standing the modality gap in clip,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Towards under- standing the modality gap in clip,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:93f7d4ad6c53c8ee392d38515660bd0d25855f036fdcfed16b7cc0a20f8ba070

Observation 4649ff1e-7a89-4042-bc86-485c1198a877 · outbound

This paper cites Confidence- aware contrastive learning for selective classification,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Confidence- aware contrastive learning for selective classification,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:a568c620ad80a5ab851ad6674f026bb7583cdf43b5ef2eef9f00645e8d2f4b80

Observation 94524f7f-cad4-40ec-80e1-a53a97a51048 · outbound

This paper cites On calibration of modern neural networks,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model On calibration of modern neural networks,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:fc29127dfce359217708122d9c91ec4906ebc8e75df5d22c3d73a864f1881145

Observation 3985df82-0674-4456-b182-22880d3916ab · outbound

This paper cites Overview of results of the muc-6 evaluation,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Overview of results of the muc-6 evaluation,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:32fe6f9768b19e669625ea4d751742c0ae660b1e4950fc5c8728e413f8f1afe2

Observation 94687791-a189-472a-b2f6-2513d944448a · outbound

This paper cites Algorithms for scoring coreference chains,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Algorithms for scoring coreference chains,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:f48424cefd4013d6d091f6a76f62a3e934f482da61d75b1babe31006835f5048

Observation a545e5a8-468d-45f1-9869-4b80637f748a · outbound

This paper cites On coreference resolution performance metrics,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model On coreference resolution performance metrics,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:39b825666271fb8b98928d7add37e0068bb45907a49df39a05dd5ed77809d4ea

Observation 540d48fd-6d2e-4879-9c2a-824318456aa7 · outbound

This paper cites Conll- 2012 shared task: Modeling multilingual unrestricted coreference in ontonotes,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Conll- 2012 shared task: Modeling multilingual unrestricted coreference in ontonotes,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:fae2b9b8f43dcd2336b143a8225d38c94607dd708a4eec5b0856bffa076ffb34

Observation d49a025f-ec5c-4fff-aede-4a4c4994262e · outbound

This paper cites Stanford’s multi-pass sieve coreference resolution system at the conll-2011 shared task,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Stanford’s multi-pass sieve coreference resolution system at the conll-2011 shared task,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:b2513c68296d83e8bfa1c2a22d2d675bffdd572e17e627984665640677fdaed0

Observation 4557615b-2d90-4143-8ef9-3ec372b114e0 · outbound

This paper cites End-to-end neural coreference resolution,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model End-to-end neural coreference resolution,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:d781e7e7ef1779a400e3acef01c56611c5326f80839bd812534f7098d0e3d50e

Observation 71694848-dfec-4640-80e3-71e6e6c82ae3 · outbound

This paper cites On gen- eralization in coreference resolution,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model On gen- eralization in coreference resolution,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:b887481716d9d1c14962134a1a5323c80a3af2b5eb35a14a8648b6a1df95ed1c

Observation aa34f6ed-e795-4aa1-9565-e073481cc37b · outbound

This paper cites Qwen2.5 Technical Report.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Qwen2.5 Technical Report

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.167500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:ef3677d8a88a51ae2df362addcc184173b98ff9b48aebff6568403f3c8884237

Observation 2cbf12cc-a88f-4bf8-b086-e8d05fac0e37 · outbound

This paper cites Maf: Multimodal alignment framework for weakly-supervised phrase grounding,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Maf: Multimodal alignment framework for weakly-supervised phrase grounding,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:c83d9ebe7b73c00c8276194c6f5a61fc6d121ad8df22e7eae10aa696086d9c90

Observation 76493984-ff25-4a23-8a76-0c0f91d1fccc · outbound

This paper cites Are language models robust coreference resolvers?.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Are language models robust coreference resolvers?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:9af57605f18ab197281237e46a4f7be16aed93ecdb1e47ed51dba3c58e6d006f

Observation f3bc6d5b-6506-4b50-9d55-8b5cf8bddcad · outbound

This paper cites Open information extraction via chunks,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Open information extraction via chunks,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:4216b4e142bc24367f5527883fccfb7a9ad5dd34dc64b4d0be87c051b48beba2

Observation 580a7e07-7997-4065-9078-2ab778e90217 · outbound

This paper cites From recognition to cogni- tion: Visual commonsense reasoning,.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model From recognition to cogni- tion: Visual commonsense reasoning,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T01:31:57.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:703ba49d2fc8d40866c1fff17f8954739b2e52014c97e5854c31b59a3334ecdb

Pith citing papers

No inbound Pith citation observations are available.