Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:31:57.305121Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2606.17950.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:31:57.305121Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6aec185-b25b-4f91-bcc8-e04bc2677e3f · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model A brief survey on recent advances in coreference resolution,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85df2914-b136-4958-9ced-fa899663cd65 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Spanbert: Improving pre-training by representing and predicting spans,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31a84b2-688e-47ea-b4ea-662f16acc808 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Coreference resolution without span representations,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850387c3-6d29-4d33-bfb8-7b8ab21b885b · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Image-based storytelling using deep learning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4396c91-40ec-4091-adef-3d11a698491f · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Pixels to Prose: Understanding the art of Image Captioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eff252fb-d3a3-4f73-8611-6927cd49c3a8 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Multi-modal self- perception enhanced large language model for 3d region-of-interest captioning with limited data,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946e652e-70b1-46cb-8d1a-ab71707412e1 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Video storytelling: Textual summaries for events,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555f3015-43d7-4709-b37a-37d812054670 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model What are you talking about? text-to-image coreference,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2807f6fa-1704-426f-a6e0-ea071d08f41c · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Who’s waldo? linking people across text and images,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a15f9a-450f-4656-801c-27edd7bf10da · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Phrase decoupling cross-modal hierarchical matching and progressive position correction for visual grounding,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9bbb425-810e-4f35-b1be-6a430b9c7c24 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Gravl-bert: Graphical visual-linguistic representations for multimodal coreference resolution,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e98a0a0-0712-4d37-ad1b-f0f5c0cc5040 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Reclip: A strong zero-shot baseline for referring ex- pression comprehension,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5c131b2-2ba4-4571-99fc-b33879a2f4ad · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model A dual reinforcement learning framework for weakly supervised phrase grounding,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a27f7227-b67a-48be-b9df-8caccb5b8f4c · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1e16611c-1532-4682-8064-eb4fcd494130 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Who are you referring to? coreference resolution in image narrations,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b744ce-c4bc-45c9-b468-ec84a428527c · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Semi-supervised multimodal coreference resolution in image narrations,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509bbd06-83e0-482d-ad73-89b6aeb0f725 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Self-adaptive fine-grained multi-modal data augmentation for semi- supervised multi-modal coreference resolution,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7ff081-24b2-49b3-a784-00923daa2c51 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Connecting vision and language with localized narratives,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4e2554-510b-40e0-8ae4-400add9bc2d8 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Revisiting multi-modal llm evaluation,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867e7710-eb73-4689-8348-a626bab1c6c9 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Knowledge en- hanced vision and language model for multi-modal fake news detection,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56c5aacc-e456-4b35-b0ee-8e77c2f4ad4f · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Learning transferable visual models from natural language supervision,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278aa20a-41bc-4d50-9e80-329fdbd20c73 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Combination of evidence in dempster-shafer theory,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3329ff-afd0-4993-82f1-b270371fdb29 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Jøsang,Subjective logic
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667bb99f-0cda-462e-b6c2-d0afa9091719 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0546ca64-1e77-49a3-8f75-ba137caf22c9 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Uniter: Universal image-text representation learning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38262897-86ce-46b6-9885-bcdadf8caa1d · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Vinvl: Revisiting visual representations in vision-language models,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04150a8a-f231-4ead-8e04-b590f669c5be · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Zero-shot referring expression comprehension via structural similarity between images and captions,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f513a35c-8d74-4331-b43f-b26c316a00cf · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Models overview - anthropic,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623531fa-1347-47e2-96a1-04176796b4ae · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ace29cde-5c94-4311-a78c-25e43e5dc1ac · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Qwen2.5-VL Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 629cdee8-bbec-4051-9148-4e44353cf548 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6b19fbf8-9d5d-418d-b852-c6dd7605ee67 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Gpt-4 in a cancer center—institute-wide deployment challenges and lessons learned,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb664dc-0067-47d6-9dfc-1fc49d468276 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model 3ur-llm: An end- to-end multimodal large language model for 3d scene understanding,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffef104-def1-4c78-8eaf-baa8c44082ed · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Hico: A benchmark for recognizing human-object interactions in images,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c1bef65-9301-4302-b3c6-0905ea8db195 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Grounded situation recognition,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12077cd-102e-4813-b29f-e1f7e7684bd6 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Visual genome: Connecting language and vision using crowdsourced dense image annotations,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359709fc-ac16-4a5a-99bd-7bb25953baf5 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761881ae-d35e-4327-86b5-74ecc1ad6af2 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Classification-then- grounding: Reformulating video scene graphs as temporal bipartite graphs,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf512b9b-a2be-4660-b2b7-2fd03e307c04 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Llm meets scene graph: Can large language models understand and generate scene graphs? a benchmark and empirical study,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfa099e-bf33-4c02-ad0a-b1a195d30443 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Information extraction,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad7fd5a-5d45-4ac9-9ca9-b470c766a780 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Open Information Extraction: A Review of Baseline Techniques, Approaches, and Applications
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d02f59e3-f637-4adf-bccf-9b2b47ff3c94 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Faster r-cnn: Towards real-time object detection with region proposal networks,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ef7bb9-c672-4f0e-9c8b-610710db22f3 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Generalization of dempster–shafer theory: A complex mass function,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294ef122-8948-4746-bb81-e5131748bd7b · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Trusted multi-view classi- fication with dynamic evidential fusion,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a2ecf6-785e-4c8d-93df-da99bba02c73 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e05d12df-0c03-4eb4-b27f-40ccdf1d9543 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Bridge the modality and capability gaps in vision-language model selection,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf64065c-2d07-49e6-a035-3f61b046533e · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Towards under- standing the modality gap in clip,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4649ff1e-7a89-4042-bc86-485c1198a877 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Confidence- aware contrastive learning for selective classification,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94524f7f-cad4-40ec-80e1-a53a97a51048 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model On calibration of modern neural networks,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3985df82-0674-4456-b182-22880d3916ab · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Overview of results of the muc-6 evaluation,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94687791-a189-472a-b2f6-2513d944448a · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Algorithms for scoring coreference chains,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a545e5a8-468d-45f1-9869-4b80637f748a · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model On coreference resolution performance metrics,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540d48fd-6d2e-4879-9c2a-824318456aa7 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Conll- 2012 shared task: Modeling multilingual unrestricted coreference in ontonotes,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49a025f-ec5c-4fff-aede-4a4c4994262e · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Stanford’s multi-pass sieve coreference resolution system at the conll-2011 shared task,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4557615b-2d90-4143-8ef9-3ec372b114e0 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model End-to-end neural coreference resolution,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71694848-dfec-4640-80e3-71e6e6c82ae3 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model On gen- eralization in coreference resolution,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa34f6ed-e795-4aa1-9565-e073481cc37b · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Qwen2.5 Technical Report
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2cbf12cc-a88f-4bf8-b086-e8d05fac0e37 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Maf: Multimodal alignment framework for weakly-supervised phrase grounding,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76493984-ff25-4a23-8a76-0c0f91d1fccc · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Are language models robust coreference resolvers?
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3bc6d5b-6506-4b50-9d55-8b5cf8bddcad · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Open information extraction via chunks,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580a7e07-7997-4065-9078-2ab778e90217 · outbound
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model From recognition to cogni- tion: Visual commonsense reasoning,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.