Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:08:48.226962Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 0 inbound Pith citation observations for arXiv:2506.11515.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:08:48.226962Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 143 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b787420d-6acc-4f61-a52a-79299a9128bc · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55f9426a-ef5c-4aa8-b023-99165f154c40 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Making the V in VQA matter: Elevating the role of image understanding in visual question answering,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc1cac8-cde5-463f-b1a8-406c2422b85d · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a06e115-c7a9-4c85-ab83-083584a3d28c · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A corpus for reasoning about natural language grounded in photographs,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e00204-edb0-4c9f-8b1e-48ec10634540 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4deea2d7-3081-4629-977c-7375b2fc0ce8 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An empirical study of training end-to-end vision-and-language transformers,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b114773-63c4-4e72-bb35-62f12f64839e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Bridgetower: Building bridges between encoders in vision-language representation learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d71738-f08d-4cd5-9398-f9bf73b4f60b · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning transferable visual models from natural language supervi- sion,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a635d12-cd8e-4d0e-a1e3-a20123271c14 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1497da84-d879-4b24-a45b-d1310f0bc138 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning deep transformer models for machine translation,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8768f4-df16-48b4-ae99-fd24685256d8 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f631e1-3a02-48f3-ad3c-0edf39ec72bb · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava- next: Improved reasoning, ocr, and world knowledge,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f78c9a6-e55e-4637-a4a3-8d5d209ae51f · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs How much can CLIP benefit vision-and-language tasks?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe89ca05-e8d3-472f-ae87-bf3686995f4a · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO-2: End-to-end unified vision-language grounded learning,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e54d8e-4a8a-4f7b-9af6-5cad87292bb6 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation of rare words with subword units,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c02c04b3-7ca1-4bca-b887-8375e818dd1f · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are unsupervised multitask learners,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db02e517-fb23-4300-a9be-3426bf0d409d · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention is all you need,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa3a712-7ab3-44c4-a016-f21eb83b3039 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127b0b7e-b728-4454-9243-9d426aeea31c · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multi-layer representation fusion for neural machine translation,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e228351-5890-459e-a8cb-e664beffcaa7 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multiscale collaborative deep models for neural machine translation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168de691-262c-4021-983a-4525d727306a · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Layer Normalization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f72292-993e-42a2-8b69-78960073aa47 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6364764-2666-4314-a977-3c149ba81cd8 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Decoupled weight decay regularization,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13501472-651d-4356-b2bc-8d4511f2f927 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilt: Vision-and-language transformer without convolution or region supervision,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d542ec-e519-46e3-a9d0-bbe5db76fa07 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Uniter: Universal image-text representation learning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4559d4da-8f84-45c6-b95c-2d17fde19ed3 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa74da3-3b4e-460c-bd91-263168634441 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Align before fuse: Vision and language representation learning with momentum distillation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ec964d-4a3b-43ff-9eb0-ac8e6074e064 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1b9acf-fcc9-438f-9f15-94e42948a820 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Simvlm: Simple visual language model pretraining with weak supervision,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d31b259-9a5c-4d4e-97d5-8f750e5c074a · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f9cc6f-3c60-48a9-8c62-b536775451df · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0fd178-91fb-40f4-8f7b-13bfa4a7129b · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Im2text: Describing images using 1 million captioned photographs,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c6ed4c-a231-4626-9caf-a5e85204bc85 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43234c42-5705-49f9-bb51-c0a6d5b9edd7 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2559ceb1-fa12-4997-bc03-069586b28428 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Improved baselines with visual instruction tuning,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f49981-e54c-4b78-a46e-cc375d64580f · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Monkey: Image resolution and text label are important things for large multi-modal models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ecea55-9c8f-4520-b022-d83c20fe3c04 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Docvqa: A dataset for vqa on document images,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537332fd-f14c-4b01-bebc-63fb4571fe5e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ed166f-4887-4415-ab39-a07af0c8721d · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Sigmoid loss for language image pre-training,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a4699c-f096-4e34-83f5-f5e54cef2ca5 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5ec591-45c1-41cb-b6d8-244f741b30bb · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring Plain Vision Transformer Backbones for Object Detection
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6490ac-a365-4e03-a021-585cda9746f2 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8158c46c-bef8-4e83-8312-21b3dc5bdd76 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OK-VQA: A visual question answering benchmark requiring external knowledge,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c7b5df-bf1e-4fdf-a001-c90e39924fb9 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs GQA: A new dataset for real-world visual reasoning and compositional question answering,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be1fb631-5e0e-400e-b7d2-6f4ae267f566 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs MM-vet: Evaluating large multimodal models for integrated capabilities,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6f0d9e-1324-4eec-b3ec-8182e7455207 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28546be3-8d16-4c27-8c88-47030076110e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Grok-1.5 vision preview
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db613073-00ff-4358-82df-5a1e2594cff4 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Towards VQA models that can read,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5ae75d-1c63-4d54-b5d4-335a3171eb90 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ChartQA: A benchmark for question answering about charts with visual and logical reasoning,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d760691-3514-4c2a-896d-d4a8a8034796 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Infographicvqa,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1350f9c-11ba-457c-9adc-0a6cd3b1d957 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A diagram is worth a dozen images,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9896961-3bb1-4c9c-bdeb-be0b8c47cb4f · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e7b22c-7d89-46c0-956b-294c214ac865 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f75a396-e90c-4cb4-90c5-f2a010df1e03 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791fa363-476e-4220-b480-1297520c747b · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava-next: What else influences visual instruction tuning beyond data?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0053e0-2ad7-40d6-9e37-341dea810c36 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfd809d-005b-4e6b-9e68-8651e9771796 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual instruction tuning,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630611ec-21aa-4254-a63b-4364f039a30e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a46b53b-59b1-4e10-a856-f00b73567ab5 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a13bfda-229f-4e88-b2f1-ff9484da7785 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation by jointly learning to align and translate,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ea5f86-41b0-4a53-9cb8-04e90e66d748 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Revealing the dark secrets of masked image modeling,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e4abef-6a5a-42c5-a7d4-42205f4297a5 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs On information and sufficiency,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11507107-d0ac-4d9d-be1d-a4aa00ff6b15 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs VL-BERT: pre-training of generic visual-linguistic representations,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21d2ddb-66ff-4d8d-a3b1-8f6fe4d4b2db · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9185042d-2240-4ff6-bab4-27b2067d0c14 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oscar: Object-semantics aligned pre-training for vision-language tasks,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e88f0c-1d92-4bd9-9257-708632af24b9 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4663b1-9867-4080-86a5-7bb1fd1cc08e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b21a005-6ee6-4ecf-9f2e-dfa787d4f2f0 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860a30ee-b50a-4881-90a6-4fac86b0df49 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring vision-language foundation model for novel object captioning,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2e4955-30c7-4ba3-b795-3ef241e06712 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unsupervised domain adaption harnessing vision-language pre-training,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9850c635-261e-4a42-996e-4a495732adf6 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Do vision transformers see like convolutional neural networks?
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e491e1b-f238-4429-bc18-38adbcdc299f · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Intriguing properties of vision transformers,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20dc7104-e1f1-4a72-ade2-e2ddc88cbc52 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dissecting contextual word embeddings: Architecture and representation,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0937db5d-0aa4-47ad-872d-540f20f90442 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Linguistic knowledge and transferability of contextual representa- tions,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d159b50-7879-4346-a4b9-4b3d6b6049ed · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs What does BERT learn about the structure of language?
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a40f94-b25b-4c85-be23-f6d1a752f0cd · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BERT: Pre- training of deep bidirectional transformers for language understanding,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ead94a-b173-4027-b219-aa7e4f40108e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Feature pyramid networks for object detection,
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5491a66c-cb13-49a4-befe-b504ac064462 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Densely connected convolutional networks,
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddf0667-663a-41c7-8681-a534e57566c2 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep layer aggregation,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81b656ed-2771-43e3-85fb-4b7858b5b7c5 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Segformer: Simple and efficient design for semantic segmentation with transformers,
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87038d9d-e3d1-497c-818f-f7eee285c3da · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Clsr: Cross-layer interaction pyramid super-resolution network,
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8239b6d0-30e7-47e8-acc0-fd8a21ec2a71 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention-based layer fusion and token masking for weakly supervised semantic segmentation,
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0638c7e-7665-4b3a-bd03-f89a1ae50506 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Artificial- spiking hierarchical networks for vision-language representation learn- ing,
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1667396-7ca9-447c-87c5-2d97a56a6271 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep contextualized word representations,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc0fba08-fdf4-4305-a1d7-a34552aad9cd · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Coarse-to-fine vision-language pre-training with fusion in the backbone,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ea500f7-b2be-4f3c-a87b-5624dbf3905b · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dense Connector for MLLMs
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a35e3cd-6c1f-489a-9ae9-bbf020cf6cb9 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are few-shot learners,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e4e945f-0795-4c44-a2fd-3049636038b5 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b98f0a0-33cd-4afc-b40b-2d8e3fbbd78e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Large Language Models Meet NLP: A Survey
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ec15ab-e7f7-4152-8e89-546a98610b0d · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f50a89ec-fb71-4975-87dd-9660679b2255 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Introducing our multimodal models,
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ffe0c48d-d30c-467d-9802-8039d28c4610 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed14fee-6450-4306-8821-431c65915b8a · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c93c76-ca12-46eb-8d29-e800355ccd76 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c172e8-53e3-4dc4-9254-b71a08b7a65e · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe259a9d-cf04-425f-93f8-ba3e61506701 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs When do we not need larger vision models?
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55ef876c-a123-4673-beb6-714e832d43db · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aac154c-821b-4d07-b583-e03cd288f6a1 · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Honeybee: Locality-enhanced projector for multimodal llm,
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0815b0f9-2125-4277-8e8e-d6211aecd15d · outbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unified language model pre-training for natural language understanding and generation,
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.