Pith. sign in

Paper Citation Record · LEDGER

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

As of 8 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 0 inbound Pith citation observations for arXiv:2506.11515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11515 v1

Coverage vector

measured 100 of 143 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:08:48.226962Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 143 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b787420d-6acc-4f61-a52a-79299a9128bc · outbound

This paper cites ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.096271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.096271Z digest=sha256:8e4773ad8db9bacb20153afd090c693b0d5b8dd90a5952c50c118f2c0e188004

Observation 55f9426a-ef5c-4aa8-b023-99165f154c40 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.202223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.202223Z digest=sha256:9baf0854402a6d851992fb86760a9e67a3ab3d132f507fecb842fde4faa76d62

Observation ebc1cac8-cde5-463f-b1a8-406c2422b85d · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.416327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.416327Z digest=sha256:37bb1082bb0e1011dfc8d9575fcf4650e40cd05d350a18bc742d39734bb993fd

Observation 3a06e115-c7a9-4c85-ab83-083584a3d28c · outbound

This paper cites A corpus for reasoning about natural language grounded in photographs,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A corpus for reasoning about natural language grounded in photographs,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.710099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.710099Z digest=sha256:180d61431503b03b66947ae22874667a0ebfe556fd95832946d2346b3e54bb39

Observation 62e00204-edb0-4c9f-8b1e-48ec10634540 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.832464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.832464Z digest=sha256:1dec854c0910bb3c4db62dfa9aa5e84b6320299c0f2d5eab0b11b0595505553f

Observation 4deea2d7-3081-4629-977c-7375b2fc0ce8 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An empirical study of training end-to-end vision-and-language transformers,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.933376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.933376Z digest=sha256:3165e2380de2f47f873cfa19e618d8378ace66449ea53a00b62f9c2e1e8e259a

Observation 6b114773-63c4-4e72-bb35-62f12f64839e · outbound

This paper cites Bridgetower: Building bridges between encoders in vision-language representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Bridgetower: Building bridges between encoders in vision-language representation learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.047268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.047268Z digest=sha256:a0c2ac67ec408860b43613538a202002ed4d5868fd7a529352bcd97292196e37

Observation f1d71738-f08d-4cd5-9398-f9bf73b4f60b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning transferable visual models from natural language supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.138970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.138970Z digest=sha256:443d157ad3ac11a04053757c0c431051046040aa8929b001c2125d59bace2367

Observation 5a635d12-cd8e-4d0e-a1e3-a20123271c14 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.269423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.269423Z digest=sha256:01b2b40f88644df1a9e7a764e73dfe98fabddc77cd7dd918105a3b1222ab58cb

Observation 1497da84-d879-4b24-a45b-d1310f0bc138 · outbound

This paper cites Learning deep transformer models for machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning deep transformer models for machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.396539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.396539Z digest=sha256:cdf30aee7ab66632e7fc7ba2ac172d82fce4fafa8e5c337b8e611ebab5f4233c

Observation 3f8768f4-df16-48b4-ae99-fd24685256d8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.510959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.510959Z digest=sha256:f292ec97048f43666596a5e36bfb1458341fa4f13a0d0f1e03f57b5567383429

Observation 17f631e1-3a02-48f3-ad3c-0edf39ec72bb · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowledge,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava- next: Improved reasoning, ocr, and world knowledge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.681245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.681245Z digest=sha256:fc7c0fdc86b83e9eb631418a54ab989bf3a07be36f07518b8ac30f502e5c5f61

Observation 2f78c9a6-e55e-4637-a4a3-8d5d209ae51f · outbound

This paper cites How much can CLIP benefit vision-and-language tasks?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs How much can CLIP benefit vision-and-language tasks?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.848451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.848451Z digest=sha256:47fb6b07c2bf20637f74e74f6f051a182ade7bc4c009dea29318f6c28da784fe

Observation fe89ca05-e8d3-472f-ae87-bf3686995f4a · outbound

This paper cites UNIMO-2: End-to-end unified vision-language grounded learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO-2: End-to-end unified vision-language grounded learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.012205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.012205Z digest=sha256:4ddb759c916a2875d215f07baa7970a2c094d7aeaf8cca763fcb860750c6b35c

Observation 83e54d8e-4a8a-4f7b-9af6-5cad87292bb6 · outbound

This paper cites Neural machine translation of rare words with subword units,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation of rare words with subword units,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.146927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.146927Z digest=sha256:afd86ee6ba507392a3311a8c8bbeb3a2294b94ab7b4a596fe8652c8ba24ee793

Observation c02c04b3-7ca1-4bca-b887-8375e818dd1f · outbound

This paper cites Language models are unsupervised multitask learners,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.256696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.256696Z digest=sha256:f3d53529d9010b757534f2f26ee9e6f5ea456a86bc4b7b7c670398f1b967f5d1

Observation db02e517-fb23-4300-a9be-3426bf0d409d · outbound

This paper cites Attention is all you need,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention is all you need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.432776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.432776Z digest=sha256:8b6c451c50e12e14d16698260eecd18e40849e5f6ceba651a42ff31e0aef17e0

Observation 2aa3a712-7ab3-44c4-a016-f21eb83b3039 · outbound

This paper cites Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.541595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.541595Z digest=sha256:5ac93bb5f8ce31b154ed935bd1403e75fe72214d6e20f6c60491f7708d267b7b

Observation 127b0b7e-b728-4454-9243-9d426aeea31c · outbound

This paper cites Multi-layer representation fusion for neural machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multi-layer representation fusion for neural machine translation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.664720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.664720Z digest=sha256:6678e3351d25764c9ade2a925b0206208c6fd883a5cb210ab4e067b1318dfb5c

Observation 6e228351-5890-459e-a8cb-e664beffcaa7 · outbound

This paper cites Multiscale collaborative deep models for neural machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multiscale collaborative deep models for neural machine translation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.819400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.819400Z digest=sha256:fa595348c47aaa86f43cdf82794967ceac6f466cb47fc965e6f38762d9cda63c

Observation 168de691-262c-4021-983a-4525d727306a · outbound

This paper cites Layer Normalization.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Layer Normalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.926027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.926027Z digest=sha256:512e2d4be5f052974cb938142a2494b819b00f94fdf639c383e45e744819bd1c

Observation d3f72292-993e-42a2-8b69-78960073aa47 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.069126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.069126Z digest=sha256:67eda83301e585a792fe6ab8f560fb7edfef0ef0eadaebc8d140592b2290844e

Observation a6364764-2666-4314-a977-3c149ba81cd8 · outbound

This paper cites Decoupled weight decay regularization,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Decoupled weight decay regularization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.245435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.245435Z digest=sha256:6f198f02c7df1196ae381ef0b3f953233312d7255f8d6fbf16db0c16aad3a2a6

Observation 13501472-651d-4356-b2bc-8d4511f2f927 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.380900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.380900Z digest=sha256:420b88dbfe02813a8e2b904584f96343be3b07404cb797b8af8f31a4c848ff1d

Observation a8d542ec-e519-46e3-a9d0-bbe5db76fa07 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Uniter: Universal image-text representation learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.538152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.538152Z digest=sha256:741d9442cdb7f75d86ad3d7cfaa0fa55f1b4b8dad76181ac17358f2d9f63737d

Observation 4559d4da-8f84-45c6-b95c-2d17fde19ed3 · outbound

This paper cites UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.679187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.679187Z digest=sha256:de35c32fe7b33fc19acc1f552343a9a186eaf676eda30389111a3bb5c73c580d

Observation 6aa74da3-3b4e-460c-bd91-263168634441 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Align before fuse: Vision and language representation learning with momentum distillation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.787717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.787717Z digest=sha256:fd93a6c15b39d36faa52ba809c53f1bb9773e0eb4a821af61ad556ca53765283

Observation 40ec964d-4a3b-43ff-9eb0-ac8e6074e064 · outbound

This paper cites Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.927145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.927145Z digest=sha256:23d15b1ab9e8e7f2535cee05546c38eadf3f3182ecb613b7e17de82bfa649e9d

Observation 0e1b9acf-fcc9-438f-9f15-94e42948a820 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Simvlm: Simple visual language model pretraining with weak supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.043824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.043824Z digest=sha256:8926db9d9272308d3045c194a38a3a37da079a12d7c307bc521bc2104ebd697b

Observation 9d31b259-9a5c-4d4e-97d5-8f750e5c074a · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.184911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.184911Z digest=sha256:cf1a3adbdf9e2a4686034a40f8f53f729419b52d089709be7647979ebe32be16

Observation 81f9cc6f-3c60-48a9-8c62-b536775451df · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.280478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.280478Z digest=sha256:830ba43ca2d718d1cfae46c2f760e69c30579e3045f83e45680a361b6d3ada31

Observation ab0fd178-91fb-40f4-8f7b-13bfa4a7129b · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Im2text: Describing images using 1 million captioned photographs,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.369147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.369147Z digest=sha256:97d32db20517e5cd76d47d6a9925d6efcfeb4f92680fbe79b1fe508d6d36c809

Observation 16c6ed4c-a231-4626-9caf-a5e85204bc85 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.503386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.503386Z digest=sha256:41893d15d5368c67bb6f34e8697f842930e925a605ca6906509c9c2548262475

Observation 43234c42-5705-49f9-bb51-c0a6d5b9edd7 · outbound

This paper cites Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.643209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.643209Z digest=sha256:9dd8a96790f5108176567c6c7ed646c8f02fabda0cfaca795bf00784f285081d

Observation 2559ceb1-fa12-4997-bc03-069586b28428 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Improved baselines with visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.753062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.753062Z digest=sha256:f173f9e5009ad0d328b8dcdc4104314f95c467c63d18d2ad5647b04a07fc56d8

Observation 67f49981-e54c-4b78-a46e-cc375d64580f · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Monkey: Image resolution and text label are important things for large multi-modal models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.870902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.870902Z digest=sha256:157e0454d34131e2d324775d453fb18d249b341f36e9356cb83986eff40c7425

Observation 54ecea55-9c8f-4520-b022-d83c20fe3c04 · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Docvqa: A dataset for vqa on document images,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.996940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.996940Z digest=sha256:194809aa677d48287ddeff1fd8b56b016fe19a77e8ab73c9c9d962428d011fe7

Observation 537332fd-f14c-4b01-bebc-63fb4571fe5e · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.138298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.138298Z digest=sha256:48377a5c8aa631d7accb2a136407baac0072d1d800ef8234623181d98732f95c

Observation f0ed166f-4887-4415-ab39-a07af0c8721d · outbound

This paper cites Sigmoid loss for language image pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Sigmoid loss for language image pre-training,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.349219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.349219Z digest=sha256:65bbe674862c81b9d89fddf5cb92d321df59ba737747deb35a575995e3673c56

Observation 44a4699c-f096-4e34-83f5-f5e54cef2ca5 · outbound

This paper cites Qwen2 Technical Report.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.535201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.535201Z digest=sha256:2ba1e414681eea0a10699b9d3399889675f387836f583bbdcbc560e233ebcdfd

Observation 9d5ec591-45c1-41cb-b6d8-244f741b30bb · outbound

This paper cites Exploring Plain Vision Transformer Backbones for Object Detection.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring Plain Vision Transformer Backbones for Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.678581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.678581Z digest=sha256:a716c5e5f6613e4c45f88f6894f3a45e2e6465a8a74f1829060633411c2b2dcd

Observation 9c6490ac-a365-4e03-a021-585cda9746f2 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.805151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.805151Z digest=sha256:33444dbc8677cc018da22b73ab0ff1c7b9685c0a8770af79d6310b19972043c6

Observation 8158c46c-bef8-4e83-8312-21b3dc5bdd76 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OK-VQA: A visual question answering benchmark requiring external knowledge,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.963496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.963496Z digest=sha256:a4ebdcbb3bd40ea6d0b9a8057de909c35353cbd56f94a6febc20e97035b449ac

Observation d4c7b5df-bf1e-4fdf-a001-c90e39924fb9 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.071175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.071175Z digest=sha256:0c5706a66d9b8043b23422a7cedd9e90b1f841a025d0f128344b12592332028f

Observation be1fb631-5e0e-400e-b7d2-6f4ae267f566 · outbound

This paper cites MM-vet: Evaluating large multimodal models for integrated capabilities,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs MM-vet: Evaluating large multimodal models for integrated capabilities,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.186859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.186859Z digest=sha256:3712eec40c44d8a981e2abac781c945d371d3c4e8ea01618450f9b5164727847

Observation 0f6f0d9e-1324-4eec-b3ec-8182e7455207 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.324730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.324730Z digest=sha256:685cde579db4b6b91469d49627c7a48711cad15559b5e21436ca6f274f3eded9

Observation 28546be3-8d16-4c27-8c88-47030076110e · outbound

This paper cites Grok-1.5 vision preview.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Grok-1.5 vision preview

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.541834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.541834Z digest=sha256:3625826c71114328495d483c369ae72c9497c1d90c84a407009ce65f7dc5c9a5

Observation db613073-00ff-4358-82df-5a1e2594cff4 · outbound

This paper cites Towards VQA models that can read,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Towards VQA models that can read,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.732658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.732658Z digest=sha256:e14be8b6ab2abaf848d41f683a6957c338ba57dcce9686e6738a113fd27ba33d

Observation fd5ae75d-1c63-4d54-b5d4-335a3171eb90 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ChartQA: A benchmark for question answering about charts with visual and logical reasoning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.947386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.947386Z digest=sha256:0a63baa25cc9143527fbd4fe080c4b1221b6257ab4019e746d31c833b2cd94c1

Observation 0d760691-3514-4c2a-896d-d4a8a8034796 · outbound

This paper cites Infographicvqa,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Infographicvqa,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.137978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.137978Z digest=sha256:7322e51a69fcdf9e44e6e08f8ead2a7fc0b80c9af554ead88e628b344a816926

Observation d1350f9c-11ba-457c-9adc-0a6cd3b1d957 · outbound

This paper cites A diagram is worth a dozen images,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A diagram is worth a dozen images,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.270291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.270291Z digest=sha256:83d443e95784c3b1e14e0335c2b563db35dc882c43a6b25aee69bc6fdced8246

Observation c9896961-3bb1-4c9c-bdeb-be0b8c47cb4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.485663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.485663Z digest=sha256:597755dc354710b1edba93c8de7ceca97a0c9747cbb29ebca8f78ef84fcf351f

Observation 86e7b22c-7d89-46c0-956b-294c214ac865 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.621152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.621152Z digest=sha256:5b37578f9b31a9060790147206a1e17e0340c05a8240325962b6df323c58e9e7

Observation 2f75a396-e90c-4cb4-90c5-f2a010df1e03 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.791974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.791974Z digest=sha256:07fc9c29cc9320184885d522586d511714e957a5fc7e1b25ee695ea6a9e5156a

Observation 791fa363-476e-4220-b480-1297520c747b · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava-next: What else influences visual instruction tuning beyond data?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.922650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.922650Z digest=sha256:d15b35b3dd0b7d7a2a6153fc673e49c34d30157d021a2708f3c864114de43c1e

Observation 1b0053e0-2ad7-40d6-9e37-341dea810c36 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.016343Z digest=sha256:de1d35eb8cdaf7a277e7b11ab8ce8a20580974aac4533082c573661863b3e12d

Observation ccfd809d-005b-4e6b-9e68-8651e9771796 · outbound

This paper cites Visual instruction tuning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual instruction tuning,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.133126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.133126Z digest=sha256:2c03a20ac6a4fded9a752761316e1063153e20fb8ed4809a42b9e4c9463d1bd7

Observation 630611ec-21aa-4254-a63b-4364f039a30e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.228490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.228490Z digest=sha256:56dd14180e51d30e3e5a18f78f28f5f772ad5fd98bb363b606a12e03886a7f2a

Observation 7a46b53b-59b1-4e10-a856-f00b73567ab5 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.334971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.334971Z digest=sha256:b65d7ff4b8328d7bd85a7e4252435709c0631d9bbd96e692e44276a8f2c3e037

Observation 1a13bfda-229f-4e88-b2f1-ff9484da7785 · outbound

This paper cites Neural machine translation by jointly learning to align and translate,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation by jointly learning to align and translate,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.454860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.454860Z digest=sha256:e42f29da4e08ced317a5d7d58049f0373a7fc0d7449f46ea8464f921dd76020c

Observation 78ea5f86-41b0-4a53-9cb8-04e90e66d748 · outbound

This paper cites Revealing the dark secrets of masked image modeling,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Revealing the dark secrets of masked image modeling,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.599170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.599170Z digest=sha256:79de8af3b729d1e778968e2ced6f172afe506e3e510091adf5e046c5574d0bb8

Observation 50e4abef-6a5a-42c5-a7d4-42205f4297a5 · outbound

This paper cites On information and sufficiency,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs On information and sufficiency,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.718447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.718447Z digest=sha256:06294be047d2098e13d0420dbc3095b542313db8437438ce3cce51530578d3d7

Observation 11507107-d0ac-4d9d-be1d-a4aa00ff6b15 · outbound

This paper cites VL-BERT: pre-training of generic visual-linguistic representations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs VL-BERT: pre-training of generic visual-linguistic representations,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.750278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.750278Z digest=sha256:d24587d5278bdddb175395a6114b55fde686e8c5c8ad8cebec0328168a6b881b

Observation c21d2ddb-66ff-4d8d-a3b1-8f6fe4d4b2db · outbound

This paper cites Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.831669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.831669Z digest=sha256:694c521a4253d0f09b05e040bcd40b27710370144294513533d46bcb66784021

Observation 9185042d-2240-4ff6-bab4-27b2067d0c14 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.945825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.945825Z digest=sha256:2d8d3bf18a3095043f30bff45b00c8a4368790d37ed0abeccc5f66c858b2df6a

Observation b0e88f0c-1d92-4bd9-9257-708632af24b9 · outbound

This paper cites OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.037290Z digest=sha256:13e6087dc86414f613a16b021265b3a2472c09b9390001ff95d6bbaeafeecb13

Observation 4f4663b1-9867-4080-86a5-7bb1fd1cc08e · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.307549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.307549Z digest=sha256:7fa897f60dec06f61cf176684a575008a62171d0157853487ad211e066d6b0c9

Observation 1b21a005-6ee6-4ecf-9f2e-dfa787d4f2f0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:43.340094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:43.340094Z digest=sha256:2daca2e1a7c6f201b198d4f573e54a52e72ebee1e59814cd23d5bf18a2cf4c5c

Observation 860a30ee-b50a-4881-90a6-4fac86b0df49 · outbound

This paper cites Exploring vision-language foundation model for novel object captioning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring vision-language foundation model for novel object captioning,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:43.896751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:43.896751Z digest=sha256:7d0c9283793746769763e18b7c7a2e98b738bcdad20b6a6e4d8fe516929f1c9d

Observation 0d2e4955-30c7-4ba3-b795-3ef241e06712 · outbound

This paper cites Unsupervised domain adaption harnessing vision-language pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unsupervised domain adaption harnessing vision-language pre-training,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.003055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.003055Z digest=sha256:10421e5fa422ef85ad6aa1c8014028f11ffac3da4bfd5b5f7333ee0a774b5e0e

Observation 9850c635-261e-4a42-996e-4a495732adf6 · outbound

This paper cites Do vision transformers see like convolutional neural networks?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Do vision transformers see like convolutional neural networks?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.141844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.141844Z digest=sha256:fdd884d1515d93dfb133f04db3b8fca43a1e10773a04d32a932218ca0efe3934

Observation 4e491e1b-f238-4429-bc18-38adbcdc299f · outbound

This paper cites Intriguing properties of vision transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Intriguing properties of vision transformers,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.267295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.267295Z digest=sha256:32c65edf8b52b75a9926249e9796776886031d5d8e7e26a7eb4c349122c4c29a

Observation 20dc7104-e1f1-4a72-ade2-e2ddc88cbc52 · outbound

This paper cites Dissecting contextual word embeddings: Architecture and representation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dissecting contextual word embeddings: Architecture and representation,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.396287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.396287Z digest=sha256:2dcd1211c050445a4cab47752b13e34da223b053f6a28d4f0b1c871339680e92

Observation 0937db5d-0aa4-47ad-872d-540f20f90442 · outbound

This paper cites Linguistic knowledge and transferability of contextual representa- tions,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Linguistic knowledge and transferability of contextual representa- tions,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.506110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.506110Z digest=sha256:80065732bd1889999f02f27c3091cd740673896f232f22d6131ef7d7463c6372

Observation 0d159b50-7879-4346-a4b9-4b3d6b6049ed · outbound

This paper cites What does BERT learn about the structure of language?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs What does BERT learn about the structure of language?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.638326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.638326Z digest=sha256:5f47c4cd391eed72c1d62a543a397f487f3e20b43cb3946fd303fe52952f3644

Observation f7a40f94-b25b-4c85-be23-f6d1a752f0cd · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.742392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.742392Z digest=sha256:21070545b2919281eed1ae8eaf634ed45622286e7485e14e341211afed70f60f

Observation 82ead94a-b173-4027-b219-aa7e4f40108e · outbound

This paper cites Feature pyramid networks for object detection,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Feature pyramid networks for object detection,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.850613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.850613Z digest=sha256:f79d31fea2e094d2e28636a0f27c23b0df130034736df0b04c237eb5f49e3b70

Observation 5491a66c-cb13-49a4-befe-b504ac064462 · outbound

This paper cites Densely connected convolutional networks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Densely connected convolutional networks,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:45.304746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:45.304746Z digest=sha256:f9e7bd796543c9e2d5027b7d0db31b7bb1ece1e091f5531f8e26789927ae5757

Observation 4ddf0667-663a-41c7-8681-a534e57566c2 · outbound

This paper cites Deep layer aggregation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep layer aggregation,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.716660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:45.913961Z digest=sha256:74483631ca226379f42a434c8164d1696d73fe81e98c8ee2892a37662fa0e233

Observation 81b656ed-2771-43e3-85fb-4b7858b5b7c5 · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Segformer: Simple and efficient design for semantic segmentation with transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.489616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.004678Z digest=sha256:338fa6a62e53d9f3ab52ada11d3037b63871e00172748010c46b384ec5e0207a

Observation 87038d9d-e3d1-497c-818f-f7eee285c3da · outbound

This paper cites Clsr: Cross-layer interaction pyramid super-resolution network,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Clsr: Cross-layer interaction pyramid super-resolution network,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.169806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.141424Z digest=sha256:c4a0826cf59feedb7b51c9d85184b6727bb47b38873da176414802bc4add4fba

Observation 8239b6d0-30e7-47e8-acc0-fd8a21ec2a71 · outbound

This paper cites Attention-based layer fusion and token masking for weakly supervised semantic segmentation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention-based layer fusion and token masking for weakly supervised semantic segmentation,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.932055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.248925Z digest=sha256:b6ba3cf9e94d0107b52419bd9f9ae234bad39ed7f97081a125a720efbf2799a8

Observation f0638c7e-7665-4b3a-bd03-f89a1ae50506 · outbound

This paper cites Artificial- spiking hierarchical networks for vision-language representation learn- ing,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Artificial- spiking hierarchical networks for vision-language representation learn- ing,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.661838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.355704Z digest=sha256:e58ecc688065fe7d06e05f2e811173ecfdf9774f392062912a3b92caaf33aaf1

Observation b1667396-7ca9-447c-87c5-2d97a56a6271 · outbound

This paper cites Deep contextualized word representations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep contextualized word representations,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.341513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.461060Z digest=sha256:e9cf81cc7ef45b6548fb17ce63ba1371d2be04caa9f66b4c7c49873ff1462c24

Observation fc0fba08-fdf4-4305-a1d7-a34552aad9cd · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Coarse-to-fine vision-language pre-training with fusion in the backbone,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.033250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.583941Z digest=sha256:04b42aec890aa35041c6f5a3e42fb39dc65cd3b76ca3c361d84b1313faa79cb8

Observation 9ea500f7-b2be-4f3c-a87b-5624dbf3905b · outbound

This paper cites Dense Connector for MLLMs.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dense Connector for MLLMs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.709519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.709519Z digest=sha256:e39095ffd1975e31eb29ef9ada46645c0b6c7315b99696dda3fe7b760b2c45b5

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:ace7d013df34cc267bfcf7a92c2e4a7051268ee9bf4f957bc8b8795fbda5cdb8

Observation 2a35e3cd-6c1f-489a-9ae9-bbf020cf6cb9 · outbound

This paper cites Language models are few-shot learners,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are few-shot learners,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.754466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:46.918627Z digest=sha256:89b3695447b23ea0dba4322eafabc5bd8523a3fc08ba5084a5bf0eca2ceb19d4

Observation 3e4e945f-0795-4c44-a2fd-3049636038b5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.068817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.068817Z digest=sha256:7f7e7b2e66220539601eb43027700be686a9a99381df095525dc277dcddea699

Observation 5b98f0a0-33cd-4afc-b40b-2d8e3fbbd78e · outbound

This paper cites Large Language Models Meet NLP: A Survey.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Large Language Models Meet NLP: A Survey

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.189430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.189430Z digest=sha256:475b190a3defc78aaca0d1cf941bc5d768d6221028216a7b28fcfc7dc9703bd0

Observation 33ec15ab-e7f7-4152-8e89-546a98610b0d · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.526388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:47.273636Z digest=sha256:3c5c85529fead5ac1db127382c16056b4c2fd3ba75d35d78095bd97c843138b5

Observation f50a89ec-fb71-4975-87dd-9660679b2255 · outbound

This paper cites Introducing our multimodal models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Introducing our multimodal models,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.298926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:47.359818Z digest=sha256:8abd880e1ecd0c6f6789b049db521213193faae0bae0af2efd60e9dc46af7d0b

Observation ffe0c48d-d30c-467d-9802-8039d28c4610 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.490864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.490864Z digest=sha256:cd181ee50a6bb1785961f83bea31e4edeb2d9e39131f33997a408f2bb6875fa0

Observation 0ed14fee-6450-4306-8821-431c65915b8a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.591621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.591621Z digest=sha256:2a4d786834bb1898455f0cbac10400f8002deff62f7a7d60228abfa956527ea7

Observation f2c93c76-ca12-46eb-8d29-e800355ccd76 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.692121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.692121Z digest=sha256:c7a3a7ddbc0fc7c0e1acabe0c54b93b81fc5b9a93ec41f88afbf5fed59ee87e9

Observation 30c172e8-53e3-4dc4-9254-b71a08b7a65e · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.795386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.795386Z digest=sha256:4751e90ebdaf52e93b1bf08f9d7ded3de7c7e922b516f943ec390a0e61e6dac6

Observation fe259a9d-cf04-425f-93f8-ba3e61506701 · outbound

This paper cites When do we not need larger vision models?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs When do we not need larger vision models?

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.061995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:47.882134Z digest=sha256:ae614594bbd4c2e0b066e567632615631dc39507a577bf647fc7671937d67819

Observation 55ef876c-a123-4673-beb6-714e832d43db · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:48.010046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:48.010046Z digest=sha256:e42c46fd14a798817385b5f7790e6869821d32238da032dad0f07be90c40d5f9

Observation 3aac154c-821b-4d07-b583-e03cd288f6a1 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Honeybee: Locality-enhanced projector for multimodal llm,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:04.816456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:48.091130Z digest=sha256:4463aec5007f5410161472d25095fd29177d9f10c19bf560001dc85d3a55ef50

Observation 0815b0f9-2125-4277-8e8e-d6211aecd15d · outbound

This paper cites Unified language model pre-training for natural language understanding and generation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unified language model pre-training for natural language understanding and generation,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:04.564950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:08:48.226962Z digest=sha256:4784c43a95dba3bf3bf09e0647f4b875fe110314e01a911de473143d3b0b8f00

Pith citing papers

No inbound Pith citation observations are available.