Pith. sign in

Paper Citation Record · LEDGER

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

As of 8 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 0 inbound Pith citation observations for arXiv:2506.11515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11515 v1

Coverage vector

measured 100 of 143 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:08:48.226962Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 143 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b787420d-6acc-4f61-a52a-79299a9128bc · outbound

This paper cites ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.096271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.096271Z digest=sha256:b26ac0fad838ccadb0a239f9b48448520ef489d5ad9e2351f5c202b2d0e2bc6f

Observation 55f9426a-ef5c-4aa8-b023-99165f154c40 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.202223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.202223Z digest=sha256:53a680eb19b67b4a5c178ac0c566c5ac5c3ad57ca5bb1c70888c0ce53edfb78a

Observation ebc1cac8-cde5-463f-b1a8-406c2422b85d · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.416327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.416327Z digest=sha256:651d90713bdc0bc89040a4f1da578274882921054ee8ccdde4654fb55661a21a

Observation 3a06e115-c7a9-4c85-ab83-083584a3d28c · outbound

This paper cites A corpus for reasoning about natural language grounded in photographs,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A corpus for reasoning about natural language grounded in photographs,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.710099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.710099Z digest=sha256:442f5eb82b24af0b399e6bb1404d882bc90ece3867475300e0bdb07fc19316f2

Observation 62e00204-edb0-4c9f-8b1e-48ec10634540 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.832464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.832464Z digest=sha256:3e61b417a1ebdb3a23baccd883858c4da0a1a105edc5327bc60542c335f86a8a

Observation 4deea2d7-3081-4629-977c-7375b2fc0ce8 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An empirical study of training end-to-end vision-and-language transformers,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.933376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.933376Z digest=sha256:d997e41b3d0d1651c56defa91abc6210e310e7f353fdd6bebbec9e2cc22306ca

Observation 6b114773-63c4-4e72-bb35-62f12f64839e · outbound

This paper cites Bridgetower: Building bridges between encoders in vision-language representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Bridgetower: Building bridges between encoders in vision-language representation learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.047268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.047268Z digest=sha256:e430929f1eee25e96dc7535d902f09e589dbf1260248ec60a09ddb7cabccd608

Observation f1d71738-f08d-4cd5-9398-f9bf73b4f60b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning transferable visual models from natural language supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.138970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.138970Z digest=sha256:439785a5f7c0150f29278f5b1c1c9f91d8f5b3e3f5c563b5d0a3d155aa3b5439

Observation 5a635d12-cd8e-4d0e-a1e3-a20123271c14 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.269423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.269423Z digest=sha256:4b071493cbc5a921696f3cd1aa4f97b9ae63f4243c28486a5cd5a3fe5663ba10

Observation 1497da84-d879-4b24-a45b-d1310f0bc138 · outbound

This paper cites Learning deep transformer models for machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning deep transformer models for machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.396539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.396539Z digest=sha256:4e03e35ac30e14d36370fe0329aa42f0a279f6f0ad22b0f40c97f4e57211c11b

Observation 3f8768f4-df16-48b4-ae99-fd24685256d8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.510959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.510959Z digest=sha256:a35aff439354bacd8780d9b7e140a726a47b6247596549a769dcce34adbc4b6c

Observation 17f631e1-3a02-48f3-ad3c-0edf39ec72bb · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowledge,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava- next: Improved reasoning, ocr, and world knowledge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.681245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.681245Z digest=sha256:edcd6b46a07eb1cd7d7d4db3161490ad6dfe602ab42697db0b68eee69db7c2b5

Observation 2f78c9a6-e55e-4637-a4a3-8d5d209ae51f · outbound

This paper cites How much can CLIP benefit vision-and-language tasks?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs How much can CLIP benefit vision-and-language tasks?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.848451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.848451Z digest=sha256:5b996ada47fa46823ad6e2d06dbcecebfc522c8a8f2b854c3da847411c3d7579

Observation fe89ca05-e8d3-472f-ae87-bf3686995f4a · outbound

This paper cites UNIMO-2: End-to-end unified vision-language grounded learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO-2: End-to-end unified vision-language grounded learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.012205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.012205Z digest=sha256:29932a506ec1e47d00fe73e623e16c1421b3782ceae2dae992528dbf3b53ff71

Observation 83e54d8e-4a8a-4f7b-9af6-5cad87292bb6 · outbound

This paper cites Neural machine translation of rare words with subword units,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation of rare words with subword units,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.146927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.146927Z digest=sha256:84e542ebb38903175f3aa1989911dee8ee9bfbe8e3ae191de7d9964db34ca3ff

Observation c02c04b3-7ca1-4bca-b887-8375e818dd1f · outbound

This paper cites Language models are unsupervised multitask learners,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.256696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.256696Z digest=sha256:2d463edcc9b31f7a9edd7997d39ae4291d4e550c2bfed30d68001d716cc938ce

Observation db02e517-fb23-4300-a9be-3426bf0d409d · outbound

This paper cites Attention is all you need,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention is all you need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.432776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.432776Z digest=sha256:4d6f8e8f18bd2313b7d579f7476e2dbe4a720c320f901c576321cd5747db4b6d

Observation 2aa3a712-7ab3-44c4-a016-f21eb83b3039 · outbound

This paper cites Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.541595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.541595Z digest=sha256:c7e2c529c5dace65407a00a19da7caa89a8cbd9a2891c9a7902335630bc06165

Observation 127b0b7e-b728-4454-9243-9d426aeea31c · outbound

This paper cites Multi-layer representation fusion for neural machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multi-layer representation fusion for neural machine translation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.664720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.664720Z digest=sha256:1197d166a0c209f1a99fa31aa3c0c8c2ba9071b58eb5e354724bc5a70f63c5c1

Observation 6e228351-5890-459e-a8cb-e664beffcaa7 · outbound

This paper cites Multiscale collaborative deep models for neural machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multiscale collaborative deep models for neural machine translation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.819400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.819400Z digest=sha256:c92a5b9fd667e3f9b3b0f899bb2921402f7a681092b79585419508fa87d0238d

Observation 168de691-262c-4021-983a-4525d727306a · outbound

This paper cites Layer Normalization.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Layer Normalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.926027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.926027Z digest=sha256:7cede2cea487a5dc6d83407d082ae70fc204cf9742409c9acbdecac7a487a617

Observation d3f72292-993e-42a2-8b69-78960073aa47 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.069126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.069126Z digest=sha256:43aed407b55cd4427d0058bcffcfc9ea57745857fcda4021d693feedfbaeef79

Observation a6364764-2666-4314-a977-3c149ba81cd8 · outbound

This paper cites Decoupled weight decay regularization,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Decoupled weight decay regularization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.245435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.245435Z digest=sha256:5cce00113a21388c6c5f6073550435acb8426b6f1968e724124cdaf5d1a05786

Observation 13501472-651d-4356-b2bc-8d4511f2f927 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.380900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.380900Z digest=sha256:c06e7b8ed9de7d88ac808ae985d50b8e2f965e32718a79432d29492fd8e6b273

Observation a8d542ec-e519-46e3-a9d0-bbe5db76fa07 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Uniter: Universal image-text representation learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.538152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.538152Z digest=sha256:f93e0a0b113dd8285baa42bec9d3da29ed5f514c65ebef425ba9b147a2b0e53f

Observation 4559d4da-8f84-45c6-b95c-2d17fde19ed3 · outbound

This paper cites UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.679187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.679187Z digest=sha256:c96502436d086b1bd06815a91c5123f7f4d17f43b3ca8cb861d2814be25a0a89

Observation 6aa74da3-3b4e-460c-bd91-263168634441 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Align before fuse: Vision and language representation learning with momentum distillation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.787717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.787717Z digest=sha256:424df3b1d53aa220c2b903d9cef9823a91c6bac84f06a18e3c01b4ad4f1eb238

Observation 40ec964d-4a3b-43ff-9eb0-ac8e6074e064 · outbound

This paper cites Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.927145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.927145Z digest=sha256:83956a189710544621f2414a7f49c6351c40cc091d402d6f67d1fad7f45ae3bb

Observation 0e1b9acf-fcc9-438f-9f15-94e42948a820 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Simvlm: Simple visual language model pretraining with weak supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.043824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.043824Z digest=sha256:10c2d07194ebc8e3b6a7eb7b499fce05557bd48905f285d55a0c8ab49c100897

Observation 9d31b259-9a5c-4d4e-97d5-8f750e5c074a · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.184911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.184911Z digest=sha256:d844312454d1141e92de34b5f444c3c4c0b4eb72b6cf850670b86eb261a012c0

Observation 81f9cc6f-3c60-48a9-8c62-b536775451df · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.280478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.280478Z digest=sha256:400195efe0b31cd25f59713478da7bcee5905e838de0327712e0464f741d72b9

Observation ab0fd178-91fb-40f4-8f7b-13bfa4a7129b · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Im2text: Describing images using 1 million captioned photographs,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.369147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.369147Z digest=sha256:79e39d81817cf281faf65037c0158754d711be7bb3678eca42ca30de661db399

Observation 16c6ed4c-a231-4626-9caf-a5e85204bc85 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.503386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.503386Z digest=sha256:de570a55eb894370a2ab3c9874bc15c494a1d9fb3e5c846fec4e8df2d4a00599

Observation 43234c42-5705-49f9-bb51-c0a6d5b9edd7 · outbound

This paper cites Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.643209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.643209Z digest=sha256:2b72945788102c58ea7d781ee63f5c2a624ea1e7081e589509043feef7ab2c65

Observation 2559ceb1-fa12-4997-bc03-069586b28428 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Improved baselines with visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.753062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.753062Z digest=sha256:12e25dd012d0e94087bfc80a72a25e2bbb0e0eb455e4662297aaf1e8528866ca

Observation 67f49981-e54c-4b78-a46e-cc375d64580f · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Monkey: Image resolution and text label are important things for large multi-modal models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.870902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.870902Z digest=sha256:937824a99339d8f28d5d18a68a4f263b97e677f9c4ee222889f073a1dbd6a28b

Observation 54ecea55-9c8f-4520-b022-d83c20fe3c04 · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Docvqa: A dataset for vqa on document images,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.996940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.996940Z digest=sha256:ccc0695776553794373bbf8e43448e906c1ac7a0ee9e0e7cdb062148b947f5a0

Observation 537332fd-f14c-4b01-bebc-63fb4571fe5e · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.138298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.138298Z digest=sha256:7ec40d3f7d3d7e7f806769cd39577ab4b07345eef466952dc2526169366fe80a

Observation f0ed166f-4887-4415-ab39-a07af0c8721d · outbound

This paper cites Sigmoid loss for language image pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Sigmoid loss for language image pre-training,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.349219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.349219Z digest=sha256:4755f95e7aebdc929b8039a978ae0f09d9724dad89d0a0555152f08eeacad1a7

Observation 44a4699c-f096-4e34-83f5-f5e54cef2ca5 · outbound

This paper cites Qwen2 Technical Report.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.535201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.535201Z digest=sha256:d3af61b9100f6424fb9f7b28486f548a2f4521c1d66aa0ada8be476cdfdcac4a

Observation 9d5ec591-45c1-41cb-b6d8-244f741b30bb · outbound

This paper cites Exploring Plain Vision Transformer Backbones for Object Detection.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring Plain Vision Transformer Backbones for Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.678581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.678581Z digest=sha256:39b99bb37acf34643e12a6fa6429cf59e1302e75cad34d9bc5c460dcdbe07689

Observation 9c6490ac-a365-4e03-a021-585cda9746f2 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.805151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.805151Z digest=sha256:e1f582706ee4c626448e52b64a0a73cf0ffe20177188acd21ee22c284df7878e

Observation 8158c46c-bef8-4e83-8312-21b3dc5bdd76 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OK-VQA: A visual question answering benchmark requiring external knowledge,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.963496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.963496Z digest=sha256:9f125c27537b3a26d20b974ab65f4adf55685ffd18a435fb3a9c2ab4f3459e82

Observation d4c7b5df-bf1e-4fdf-a001-c90e39924fb9 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.071175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.071175Z digest=sha256:096d38c9be756f8798e657aa688c23e7663c821f08c689995344d03456913f1d

Observation be1fb631-5e0e-400e-b7d2-6f4ae267f566 · outbound

This paper cites MM-vet: Evaluating large multimodal models for integrated capabilities,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs MM-vet: Evaluating large multimodal models for integrated capabilities,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.186859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.186859Z digest=sha256:2f4d7516ff60d304195e717fbcf5f773b0d3aa06768d8a4927877c58479d6f71

Observation 0f6f0d9e-1324-4eec-b3ec-8182e7455207 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.324730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.324730Z digest=sha256:b425dea1cfd03ee1f262814464f836aecccc5fb2094f47d9c2f9887b568e08b3

Observation 28546be3-8d16-4c27-8c88-47030076110e · outbound

This paper cites Grok-1.5 vision preview.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Grok-1.5 vision preview

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.541834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.541834Z digest=sha256:158e99861d1328910b841354059efb68d7bc617ee44e74d7fe477da6dc90261e

Observation db613073-00ff-4358-82df-5a1e2594cff4 · outbound

This paper cites Towards VQA models that can read,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Towards VQA models that can read,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.732658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.732658Z digest=sha256:d7f67c673910c52e32089ffdf50b08261980c5f94771b7fe86fb9154316e1707

Observation fd5ae75d-1c63-4d54-b5d4-335a3171eb90 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ChartQA: A benchmark for question answering about charts with visual and logical reasoning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.947386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.947386Z digest=sha256:b0357f298dd8ceaf38962c97a33d8dc5dba88b533dd70f69846ba5e8ea59931c

Observation 0d760691-3514-4c2a-896d-d4a8a8034796 · outbound

This paper cites Infographicvqa,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Infographicvqa,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.137978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.137978Z digest=sha256:91bf0ce406bd13b437ca1018c0ee233ac40aa7e58b4e7cf82622a5ed25e5ac0f

Observation d1350f9c-11ba-457c-9adc-0a6cd3b1d957 · outbound

This paper cites A diagram is worth a dozen images,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A diagram is worth a dozen images,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.270291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.270291Z digest=sha256:b27dd0598fccfe4a6aab83b80837f848c6c612663b39bf87780f637a8c8098b6

Observation c9896961-3bb1-4c9c-bdeb-be0b8c47cb4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.485663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.485663Z digest=sha256:0f1c0e9395f3b28775b14aa6e88319c0c11b5d37d73dc8ceef7014109f1b7e5a

Observation 86e7b22c-7d89-46c0-956b-294c214ac865 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.621152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.621152Z digest=sha256:781cc465429e746d3ad0e78c43daf05ea5de77ad40817172d454139ae1ce3c32

Observation 2f75a396-e90c-4cb4-90c5-f2a010df1e03 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.791974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.791974Z digest=sha256:0661c741a38b9e62f68c56a20a4e3248dd0007b4a07ed6f10e8ec82b29713492

Observation 791fa363-476e-4220-b480-1297520c747b · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava-next: What else influences visual instruction tuning beyond data?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.922650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.922650Z digest=sha256:73468523fb239292c27a1980e9afe62f373c7ff9dbc8c9fc7fbe7d76d271e816

Observation 1b0053e0-2ad7-40d6-9e37-341dea810c36 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.016343Z digest=sha256:b8968c87c1819a862132a14bfb95a1d7667467b20ba8f23a4ea831ff2d7f505e

Observation ccfd809d-005b-4e6b-9e68-8651e9771796 · outbound

This paper cites Visual instruction tuning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual instruction tuning,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.133126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.133126Z digest=sha256:aee788686d1f2786aa095a49ca3fa427c4579806f0e16933eca565ffb0d4ff79

Observation 630611ec-21aa-4254-a63b-4364f039a30e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.228490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.228490Z digest=sha256:24961285ac84b65787cffb335f8d7bb275a1a25684a14d02041f6280bd44f31b

Observation 7a46b53b-59b1-4e10-a856-f00b73567ab5 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.334971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.334971Z digest=sha256:f2fd01907eb3b9cdd8fe50a48ba0cfcc51893af61a3fc0868e7bf9f305f2fb4a

Observation 1a13bfda-229f-4e88-b2f1-ff9484da7785 · outbound

This paper cites Neural machine translation by jointly learning to align and translate,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation by jointly learning to align and translate,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.454860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.454860Z digest=sha256:8a83c5e39660b66d350566cb43417d052fc56e5f0d3849f54be7719de91f8052

Observation 78ea5f86-41b0-4a53-9cb8-04e90e66d748 · outbound

This paper cites Revealing the dark secrets of masked image modeling,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Revealing the dark secrets of masked image modeling,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.599170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.599170Z digest=sha256:e26420cca520c3d24e22b2aff9fcdfa699be8b7c0e8216b983cd17ca9eddb0df

Observation 50e4abef-6a5a-42c5-a7d4-42205f4297a5 · outbound

This paper cites On information and sufficiency,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs On information and sufficiency,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.718447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.718447Z digest=sha256:4f8b26d56c4b1345ba19631d3be29a8185b838388be89a2e8c4a773d19353b1a

Observation 11507107-d0ac-4d9d-be1d-a4aa00ff6b15 · outbound

This paper cites VL-BERT: pre-training of generic visual-linguistic representations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs VL-BERT: pre-training of generic visual-linguistic representations,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.750278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.750278Z digest=sha256:1bdb07974e991956f4637cbb0398a91324408d0565ba00f421de6cfde60582d4

Observation c21d2ddb-66ff-4d8d-a3b1-8f6fe4d4b2db · outbound

This paper cites Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.831669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.831669Z digest=sha256:282aa8b6598e90df13ac4bec39af76e3ba3ef99cfb3a14360e1949b2e1af7154

Observation 9185042d-2240-4ff6-bab4-27b2067d0c14 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.945825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.945825Z digest=sha256:f9eb345fe6db90c6147f9d42f63c3355e5d66b55e0460a4d792560f41cee1e21

Observation b0e88f0c-1d92-4bd9-9257-708632af24b9 · outbound

This paper cites OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.037290Z digest=sha256:475adf97fe9bff2e0b093bf4dca676795ab21fcb5c7625d72397ec88be476a61

Observation 4f4663b1-9867-4080-86a5-7bb1fd1cc08e · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.307549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.307549Z digest=sha256:51ca373ad6863a22331d28b6823411d45eecbc70a87f3bef5d8694f330ec65ee

Observation 1b21a005-6ee6-4ecf-9f2e-dfa787d4f2f0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:43.340094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:43.340094Z digest=sha256:689da6157f7c8eac9afa443d5dbe7b65b6af35bcd70dcb351470e93f8998f612

Observation 860a30ee-b50a-4881-90a6-4fac86b0df49 · outbound

This paper cites Exploring vision-language foundation model for novel object captioning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring vision-language foundation model for novel object captioning,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:43.896751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:43.896751Z digest=sha256:82603af52efdb56cd3a3e507d066775da14daf7b37aeee007a262e14173575b0

Observation 0d2e4955-30c7-4ba3-b795-3ef241e06712 · outbound

This paper cites Unsupervised domain adaption harnessing vision-language pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unsupervised domain adaption harnessing vision-language pre-training,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.003055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.003055Z digest=sha256:b39a8fab8ca13c09c5ac7651874a5ff65e2ba24163d4ec2090d0503f02c11235

Observation 9850c635-261e-4a42-996e-4a495732adf6 · outbound

This paper cites Do vision transformers see like convolutional neural networks?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Do vision transformers see like convolutional neural networks?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.141844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.141844Z digest=sha256:b76ce3b4e149f3a36df2cc0fc5881a24dec689c7b8ef83d308eda51fd8ec0a73

Observation 4e491e1b-f238-4429-bc18-38adbcdc299f · outbound

This paper cites Intriguing properties of vision transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Intriguing properties of vision transformers,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.267295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.267295Z digest=sha256:f21c0072214ced800f55f3e7ffa32262345b880e47dbd742f24926569e607848

Observation 20dc7104-e1f1-4a72-ade2-e2ddc88cbc52 · outbound

This paper cites Dissecting contextual word embeddings: Architecture and representation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dissecting contextual word embeddings: Architecture and representation,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.396287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.396287Z digest=sha256:0fb5f4fca904290555acc6be0cf3ec921a2d828cbcd915c0647ff03efaaba961

Observation 0937db5d-0aa4-47ad-872d-540f20f90442 · outbound

This paper cites Linguistic knowledge and transferability of contextual representa- tions,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Linguistic knowledge and transferability of contextual representa- tions,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.506110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.506110Z digest=sha256:007b7ce1d727c61d4619fe27306691d80ca273bfa375b0e47374447895c417df

Observation 0d159b50-7879-4346-a4b9-4b3d6b6049ed · outbound

This paper cites What does BERT learn about the structure of language?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs What does BERT learn about the structure of language?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.638326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.638326Z digest=sha256:f362c7596818100c07ad136eb4176bb0801c7a35974750b3219914baf7727d10

Observation f7a40f94-b25b-4c85-be23-f6d1a752f0cd · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.742392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.742392Z digest=sha256:ee9c97bf64a683b55bb213c5d8d75d13e4f45f284ed09fbacc847e0618845334

Observation 82ead94a-b173-4027-b219-aa7e4f40108e · outbound

This paper cites Feature pyramid networks for object detection,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Feature pyramid networks for object detection,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.850613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.850613Z digest=sha256:d0a3ef9dbf24980ea845135cbc3c9019bdc0bfb3f45c321a6a80ad7c984179ba

Observation 5491a66c-cb13-49a4-befe-b504ac064462 · outbound

This paper cites Densely connected convolutional networks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Densely connected convolutional networks,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:45.304746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:45.304746Z digest=sha256:a702b24e07db322993151e7c61468bfe406dd1a55f89475fdaa4d26bb784cda2

Observation 4ddf0667-663a-41c7-8681-a534e57566c2 · outbound

This paper cites Deep layer aggregation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep layer aggregation,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.716660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:45.913961Z digest=sha256:ceb2ec35062110510e24844af51b03dad298efc3ebc4df31ddc1ce205a1ca97d

Observation 81b656ed-2771-43e3-85fb-4b7858b5b7c5 · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Segformer: Simple and efficient design for semantic segmentation with transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.489616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.004678Z digest=sha256:5e333c685d0a89f3141aea3bff3b18c110fe4e62644a6fedc4ed06df69eedbc9

Observation 87038d9d-e3d1-497c-818f-f7eee285c3da · outbound

This paper cites Clsr: Cross-layer interaction pyramid super-resolution network,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Clsr: Cross-layer interaction pyramid super-resolution network,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.169806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.141424Z digest=sha256:63e5ac4ad2045671b7cede7052e80d172d2f55409b2fad32181afc1b92c2ffd9

Observation 8239b6d0-30e7-47e8-acc0-fd8a21ec2a71 · outbound

This paper cites Attention-based layer fusion and token masking for weakly supervised semantic segmentation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention-based layer fusion and token masking for weakly supervised semantic segmentation,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.932055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.248925Z digest=sha256:ebd4daee164c97cd9db35e6ed21a0331fbf2a6d112768157966d33e57e7430a8

Observation f0638c7e-7665-4b3a-bd03-f89a1ae50506 · outbound

This paper cites Artificial- spiking hierarchical networks for vision-language representation learn- ing,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Artificial- spiking hierarchical networks for vision-language representation learn- ing,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.661838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.355704Z digest=sha256:198d80e5205de8b20b6a3591d475750ce63c64620a295dd72bf087dc4edb52ba

Observation b1667396-7ca9-447c-87c5-2d97a56a6271 · outbound

This paper cites Deep contextualized word representations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep contextualized word representations,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.341513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.461060Z digest=sha256:7826d2d491b130026da5bcac1b577bfdefd5fe85a3826beb6073e84c881cd81f

Observation fc0fba08-fdf4-4305-a1d7-a34552aad9cd · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Coarse-to-fine vision-language pre-training with fusion in the backbone,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.033250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.583941Z digest=sha256:fb7e7f83e7e3b711f2d64ff0f309a25d611187fc736177c5691c0fe393808d9c

Observation 9ea500f7-b2be-4f3c-a87b-5624dbf3905b · outbound

This paper cites Dense Connector for MLLMs.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dense Connector for MLLMs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.709519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.709519Z digest=sha256:62b4a2ae588090ce9c16a6440d780ab66549bdb818776da2b796477b0d611141

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:7dc6b9d10f5f74c8afe26c49f8440d4393b9a6045dfcde8db84c738a5594d9d6

Observation 2a35e3cd-6c1f-489a-9ae9-bbf020cf6cb9 · outbound

This paper cites Language models are few-shot learners,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are few-shot learners,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.754466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:46.918627Z digest=sha256:d146e01bc10f413f61fdacafcbb0824664a474e9fed7e4a095e62d306760d97d

Observation 3e4e945f-0795-4c44-a2fd-3049636038b5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.068817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.068817Z digest=sha256:fb7d40e95a8790279e0e9e84ad21cbd4345ef7643e6528dfa488ebd2c7781154

Observation 5b98f0a0-33cd-4afc-b40b-2d8e3fbbd78e · outbound

This paper cites Large Language Models Meet NLP: A Survey.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Large Language Models Meet NLP: A Survey

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.189430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.189430Z digest=sha256:d0aa726ce7364d31afa0cb9f0bc028897f1db123873b56d115cebd63fc5f3ce0

Observation 33ec15ab-e7f7-4152-8e89-546a98610b0d · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.526388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:47.273636Z digest=sha256:5bace24bd869b4abfa0f3cddbd003b8b41e24fce6bda7a02806b78b5cd53cb6a

Observation f50a89ec-fb71-4975-87dd-9660679b2255 · outbound

This paper cites Introducing our multimodal models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Introducing our multimodal models,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.298926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:47.359818Z digest=sha256:13da015a757b795eeb88d6b36e433984d02416ab7a997c7b04d5e3f7a562a38f

Observation ffe0c48d-d30c-467d-9802-8039d28c4610 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.490864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.490864Z digest=sha256:861c2a246d339d6c020482e8c100a0844b6d8e6a3e1f6db52bb747df86ff56bd

Observation 0ed14fee-6450-4306-8821-431c65915b8a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.591621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.591621Z digest=sha256:0cc2434c895c5caa8ca727414b4721d046c8e8a591f7c84e0417284f95ec964b

Observation f2c93c76-ca12-46eb-8d29-e800355ccd76 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.692121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.692121Z digest=sha256:71267369d449e12eeef7647a0f81763bdbc8b0c344098ab86bb3a3b9c29176ae

Observation 30c172e8-53e3-4dc4-9254-b71a08b7a65e · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.795386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.795386Z digest=sha256:065b4c70e481dec94457d9bb43ead3759f236cc02b09640d0021e7721e617ffa

Observation fe259a9d-cf04-425f-93f8-ba3e61506701 · outbound

This paper cites When do we not need larger vision models?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs When do we not need larger vision models?

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.061995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:47.882134Z digest=sha256:eb8c904519956fb8294bdbd5137f2c143612bae9dc9c80e2d3befee5a14f2961

Observation 55ef876c-a123-4673-beb6-714e832d43db · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:48.010046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:48.010046Z digest=sha256:ff503807b0c3a48dd6c3e3b5766e8bf1d82de9a2335b0da1c44c223f2917e0de

Observation 3aac154c-821b-4d07-b583-e03cd288f6a1 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Honeybee: Locality-enhanced projector for multimodal llm,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:04.816456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:48.091130Z digest=sha256:acf6b82d4a12ad67b8b577f089d302991593738c2e1e6995edae5c18dbb96b06

Observation 0815b0f9-2125-4277-8e8e-d6211aecd15d · outbound

This paper cites Unified language model pre-training for natural language understanding and generation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unified language model pre-training for natural language understanding and generation,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:04.564950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:08:48.226962Z digest=sha256:8366b2f8401e98034e810483620c3a30031f131def13dd5103827036283d92eb

Pith citing papers

No inbound Pith citation observations are available.