Pith. sign in

Paper Citation Record · LEDGER

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

As of 23 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2603.24690.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.24690 v3

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T18:42:33.178622Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-07T19:32:20.095114Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T19:34:06.409175Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa471f00-4758-4517-ad3c-00dad4735a79 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:7776c8a64836cd47cb218608a21d87e424141d1812c42c03fb0c152ee1167e70

Observation aa6fcbb4-dfd2-4cef-b9fa-b121ec75b9e9 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:82faacc3a694bdd354cc2dda136c67f66cbc0554bb3603dad3c811bd3fc2af66

Observation 9ce49761-6b9a-41eb-86da-532f94f675f2 · outbound

This paper cites Qwen3-VL Technical Report.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:5e5fc4d45bb18b2410dd48c74200d8da822066d3f55b0762a8f1285702813aa7

Observation af49630d-8532-43b9-867c-ebc8731dcfdc · outbound

This paper cites Language models are few-shot learners.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:14668a3f446f0bdc9e013f4f94ca8109297ea506d78bf7f8e897d63d30db8a17

Observation 4e4c3c86-497f-4d38-8dde-9db7334448fd · outbound

This paper cites Can multimodal large language models truly perform multimodal in-context learning? In2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), pages 6000–6010.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Can multimodal large language models truly perform multimodal in-context learning? In2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), pages 6000–6010

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:d18d4dc263e6b5d6934168e74e3321b976616d60e468ef6c9ecab6c991bcf318

Observation 7d7656c1-c31b-4bf7-9a25-fdb16806f88a · outbound

This paper cites Levels of processing: A framework for memory research.Journal of verbal learning and verbal behavior, 1972.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Levels of processing: A framework for memory research.Journal of verbal learning and verbal behavior, 1972

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:367281b0dd3d010b8f6de01e287b060be0d11de91fac37a08bfe0b13742e1dde

Observation 8b2d24ca-37a9-4fb6-aa3e-e5a509426dfb · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:3147663540a431774674fc56a0fa47bc24b50449a9a347699c9bb1d5ade45761

Observation cde3d691-1321-435b-a7c4-5c13014a06ae · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:d3400ac37cb583a922b0685d22c1033c606e6b8c6fc3b017d0231e952b27d0e5

Observation 22658d71-2463-4b6f-95a0-a2c44cdcec12 · outbound

This paper cites Distributed hierarchical processing in the primate cerebral cortex.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Distributed hierarchical processing in the primate cerebral cortex

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:50b47842089f89424ba46019fb1c26406fb4674c7e2f4c4754bcda5c54fb0458

Observation ea590123-6760-4a52-851c-594a4a392e70 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 2023.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:791a0bcb490f52ab616e5dbc6454609b912fb84c1ea2db26134b9abed84c32b4

Observation 175285c7-56dd-4630-b3aa-8ba9c27fa61a · outbound

This paper cites Processing capacity defined by relational complexity: Implications for comparative, developmental, and cognitive psychology.Behavioral and brain sciences, 1998.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Processing capacity defined by relational complexity: Implications for comparative, developmental, and cognitive psychology.Behavioral and brain sciences, 1998

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:a4a4fdfd79a529f714991c01c2d70c9877c0a0d349018b4f7c22e7ec88976e74

Observation 9e3b0ec1-8fb0-4d8f-bc0b-86754c7b0ac4 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Clipscore: A reference-free evaluation metric for image captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:064125b223cbc15a627211318761e98b11a6c7d0f6995e734c02d95bff3b34b2

Observation 0829e7d9-7d9b-45d9-ae28-c5f2f7dca782 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:ad75d4c14d097b67997dc97cf7f59ec625511c9cc61301b956fd2382f9a27e9d

Observation 50cf13b9-172d-4447-a5c4-71bed2421ef8 · outbound

This paper cites Detect anything via next point prediction, 2025.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Detect anything via next point prediction, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:94989c0c67bae31bc07c42110b8f48d8076a27955a868979dad249c571128b42

Observation cb655419-7889-4d8a-8c6b-a8cf82f86e78 · outbound

This paper cites Generating images with multimodal language models.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Generating images with multimodal language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:7d2a7208fcbb265a9cf2862777b7bc90fae061803159fdf0357734464a12e215

Observation c4e35e50-9b82-4d56-8af8-771747724d36 · outbound

This paper cites What matters when building vision-language models?NeurIPS, 2024.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy What matters when building vision-language models?NeurIPS, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:32f9247a2c805f9d27de26876613c5e3607253576eb02df164126ff1deff00df

Observation 97b6a8ae-8cff-4efe-98ed-4afc998a3148 · outbound

This paper cites Mimic-it: Multi-modal in-context instruction tuning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Mimic-it: Multi-modal in-context instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:2d727f3ab890a49ec5296d11896be59d339ac9b03d088e1cacfc99d6c1480917

Observation 3646c887-112f-4be0-81c2-320e1a1840ae · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning.T-PAMI, 2025.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Otter: A multi-modal model with in-context instruction tuning.T-PAMI, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:4f2eab106a463d70242e92e38c03b68d7f64c3e1f484bd3d1904765e3a82cff8

Observation 550ffadb-107e-407b-b96e-bdf27594fbd8 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:19dd679cdfd174c1ae617864215c7ddabfbff5d6fb0327846192281aa7068a38

Observation dc169f5e-f476-49e4-80e2-ffd041c20d23 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:9d49b2cc55400af68c52c061b477fe7ecdc4ae428adc5f5cf0dcb8571558df51

Observation f23a9da6-1c21-4915-937b-c76d173fd673 · outbound

This paper cites M2iv: Towards efficient and fine-grained multimodal in-context learning in large vision-language models.arXiv e-prints, pages arXiv–2504,.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy M2iv: Towards efficient and fine-grained multimodal in-context learning in large vision-language models.arXiv e-prints, pages arXiv–2504,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:d301a72fe98a9fc66377f2557bf10fdd506d3e98d2f1bd234a78bab1301bfc5e

Observation cf74c5d0-24db-4404-a096-300d1c350b18 · outbound

This paper cites Visualcloze: A universal image generation framework via visual in-context learning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Visualcloze: A universal image generation framework via visual in-context learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:fae3f9a4dfa996804cecd238d268e85068a53516352aa32ded717d20157c5bc1

Observation ac02e138-4057-49c5-99f7-f192d580cac3 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:252cbeeaabf5aec23254f146a19a3dea7b93db4c3369e5d34fdbbd94212e4549

Observation 98a085e4-e67c-4371-b71f-e5a00d8faa15 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV, 2024.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Mmbench: Is your multi-modal model an all-around player? InECCV, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:263c5b43b0dc13294b089e3ca868eb50fbdcbf2c48a50fd7254c81ec8a4c6db3

Observation 1ede21f4-6e36-41cd-b56e-703fbd024205 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy The flan collection: Designing data and methods for effective instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:710dc6a53a375f0acb81aef3429d962e6c8bbe56fcbb90c86813f7696460f9d4

Observation 821d1b88-4924-4ece-ad05-93aad78325b6 · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:7ff76498e1a11bf145a0926696fdc0541ed691d9e55dd678019c4a1aa9baa752

Observation ad177161-bbcc-4adb-bdac-f0a44642bd0e · outbound

This paper cites Hpsv3: Towards wide-spectrum human preference score.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Hpsv3: Towards wide-spectrum human preference score

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:35960b2c444d16eb1257a7c7a81f74907de7698c34acfd3967140e524f13dc0a

Observation dd10b1a6-d4a9-4cf0-a7c0-b93cbb99a353 · outbound

This paper cites In-context Learning and Induction Heads.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy In-context Learning and Induction Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:5f5f7d32b2fa5201bc355c72d8cd4291aa7de9d885d894d81dd57f2f46c04d86

Observation bf778a14-ca72-4c44-b9ec-0122a1de2d74 · outbound

This paper cites What factors affect multi-modal in-context learning? an in-depth exploration.Advances in Neural Information Processing Systems, 2024.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy What factors affect multi-modal in-context learning? an in-depth exploration.Advances in Neural Information Processing Systems, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:ad8f34aa539b24deb99c85bf55350ed8a711781a4ac0499820dddcc1f6ea5358

Observation 9f1ce7c2-8d07-4dcd-9910-bad04b5284ec · outbound

This paper cites Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:0b86851cb48aa6b7eac97c81844eed25280e8b231c5427c179962e95cb85e8a1

Observation bd889ec9-24ed-4b6b-adc3-583c5bed9414 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy SAM 2: Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:d2ef824cad54f251cb6af8882fbe383112745cc641ab1564fb2818de09f9fa25

Observation 9a148b61-a768-420d-9ce5-8937b57faeeb · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.NeurIPS, 2022.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Laion-5b: An open large-scale dataset for training next generation image-text models.NeurIPS, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:7e8a64b8a24868369f655a5d5b5dc68938260bb93779d137efdc369290b1582e

Observation 57a8c9a8-455f-4352-bdb8-57d8ab4a7aec · outbound

This paper cites DINOv3.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy DINOv3

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:241a6dbb1e43684aa8615c291d99082cf3c713395dd2fa2f5d62f1bfd89bac28

Observation 4f99a9fe-14e3-459c-bd7e-33a2a93876ea · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Emu: Generative Pretraining in Multimodality

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:f3c04d6341a81fb341a8510a0f42c1b393304ec0a2b26d5ffc97f7661cdffa53

Observation f4cff1d0-fb9e-453f-8d6d-7aa21ac13e36 · outbound

This paper cites Generative multimodal models are in-context learners.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Generative multimodal models are in-context learners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:ce768391d5a33d582b912f3c887672b47537ef82cc535b1828c00c501d0148b3

Observation d672e8d4-fe0b-4ecb-9978-63ef903fbd5b · outbound

This paper cites Codi-2: In-context interleaved and interactive any-to-any generation.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Codi-2: In-context interleaved and interactive any-to-any generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:69efbc3c499aacd5119f281bb58e49f18831711192867fe9d63c178b0bcca46f

Observation 6619c410-9a38-4de8-8ace-5eb04f19addd · outbound

This paper cites Determinantal point processes for machine learning.stat, 2013.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Determinantal point processes for machine learning.stat, 2013

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:37b6e3261482b6d2f9368b73b0a55b353dae05fe0a2e8d37d13e04fccff25cd0

Observation a469517b-ad58-47a2-ad05-d6706c7ae686 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:0c4ff3f6a21dfd2768452ca47fa0f10b6632ccd66c82324171f580800c2f9499

Observation a0c5863c-4ab9-49d5-b9e8-50aa99d2a707 · outbound

This paper cites Qwen2.5-vl, January 2025.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Qwen2.5-vl, January 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:3d72ff150157617566a145e39079a9a039457d869b532d127ba2f716eece5052

Observation bdbf2c55-925e-46f8-8b79-6b9f46f1df32 · outbound

This paper cites Transformers learn in-context by gradient descent.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Transformers learn in-context by gradient descent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:ce79d40a3494520e2cc30918b378f78cb17f9324a4b5db07e806c4437449e95a

Observation 72272e35-bc7a-4e93-98b7-74e476b2700d · outbound

This paper cites Ovis-U1 Technical Report.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Ovis-U1 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:b38a00b68c9e924d832667ee51b6dee969f68ac2adf1da15138eb06fbad90696

Observation a3fda085-5bfd-478c-b002-20cc6cce6f78 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:1bae4f11d143d308e62a197fc0dd2ad7061ac887db83434239db8f04976f7708

Observation 6383ef36-0853-428f-987d-c3f59c9dba6f · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Emu3: Next-Token Prediction is All You Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:e5453511384e13a10231851a6dff78a03caa578b3e6ad581237947b810baeb50

Observation 73c2d116-16b0-4735-81f8-8977947c9cf2 · outbound

This paper cites Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:59b259ed5d8ee85b575c1562edd754545cc57c70703b04ee3ef32ecb7a5a64e9

Observation 62a2d729-a2de-4553-aff2-31f78d05960f · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:1d96314b59842953d269cfb52ceba90b0e919c6bf43b4988e72d22cf41db5850

Observation f3fd0217-5e95-4c5c-8797-377cb7c6a6dc · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:03953adf4c5a0b4574cfac6fa766b9e5cfeb9e1e711feebc3e9ced607ff9bcdf

Observation 13d7c64c-63e9-42c8-bb53-4e2a0a1a3ad2 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:42f00729f55a2f4bba4362fbd6db3a60733657fe752239cc7dc61a6531fc1a0a

Observation 937d8872-13d5-494c-b4d3-060343842ae0 · outbound

This paper cites Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:db323fbd1de685ac4c6bad8202630569acac27c3701e1a6b51b326bbfc38e9b1

Observation 02ec4337-884a-41a9-a60d-fba9b641e025 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:c8797ebcf720900d94a65d2d0319ba19a5b933bef5a6cbc9defaa6a074600be4

Observation 69aa318b-a947-4cbe-a5c6-4cea9f9c0510 · outbound

This paper cites Recognize anything: A strong image tagging model.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Recognize anything: A strong image tagging model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:e4a6239e160931a35278d9be3f9fac4c8d11ed1818980219a531fe70ce1f4f15

Observation cc5a4c69-2201-4625-91c6-b326f0c9c601 · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:8e32af799e5a21a73aa2531ac5e1e94a594fcb3a2250a6d8971a7590643de34f

Observation 04efb5d9-1e1d-49c3-ad90-63191f9bfbb5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:4cc530b65a7a4c30874bf86f003c93799debc5937ea16e3b1a6ea18bf74a4cba

Observation ced217b6-825d-472f-9eb2-4d0ad76de108 · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:c572cbfeaa80b404eb6ea1deef315b66e2d90d397ab5d1e2602dacb622459d0a

Pith citing papers

Observation c075f7a1-ad10-4891-9855-286a7582edd3 · inbound

ChatImage: Navigating Long-Form LLM Answers through Interactive Images cites this paper.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.411651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:e54eda81ffb110d7b1250cf9e436fcf2ea0e1afa80f8cc505f8f3e5e740d8cc6