Pith. sign in

Paper Citation Record · LEDGER

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

As of 5 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 34 inbound Pith citation observations for arXiv:2309.15112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.15112 v5

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T13:48:48.661566Z

measured 134 of 134 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:01:21.945592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:29:56.656404Z

Reference resolution

100 of 119 outbound references displayed

  • verified exact17
  • verified fuzzy78
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 979ac81f-b2a3-4aa8-b3d3-7d94597bb6b2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.070679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:cec361c7638279856fe6754724d9416f290c7b0fd51924a4f041d73e6dfcf009

Observation 1b1e2b18-c99e-4ada-8b85-8b5eaa19535e · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.076569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:238e424eef8f111ed7e286ae905daf7f21d0cb51b045816d11204438ddaa65d7

Observation 3bc5c20c-2329-487a-861b-bd676c3c15ce · outbound

This paper cites Openflamingo: An open- source framework for training large autoregressive vision- language models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Openflamingo: An open- source framework for training large autoregressive vision- language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.079472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:3b5b2b40514760a70d1246ea44d2344b813a1a697f6ef6358446127c2312e198

Observation 9a82d99e-7971-4650-b024-a104b019e64b · outbound

This paper cites Qwen-vl: A frontier large vision-language model with versatile abilities.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Qwen-vl: A frontier large vision-language model with versatile abilities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.082437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:167120b1dc4cfa450316c81561e78ee860f0c33ea7b6af782602d37ae3de196d

Observation aa52d3b9-74f2-42d9-8f55-4e4091b28454 · outbound

This paper cites Baichuan 2: Open large-scale language models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Baichuan 2: Open large-scale language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.085492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:5f9d88dea5c1f80eb2e9dbbe5f64b1e46df485c684d56198af44e5ad64408d7c

Observation 5d477ac7-8eba-4031-899c-1d6f4b87c53c · outbound

This paper cites Improving image generation with better captions.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Improving image generation with better captions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.088354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2b635891be088775b499217961fb81955ffefe2d0413187251ec06de033a1a71

Observation d116c809-d25b-49bf-a958-e8d67c2e2052 · outbound

This paper cites Language models are few-shot learners.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Language models are few-shot learners

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.091338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:95164070e42ba0ef5358d4a2a7d5af5ebe6266e10b6290de5098b4cc924cb1ff

Observation bf8ca6a2-fff9-41f2-b001-b7ab7b31bb08 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.094827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:730171548add009e98d6912053de0821bdeb0c39aacda812a207c3a5f70bb241

Observation a106cb8b-f31b-4268-ac7e-ebce22e1fe7c · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.750356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:431da4d817b1a8deb530e57aa6be1a67275a974cc3a4b9f7da101d796cc31d94

Observation 65bf8951-f540-47ec-aa0a-057bbe4d46ca · outbound

This paper cites Shikra: Unleashing multimodal llm’s referential dialogue magic.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Shikra: Unleashing multimodal llm’s referential dialogue magic

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.098339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2ef86f394e13b4db2450be703c868586c4ac03be1383d94c09cc710d3392db26

Observation f2e7a491-8067-46e0-9f26-b3213f1cbe3a · outbound

This paper cites Pali-x: On scaling up a multilingual vision and language model.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Pali-x: On scaling up a multilingual vision and language model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.102232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:b7f9f6ef52b17442d19c4cdb0b9f26da6a8a33d97750b12532a24e52206beab4

Observation 9cc53447-c692-4081-8737-a6fdae6c30c0 · outbound

This paper cites Lawrence Zitnick.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lawrence Zitnick

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.105682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2b7004503d8ecee1badb9224e6e34fb03b374ba4f4b2eb1c4d5e9c31d1b20eda

Observation 2d8b29a9-5d29-4d34-b425-69f40dfe8730 · outbound

This paper cites Pali-3 vision language models: Smaller, faster, stronger.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Pali-3 vision language models: Smaller, faster, stronger

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.109053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:a8b2dd6201d6df4b90747e7e838250adc1dcc34dfe629d65d8a75f10077b9fba

Observation 6628fc6d-209c-40c1-bbd8-a684ded97381 · outbound

This paper cites Pali: A jointly-scaled multilingual language- image model.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Pali: A jointly-scaled multilingual language- image model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.112834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:5d150f9cd61747166af2b117e86a081d563631e2ee29bcf8670ace9be24d7cd9

Observation 240c8446-ec7d-4ee6-a870-8c7a96db8c30 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.116256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:a9d4491a5b378854bd5b23cbc80754bffac1905646d234f4f5a55e6cbbd0583a

Observation 9a075ae1-0a14-4d76-aeb2-9436be4f728a · outbound

This paper cites Palm: Scaling language modeling with pathways.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Palm: Scaling language modeling with pathways

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.119703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:8de83a1ca1bcb96a08081d346bec2d607ab8326e294c6095a1987718f6ca4b8f

Observation 5e457e9b-7027-4b12-98ca-a56c94fa0d3b · outbound

This paper cites Class-balanced loss based on effective number of samples.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Class-balanced loss based on effective number of samples

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.123240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:0bb15b2971d6f47db46599a6043b37727aa302153c80374d343006b5401251ef

Observation 4ccdc6ea-2aaa-438f-a13f-b83c810ac715 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning, 9.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Instructblip: Towards general- purpose vision-language models with instruction tuning, 9

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.126859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:10a7e6fd26991fac01a1c47d0cf575fb38cdb460e68f2703e4121c48a775937a

Observation 9d3a463c-b9f5-4c0e-bd37-b83c5f788c78 · outbound

This paper cites Visual dialog.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Visual dialog

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.129967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:ce38e4115416e0b1f4bd3657137d68bb1b53060af61c0f9e1476b423ec72a0e6

Observation 2cb3bbd5-a823-4be9-8abb-9f23de587b74 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Imagenet: A large-scale hierarchical image database

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.132759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2c31f5143dabd6c31c521e39bc37ad22925672643796c540af35f39e0bdb9c76

Observation 306ad5da-4ad3-4f14-a386-99694741f938 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.135664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:499b0b08f7922992af842f50deeca7540252dafa1f90b56fb09e96edb52f3f8f

Observation 8d9f494b-7264-478c-8a43-ee5784727b85 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.760046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:13d35cab4b1613efa70e9941fc334af86acc4d986416a011033bfca63f7fb805

Observation f0331151-8237-4e7d-9eca-407cad275eaa · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.724699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:4f43d13551e9689e44cbfcb34dcb833698ee62c1d78b27438ecc94ed350c9037

Observation 7bdeef09-6e21-47aa-992a-e417f0ef1e98 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Glm: General language model pretraining with autoregressive blank infilling

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.824363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:77727f0df174e8568e802c99cca10160a507e7cc99acf7b6a58a3e1e47318ace

Observation c44bc926-145a-4d21-8577-4a85ee158fb9 · outbound

This paper cites Eva: Exploring the limits of masked visual represen- tation learning at scale.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Eva: Exploring the limits of masked visual represen- tation learning at scale

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.828653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:76909ccffa2d54a778c7f9915749e128241ae8656e3d591714e55ebec85cf020

Observation 0bd4293c-b96c-4cf6-a210-db1a429d4110 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.799970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:692b0f685164861a1c2a4914d2206692527e09760e473d78589dcfde287f6d3c

Observation 4942f170-a105-4538-9095-eb8df5913349 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T13:48:48.819475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:ae0293d9991c1a7d974aacc1f193d14fe2dab0a6605819126d4de0cff386f5bf

Observation aeb05894-a315-4054-8f73-026026ddce2e · outbound

This paper cites Planting a seed of vision in large language model.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Planting a seed of vision in large language model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.832607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:c76888039b6d9e27afaa0d91a915cb84cdac5b1e43d39a2c66dc0e1e0c71f5a9

Observation d4495152-225e-4c73-aaed-e42d3fb5d321 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.836552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:9d177753e4910cef3afb4a9cb1eb5340059409dbb28d82a43ef7ed6a03945584

Observation ab9cdb45-d6db-471a-8850-4b4cf2a48598 · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.755358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:58b2a3a07fc8cf5362284275263e41b839c9fd22766b7f7de8587bd01c7ae82f

Observation 42107cc3-71a5-4c4f-8229-5e3c617f39b3 · outbound

This paper cites LoRA: Low-rank adaptation of large language mod- els.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition LoRA: Low-rank adaptation of large language mod- els

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.839983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:8a359b8cd7c0b0b370215194636efd6f9041865dcfd77b041523bca4011c3d40

Observation b3226ca6-54ca-4841-a2d1-7071849a604a · outbound

This paper cites BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.789741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2b869a3b613ab2bfe6b967367513ed4940151e984df32ca2e87c0a38c3d7810a

Observation 09f4f62f-9329-470b-b15d-d9690885be80 · outbound

This paper cites Reveal: Retrieval-augmented visual- language pre-training with multi-source multimodal knowl- edge memory.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Reveal: Retrieval-augmented visual- language pre-training with multi-source multimodal knowl- edge memory

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.843446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:48fd94034620a08c3c2512fa40e2d36312c26e8579594a9f10343f92b408fa9e

Observation 35753301-96f4-4cca-a9fe-002ae1d88fcc · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.846622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:f60fcce68b015162eb1717de509f9488e43f6baaed4a7b461505eef1d5d1fbc8

Observation c1ad230f-791d-4908-b51e-dc7f69c30460 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Scaling up visual and vision-language representation learning with noisy text supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.849897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:0bd7742a120b7d80aedf464564a560a8c6b4d3b3225e4eb610b9cf85d36bf55f

Observation c7c10019-60aa-4fd4-9b2e-90bbd3f0ef94 · outbound

This paper cites an unresolved cited work.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:48:48.853045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:c27a7a56d6788a7044627853c09f321769488f226886f4c5ea12d00f88b1310d

Observation 0013b0c1-fb74-406c-8df2-273eca0d701b · outbound

This paper cites Grounding language models to images for multimodal in- puts and outputs.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Grounding language models to images for multimodal in- puts and outputs

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.856390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:87ab653a3d42a93c451a0f50c6468b53a2b5dd4ef9da0685d9dd01342bbb0d4b

Observation a6e581ea-b5bf-4ea6-b328-ea90fc000460 · outbound

This paper cites Openassistant conver- sations – democratizing large language model alignment.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Openassistant conver- sations – democratizing large language model alignment

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.859818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:ec6550c89932ccd1e028a4f1d169177950999a7496add8b1119c67dd0514dc73

Observation 9cceac28-acfe-462d-91c8-b50accb0a360 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.862966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:db2c0624811b2331ee6a34d92d8eca140605cd9dcc07eb6344f32d0ee2e83dd2

Observation 0974ee3a-f490-44d4-b6eb-13bc4942c818 · outbound

This paper cites Seed-bench: Benchmarking multi- modal llms with generative comprehension.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Seed-bench: Benchmarking multi- modal llms with generative comprehension

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.866507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:35d92e1c56d76c8a0606aca9ec2908a038ae28876a52a3c1ea7598640363628b

Observation 53728e28-7a43-4baa-b6a9-8c6360441e7d · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Otter: A multi-modal model with in-context instruction tuning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.869651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:160a23bf37bdc8ccb555dc64b17817eea86259b9629fc583ccbb33ecc2e6e61d

Observation 67bebcf0-e3c0-42ab-93c0-b5d0b26c5e44 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.719613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:078191834af21ec6410f002df155eafac102cd1baabd9471f2fc018d3e2315f7

Observation 558b4f9b-ca87-4337-956b-b3a34327e36f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.873220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:f2ffa627ce8e4e0cf80290f8278b6955cecaac33443f1a4878f1ed81cba9613e

Observation b2a9dea2-6d23-4c30-ae80-9c3a6e17766e · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.740728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:904406ab776565d96233da531656a791c4376000a4fbe98d7fa243a2d9cb7ff0

Observation 0a4a90c5-f97b-4748-bec3-78be7ddde092 · outbound

This paper cites Grounded language-image pre-training.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Grounded language-image pre-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.876451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:c0673be2a22d303d65fee114016df6f6757c9077ebe0e73c76e867892451f007

Observation ef49e8f3-1738-457c-bec2-fa450ba70519 · outbound

This paper cites Lmeye: An interactive perception network for large language models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lmeye: An interactive perception network for large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.880117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:0554cc171cc5ab31314015a2f88c0f1a39a9ca13dfe61ba24794a9cdb7c36d44

Observation 7c529356-af24-4d2b-a476-bc1d725ceaca · outbound

This paper cites Visual spatial reasoning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Visual spatial reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.883041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2854ccca20c35ada5199f4f83e713eb373f9fbeeea920820c196cd1a8c66c283

Observation 1248191f-170f-480f-9ce9-c4c46bdc728f · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.764410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:ddf570b192d0655e2749c690c2526bc6e0d531b6a30fa483ac7295e2b9c052fa

Observation 7dddacd5-b221-4ede-87c6-825736dfc500 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Improved Baselines with Visual Instruction Tuning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.778256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:be98da6c9325c202157e8e39b86c2d435c78361aedbafd75e87a68678679ea6c

Observation 1bc0adf0-50ed-440c-bf6c-a27fddaf0053 · outbound

This paper cites Visual instruction tuning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Visual instruction tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.886464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:8b09c38a8cae0e5fe3ad97407d3108c9703cc1e180ac3ed855a731d3cfb34be6

Observation deb96abc-cb5f-4a0a-92fb-4568e4db9f58 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.889874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:bb695955da00e650bba7b5fd71c240779505a0be3a4b2bdffa16529af4d48189

Observation abb5e0db-e7f9-471f-8b44-807025396787 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MMBench: Is Your Multi-modal Model an All-around Player?

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.805997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:75850f01c0f1999d13a9cfa535092abacb133d0ee2176b8681cccb6ea324114c

Observation 483b757c-3531-486e-b552-d69b1503a51d · outbound

This paper cites Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.893018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:ac2d45f1c6d760501b973918ac3c3b6b57f26fd99f8086883f2f5bd8265ff5d5

Observation 7ff8e631-72c4-41b7-bf1a-1517384975c4 · outbound

This paper cites an unresolved cited work.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:48:48.896093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:af36c51ffe9a549ef3d755c7e0ac843185af217f42d98cab961b45e4d0b3dcc0

Observation 094cfbac-c5c1-4a97-975b-72624e879d10 · outbound

This paper cites Learn to explain: Multimodal rea- soning via thought chains for science question answer- ing.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Learn to explain: Multimodal rea- soning via thought chains for science question answer- ing

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.899868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:cadc25382c753d0591e8006d7dc09fe12b8aaefa91acaa67c69bbab7eaa20acd

Observation de83e9c4-7deb-44e2-a114-bb3af54e26bf · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.735693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:747918b906c0a538fb842448a27646b221627d611b4214f8d51b4490567ec92b

Observation b6a51c30-603b-4857-b1ab-59ecb00fced3 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.903761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:04f0076448022fe0391e236e0c286dcc6c06eae6636e564fe8aab7669d1d2a3d

Observation 1357f108-2ed9-47a6-8d0c-b207cd649896 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Ocr-vqa: Visual question answering by reading text in images

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.907186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:bb8bb7998b3779754102f514869c12e62fdb96247336e446a5974e0400692a5e

Observation 714a7a29-c6bb-4a50-9371-9194d5c3d40c · outbound

This paper cites Power laws, pareto distributions and zipf’s law.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Power laws, pareto distributions and zipf’s law

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.910454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:f0606e8708f24c5d02ee012fa7db15a2f7d70d6b3de7a81a91d9c21c07c3a463

Observation 12b653c1-7f5d-416f-bbe2-d40020490428 · outbound

This paper cites an unresolved cited work.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:48:48.913531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:cc4af32e12d2102497c4b8ff830927eaf9ed47455796122c26b55dc5f00de8a9

Observation 52812451-604e-4259-b71c-827a21b764d5 · outbound

This paper cites Gpt-4 technical report.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Gpt-4 technical report

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.916089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:19b523d5ae40c6907cf7668b27fe8986ccc9a142e30d62db92ed37bfa5096b91

Observation f02c3809-f14c-47e4-aa04-c1c94118f8ec · outbound

This paper cites an unresolved cited work.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:48:48.919192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:a97cca24091e38ec6be9298724c683d652f05298c8876168268c45854142cd86

Observation 0fc9a6b4-8aae-4488-867b-2e1d54aae0a7 · outbound

This paper cites Training language models to follow instructions with human feed- back.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Training language models to follow instructions with human feed- back

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.922532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:0f9705e18955afdcbc9b520774778dca6eeb189dd6673be9a3961fcad06ee8d6

Observation c3a2911f-3b2b-4a1d-a5c9-9bcdacfcf6e1 · outbound

This paper cites The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.925958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:d407ff251c473021fc2618be5a7ffe3accd02b0c4236b6ecc5d2911f866e070b

Observation 0fd2ac72-cec9-4eb9-a53e-584430c1c59b · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Kosmos-2: Grounding multimodal large language models to the world

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.929741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:968bcc22119c9984d3b0ec56b713400be2a08e067335295ea5e7477a412f0a5b

Observation c661bb59-eb20-409e-9679-8dd8da5118eb · outbound

This paper cites Introducing qwen-7b: Open foundation and human- aligned models (of the state-of-the-arts).

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Introducing qwen-7b: Open foundation and human- aligned models (of the state-of-the-arts)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.932885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:20f4f01f3a2e98e93fbeff5687bbb21bf07cdfde33ce7d132706526663ad8d1b

Observation 95025b0e-807f-43d9-abcf-7f0912bd15d1 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Learn- ing transferable visual models from natural language super- vision

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.936067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:3cbb6f69cee330d5c7366b8202723cb7d4793f8b3471f552ed25e3e4e5647028

Observation 6f3897a2-8273-43ee-af0c-2d62ef8fbcad · outbound

This paper cites Improving language understanding by gen- erative pre-training.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Improving language understanding by gen- erative pre-training

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.939311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:1fab65bc31391d5e09ecdedb919d7baa38080ca542d925094a79119a11dc7df2

Observation 5c9e6575-d48c-491e-8772-dd8e9a7a56a9 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.942400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:fc3546a78239cc1493f4eaf63d3a2179500995178269f16db9e8f048e3f4b1de

Observation 26d41905-6dae-45eb-a36f-e7390fefbe9b · outbound

This paper cites Hierarchical text-conditional image generation with clip latents.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Hierarchical text-conditional image generation with clip latents

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.945540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:87d606470c3da2f9916c01e0a7e55f1d7a13865fe4d867986d1b62b4575dae8d

Observation 6ffd6cc8-9a02-4299-8865-b0c0d8661940 · outbound

This paper cites Zero-shot text-to-image generation.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Zero-shot text-to-image generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.948395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:118de1b378593d49ac2ce7dc497b802e531704286363e86945f8b4c455cb8038

Observation aa31804e-c4bb-4990-b58e-acccba4a8a2d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition High-resolution image synthesis with latent diffusion models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.951146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:455be65b402fa97b640bdeb79798660974ef252f5109a099b86aa75c1ac9426b

Observation 01fee63c-ca7c-46e3-9cf4-1104040ff200 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Photorealistic text-to-image diffusion models with deep language understanding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.954258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:d14b4bc96147e72c4a031a2ff63c20bfa6d462e06f9d59ebb24d6ed0bdc3fcf9

Observation 66fd9a01-8951-4007-b3b0-50a8d000ed2c · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.957596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:ccdc7c60bdb7737d311bfde44aefe39e7eee3d3e3c4423524d3fc28e46bb343c

Observation 8bbedc54-57b5-4c8b-a0a6-e79a7a281bfa · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.773235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:1e1664064ab922a92ca419f3d3db126b12fa3790872318adb40ec3924663919a

Observation af4f7762-243c-4e28-823a-00c59e6c54ac · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition A-okvqa: A benchmark for visual question answering using world knowledge

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.960728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:462c31da288a1c96551b47773b2c4ac33880c55efcba925b733ccbf6513a4fdc

Observation 55885b6a-4cd0-4fd8-aa7f-2329eabc3bdb · outbound

This paper cites TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.783620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:69d709bde5a490d728320437056c0c1c4c97e70ae4a60799f4ba7b2ffd89589d

Observation a75387cf-17ef-4b96-ab27-3fbe80b31663 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.963747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:5372535da11986ff1c62ba1287f08ae0a61b3ba62ea6f685e3ebe502df0127f2

Observation 28a41c36-b47a-42c0-8474-930611cd9e8a · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Textcaps: a dataset for image caption- ing with reading comprehension

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.967061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:7269d934c4ae8e98e90e9c351ac0fc9850334f569d1da9d19d9c4e1a8261e42b

Observation 0c7632a7-ec81-49e4-8764-a1c8ded7ebd0 · outbound

This paper cites Towards vqa models that can read.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Towards vqa models that can read

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.969975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:432332b916b9fdd5ae531d6c368c16abe2f5190c38fe9c10bb8938893be92fe5

Observation 4fc8839e-b62a-4758-a16a-a5bd277471ad · outbound

This paper cites Wit: Wikipedia-based im- age text dataset for multimodal multilingual machine learn- ing.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Wit: Wikipedia-based im- age text dataset for multimodal multilingual machine learn- ing

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.972950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:b0152779f6a475c8bb392899a8cdb5bd7b83e40525ae9ea63d808849595c0dc2

Observation 13b1b070-5143-4668-9232-24ffe6aedb14 · outbound

This paper cites Generative pretraining in mul- timodality.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Generative pretraining in mul- timodality

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.975697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:4491cced7f50da236ccc65ac188364f8307e36c626073650a9b9626b69dd5ae7

Observation e58cdb82-0f2b-43e1-b65a-7f6731ad4bd6 · outbound

This paper cites Hashimoto.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Hashimoto

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.978299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:3408939caf2dd41099203d81114bc5e945fa4d75abb370e77e2d52c4c6a97f54

Observation 1d3a0a36-a199-4142-b6a0-16975e15de1a · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Internlm: A multilingual language model with progressively enhanced capabilities

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.981540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:98b0367f98a13ae964f1fd63b512bd904d382d7cc748dd5bcd43e740d0a80b50

Observation f680910f-12db-4574-a02e-c04fb73f51a5 · outbound

This paper cites Llama: Open and efficient foundation language mod- els.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Llama: Open and efficient foundation language mod- els

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.984344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:fd6183adf4f3333d44c64d3c4d92133776ae58b832488f6a3c13484e7e3d485d

Observation 030b4e6a-a13f-45ab-bdb6-60ceca6f301e · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Llama 2: Open foundation and fine-tuned chat models

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.987669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:da75799855593edaedd32225e19a6c4898ea4849233bda332aa5d188872b5ee5

Observation d68bd975-71a9-4908-a508-5a9c359f40a6 · outbound

This paper cites Attention is all you need.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Attention is all you need

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.990812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:0ecc296f86260fc4f749e0945335d5d4b8d28709dcd13df896c26de57349950b

Observation 76633ef0-68dc-4d2b-b1ae-b20a5ac33c0c · outbound

This paper cites Vigc: Visual instruction generation and correc- tion.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Vigc: Visual instruction generation and correc- tion

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.993773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:c757fbd36c5fe4e0112eb98861b2cc9102cc86ca7fb1ea35b51085218477c265

Observation bacd1d0f-f64f-4c71-b0c3-6135f6b76ea4 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Cogvlm: Visual expert for pretrained language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.996855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:16fee1ad0a6c812e55fbc155d12e7b34fb8f9100801afff34c22f50c77b4462d

Observation ed5a36fc-22a4-429b-8d46-d70709c2e554 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.768680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:c2a3b3244e7c1aaf879d88a13cfb08d9a65c2b52282b8c99973701b110f61398

Observation 65a556fd-8bd7-4d22-b4c9-0cf6d7491027 · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:48.999879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:4527c84cc44e5f492ad5ff9173e21164c5ef476fb185219a70814aba855bdfd4

Observation c26c6150-1754-4a06-825a-d71e26062f6c · outbound

This paper cites Chinese clip: Con- trastive vision-language pretraining in chinese.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Chinese clip: Con- trastive vision-language pretraining in chinese

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.002905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:39c459625c2c6affcd6f0a42ac193e1f057157b2ae8724560e7937d5477e82cd

Observation a1872810-2a25-4109-8b6d-c170ea82e671 · outbound

This paper cites Retrieval-augmented multimodal language modeling.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Retrieval-augmented multimodal language modeling

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.006024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:c5997066cf0e9490f40688fb1e0cd0f97ab16936e498ac39cece2584b8b309ce

Observation 6c67a29b-7c48-4e89-a9c4-13ae1b736c24 · outbound

This paper cites mplug-owl: Modularization empowers 12 large language models with multimodality.arXiv.org.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition mplug-owl: Modularization empowers 12 large language models with multimodality.arXiv.org

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.009201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:2c5d623092b3c11367e1bc1df318e50f065265fb8f6c11186d2621553c7a582d

Observation 8156aee6-b513-4fed-bda0-b098f66b8abe · outbound

This paper cites A Survey on Multimodal Large Language Models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition A Survey on Multimodal Large Language Models

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:48:48.794372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:63054be6bbe3de0c1241f04268ff55e3e8f1af5c2e6f5670f3f83cd3c71237ed

Observation ce49f9a2-afe8-4249-888b-db8cd7fea919 · outbound

This paper cites Scaling autoregressive multi-modal models: Pretraining and instruction tuning.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Scaling autoregressive multi-modal models: Pretraining and instruction tuning

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.012635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:fca6e9071c4e99f1b3c28e5d0c3ceef31ae5749e2b748eef133367a1843cf442

Observation 8d5e6de5-c702-4abb-ae5b-0b20d0883147 · outbound

This paper cites GLM-130b: An open bilingual pre- trained model.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition GLM-130b: An open bilingual pre- trained model

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.015632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:d08ce7f6c66bb793db01ab4c8eb6a2ef00f6998aff23445bae4529be8c744424

Observation 5e3faddd-0351-409c-b533-c105526e204d · outbound

This paper cites What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.813327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:eb3aeda9a635214a415bee58e88ba8f1232df611f547cc0fe26b961d5724b0ac

Observation fa1c1330-c9f8-4663-9329-352b5b999808 · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Glipv2: Unifying localization and vision-language understanding

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.018756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:289f069daa5f122692a3ff22c751d3107052a28720d514cd58b6732292e60b9f

Observation aaba5360-a328-4f28-97bb-814f6f1be40a · outbound

This paper cites Opt: Open pre-trained transformer language models.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Opt: Open pre-trained transformer language models

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:48:49.021637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:a93ef105d82c3ff7aff531873201ec15bb3e2ad21ff1a959a1db2884ae1ef8cd

Pith citing papers

Observation b3fc066f-f3ac-4c3b-8ef4-362c3f307064 · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:5dbc9c30fdb6a0e74ed2cd3058fc1b7f2149f8c323558622280b55876f9213d0

Observation 41e794b6-71a1-4de2-9f4c-ac9a97b67f59 · inbound

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions cites this paper.

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:08:12.727773Z digest=sha256:b6485eae0c4ffd2f649816c559a07e70da00f7583483ab379d3ada5505671587

Observation d37652f5-1113-400d-b4ba-ee967f234b94 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:037a5ca35e3a2799ed4d32455bc3c16d972392327d101e7f5ce44c64a3492aec

Observation b2c9972c-ba2a-4b05-90d0-110dcecc03af · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:f8291c4919fccb91de2b9cc475633658a72cb70a1f6944afd5abfc0d3c5d185f

Observation cc57617f-34ea-4e67-a99d-838e2a7e39d9 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:42cea50a3671daa0ee7b4076841f6c486f721d3e45a2681fbe8fd75c6f7bcaed

Observation d214b145-3aa1-4b97-a10c-813ef3ddd0ec · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:b44b989813c2651cb8060bf059b603d8f7be5623798e75ce272c71d1e609825d

Observation d39d6b85-5063-48f2-8329-4a77b1a1ffbf · inbound

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition cites this paper.

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T03:33:50.806971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T03:31:58.848180Z digest=sha256:38df0ebc8754d9c14c41717927967548db8dd984a5ad24415342713c5c0a3f6e

Observation 900ac875-69ac-4afa-b294-7c739943f0a8 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:0bd7756d18699bf8a0bddb7cc88a9a2f081c0bbbeb6f23fe755b18996cabc44e

Observation f77d35f8-7e2d-43b0-b412-be587fbc5f5b · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:3e8a890d119c878051ad7de44ef93a60a18b4a637ec22c807ccd5dec3e0909ac

Observation 7d393e68-5f8b-46e1-81f5-3e0696907aba · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:2649d34271a1194c0a91bf097442940f36cbadd0577b5166cf1ba90ce65d96b3

Observation 4f73a51a-041b-4ea7-ae15-e33f1910efce · inbound

VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model cites this paper.

VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:38:37.332882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:37:12.985780Z digest=sha256:68b3e14b47c6ed5b663682d24db3ad3861658eec73b08cd2b5f73938c72dfa16

Observation 1a9c0f0d-1350-4044-8548-136d7ea41349 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:2ec474e07b1f9a539f67e38916973982f8927c8e1e25d19ddcf0238b8a7f0282

Observation 2159a0b1-d809-435c-9f0e-4c1390f85385 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 266

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.575164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:6e83e013de8a345f3e284f875fb9963ade93847eb65c2b64617b3bfdba5e34fd

Observation 3461b7a8-4d8d-42da-a4f2-936186663aaa · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:028c79ae11291c62c838c79fbb27c517b4490994e4523f177275bb878e1d06dc

Observation fa85359e-ae8e-4aa4-8938-f8392019cbaa · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:08:19.721250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:7600d200bb372fc120ff1122edcb010429f794fc0f9d542b912a90bcda54ff11

Observation 1bbac56b-66a6-443a-b7eb-0057d7443e67 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:3f4f0ad2fb3560011f3dc6b77fada4b1eff13376b62c4a6ba59b035db454e5a2

Observation 86149ad6-9a2a-4921-a378-1339e2d11331 · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:b8eb0d364dd2f1decb41bf4584f35e0c65485abd264e225c0739c34c2ed96344

Observation 0c4b6d2b-307f-4182-91e6-6717ded38838 · inbound

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs cites this paper.

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:21.945592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:21.945592Z digest=sha256:ae19be01f3b5d68423bb3402c1909561a392bbe66ea5b370a97c8f9ada2e6e1d

Observation b76dc277-2c99-4e8b-95b7-b1b7602bae7c · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:12:54.088430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:7b93b7bc63c310a784d439c33a6b796d758048451e2bb74595cc3a38e968e942

Observation af4b3667-a3d1-4b94-ab9e-0f9efcae4f99 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:05:30.607459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:bf2c2d100236320c8d02a2607d26019f40602dfab5de09716c0d8ee0d38d27f4

Observation 96817b62-3602-4b8e-9c65-4e165c82f11b · inbound

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration cites this paper.

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:17:18.089212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:17:18.089212Z digest=sha256:bef020a7395c0ebb586b2364f0efce7bcee403f94565b00df8775935b0f613bb

Observation 0aed436a-99e6-4183-9ff0-fdfcb493c47f · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:27eb381c1fe2bbe874f933d0406cc6260432bb422bcb0e878beb699c89143411

Observation 81a8e0c5-e64d-4f5a-8818-d86ef705a326 · inbound

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models cites this paper.

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.228898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.228898Z digest=sha256:a2d8b252698eaffa200518bb88317e14a823d363a312cb264d311587f7001ccc

Observation a0e51aab-85da-467e-848f-b9d5f189b9ce · inbound

Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models cites this paper.

Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:58.668589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:45:58.668589Z digest=sha256:65fe151529efc70d9f9ee88a196b222f2e4df6e038809637d78e03206d966191

Observation a6772b53-7dab-4312-a7c2-a4915253a60e · inbound

CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning cites this paper.

CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:18:38.658857Z digest=sha256:e040eb135a7fd89821ef1b1bfc74197ffcc2fcee3a528ec236137116f73b431d

Observation 9317b03e-b415-4d34-9062-8c89946060ef · inbound

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing cites this paper.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:9c2ae6f58b00652107e7782a0c5a10e60d9a93b0d1da894896dd64c9c370f958

Observation d71e3551-5474-4e3b-a7e9-f4a16508c56f · inbound

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring cites this paper.

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:34:40.560924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:28:01.889147Z digest=sha256:62512773215a8fe8049d9ce02ca972574c35c5c265bb2abc036426ea7f8c52a9

Observation f74f03d8-1d1f-4ae4-b059-f053af3e914d · inbound

Linear Scaling Video VLMs for Long Video Understanding cites this paper.

Linear Scaling Video VLMs for Long Video Understanding InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:02:46.234160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:00:11.246232Z digest=sha256:84a4987e0136fd328bc45f4a99087f55b9420b41a323329ed47fd1fe5b9488c5

Observation d6ba6cb0-0a00-42dc-9558-7ee5b736da93 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 170

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:18:43.799277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:9241baa97c7f672d42796dbca07606942af88b400e9dd03ff8263e0b64172fa0

Observation c75803fd-d409-4189-81bb-148dcce7a61c · inbound

Curvature-Guided Mixing for MLLM Adaptation cites this paper.

Curvature-Guided Mixing for MLLM Adaptation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:29:56.658035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:42:56.323034Z digest=sha256:bde89b51242bb7b40edf5963f8509e497412758a64357156ddf745de1c162419

Observation 75f01b8b-f14a-4cc0-891a-00fd6a4f93d3 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 183

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:b24ad6e40ac871926f7f30e44812d6641b0817b2fca81a33b864baf96e92bedb

Observation dd65781c-dcdf-4400-a77b-36a2da1576d7 · inbound

Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model cites this paper.

Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T13:51:21.256489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:51:21.256489Z digest=sha256:047476e046e72ba6befc51283e982842cddef11795806ec8e9dbbf8fc2b3df45

Observation 0d564aa3-ac52-4aef-b627-9e34a4238f1d · inbound

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose? cites this paper.

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T10:20:57.609350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:20:57.609350Z digest=sha256:533635b01a6d3061a6300663c0ec20d6252309d6a1238a8f6f61d07e7f076d2b

Observation edcfe875-9729-4fd4-b2a6-521312c36c1c · inbound

VIG-RL: Learning to Search and Insert for Verified Image Grounding cites this paper.

VIG-RL: Learning to Search and Insert for Verified Image Grounding InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T19:16:58.483265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T19:16:58.483265Z digest=sha256:16a54735c7c991f53b7285889a2437e028b24c3815ea9395965898b88cb2c03f