Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T13:48:48.661566Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 34 inbound Pith citation observations for arXiv:2309.15112.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T13:48:48.661566Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:01:21.945592Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:29:56.656404Z
100 of 119 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 979ac81f-b2a3-4aa8-b3d3-7d94597bb6b2 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b1e2b18-c99e-4ada-8b85-8b5eaa19535e · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lawrence Zitnick, and Devi Parikh
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3bc5c20c-2329-487a-861b-bd676c3c15ce · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Openflamingo: An open- source framework for training large autoregressive vision- language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a82d99e-7971-4650-b024-a104b019e64b · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Qwen-vl: A frontier large vision-language model with versatile abilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa52d3b9-74f2-42d9-8f55-4e4091b28454 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Baichuan 2: Open large-scale language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d477ac7-8eba-4031-899c-1d6f4b87c53c · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Improving image generation with better captions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d116c809-d25b-49bf-a958-e8d67c2e2052 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Language models are few-shot learners
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf8ca6a2-fff9-41f2-b001-b7ab7b31bb08 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a106cb8b-f31b-4268-ac7e-ebce22e1fe7c · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65bf8951-f540-47ec-aa0a-057bbe4d46ca · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Shikra: Unleashing multimodal llm’s referential dialogue magic
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2e7a491-8067-46e0-9f26-b3213f1cbe3a · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Pali-x: On scaling up a multilingual vision and language model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9cc53447-c692-4081-8737-a6fdae6c30c0 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lawrence Zitnick
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d8b29a9-5d29-4d34-b425-69f40dfe8730 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Pali-3 vision language models: Smaller, faster, stronger
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6628fc6d-209c-40c1-bbd8-a684ded97381 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Pali: A jointly-scaled multilingual language- image model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 240c8446-ec7d-4ee6-a870-8c7a96db8c30 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Gonzalez, Ion Stoica, and Eric P
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a075ae1-0a14-4d76-aeb2-9436be4f728a · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Palm: Scaling language modeling with pathways
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e457e9b-7027-4b12-98ca-a56c94fa0d3b · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Class-balanced loss based on effective number of samples
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ccdc6ea-2aaa-438f-a13f-b83c810ac715 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Instructblip: Towards general- purpose vision-language models with instruction tuning, 9
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d3a463c-b9f5-4c0e-bd37-b83c5f788c78 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Visual dialog
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2cb3bbd5-a823-4be9-8abb-9f23de587b74 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Imagenet: A large-scale hierarchical image database
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 306ad5da-4ad3-4f14-a386-99694741f938 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d9f494b-7264-478c-8a43-ee5784727b85 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0331151-8237-4e7d-9eca-407cad275eaa · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition PaLM-E: An Embodied Multimodal Language Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bdeef09-6e21-47aa-992a-e417f0ef1e98 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Glm: General language model pretraining with autoregressive blank infilling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c44bc926-145a-4d21-8577-4a85ee158fb9 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Eva: Exploring the limits of masked visual represen- tation learning at scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bd4293c-b96c-4cf6-a210-db1a429d4110 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4942f170-a105-4538-9095-eb8df5913349 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aeb05894-a315-4054-8f73-026026ddce2e · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Planting a seed of vision in large language model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4495152-225e-4c73-aaed-e42d3fb5d321 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab9cdb45-d6db-471a-8850-4b4cf2a48598 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42107cc3-71a5-4c4f-8229-5e3c617f39b3 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition LoRA: Low-rank adaptation of large language mod- els
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3226ca6-54ca-4841-a2d1-7071849a604a · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09f4f62f-9329-470b-b15d-d9690885be80 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Reveal: Retrieval-augmented visual- language pre-training with multi-source multimodal knowl- edge memory
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35753301-96f4-4cca-a9fe-002ae1d88fcc · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1ad230f-791d-4908-b51e-dc7f69c30460 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Scaling up visual and vision-language representation learning with noisy text supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7c10019-60aa-4fd4-9b2e-90bbd3f0ef94 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0013b0c1-fb74-406c-8df2-273eca0d701b · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Grounding language models to images for multimodal in- puts and outputs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a6e581ea-b5bf-4ea6-b328-ea90fc000460 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Openassistant conver- sations – democratizing large language model alignment
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9cceac28-acfe-462d-91c8-b50accb0a360 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0974ee3a-f490-44d4-b6eb-13bc4942c818 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Seed-bench: Benchmarking multi- modal llms with generative comprehension
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53728e28-7a43-4baa-b6a9-8c6360441e7d · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Otter: A multi-modal model with in-context instruction tuning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67bebcf0-e3c0-42ab-93c0-b5d0b26c5e44 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 558b4f9b-ca87-4337-956b-b3a34327e36f · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2a9dea2-6d23-4c30-ae80-9c3a6e17766e · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a4a90c5-f97b-4748-bec3-78be7ddde092 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Grounded language-image pre-training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef49e8f3-1738-457c-bec2-fa450ba70519 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lmeye: An interactive perception network for large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c529356-af24-4d2b-a476-bc1d725ceaca · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Visual spatial reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1248191f-170f-480f-9ce9-c4c46bdc728f · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7dddacd5-b221-4ede-87c6-825736dfc500 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Improved Baselines with Visual Instruction Tuning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bc0adf0-50ed-440c-bf6c-a27fddaf0053 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Visual instruction tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation deb96abc-cb5f-4a0a-92fb-4568e4db9f58 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation abb5e0db-e7f9-471f-8b44-807025396787 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MMBench: Is Your Multi-modal Model an All-around Player?
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 483b757c-3531-486e-b552-d69b1503a51d · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ff8e631-72c4-41b7-bf1a-1517384975c4 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 094cfbac-c5c1-4a97-975b-72624e879d10 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Learn to explain: Multimodal rea- soning via thought chains for science question answer- ing
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de83e9c4-7deb-44e2-a114-bb3af54e26bf · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6a51c30-603b-4857-b1ab-59ecb00fced3 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1357f108-2ed9-47a6-8d0c-b207cd649896 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Ocr-vqa: Visual question answering by reading text in images
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 714a7a29-c6bb-4a50-9371-9194d5c3d40c · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Power laws, pareto distributions and zipf’s law
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12b653c1-7f5d-416f-bbe2-d40020490428 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52812451-604e-4259-b71c-827a21b764d5 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Gpt-4 technical report
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f02c3809-f14c-47e4-aa04-c1c94118f8ec · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0fc9a6b4-8aae-4488-867b-2e1d54aae0a7 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Training language models to follow instructions with human feed- back
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3a2911f-3b2b-4a1d-a5c9-9bcdacfcf6e1 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0fd2ac72-cec9-4eb9-a53e-584430c1c59b · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Kosmos-2: Grounding multimodal large language models to the world
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c661bb59-eb20-409e-9679-8dd8da5118eb · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Introducing qwen-7b: Open foundation and human- aligned models (of the state-of-the-arts)
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95025b0e-807f-43d9-abcf-7f0912bd15d1 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Learn- ing transferable visual models from natural language super- vision
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f3897a2-8273-43ee-af0c-2d62ef8fbcad · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Improving language understanding by gen- erative pre-training
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c9e6575-d48c-491e-8772-dd8e9a7a56a9 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26d41905-6dae-45eb-a36f-e7390fefbe9b · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Hierarchical text-conditional image generation with clip latents
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ffd6cc8-9a02-4299-8865-b0c0d8661940 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Zero-shot text-to-image generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa31804e-c4bb-4990-b58e-acccba4a8a2d · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition High-resolution image synthesis with latent diffusion models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01fee63c-ca7c-46e3-9cf4-1104040ff200 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Photorealistic text-to-image diffusion models with deep language understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66fd9a01-8951-4007-b3b0-50a8d000ed2c · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bbedc54-57b5-4c8b-a0a6-e79a7a281bfa · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af4f7762-243c-4e28-823a-00c59e6c54ac · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition A-okvqa: A benchmark for visual question answering using world knowledge
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55885b6a-4cd0-4fd8-aa7f-2329eabc3bdb · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a75387cf-17ef-4b96-ab27-3fbe80b31663 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28a41c36-b47a-42c0-8474-930611cd9e8a · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Textcaps: a dataset for image caption- ing with reading comprehension
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c7632a7-ec81-49e4-8764-a1c8ded7ebd0 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Towards vqa models that can read
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fc8839e-b62a-4758-a16a-a5bd277471ad · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Wit: Wikipedia-based im- age text dataset for multimodal multilingual machine learn- ing
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13b1b070-5143-4668-9232-24ffe6aedb14 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Generative pretraining in mul- timodality
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e58cdb82-0f2b-43e1-b65a-7f6731ad4bd6 · outbound
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d3a0a36-a199-4142-b6a0-16975e15de1a · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Internlm: A multilingual language model with progressively enhanced capabilities
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f680910f-12db-4574-a02e-c04fb73f51a5 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Llama: Open and efficient foundation language mod- els
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 030b4e6a-a13f-45ab-bdb6-60ceca6f301e · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Llama 2: Open foundation and fine-tuned chat models
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d68bd975-71a9-4908-a508-5a9c359f40a6 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Attention is all you need
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76633ef0-68dc-4d2b-b1ae-b20a5ac33c0c · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Vigc: Visual instruction generation and correc- tion
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bacd1d0f-f64f-4c71-b0c3-6135f6b76ea4 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Cogvlm: Visual expert for pretrained language models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed5a36fc-22a4-429b-8d46-d70709c2e554 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65a556fd-8bd7-4d22-b4c9-0cf6d7491027 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c26c6150-1754-4a06-825a-d71e26062f6c · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Chinese clip: Con- trastive vision-language pretraining in chinese
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1872810-2a25-4109-8b6d-c170ea82e671 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Retrieval-augmented multimodal language modeling
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c67a29b-7c48-4e89-a9c4-13ae1b736c24 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition mplug-owl: Modularization empowers 12 large language models with multimodality.arXiv.org
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8156aee6-b513-4fed-bda0-b098f66b8abe · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition A Survey on Multimodal Large Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce49f9a2-afe8-4249-888b-db8cd7fea919 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Scaling autoregressive multi-modal models: Pretraining and instruction tuning
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d5e6de5-c702-4abb-ae5b-0b20d0883147 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition GLM-130b: An open bilingual pre- trained model
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e3faddd-0351-409c-b533-c105526e204d · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa1c1330-c9f8-4663-9329-352b5b999808 · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Glipv2: Unifying localization and vision-language understanding
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aaba5360-a328-4f28-97bb-814f6f1be40a · outbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Opt: Open pre-trained transformer language models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3fc066f-f3ac-4c3b-8ef4-362c3f307064 · inbound
MMBench: Is Your Multi-modal Model an All-around Player? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 41e794b6-71a1-4de2-9f4c-ac9a97b67f59 · inbound
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d37652f5-1113-400d-b4ba-ee967f234b94 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2c9972c-ba2a-4b05-90d0-110dcecc03af · inbound
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc57617f-34ea-4e67-a99d-838e2a7e39d9 · inbound
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d214b145-3aa1-4b97-a10c-813ef3ddd0ec · inbound
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d39d6b85-5063-48f2-8329-4a77b1a1ffbf · inbound
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 900ac875-69ac-4afa-b294-7c739943f0a8 · inbound
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f77d35f8-7e2d-43b0-b412-be587fbc5f5b · inbound
Are We on the Right Way for Evaluating Large Vision-Language Models? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d393e68-5f8b-46e1-81f5-3e0696907aba · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f73a51a-041b-4ea7-ae15-e33f1910efce · inbound
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a9c0f0d-1350-4044-8548-136d7ea41349 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 174
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2159a0b1-d809-435c-9f0e-4c1390f85385 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3461b7a8-4d8d-42da-a4f2-936186663aaa · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa85359e-ae8e-4aa4-8938-f8392019cbaa · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bbac56b-66a6-443a-b7eb-0057d7443e67 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86149ad6-9a2a-4921-a378-1339e2d11331 · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c4b6d2b-307f-4182-91e6-6717ded38838 · inbound
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76dc277-2c99-4e8b-95b7-b1b7602bae7c · inbound
Training-Free Multimodal Large Language Model Orchestration InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af4b3667-a3d1-4b94-ab9e-0f9efcae4f99 · inbound
Training-Free Multimodal Large Language Model Orchestration InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96817b62-3602-4b8e-9c65-4e165c82f11b · inbound
Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aed436a-99e6-4183-9ff0-fdfcb493c47f · inbound
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81a8e0c5-e64d-4f5a-8818-d86ef705a326 · inbound
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e51aab-85da-467e-848f-b9d5f189b9ce · inbound
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6772b53-7dab-4312-a7c2-a4915253a60e · inbound
CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9317b03e-b415-4d34-9062-8c89946060ef · inbound
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d71e3551-5474-4e3b-a7e9-f4a16508c56f · inbound
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f74f03d8-1d1f-4ae4-b059-f053af3e914d · inbound
Linear Scaling Video VLMs for Long Video Understanding InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6ba6cb0-0a00-42dc-9558-7ee5b736da93 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c75803fd-d409-4189-81bb-148dcce7a61c · inbound
Curvature-Guided Mixing for MLLM Adaptation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75f01b8b-f14a-4cc0-891a-00fd6a4f93d3 · inbound
Qwen-Audio-VAE Technical Report InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 183
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd65781c-dcdf-4400-a77b-36a2da1576d7 · inbound
Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d564aa3-ac52-4aef-b627-9e34a4238f1d · inbound
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose? InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edcfe875-9729-4fd4-b2a6-521312c36c1c · inbound
VIG-RL: Learning to Search and Insert for Verified Image Grounding InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.