Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:31:26.112925Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.21741.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:31:26.112925Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b290fad-7789-49b4-81c4-471ea9f4f717 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6158cf1-a03d-4fa7-82d0-0b2f80fb5002 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Uniter: Universal image-text representa- tion learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19b30b09-1918-4e97-9c92-8b68e719fc18 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cbb349-b375-4e07-979b-b13e43915fe4 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Gonzalez, Ion Stoica, and Eric P
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1c8e38f-9a7a-46f5-88c9-4a4c482968e3 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Instructblip: Towards general-purpose vision-language models with in- struction tuning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20f951cd-4963-4261-a512-14d0829a679b · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces NVLM: Open Frontier-Class Multimodal LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68acd8d9-cf1d-4eb4-b295-f68536314c2a · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Mme: A comprehensive evaluation benchmark for multi- modal large language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a444735a-947f-4d70-b860-00624c37a323 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4015ee-c20f-46d4-ab7a-afa9f0be0a91 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LoRA: Low-Rank Adaptation of Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2704b91-d00a-4a38-96a5-37ccec7f7ff1 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Dvqa: Understanding data visu- alizations via question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ba075e9-3b99-40b1-8010-230270d6c175 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c167ae1-5c2a-43d5-825b-1f490f60785a · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ba7ca47-eae9-479e-a316-ec72aeb044c8 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76aa291a-bc34-46bf-a010-0976aa03728c · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Evaluating Object Hallucination in Large Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede1ea7f-a126-498b-8996-33326c899b63 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd6caea-b43d-4aef-9ce0-1164a14fa382 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Docvqa: A dataset for vqa on document images
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0249a39e-5b92-43f2-8dc0-9867770aaaa5 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Hello gpt-4o
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1136b45-0eb7-4423-b9cf-45ab13f36844 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces [Radford et al., 2021a] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36f9e785-7302-4ec7-a97b-9e954a01647e · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3fab29e7-4af2-4062-99ec-f3d863f35600 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Gemini: A Family of Highly Capable Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a37acac-9ece-45ac-b4fd-db999bb3fd8b · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Attention is all you need
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af5e91bc-0260-4800-a078-d8b76e320927 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Introduction to convolutional neural networks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 755eb3ef-c83d-46cc-b854-83e06da943dc · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601549d3-cebe-4ac9-940d-289cc83354ca · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b1656b-993a-4acb-a062-d5813e053d15 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Florence: A New Foundation Model for Computer Vision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f0bfe4-120a-48ef-aa2e-b2bde586493c · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Easygen: Easing multimodal generation with bidiffuser and llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 832fbd14-48b9-4ce9-a888-c694d51f440d · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems , 36:46595– 46623, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67244016-c8d3-4d9f-9475-48c2196fe6d2 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39d8168-5047-4033-9f8d-a1275fdd40f3 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Ocr-free document understanding trans- former
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e640e00d-ee1a-49ad-9e67-2782314a0c7c · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Honeybee: Locality-enhanced projector for multimodal llm
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc5c1853-1c5a-4793-94e1-21bd4be5d566 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a606dd0-01a3-457e-987e-de605d458991 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Claude 3.5 sonnet
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9781c94c-d3a5-4783-b2cd-f10bc6e367d7 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Language models are few-shot learners
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fabbac7f-b0b7-49b4-a84c-845b2a2754cc · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a858a163-883c-4807-a054-80b7481666a0 · outbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.