Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:29.201224Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2411.17794.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:29.201224Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:24:57.169737Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:36:05.014470Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9c1ff511-65d0-40df-b468-58e372740f50 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Flamingo: A visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ebb4e11c-4f2b-4fb0-8171-5f3d23a204d6 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d88c780-105f-482c-bf8a-810709dde470 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28130577-7c02-4337-8cc3-d06c740ee696 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c01332-1ba4-4faf-9734-98138069ce28 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM2 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c78c158-3ead-4c2d-bc7a-a01149fe323d · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1919acb8-38ab-48f9-abdf-90f621f7af2a · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00deef47-f46a-4d61-862a-2d659f2a2e3d · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ffb52276-4dbf-42a5-a8a7-6188c612ccf3 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InstructBLIP: Towards general-purpose vision- language models with instruction tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74bc106e-3051-457d-bd6b-6fdf84bf9da0 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ImageNet: A large-scale hierarchical im- age database
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation acce52c3-5ad4-4f5d-ab4d-d6ec5a633617 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Scaling laws of synthetic images for model training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6958ca54-6c05-485d-bb6d-d0752481924a · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EV A: Exploring the limits of masked visual represen- tation learning at scale
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f028fc6e-4b57-478f-87fb-90e4505b9d39 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e81c1a8-f204-4538-aa07-3588ac598431 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gemini., 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6413d14b-cc83-41c8-b741-b53012ff8f07 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c0226eb-72d7-4783-b93b-32da6202be57 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench-2: Benchmarking Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf354ab2-46c2-4c6d-bda0-fce9785b7fb3 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench: Benchmarking multimodal large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e0f7929-dcc2-4b94-94b7-2719b23accc1 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff023b8e-d03c-49d6-b83c-0884b3184ae7 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46ab11c0-667f-421e-85ff-310e912d9999 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8171e875-5c42-4ad9-b087-14455198dcc0 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e336f4b7-b409-40ee-b9c5-38effed72c9f · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Evaluating object hallucination in large vision-language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa38c0a9-84f6-4197-82be-97650e95f758 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual instruction tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fcbe0d19-002d-48d1-b745-40cae90a2cd5 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMBench: Is your multi-modal model an all-around player? In Proc
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f3b36e6e-ba15-4cde-b70b-14c99e052a24 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? From here to human-level AI
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a8c0e8a-bdf5-41f1-a0cf-5c0bd195d76b · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49d540f-1f35-4ffc-9263-f5213194d74c · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Position: Levels of AGI for operationalizing progress on the path to AGI
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ffeb8b1c-da57-40a3-b4ff-3d21227f552e · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? GPT-4o, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b80c38a8-aee4-464f-a679-dbe8ac3a27d2 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Learn- ing transferable visual models from natural language super- vision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 47900640-6151-4743-9d30-7e73eb74b853 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Zero-shot text-to-image generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2ecd108c-8b53-46b6-bdcb-7417fe87ef9e · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed030b23-9307-49ba-a19a-715009b5ff27 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Link- context learning for multimodal LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aace756d-93a0-47ea-b5ba-7a43b3f73cb4 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Label Studio: Data labeling soft- ware, 2020-2022
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 79b3aa1f-b5ff-434c-9691-f332dc9bf7d9 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7dc8d40-bde3-44b3-b14b-90fd04aab1b4 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a603c976-d34d-4b55-9bf0-0b714c5c1afc · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaMA: Open and Efficient Foundation Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a914c956-324d-478b-ac7e-0e1a1f32ba7d · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Recent advancements in fruit detection and classifica- tion using deep learning techniques
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b55d1ac3-ea59-4cc8-94fb-06298f4b59f8 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Le, Thang Luong, and Golnaz Ghiasi
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation faa5d3f1-5b31-417b-95bf-b4231f2d58ca · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? xGen- MM (formerly BLIP-3): A family of open large multimodal models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf36864-c24d-45ee-b855-05555cb4c46b · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9af856-ff51-4cec-a2a9-398e3f64cbae · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Yi: Open Foundation Models by 01.AI
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f2d605-937b-4203-a357-141d63b74a09 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81cedbb2-41d0-4939-9e2c-3a3ce387579c · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8aa4ff3-da1c-4cf7-81b8-027479da32d1 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Sigmoid loss for language image pre-training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cadb536a-be2d-4ff9-b053-f671129816ee · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad8ed8e-a71d-4e96-8c3e-28895f934dcc · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e0c93ba-4e6b-4e1c-ba84-fb4350a2a774 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? On evaluating ad- versarial robustness of large vision-language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3dfd032-6ed0-46f1-8e59-f7b0b56f4b73 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 84c5baff-60b7-4802-a6df-9bcfe455bf38 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7053f530-aa0c-49e2-b1da-d1771aef3df6 · outbound
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? dall-e-3
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12c5c272-28de-4f42-aea7-07925d6fb9d8 · inbound
Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.