Pith. sign in

Paper Citation Record · LEDGER

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2411.17794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17794 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:29.201224Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:24:57.169737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:36:05.014470Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c1ff511-65d0-40df-b468-58e372740f50 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Flamingo: A visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:30.011651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.007620Z digest=sha256:976c5ab5af54e07c7cbe0e5a12a0b678c3206eb561469b488b70c046fe52478c

Observation ebb4e11c-4f2b-4fb0-8171-5f3d23a204d6 · outbound

This paper cites Qwen Technical Report.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.012130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.012130Z digest=sha256:cac07a8e3234f660e7942ac84d8541ac284b70d854db282f13ec172686edeab3

Observation 4d88c780-105f-482c-bf8a-810709dde470 · outbound

This paper cites Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.999362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.016492Z digest=sha256:365ce4fcc23e83dea902fedd1622289354898311425d81da09cd7e9ede4d0988

Observation 28130577-7c02-4337-8cc3-d06c740ee696 · outbound

This paper cites Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.020551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.020551Z digest=sha256:0cc40878a418dd27827d0d12a9ac2a88590ad43e091d8eab3e7a78e80f45dbba

Observation 35c01332-1ba4-4faf-9734-98138069ce28 · outbound

This paper cites InternLM2 Technical Report.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.024930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.024930Z digest=sha256:e8bd662d5af16d9533a377e2ff1aae6fefba79cd88f6b6c81a4f7f27f2ffd124

Observation 1c78c158-3ead-4c2d-bc7a-a01149fe323d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.029279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.029279Z digest=sha256:01a6614e8bde5e00300f5e5f80f5883c991533ec8e1f327b4561552bc475ef26

Observation 1919acb8-38ab-48f9-abdf-90f621f7af2a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.033836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.033836Z digest=sha256:773a569b2ad81fcbcf947c9a420f7f912f9337aa8215583168853dca9da34224

Observation 00deef47-f46a-4d61-862a-2d659f2a2e3d · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.986905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.038639Z digest=sha256:6c0bea1a4337152a1b71ef5f1325653a54a4dc6c05f916a58668c4575b7c37a3

Observation ffb52276-4dbf-42a5-a8a7-6188c612ccf3 · outbound

This paper cites InstructBLIP: Towards general-purpose vision- language models with instruction tuning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InstructBLIP: Towards general-purpose vision- language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.974761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.042409Z digest=sha256:96b5fa15881291e0bf7f6c0a7aa2f02ab345c1c929386eedaffefeb6e367a3f4

Observation 74bc106e-3051-457d-bd6b-6fdf84bf9da0 · outbound

This paper cites ImageNet: A large-scale hierarchical im- age database.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ImageNet: A large-scale hierarchical im- age database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.962422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.046295Z digest=sha256:af92820d594a08f636ecf3148e3e1d6977f33a64f8d57e790036bd586b622af0

Observation acce52c3-5ad4-4f5d-ab4d-d6ec5a633617 · outbound

This paper cites Scaling laws of synthetic images for model training.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Scaling laws of synthetic images for model training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.948764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.049953Z digest=sha256:110f2d6257bd8dc588e9d149c6914842ddcae4d557d5eafe41510b8e8f18adfb

Observation 6958ca54-6c05-485d-bb6d-d0752481924a · outbound

This paper cites EV A: Exploring the limits of masked visual represen- tation learning at scale.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EV A: Exploring the limits of masked visual represen- tation learning at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.936580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.053556Z digest=sha256:30e6a8b1098736cb2d71abb4cb2e8be9044ac59701945a701ca0f08898c46044

Observation f028fc6e-4b57-478f-87fb-90e4505b9d39 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.057433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.057433Z digest=sha256:9bc53515ce1abdbe90be0fe1ae6db4661cea764fdcf3abaec6f352bd8f0cbe79

Observation 8e81c1a8-f204-4538-aa07-3588ac598431 · outbound

This paper cites Gemini., 2023.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gemini., 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.924269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.061609Z digest=sha256:5b0eb126139fe5229038cda9a313b4014044aa658e348dc95f6e5587ec487df6

Observation 6413d14b-cc83-41c8-b741-b53012ff8f07 · outbound

This paper cites Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.912168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.065609Z digest=sha256:071619d158ce18df9adbedb2858a9f762f1d0979d7a0ca59d71b066b54f66f30

Observation 3c0226eb-72d7-4783-b93b-32da6202be57 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.069305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.069305Z digest=sha256:ad9aa5e01ad4c400821516998d9becbc0d7c6132f5c0ee0b92a2cf4c1d1d0641

Observation bf354ab2-46c2-4c6d-bda0-fce9785b7fb3 · outbound

This paper cites SEED-Bench: Benchmarking multimodal large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench: Benchmarking multimodal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.898939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.073147Z digest=sha256:24f65844893f1b4a045d9190fc11ae4fd9aa151041590d4c50555c0e20a81998

Observation 9e0f7929-dcc2-4b94-94b7-2719b23accc1 · outbound

This paper cites Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.885960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.076856Z digest=sha256:a8d78460793f505cc4e92007bd7a59dd886dea4e84db0ba50f37925bb4a45fa2

Observation ff023b8e-d03c-49d6-b83c-0884b3184ae7 · outbound

This paper cites LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.873904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.080548Z digest=sha256:5b75e8eb8b243a9b9f27fff05546695416c955d6750ba8ced5c5002508b404ed

Observation 46ab11c0-667f-421e-85ff-310e912d9999 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.861617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.084349Z digest=sha256:13f5ee0d6b8096f68c448302cc8ca9eeadad34d2bd3588a581226fbf0102de9a

Observation 8171e875-5c42-4ad9-b087-14455198dcc0 · outbound

This paper cites FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.848224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.088284Z digest=sha256:af1d2d48bd9783bd042ed47eaba640a5ce5fc5d6eae5a4aefe55a1ddc81457a6

Observation e336f4b7-b409-40ee-b9c5-38effed72c9f · outbound

This paper cites Evaluating object hallucination in large vision-language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Evaluating object hallucination in large vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.834662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.092028Z digest=sha256:db2aed0e99083c8676962fbae8d487230378d3e29cb9c7bf901f6054f7d41ca0

Observation aa38c0a9-84f6-4197-82be-97650e95f758 · outbound

This paper cites Visual instruction tuning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.725556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.096069Z digest=sha256:43d19318f59ea918c3769a1e9eee1ab92ca1d88242131cc355719611c890a593

Observation fcbe0d19-002d-48d1-b745-40cae90a2cd5 · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? In Proc.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMBench: Is your multi-modal model an all-around player? In Proc

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.712165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.099764Z digest=sha256:9a7cf39e684edaa58de65cfcc3d01fcf0e4e605cba5c6edf5a24924a28d51322

Observation f3b36e6e-ba15-4cde-b70b-14c99e052a24 · outbound

This paper cites From here to human-level AI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? From here to human-level AI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.699235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.104119Z digest=sha256:4a78215efd465c6cfbdbe2a529faac3b1dfda306931bab0f9db1cc8748f9cb84

Observation 1a8c0e8a-bdf5-41f1-a0cf-5c0bd195d76b · outbound

This paper cites Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.107976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.107976Z digest=sha256:5f6096e6dc1fca49d3350d65471ba8a8d2850c8e9e6cffd2b8acfbf1dbcb7b5d

Observation c49d540f-1f35-4ffc-9263-f5213194d74c · outbound

This paper cites Position: Levels of AGI for operationalizing progress on the path to AGI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Position: Levels of AGI for operationalizing progress on the path to AGI

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.687401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.112023Z digest=sha256:7d698c97e48b48ba263a1102abea9daccb322a9f08c0f6451467b4fd80eacac3

Observation ffeb8b1c-da57-40a3-b4ff-3d21227f552e · outbound

This paper cites GPT-4o, 2024.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? GPT-4o, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.675430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.115794Z digest=sha256:018c74b6800f5709ba085a4fe99f913ff0a27ede281648994db75c4ed84c7cbc

Observation b80c38a8-aee4-464f-a679-dbe8ac3a27d2 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Learn- ing transferable visual models from natural language super- vision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.663227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.119566Z digest=sha256:81968af498895a6ec9669ee054ff4bfe65fa1449c9c0ed63fd7b555dd27d6a1e

Observation 47900640-6151-4743-9d30-7e73eb74b853 · outbound

This paper cites Zero-shot text-to-image generation.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Zero-shot text-to-image generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.650633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.123316Z digest=sha256:cd921b4f50c07d5dd3d3181aea87eb7355492e357acd518c7af9be2df99cc2ab

Observation 2ecd108c-8b53-46b6-bdcb-7417fe87ef9e · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.127230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.127230Z digest=sha256:a558e21ec74e7873e849a321e21873e3ea6002aba50ada7e2022a09ba9fac1aa

Observation ed030b23-9307-49ba-a19a-715009b5ff27 · outbound

This paper cites Link- context learning for multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Link- context learning for multimodal LLMs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.638505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.131172Z digest=sha256:c065932b0703f6d702bbfa8b487d818230eb6917b3ee8429947c3d573d6bedc7

Observation aace756d-93a0-47ea-b5ba-7a43b3f73cb4 · outbound

This paper cites Label Studio: Data labeling soft- ware, 2020-2022.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Label Studio: Data labeling soft- ware, 2020-2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.625949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.134990Z digest=sha256:4cd061aa674f149debd8d065358197c64d075f1c4464b924bb3e314825f999b1

Observation 79b3aa1f-b5ff-434c-9691-f332dc9bf7d9 · outbound

This paper cites Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.613670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.138719Z digest=sha256:3eda37d75a045d69cdb22b2655f6ae398f1dcf6b04ca5defb7d0597a6cffec25

Observation c7dc8d40-bde3-44b3-b14b-90fd04aab1b4 · outbound

This paper cites Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.600539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.142487Z digest=sha256:d5b29f44fdb89f459c37e31d97d5b79f67935c6ceb43b56396fd778e1b6ee58a

Observation a603c976-d34d-4b55-9bf0-0b714c5c1afc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.146242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.146242Z digest=sha256:646f12b972de9850dacbfd5b0c8078dc931108aa359c311d4bf4deb46160f9d3

Observation a914c956-324d-478b-ac7e-0e1a1f32ba7d · outbound

This paper cites Recent advancements in fruit detection and classifica- tion using deep learning techniques.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Recent advancements in fruit detection and classifica- tion using deep learning techniques

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.588119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.150418Z digest=sha256:4d8747ae2790e2a72a33393ea4e54fe1a7c079204a2f040741cfcb4549235461

Observation b55d1ac3-ea59-4cc8-94fb-06298f4b59f8 · outbound

This paper cites Le, Thang Luong, and Golnaz Ghiasi.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Le, Thang Luong, and Golnaz Ghiasi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.575746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.154142Z digest=sha256:2e5ace3bfd466f0db34d1e17140d79ca2be36206b3ee405c750a6f3e64fd7668

Observation faa5d3f1-5b31-417b-95bf-b4231f2d58ca · outbound

This paper cites xGen- MM (formerly BLIP-3): A family of open large multimodal models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? xGen- MM (formerly BLIP-3): A family of open large multimodal models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.158345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.158345Z digest=sha256:4f5912a4cbb0c77c8c964faa6b47f0e87a2b01bb7958d23ab2f7e1b0c4583d45

Observation 1bf36864-c24d-45ee-b855-05555cb4c46b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.162061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.162061Z digest=sha256:bafcbfecef26559430f78cafdc6cb644bc80cacb3c782756672323c42cd9323d

Observation 5a9af856-ff51-4cec-a2a9-398e3f64cbae · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Yi: Open Foundation Models by 01.AI

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.166131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.166131Z digest=sha256:ba0ce5a1906bb7d115661139e5c74031823ba2214ea1e960d4921c036b070c77

Observation 74f2d605-937b-4203-a357-141d63b74a09 · outbound

This paper cites MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.562615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.170068Z digest=sha256:82e63242968a07e5a734ed0fb6b1dcf23717ed8870dc02ac0a01fd939c9da8bb

Observation 81cedbb2-41d0-4939-9e2c-3a3ce387579c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.173837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.173837Z digest=sha256:2e0326609058198a91fa8efd1a31c2ca6d73cb34a5d449d0b9009f3abd520824

Observation b8aa4ff3-da1c-4cf7-81b8-027479da32d1 · outbound

This paper cites Sigmoid loss for language image pre-training.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Sigmoid loss for language image pre-training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.549109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.177834Z digest=sha256:696abada4f6181a61467cbd3de05d4b39b912d16d11793652095fa3aaf47801e

Observation cadb536a-be2d-4ff9-b053-f671129816ee · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.181590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.181590Z digest=sha256:d8845221bf8337fb93c16ba52002b755686974314cde788c95029cbba8fe47b7

Observation aad8ed8e-a71d-4e96-8c3e-28895f934dcc · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.185644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.185644Z digest=sha256:51a507fbff73a54dcde193458ecf43d0ed674935b60fd74934ef1261bc20484e

Observation 2e0c93ba-4e6b-4e1c-ba84-fb4350a2a774 · outbound

This paper cites On evaluating ad- versarial robustness of large vision-language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? On evaluating ad- versarial robustness of large vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.536158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.189782Z digest=sha256:cb4583597a927f32949664097f1d899781d0b8693ba10f111f46c7dda32a5d54

Observation c3dfd032-6ed0-46f1-8e59-f7b0b56f4b73 · outbound

This paper cites ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.523316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.193661Z digest=sha256:1d4b18ed4532adf00ccc693384f16643f365eb3d22b085acd39be3387fcb63fa

Observation 84c5baff-60b7-4802-a6df-9bcfe455bf38 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.510486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.197392Z digest=sha256:10e56c6f386ec3ba6191384c811348fbc2fb779d6648d12be980c83a1c0dddae

Observation 7053f530-aa0c-49e2-b1da-d1771aef3df6 · outbound

This paper cites dall-e-3.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? dall-e-3

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:57:29.496650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:57:29.201224Z digest=sha256:babe9a152783e54ae71ff10eb47783401b77a9fae5b9760dda7cb53455c8356a

Pith citing papers

Observation 12c5c272-28de-4f42-aea7-07925d6fb9d8 · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.021149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:896b0c8647dc492d034fd65e1920ab1f9c8c8907bcbcd48c558c81ef324dca0c