Pith. sign in

Paper Citation Record · LEDGER

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2411.17794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17794 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:29.201224Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:24:57.169737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:36:05.014470Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c1ff511-65d0-40df-b468-58e372740f50 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Flamingo: A visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:30.011651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.007620Z digest=sha256:13aad72ec8ac32ea4b2fdde5c6478d077f1106e604b43f6cb040cb0eb7ebea7b

Observation ebb4e11c-4f2b-4fb0-8171-5f3d23a204d6 · outbound

This paper cites Qwen Technical Report.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.012130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.012130Z digest=sha256:7c4ed4aa83e77642eaa451aab1b8636df8b039bc638a1419cd5057d498a7c84f

Observation 4d88c780-105f-482c-bf8a-810709dde470 · outbound

This paper cites Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.999362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.016492Z digest=sha256:5cda37855ee2627b6033061cb01bf544829396b94359fc4a852c43375b6deb64

Observation 28130577-7c02-4337-8cc3-d06c740ee696 · outbound

This paper cites Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.020551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.020551Z digest=sha256:a757c284e10fdc34e2d7c00e588e3bcf1fd12262a1a1fc4a30f8588c09482ea3

Observation 35c01332-1ba4-4faf-9734-98138069ce28 · outbound

This paper cites InternLM2 Technical Report.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.024930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.024930Z digest=sha256:09a136627b4eac16b0db7f1e3254efdaac00651516bbeecbe96b2310fedf33e3

Observation 1c78c158-3ead-4c2d-bc7a-a01149fe323d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.029279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.029279Z digest=sha256:ffc59dcdb4fdedfd6794642b5fc54ee149672bf9f2aeb0e9e758c70ab9ced7f7

Observation 1919acb8-38ab-48f9-abdf-90f621f7af2a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.033836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.033836Z digest=sha256:4895d97728cb91543579a1ee751e7c85ac4532a71db8238205898e09174975ee

Observation 00deef47-f46a-4d61-862a-2d659f2a2e3d · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.986905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.038639Z digest=sha256:40a7525ee1e709494832f86da91857869a64631d52d2f528b208be8d261dd82f

Observation ffb52276-4dbf-42a5-a8a7-6188c612ccf3 · outbound

This paper cites InstructBLIP: Towards general-purpose vision- language models with instruction tuning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InstructBLIP: Towards general-purpose vision- language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.974761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.042409Z digest=sha256:4dd89e28ed1ae8473c204f6177e1122b3153f78696041fde64e5b8f700c4ff5e

Observation 74bc106e-3051-457d-bd6b-6fdf84bf9da0 · outbound

This paper cites ImageNet: A large-scale hierarchical im- age database.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ImageNet: A large-scale hierarchical im- age database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.962422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.046295Z digest=sha256:0e0ab3785ae86df156efafc5f2b59d3697753aa0e6157786bbe4577dae6ee2cd

Observation acce52c3-5ad4-4f5d-ab4d-d6ec5a633617 · outbound

This paper cites Scaling laws of synthetic images for model training.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Scaling laws of synthetic images for model training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.948764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.049953Z digest=sha256:1135c4fa83828dbc9741497db6570ccc7e8a5799b4044ab6d789b74495f6da19

Observation 6958ca54-6c05-485d-bb6d-d0752481924a · outbound

This paper cites EV A: Exploring the limits of masked visual represen- tation learning at scale.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EV A: Exploring the limits of masked visual represen- tation learning at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.936580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.053556Z digest=sha256:9ba8aaea67e191d7a62db238564f7a6952614034f9401b5be9552c63d13dff43

Observation f028fc6e-4b57-478f-87fb-90e4505b9d39 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.057433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.057433Z digest=sha256:31b46838b3377cb7336d41a7c97b49d149c0442e19e587465d6f8604f268acee

Observation 8e81c1a8-f204-4538-aa07-3588ac598431 · outbound

This paper cites Gemini., 2023.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gemini., 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.924269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.061609Z digest=sha256:af0c965d2173e6cd5bfa8de78b06369de117c31f2894d21483bba5955cce6f21

Observation 6413d14b-cc83-41c8-b741-b53012ff8f07 · outbound

This paper cites Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.912168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.065609Z digest=sha256:ce9395c518cb796016f25475af3cd79e69ee3e9f120ba67b502dcaf529fd01ef

Observation 3c0226eb-72d7-4783-b93b-32da6202be57 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.069305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.069305Z digest=sha256:f28c726c0fa16c394225484a6373cf0cedfd26835334c9fffdfc23eb9a8d9ee3

Observation bf354ab2-46c2-4c6d-bda0-fce9785b7fb3 · outbound

This paper cites SEED-Bench: Benchmarking multimodal large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench: Benchmarking multimodal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.898939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.073147Z digest=sha256:f7e43db951ed09ae5227f679ac38f69791fe4e41f83892f30ce978131bbfe006

Observation 9e0f7929-dcc2-4b94-94b7-2719b23accc1 · outbound

This paper cites Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.885960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.076856Z digest=sha256:cfd31b097d9500044a16d85134301bd5d0d21c089baf259b2314fbe81fccdeac

Observation ff023b8e-d03c-49d6-b83c-0884b3184ae7 · outbound

This paper cites LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.873904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.080548Z digest=sha256:04c2af3731635c148d6865e56104185a7b0c8bd118968925ad7cff9827f3203f

Observation 46ab11c0-667f-421e-85ff-310e912d9999 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.861617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.084349Z digest=sha256:cd815652ed07a525d5778075bf1b454e8cc48baefebfd14939c85919733d918e

Observation 8171e875-5c42-4ad9-b087-14455198dcc0 · outbound

This paper cites FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.848224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.088284Z digest=sha256:4878898cefb1437424afcc9f5157a7c5853a411e59b3ea45e1d8416493ea6535

Observation e336f4b7-b409-40ee-b9c5-38effed72c9f · outbound

This paper cites Evaluating object hallucination in large vision-language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Evaluating object hallucination in large vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.834662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.092028Z digest=sha256:b472904e00fd9310256861d82cefd2d0ec1b4398c8e2060d1c4a2d2b200f6abc

Observation aa38c0a9-84f6-4197-82be-97650e95f758 · outbound

This paper cites Visual instruction tuning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.725556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.096069Z digest=sha256:d8967809bc71deaa46b7b5f8d2310f57d12073c01b5fd4eca2397fb5c333e2f9

Observation fcbe0d19-002d-48d1-b745-40cae90a2cd5 · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? In Proc.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMBench: Is your multi-modal model an all-around player? In Proc

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.712165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.099764Z digest=sha256:e4f2ae4ea624be8a409fe54993a2ac95d7dca70de56ab0e933d5855a8d3ef2e4

Observation f3b36e6e-ba15-4cde-b70b-14c99e052a24 · outbound

This paper cites From here to human-level AI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? From here to human-level AI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.699235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.104119Z digest=sha256:2791cdad213f85e0f0091ac92a8f581e19b7c5210706684deb404f8f09ec3b04

Observation 1a8c0e8a-bdf5-41f1-a0cf-5c0bd195d76b · outbound

This paper cites Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.107976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.107976Z digest=sha256:ead367b93cc81fef7421978f516ef10400bc89dc8ffac18437727524bf1b488d

Observation c49d540f-1f35-4ffc-9263-f5213194d74c · outbound

This paper cites Position: Levels of AGI for operationalizing progress on the path to AGI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Position: Levels of AGI for operationalizing progress on the path to AGI

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.687401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.112023Z digest=sha256:adb0d27e215a53bd9bf91191f1c0e5a94a9b40bd4e39639607b8b9c6b0cf9cc7

Observation ffeb8b1c-da57-40a3-b4ff-3d21227f552e · outbound

This paper cites GPT-4o, 2024.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? GPT-4o, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.675430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.115794Z digest=sha256:f87809fd2c896b0f721107096a116b0a5a1810c11a6086b19e1f3667b341481d

Observation b80c38a8-aee4-464f-a679-dbe8ac3a27d2 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Learn- ing transferable visual models from natural language super- vision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.663227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.119566Z digest=sha256:9c1ca94fba5de2a643fbcfef045ed26660ff82aeef30b8c4ecc5b2a1388231f4

Observation 47900640-6151-4743-9d30-7e73eb74b853 · outbound

This paper cites Zero-shot text-to-image generation.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Zero-shot text-to-image generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.650633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.123316Z digest=sha256:860ad1bde650eada3edf6126a21b93fbf5cbdac5da191a9cdb32b2436a509a3b

Observation 2ecd108c-8b53-46b6-bdcb-7417fe87ef9e · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.127230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.127230Z digest=sha256:afe5efdebc0eefde0e5bb76bc542508eec5e284e1d59d6a2e9359bb5fe840d13

Observation ed030b23-9307-49ba-a19a-715009b5ff27 · outbound

This paper cites Link- context learning for multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Link- context learning for multimodal LLMs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.638505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.131172Z digest=sha256:c8a793828bda80acf0e7af512498c1d25b7e622eed37e434c403ca4ea7568512

Observation aace756d-93a0-47ea-b5ba-7a43b3f73cb4 · outbound

This paper cites Label Studio: Data labeling soft- ware, 2020-2022.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Label Studio: Data labeling soft- ware, 2020-2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.625949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.134990Z digest=sha256:1c98fa2a118f0c7cc38e585be7c84fdfde459ffe02698f7e26dc15cfea994538

Observation 79b3aa1f-b5ff-434c-9691-f332dc9bf7d9 · outbound

This paper cites Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.613670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.138719Z digest=sha256:749376f94d158e0352ac1dcad4022984c76792b3991c42686282ec3468ddcacf

Observation c7dc8d40-bde3-44b3-b14b-90fd04aab1b4 · outbound

This paper cites Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.600539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.142487Z digest=sha256:922aeca95bd0e226446b3d146a6cab04fda29b4b38520e09b4ba766182b5627d

Observation a603c976-d34d-4b55-9bf0-0b714c5c1afc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.146242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.146242Z digest=sha256:f54cc340950b8de8b027a1893140c8a77fa8781d7e6f154c97ef4f9796577b16

Observation a914c956-324d-478b-ac7e-0e1a1f32ba7d · outbound

This paper cites Recent advancements in fruit detection and classifica- tion using deep learning techniques.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Recent advancements in fruit detection and classifica- tion using deep learning techniques

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.588119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.150418Z digest=sha256:f2035942f819b7f19a09cbd5405cc064f4767c5c089580b680070d0c1cbfd0da

Observation b55d1ac3-ea59-4cc8-94fb-06298f4b59f8 · outbound

This paper cites Le, Thang Luong, and Golnaz Ghiasi.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Le, Thang Luong, and Golnaz Ghiasi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.575746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.154142Z digest=sha256:29c39458254981202932f1a3ba03e11e0ebf9ca2064fce1bb5316d6c7c222389

Observation faa5d3f1-5b31-417b-95bf-b4231f2d58ca · outbound

This paper cites xGen- MM (formerly BLIP-3): A family of open large multimodal models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? xGen- MM (formerly BLIP-3): A family of open large multimodal models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.158345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.158345Z digest=sha256:9dca61964e6a8d19eea594acf221452462e12dcea8ef3df2e50bc4c6a5c7571f

Observation 1bf36864-c24d-45ee-b855-05555cb4c46b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.162061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.162061Z digest=sha256:086db5af00d68f9e361db7a9f2da7cf5c4107c551441c10e929e8a6d901d2cbe

Observation 5a9af856-ff51-4cec-a2a9-398e3f64cbae · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Yi: Open Foundation Models by 01.AI

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.166131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.166131Z digest=sha256:bde5333e48b3ed8840b09031e13e4ca80386733eb1e259d9a58b613b4d03a59a

Observation 74f2d605-937b-4203-a357-141d63b74a09 · outbound

This paper cites MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.562615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.170068Z digest=sha256:9d582f78136f60b230c4a7d01f1237efe412c4653dc85352c27c62e80994cc3c

Observation 81cedbb2-41d0-4939-9e2c-3a3ce387579c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.173837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.173837Z digest=sha256:53fcf0f94e3d9a8cadf093790eed3c681a6e4779f354c66b73e8847625f1da44

Observation b8aa4ff3-da1c-4cf7-81b8-027479da32d1 · outbound

This paper cites Sigmoid loss for language image pre-training.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Sigmoid loss for language image pre-training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.549109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.177834Z digest=sha256:84964f47dc783012c6db2a7a1a99c0216a166cf2f0a318118bf62255b4e302e5

Observation cadb536a-be2d-4ff9-b053-f671129816ee · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.181590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.181590Z digest=sha256:e53c25d2d91603a00d7f5dab984e19624042fcdf14cdc6aaf603d31e28a9bcf5

Observation aad8ed8e-a71d-4e96-8c3e-28895f934dcc · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.185644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.185644Z digest=sha256:61f7484e7a9de493ad8ffccb93721815f0e72a97d7698f44a53938c8cd4be3d2

Observation 2e0c93ba-4e6b-4e1c-ba84-fb4350a2a774 · outbound

This paper cites On evaluating ad- versarial robustness of large vision-language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? On evaluating ad- versarial robustness of large vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.536158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.189782Z digest=sha256:72ab9cf749d7fa868ded22db8956cd4ccd43ed01fa24fab02166ae930c9233dd

Observation c3dfd032-6ed0-46f1-8e59-f7b0b56f4b73 · outbound

This paper cites ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.523316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.193661Z digest=sha256:9bcc1257c9b44378c442c17128516b3e4a4aae45a9e6dd96be1a7486324fe7ca

Observation 84c5baff-60b7-4802-a6df-9bcfe455bf38 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.510486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.197392Z digest=sha256:5af809bf05d691ceff09c6d25985bf2fe69445600beec376a0622ef617853b98

Observation 7053f530-aa0c-49e2-b1da-d1771aef3df6 · outbound

This paper cites dall-e-3.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? dall-e-3

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:57:29.496650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:57:29.201224Z digest=sha256:a063d2bafb8c5b4fb37afc5cc4eb6dc3efb68db832bcb8c8caf4603feb990bec

Pith citing papers

Observation 12c5c272-28de-4f42-aea7-07925d6fb9d8 · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.021149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:1178a5cef99a194ed49829ade2a1eabd6f6a78920f90040bb79c719dca777523