Pith. sign in

Paper Citation Record · LEDGER

Region-Level Context-Aware Multimodal Understanding

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2508.12263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12263 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:39:48.245781Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:09.424972Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6aea20c-74fd-41b4-953c-32687809851b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Region-Level Context-Aware Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.022168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.022168Z digest=sha256:68bd43df8dbbb53b5cd7fc26f255226117fb1e1ba200d99e63686a457b8f63f3

Observation 7df1ea6a-cf53-4abe-8f61-2d24ae8de02a · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Region-Level Context-Aware Multimodal Understanding Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.032858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.032858Z digest=sha256:ca85af08704a79673afa3869a8c192a1f3d60b10aa1169fda9168a1a9f86d98c

Observation 8424158f-e0c5-4b3d-9f5f-c5aab1d835ac · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Region-Level Context-Aware Multimodal Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.041259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.041259Z digest=sha256:93ff978d42e6a4574385d17aa673d4589fde0149bcfd81fbce2a3766871258af

Observation 6a73e7ef-8b2c-40c4-b159-6154601ab78b · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Region-Level Context-Aware Multimodal Understanding InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.049375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.049375Z digest=sha256:0ceb3ec5d12b94795340518687b47085e0c16dd9096461f9ef578b9c53115785

Observation ec805674-2e99-49b6-8ff8-56e581a3b0f5 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Region-Level Context-Aware Multimodal Understanding DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.055078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.055078Z digest=sha256:a81c0ca8e7397c0310f94596569b49a174703eb4bc02f29280cdc659f8e8c214

Observation f5b6ccda-328f-47e0-895c-d34039522811 · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

Region-Level Context-Aware Multimodal Understanding MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.059765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.059765Z digest=sha256:4f7e5fbbe918c988aba9b665face90e6f92420b07e28ab8c883c8e1933a4a1eb

Observation dc0acddf-0452-4542-8916-34dbb5c29744 · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

Region-Level Context-Aware Multimodal Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.064324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.064324Z digest=sha256:26ecd1b99a2d6c5f800e896723b7ccdcf6915fd7f041ea0cfa33a73e723232db

Observation 16e4f698-fc0d-4b37-80d7-dc44182c46ea · outbound

This paper cites CaMML: Context-Aware Multimodal Learner for Large Models.

Region-Level Context-Aware Multimodal Understanding CaMML: Context-Aware Multimodal Learner for Large Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:39:48.648811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.069121Z digest=sha256:06ed9f0860097dd9a1fa521e747ad693c87a24502750030a92692a9a66022e0e

Observation 8a816166-f40f-4779-b0a8-40d5702a3811 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,.

Region-Level Context-Aware Multimodal Understanding Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.075111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.078598Z digest=sha256:67731a0488c55a5289b6d323bdf4283a1a36c6ba42cd863da0e3a3e2ed02d9b2

Observation 064f0dad-04c1-40ce-a4be-87c156d73b20 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Region-Level Context-Aware Multimodal Understanding Referitgame: Referring to objects in photographs of natural scenes,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.062001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.082164Z digest=sha256:70f18ad41d62a62d85c83e497f3ced22b89289670ff107dcbe39b3a20e0d86b2

Observation 7da6b354-52cc-4587-a071-ef4c2aebda7c · outbound

This paper cites Panoptic scene graph generation,.

Region-Level Context-Aware Multimodal Understanding Panoptic scene graph generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.049397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.087505Z digest=sha256:c9d41ca06d105b17f6d9ce9bac1a998f883bc78a94cafa083a977526bda2b74f

Observation 0f216edd-f8c6-4466-954b-3e5636654f35 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Region-Level Context-Aware Multimodal Understanding Glamm: Pixel grounding large multimodal model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.037852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.092365Z digest=sha256:1d43bd971c2780912869232a75e4e7033a752c1b62e02f5fd36aa1c0538cad51

Observation 2157a260-b549-480a-8763-0fa70a8c963e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Region-Level Context-Aware Multimodal Understanding Bleu: a method for automatic evaluation of machine translation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.025900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.097391Z digest=sha256:81126127dfb32175840b578a75e5d981490574b8c6209ff36e13898b5b96285e

Observation 56ae1682-2407-4411-8588-67915790807f · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Region-Level Context-Aware Multimodal Understanding Rouge: A package for automatic evaluation of summaries,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.100736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.100736Z digest=sha256:bd7cc4f8ee378a1f391d9177172bbf455e56cfc9e2761ce7ed73fc1b1f08f84f

Observation d28d5f7c-fb8d-454a-a840-b822c9a5e598 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Region-Level Context-Aware Multimodal Understanding Cider: Consensus- based image description evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.005462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.104232Z digest=sha256:6e7ad86e9296745d7f11ac508deb03a9fd0c573863224693b78e2ab1d1f506cc

Observation a41a9f92-bcd7-46f8-a4dd-f57476b6cbe9 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Region-Level Context-Aware Multimodal Understanding Spice: Semantic propositional image caption evaluation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.984661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.108957Z digest=sha256:3cf3d8f561a0ea0196921911ee01b4f3a5d7b3ee01cc4639d095de157b020190

Observation ba4e72e0-11cf-4a11-a933-129254c97da0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Region-Level Context-Aware Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.971525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.113607Z digest=sha256:22741d6f29431118ee1ed29df50b3847f17197f99e16384ad5a5b8ae999f58cb

Observation a70b2165-d355-46ca-91c3-e18b54b1079e · outbound

This paper cites Visual Instruction Tuning.

Region-Level Context-Aware Multimodal Understanding Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.121173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.121173Z digest=sha256:79a69a0c930ea7cd3b82be105438943c968504bb6f73ad187f70a13a726c72fe

Observation 069b9a4b-b512-452b-ab20-7cb3fc479039 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

Region-Level Context-Aware Multimodal Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.124792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.124792Z digest=sha256:4deab0b17bf9268c4f51c283710b6fb3db53ce8d06a041e962f59d3622be649f

Observation 4e50f82b-d859-4e4a-b256-471307846976 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Region-Level Context-Aware Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.130365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.130365Z digest=sha256:035b515e38d261c9486d526b81e5ba3ae22e79cb9ec077b5adfe9bff9fc2d1e8

Observation 5a10a371-7cdf-43c3-bbfa-86c72a2a00cc · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 256390509.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 256390509

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.958837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.117178Z digest=sha256:440ca4626d6dff3d5d82b1ca728dd03d381dbf8d7f97912876bef2a475b1a5c3

Observation f4621501-30fc-44ef-897a-3d9c6edede2f · outbound

This paper cites Mmict: Boosting multi-modal fine-tuning with in-context examples,.

Region-Level Context-Aware Multimodal Understanding Mmict: Boosting multi-modal fine-tuning with in-context examples,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.948011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.137165Z digest=sha256:345dcc757a633cc06b5f72a6db3e634ebfd73b4cec22b5353a79d7638fa8b643

Observation 5ed4d786-02e6-4e98-bbb7-f525b886ef24 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Region-Level Context-Aware Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.936899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.140247Z digest=sha256:3d135e8291303126a8e8bd6bcf9f1b6c22422c791a1bc8efba9db876122ea2f7

Observation ac638776-fc03-48f3-85c0-9b918924fde1 · outbound

This paper cites Cantor: Inspiring multimodal chain- of-thought of mllm,.

Region-Level Context-Aware Multimodal Understanding Cantor: Inspiring multimodal chain- of-thought of mllm,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.916037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.144298Z digest=sha256:add6b77d4a60d2bd3ddcfec9717303b2b5429fb3722efc6ce74d9a79ddda6f0c

Observation 5a2cff7f-f92a-4cfd-9d31-a79215644a5a · outbound

This paper cites INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model.

Region-Level Context-Aware Multimodal Understanding INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:39:48.570485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.133771Z digest=sha256:bfc41cd4bd5db0350df56e21068f10eccbcca48a77182f0a7df391da668ab661

Observation 2fcc1d4d-6c3a-42db-a9a9-0668919903c7 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

Region-Level Context-Aware Multimodal Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.874397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.159498Z digest=sha256:20765584429d738c2b0262863a7d6e6b6068d5eb8df336bf98bcaa2c4cfaf675

Observation cf21f313-22a6-403a-b487-4054478e175c · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension,.

Region-Level Context-Aware Multimodal Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.163553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.163553Z digest=sha256:a92a9e845daed3877282f3d9f4fc626a8dd4d89c37eb14b250b33f342e7979be

Observation ec322ad7-3d90-4a70-ac13-c2e78f070e97 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Region-Level Context-Aware Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.167701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.167701Z digest=sha256:97ed67689acb1ae62a22c53b4bc9a22c2c5cf14e75badf9e937becb0347b85dd

Observation c4866031-3d7d-46ce-968e-a6d1cb721dd1 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection,.

Region-Level Context-Aware Multimodal Understanding Video-llava: Learning united visual representation by alignment before projection,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.902281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.149659Z digest=sha256:e306537f151d542866ce436ea040a8806467ec3e23f981cbd24034fc04cffcb2

Observation c0506359-bb7d-41f3-a44d-c2e962694888 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 265281544.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 265281544

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.888958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.153196Z digest=sha256:91e0ebd1e945354f7ea94dabe37663e983da679cdeb5f88b610d86bf1365631a

Observation 66411ede-fb32-46f8-9fc1-6074fe3a8c78 · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,.

Region-Level Context-Aware Multimodal Understanding Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.834882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.181046Z digest=sha256:549e0ed65ff4cd5fffbd0ee4563cecb8d7028bf6a75afa26a8c4ce6af4534702

Observation 7d7eda24-ed12-4dfe-90c5-206349a46147 · outbound

This paper cites Rap: Retrieval-augmented personalization for multimodal large language models,.

Region-Level Context-Aware Multimodal Understanding Rap: Retrieval-augmented personalization for multimodal large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.804984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.188905Z digest=sha256:b359f6c5582610bfc1165196fc8354da310f867f22433c1d8346ed95480b7376

Observation 863d828b-bcf9-4ba0-af2f-ac13b716c004 · outbound

This paper cites Yo’llava: Your personalized language and vision assistant,.

Region-Level Context-Aware Multimodal Understanding Yo’llava: Your personalized language and vision assistant,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.790306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.191574Z digest=sha256:9db66ed0d8f894cfc660595dc3fb8bb613a83f17421376eeb1bd62f93d7ce7fa

Observation db9dc312-26fa-4102-9fae-f82b7ab730a3 · outbound

This paper cites Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,.

Region-Level Context-Aware Multimodal Understanding Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.861748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.171506Z digest=sha256:522226b854f36ce5cbbd33bc6237dc9ca87401e21151786b6df36c1f6c10a91a

Observation 2631a3bc-7b7e-4b32-bb7d-bb594ec41218 · outbound

This paper cites Palm-e: An embodied multimodal language model,.

Region-Level Context-Aware Multimodal Understanding Palm-e: An embodied multimodal language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.850483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.176293Z digest=sha256:3569e2d54b60780c4ee125b678548e2ded16e7b49423449103b74f983bdd53cb

Observation 36873bb4-dd1a-45ee-940a-2d65d5c98aa0 · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation,.

Region-Level Context-Aware Multimodal Understanding Llm2clip: Powerful language model unlock richer visual representation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.207295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.207295Z digest=sha256:3f3fc5bd21b44731b08bdbc4594b1873fb7baf4dc2a4f1ae0a93a9d29bc7c38d

Observation 0ae8dd8e-75a3-4d8d-9bc0-668bf4a658cd · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 266573457.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 266573457

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.819456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.185416Z digest=sha256:d00939f11abcca8dac26bf40523c60015f22317b592e01edccf63ab37d13a5be

Observation 5f8819b5-6c47-44eb-a65c-fbf2719e09ef · outbound

This paper cites GPT-4o System Card.

Region-Level Context-Aware Multimodal Understanding GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.214532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.214532Z digest=sha256:3cc0443b9bc37d76becc000c000c5e653ac0c0c14e86f21438f8cb480d169d7f

Observation ea84cffa-cf95-4d85-b959-fcb037b10bb2 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Region-Level Context-Aware Multimodal Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.225203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.225203Z digest=sha256:bdae87856bfd9baa6eaf525f7c4a60395d52f9e9f7d235396b4cc4897eaec480

Observation 06819efa-51cd-4240-b180-13c4f8121ed3 · outbound

This paper cites Myvlm: Personalizing vlms for user-specific queries,.

Region-Level Context-Aware Multimodal Understanding Myvlm: Personalizing vlms for user-specific queries,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.778005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.198688Z digest=sha256:1491ba11e8b4dfa66ec925a5e8526d55e2891b4d75e8be47674d41d6914248d6

Observation 5e8108fc-7083-4939-85cd-9dcfe042a13d · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Region-Level Context-Aware Multimodal Understanding CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.202916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.202916Z digest=sha256:efcf1bc0504c630c884839792667be093535ce776ea8364f52778f29fb6821c6

Observation 80b1a083-761c-416a-9ece-e9b080dc6912 · outbound

This paper cites DeepSeek-V3 Technical Report.

Region-Level Context-Aware Multimodal Understanding DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.237682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.237682Z digest=sha256:1de815a8eeb52f3162af632d082418ba5de59b6afca6e89b95627efa4be0a0f0

Observation aacd94b3-7518-4400-b327-feb67497b463 · outbound

This paper cites Gemini 2.0: A new era of multimodal models,.

Region-Level Context-Aware Multimodal Understanding Gemini 2.0: A new era of multimodal models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.766616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.210718Z digest=sha256:8045dd4a834d5127de338553ca4b783cc71b213443c30f4ac17289301ce385ed

Observation 7a31e000-7b9a-412b-9ce5-ab7d02769624 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Region-Level Context-Aware Multimodal Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.245781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.245781Z digest=sha256:c76854ce09d09a559fccbf2f03a1d4e0d82e0962ac6c03767bf1452257ba4b42

Observation 7caa56f4-8a90-467c-939b-7101d79bffb0 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Region-Level Context-Aware Multimodal Understanding Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.756620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.229795Z digest=sha256:4462b0d139ab2348125a133404dc63efa036a3446bf6adb7415e9c4f4730dd13

Observation 5f8d553e-db90-487d-b911-114ff3e7f983 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Region-Level Context-Aware Multimodal Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.233486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.233486Z digest=sha256:a089862a69ae899e62e37bacf969a0f35b806d16ba0c575f82e017ed1108bfe8

Observation eb47bbe7-f3a4-47da-b8e7-ab3b418008ce · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

Region-Level Context-Aware Multimodal Understanding Enabling Large Language Models to Generate Text with Citations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.241070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.241070Z digest=sha256:19919bc8b7566c4f1ea0563ef6cfb521e8468e5a079096f8eec673d8652771ca

Observation be0e5648-70cf-4467-88ea-e2d6934d2848 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 248476411.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 248476411

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.242414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.037500Z digest=sha256:7f97d1fa539a9ea5c10cf2415095a2e3cd76f292eb3c67b81a254639e27ac1b0

Observation af0d61fe-6251-47e2-b377-ccfecb726f3d · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 258615266.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 258615266

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.168952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.052310Z digest=sha256:e6913ec9c6b98fd9aaf941da03516a1f32272028d0e0a2a4251b8f11cae58da0

Observation eba9c55e-cb75-44ff-adae-69bb5301e5f3 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 266844925.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 266844925

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.110845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:39:48.074850Z digest=sha256:d7d3fb641065486aa5a831169dd03ac49e690fe3177ba64436fd69f155d77229

Pith citing papers

Observation 9865e114-cb2d-446a-9799-6fb45b263c7f · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Region-Level Context-Aware Multimodal Understanding

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:09.472995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:4c0a105121a5a700a5f5b28cc0fa35c680fbe61e6815639f264c2cff375c7099