Pith. sign in

Paper Citation Record · LEDGER

Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2311.06607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.06607 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:58:19.334924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.479106Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cb96e4f-8e4b-4928-83dd-191236d49a52 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.471878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:610285715d9287c75df434c47289686e6706cddf68c9c2c5bdc6a794acd53f12

Observation 77798ffe-5e13-4a9e-9b18-b33c9ab5be2a · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:53.896871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:b618e77d0dd6cf3db4a87f5c00172d434397a3231068279379442b309bd96eb9

Observation 2dc1f223-840d-47b3-83b6-30d50bd8bf64 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.003085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:c7c0754757468a555e66db27a58cd5725b9cf55c73e84783c8826c2e9f0c7686

Observation 7194c1c2-9fc8-4a1f-91ef-9ce83b2e3d93 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.732796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:4d118abde02adf8403a74e7aafb3ec0f44baea0696a7f8068b1121cd753efb7a

Observation d49e1eb5-1728-405e-ba7e-138d7f79f98a · inbound

A Survey on Hallucination in Large Vision-Language Models cites this paper.

A Survey on Hallucination in Large Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.326163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:10:10.186950Z digest=sha256:0cca9181be15efc94fc29cf3fab1fdeafb4ed0aa2f22e3a6ecfca62815732398

Observation 6b69ddb0-f51d-49e8-ad2a-daef9811f7b1 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.493562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:0209b577757592d94639caa070ed558de1efd4d3a0ec278843b90037c75a2c21

Observation 245f14d7-b4f0-49f3-ad45-72fc13368c17 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.535674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:fd404cb84072ad597742b06c8efc9a0f80ffb6e93ea4467de33440aa3228bbee

Observation 1ed97519-ff46-4555-8a73-4419df871bfe · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.053339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:f43264ff11f6475b06a3c91f5b1009e742cef770f785428b129cf163ef3bc7e3

Observation 013f2f03-743b-49cf-bbf8-339868f8f464 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.688589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:4d1a01f7198574c9a674b70ee4ef6326c06264cc540c513f89ae139abe1cc45d

Observation e2c0c614-df4b-4e47-8fa7-185b0efbc18e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.808098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:7ce32e7d749876aa4f6f20d9eb756d046d64b6274e97b8861527d932cfe7eb0a

Observation 9d3a7fc1-83a7-4d6c-9880-70a9f42332fa · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.703956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:25f194f535a8be5c35d5b50b2414c44a0b2439442b13228b9b92fb33a9f8ce21

Observation 27a21b8e-3b2d-4ad6-8b6b-c8b3cfe7171e · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.253757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:2e455839cc2134c11f9f4d01bf0275ed78b394de023630c9730a725f1c8ee0dc

Observation ea2a5ef6-e5f0-4db7-bc73-a67dbf644f8d · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.854770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:09fbd7c24bb42ee0c460f7572b10ef42ee5a8caab8fe9af941dd5c5e7ec7b1db

Observation 3984b339-064a-43db-9912-503af55cc6e9 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.362333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f03a3f186c2ae5dd63ef287a434bb0ebd12fb503adbcfa0c3a631c62359ed9dc

Observation 88447597-4568-469d-a634-9e20ed3c5b47 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.931321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:891219de9ba69c931beff9e414c82eeb3ebd2cfed1c927a08da2d9b73c81027c

Observation 57eecde9-6e1e-4448-9fe3-a94055097536 · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:19.334924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:19.334924Z digest=sha256:2a4df78160186f59695a773db3d431ec4687888fe48da6dbc3cb353dbaf70c2b

Observation e15e1f92-b278-4d04-89e0-b91bb83ed2e0 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.922398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:91b818dd011c51810840c7b7ddc9dd19b8952a143e9a0036f10713b6b2dfc62b

Observation 0ed651f1-a7c3-44ae-8b8a-c60a789c91ac · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 279

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.480709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:b8c97564520431d62b19c256a2f0fd76568c5202b44856a6f949bef863175607

Observation a6cd8e01-3f80-4ee3-86ce-87e3715168a5 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.426378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:a6b7557acb3fa7db07a5948fad93bcc21c4943dc78a57ea19013ac5220de9702

Observation f03fc67c-d05c-4124-ba0c-b89f0ee94c29 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:b334c8a7c92bea43162c65d348897a0a5eb001efe7d0980782bb7cae84c15d1c