Pith. sign in

Paper Citation Record · LEDGER

Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2311.06607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.06607 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:31:36.855085Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.479106Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cb96e4f-8e4b-4928-83dd-191236d49a52 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.471878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:b2b8e9ed786f6cf5e44e4546860467f76c1cf612a3e54bb632003bddd7bfd47c

Observation 77798ffe-5e13-4a9e-9b18-b33c9ab5be2a · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:53.896871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:708bef86a8c9c9814080482da646b571d4d2dd49024456287bf84967bf6cdbf4

Observation 2dc1f223-840d-47b3-83b6-30d50bd8bf64 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.003085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:e42a2969695d6a082e6d39a34bb3c1f29b199d860378b2018b11c00138d3f0d2

Observation 7194c1c2-9fc8-4a1f-91ef-9ce83b2e3d93 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.732796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:1ad9fec91cbb21d56a8f95db1a6f53785836b6432eb88aae765fd693a1bbc84b

Observation d49e1eb5-1728-405e-ba7e-138d7f79f98a · inbound

A Survey on Hallucination in Large Vision-Language Models cites this paper.

A Survey on Hallucination in Large Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.326163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T22:10:10.186950Z digest=sha256:c801dce2ac0a6e4264950aa6945006f3df628ccedacccc6bf4f932e4c792b462

Observation 6b69ddb0-f51d-49e8-ad2a-daef9811f7b1 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.493562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:077e7e427a6dc444594f4ddbab733868088ec90c6ddd72cdc898bdd133742bad

Observation 245f14d7-b4f0-49f3-ad45-72fc13368c17 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.535674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:cb43329f97412aaa8ea1021a611814ced7637a5695f42d27270f18ab76e4d8ed

Observation 1ed97519-ff46-4555-8a73-4419df871bfe · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.053339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:1df5f532d5afc65ca5cc180eef78e726d9cdbe1b322a56112cbde49fa9501939

Observation 013f2f03-743b-49cf-bbf8-339868f8f464 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.688589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:28e006639d82b3d6d908c82ce02c73b48c8e6e315df5f0b3e186e48056774ae2

Observation e2c0c614-df4b-4e47-8fa7-185b0efbc18e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.808098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:64b7eb43874f267d8c64f4db95a994da100d00cbba0b6334a7e44bb66ff16cae

Observation 9d3a7fc1-83a7-4d6c-9880-70a9f42332fa · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.703956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:5559c609e84097073e86c1a2ba83afce98145a9f9222b5daabfd821a5062d910

Observation 55873e8b-ff74-4978-b429-92c8535b0f80 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.619612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.619612Z digest=sha256:e3b66a756035dc5d28cb3554779ac1214cbede848e0097653bd8b4e574405d07

Observation d5d72a91-35b8-4408-9c4c-628b2ff4fabc · inbound

DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models cites this paper.

DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:31:36.855085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:31:36.855085Z digest=sha256:be13c45024124e7e53e39619df3ca34f1b0e0e750058560d1908ecea84139d03

Observation 27a21b8e-3b2d-4ad6-8b6b-c8b3cfe7171e · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.253757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:24eee02cfd5fea418ecb543a8f3801a91b8ffafc457fbd3d80fa0f2e648f49f6

Observation f99a26e0-6a84-471e-9778-0f011d2d64ae · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.662855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.662855Z digest=sha256:69dc11616cd8cad9d9fdc1315d20a76faaefa04d2451383c0ef648906d1bab03

Observation 42c0552b-b785-414c-8216-a17b76ff1262 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.790703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.790703Z digest=sha256:1f504cf8317054918aaa48adc18850f5e3aea47fc49e39b4c709be59d7362845

Observation aad61dc5-5b86-47f0-9cac-5ddc9a501f25 · inbound

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking cites this paper.

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:11.735753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:11.735753Z digest=sha256:0dba34a2a2036199d666bf77fd8eb594cf02cc66a32d9d0d83fed5449ffdad02

Observation 42bc1c55-1abe-433d-827b-4849628fe22f · inbound

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models cites this paper.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.758465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.758465Z digest=sha256:0b15a8229bdd1ae9c7b210f4aef94ca3cc7d46e1def7008089334b61151411fb

Observation ecff9ba3-20f8-41fd-9c68-3dd9d1ecab2e · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.954939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.954939Z digest=sha256:89cb74f36f433d2ff826cea0b7c22c14378f68f28c952ce97e045f0c2749a55f

Observation c040ae44-2846-4573-ab31-c0b068f5edca · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.709273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.709273Z digest=sha256:36be8a3d17dcec87f5b0e367152f1f593a4613440f577c23d45d819a4b4acdad

Observation 31365a7f-c765-4263-a232-b30f7ed0088d · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.752574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.752574Z digest=sha256:df845e5220dcf8d28e8da8532e94dd9730189378e9325bcc509c6b4fba10e5f0

Observation ea2a5ef6-e5f0-4db7-bc73-a67dbf644f8d · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.854770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:af37627e705fb1848ce89fc14894bce7e155315ce445583ad7e1dcb4c672f9a8

Observation 3984b339-064a-43db-9912-503af55cc6e9 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.362333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:23a6132f6b28e212cf98c053b67a06530b69b0d1433a82a162fcede12c9bf761

Observation 810e2824-ad71-43d2-b5fb-4cbea895bf9b · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:21.845958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:21.845958Z digest=sha256:09752e472ffb78fc8c2cba72bdc5058efa37ec807d4dff54afcf094ce385d8f1

Observation 8c015068-9151-440b-b307-0fba713fcd06 · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.979543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.979543Z digest=sha256:f8174be56f0015e8ca805729224077e9a6c646e990ad8772bdcb116d91736e92

Observation cc1567fc-3ba2-4fc9-9c42-1ccde88ab3d8 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.542586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:52.542586Z digest=sha256:d3bca4a5f7f0bf3e98a580aab176eb84da531e4eaaf989eaa7b1ce4cd21e6d26

Observation 543ae786-de83-4dbb-b149-4ecbc8f6fb43 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.134795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.134795Z digest=sha256:1edb37da0cc6934c33efb757340e9f59464197d39783576f3dddc01ef207b0af

Observation 88447597-4568-469d-a634-9e20ed3c5b47 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.931321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:9402b08fea69a2437039aec9485e914aec76d0d82cd3920250d60cf60c5d0bed

Observation 57eecde9-6e1e-4448-9fe3-a94055097536 · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:19.334924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:19.334924Z digest=sha256:d0148081e166843eea205f50ff058a8ebcf1451828ebce56c227b00855163301

Observation e15e1f92-b278-4d04-89e0-b91bb83ed2e0 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.922398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:1f54e4856b2cb5b33af12777c41f9c658498c045a2a2b9cb7421014ee74b42f2

Observation 0ed651f1-a7c3-44ae-8b8a-c60a789c91ac · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 279

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.480709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:a84af438c8e604d7b8c87096a28a790fc80f9570bc0431920f17709d582c3192

Observation a6cd8e01-3f80-4ee3-86ce-87e3715168a5 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.426378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:3a882d81b61b65a1ff3272850b469f9b9917d2a4ce812650f32ed7a24e9306ba

Observation f03fc67c-d05c-4124-ba0c-b89f0ee94c29 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:be54d8396040d64e5566afa67bd59d15974be394ac76fb4a79d2ef7661869077