Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:03.038918Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2507.06272.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:03.038918Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.109470Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 07667d61-3b90-4a50-90bc-9f7a8a79023a · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f30626-190d-43bc-819e-4e08e8f582bc · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d23a0cfd-ddfe-4b24-8f4c-4404ada3607a · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Internlm2 technical report, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 123fb9d3-7959-44c2-9874-ddae3b8b1325 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3620d91-5c04-46bf-893c-40517ccf4caa · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946f5be3-94c9-4f85-bc6b-7c95eeeabba7 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81d9a75-d6a1-44b0-939c-e8ea40a342be · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0dd2dbc-8952-4815-99c9-5af96197e5de · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a9d86c-06d2-405c-bf13-9c82085b0b48 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15eff456-804c-4064-8188-8081a0d0bf2d · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c9bccd3-d49d-4815-ac94-3b5dbda9a750 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Vizwiz grand challenge: Answering visual questions from blind people
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81508aa-362c-4002-8d5a-19dc7109a3c5 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cogagent: A visual language model for gui agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7d658121-544e-4f2b-8fa0-5497626ca68c · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lora: Low-rank adaptation of large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2fd3d24c-5388-4590-a076-c05969ca92b0 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92808272-303a-4110-968f-bc3d929e9205 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d63ed73-3ef4-4c6e-8065-d8fcd8893485 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Dvqa: Understanding data visualizations via ques- tion answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be61820e-baf0-4138-88c9-9fe5c3de2be3 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A diagram is worth a dozen images
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f422613-5648-437a-a5b1-21c108970889 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Segment any- thing
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9151211-5ed7-40c3-a995-6a6ec82e37ca · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lisa: Reasoning segmentation via large language model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9912157e-0ba8-4078-98f6-3d9d0e1c2f97 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Text4Seg: Reimagining Image Segmentation as Text Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 438b3181-0746-4800-8651-c1a3abe12590 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96330e5-cbf0-4c07-b9e9-957de476c346 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance OMG-Seg: Is One Model Good Enough For All Segmentation?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3b72e5-80cd-47c4-ade6-52e4b22c0a6a · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Evaluating object hallucination in large vision-language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1effccf-7352-46f9-ac3a-7c10cf449f54 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776f9d73-49bf-4390-8d13-8be844bf18ac · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66c9b9c1-9e14-4c63-98cd-3117bb196e78 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gener- alized referring expression segmentation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1309b116-be41-4492-8150-93af04f8a6ef · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gen- eralized referring expression segmentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 25cbb403-d745-40c7-9d46-33681c38d58f · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a6f45e1-600a-4186-b0b4-b66d8c1959aa · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visual instruction tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6798a7d-cf83-4917-bc88-4894cf994b78 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608ab5c5-c239-4760-b351-dfcd833a97c9 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b577851-b868-4a87-a58d-e891afe63e52 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4a3485-1b56-4e99-860b-5402c90fa88f · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5cb8d7-a020-43c8-a85c-1bb846236cd4 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c64fb980-aa3b-45df-be78-a21e7f0a2fe3 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327d3d31-edab-4095-808e-0eb026ee171d · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5625cbc4-79a1-42c8-8255-c8f792690acb · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446b98f5-a107-4b23-8d14-f98ce24d6a7b · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mod- eling context between objects for referring expression under- standing
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79209ef-18ee-4598-bffc-0ed393c22227 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087fa912-cbf8-437d-bf28-1d9225376729 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce61a4c-e2d4-4cfb-9076-cd599d3f71df · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Reasoning to attend: Try to understand how¡ seg¿ token works
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2eca709-119d-497e-88a1-f6efe28401e5 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance 10 Glamm: Pixel grounding large multimodal model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 321b42bd-a7ad-405b-bdc7-4d7fd7a93ac3 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Pixellm: Pixel reasoning with large multimodal model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ef42190d-5df3-4e58-8f53-d637060f1479 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Object hallucination in image cap- tioning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f1824d-f5a8-4416-970e-15639a0b1d05 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A-okvqa: A benchmark for visual question answering using world knowl- edge
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70e555ba-6fe2-4a70-b233-858ccd208a78 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1553ffa-e220-4930-84d1-3e7668aa78e4 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Towards vqa models that can read
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 951a132c-acf1-4cbd-b158-19561d0a73d8 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 34dc3d2c-51ae-41a9-a577-67bf4810d1a9 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1834aff1-5fff-49d6-8c83-2fbb9966a293 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogVLM: Visual Expert for Pretrained Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6664a6e4-dea9-464e-bce3-8631bbe9691b · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visionllm: Large language model is also an open- ended decoder for vision-centric tasks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7727ef6d-8e2a-4681-add6-7b3405fa1997 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Hierar- chical open-vocabulary universal image segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a118ef6e-a4e0-4a5f-a839-c70b4d9bbcac · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance SegLLM: Multi-round Reasoning Segmentation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079b2c5b-371f-4f8f-8e52-c5cff871c6e6 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LaSagnA: Language-based Segmentation Assistant for Complex Queries
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b659962-2df3-4eea-b750-fe0c9decdd8a · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance General object foundation model for images and videos at scale
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8036437-e011-48ec-9078-abfa965eb13e · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59987610-2c72-4322-9dd0-1d553565e551 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance V*: Guided visual search as a core mechanism in multimodal llms
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30aae116-914d-4b0e-8ae2-252453f7df71 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gsva: Generalized segmentation via multimodal large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991fa207-564b-42eb-b34a-12de5323ffce · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2 technical report, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae5f731d-8a5b-4e55-8359-535d567b62d7 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b3102a1-ab8f-4168-8805-3f6862cb943c · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f156745b-ab03-44ca-becf-967ee861c55b · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2182922a-1b65-4d5c-8bf2-36da9a7155b1 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836ac2fe-de2b-4583-aa04-dd43e5df3ee7 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe05fdce-4e5e-4b39-bd41-193b88ab0b47 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ferret: Refer and ground anything anywhere at any granularity
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19788345-0a4a-453f-8243-d4300b29a1ec · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Modeling context in referring expres- sions
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9a6494e-4048-411a-9dcd-9946a68d8d5b · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0bda23-d1d4-4f18-ab4c-7c1957fad63b · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5faef72c-b81c-4aea-8c07-04ecd20f71e1 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d06d3946-b1e1-46c6-82e1-dcc182433d74 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Psalm: Pixelwise segmentation with large multi-modal model
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 068d6238-60ab-4185-b6b0-29f180f24a04 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffabad69-66d6-401c-9b0b-cfdb05f446f3 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance When the provided information is insufficient, respond with ‘Unanswerable,’
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd18e0b8-357d-4092-a28a-536a65f6a9a0 · outbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Unresolved cited work
Reference 251
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa7ad16a-b5f6-4ed2-bd67-b6d1ce5d30c6 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
Reference 144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.