Pith. sign in

Paper Citation Record · LEDGER

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

As of 7 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2507.06272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06272 v3

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:03.038918Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.109470Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved40
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07667d61-3b90-4a50-90bc-9f7a8a79023a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.560688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.560688Z digest=sha256:605bf011539f2c287e5c482ddaf587e3a253732084ea0afad9f00d5152c98069

Observation c1f30626-190d-43bc-819e-4e08e8f582bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.566717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.566717Z digest=sha256:91941a3aa6a8e9d7c516dae22935e904583f56abbf0456b7a2f9a011e21926c7

Observation d23a0cfd-ddfe-4b24-8f4c-4404ada3607a · outbound

This paper cites Internlm2 technical report, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Internlm2 technical report, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.472428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.573028Z digest=sha256:4f5689c3b80c98ccb323e64a48346765b17383008391bd5ef568c17f41daf30a

Observation 123fb9d3-7959-44c2-9874-ddae3b8b1325 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.578412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.578412Z digest=sha256:64d0cdef93eb8dda6eb0bba6e986e94f6296536312d8462e4e0d85815ae77545

Observation c3620d91-5c04-46bf-893c-40517ccf4caa · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.585377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.585377Z digest=sha256:b8b28978447b623dd18a1fb0bec771ba6fd6651f5bbb11508770d622d48d5b42

Observation 946f5be3-94c9-4f85-bc6b-7c95eeeabba7 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.592591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.592591Z digest=sha256:a8e9e1a9e5b945251b716aec87b9469879f6a810250c5b4c4c61aa4e8154701a

Observation d81d9a75-d6a1-44b0-939c-e8ea40a342be · outbound

This paper cites Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.454166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.600401Z digest=sha256:f7c3a30c28d42bfecac563c0eec809934eae2f32796aae5e168fe48d8d8aeaae

Observation b0dd2dbc-8952-4815-99c9-5af96197e5de · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.612441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.612441Z digest=sha256:2835066d439ca38c5b54b51af1e38370f017385e8a18488405c70e0e4f775080

Observation c8a9d86c-06d2-405c-bf13-9c82085b0b48 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.620079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.620079Z digest=sha256:874197e0e6d5d7defb53749e0b49188948e878271e110a32dfb9d287ffb8de3c

Observation 15eff456-804c-4064-8188-8081a0d0bf2d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.433594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.626331Z digest=sha256:ae9b7930fdeb3151ceaec028464359af4945a3a5de17eace91bd2ac257608d2a

Observation 1c9bccd3-d49d-4815-ac94-3b5dbda9a750 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Vizwiz grand challenge: Answering visual questions from blind people

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.632584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.632584Z digest=sha256:5f1431952192e2709b62d38610bbf2a1275c9fe3d99ccd74762989002ae381b2

Observation d81508aa-362c-4002-8d5a-19dc7109a3c5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cogagent: A visual language model for gui agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.402511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.639805Z digest=sha256:43430068ff5cbcf0f4caa552249c8b803b9a1d1ae88803d54424254fa66537d3

Observation 7d658121-544e-4f2b-8fa0-5497626ca68c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lora: Low-rank adaptation of large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.385021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.645404Z digest=sha256:8f350a10139fb1f3d3b7ae45f15680cdd03227190c09cfb23fade36692834417

Observation 2fd3d24c-5388-4590-a076-c05969ca92b0 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.650218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.650218Z digest=sha256:5085198d4cc166047041b6384eb300010b466ce60e7c9658b8d41fe38e7eec9e

Observation 92808272-303a-4110-968f-bc3d929e9205 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.367323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.655642Z digest=sha256:3a935e9d4c90b9833f0252e3fd75d44190eb9edbf197aa26ff5ee66af732748f

Observation 5d63ed73-3ef4-4c6e-8065-d8fcd8893485 · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Dvqa: Understanding data visualizations via ques- tion answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.660377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.660377Z digest=sha256:24b028057f77367a1ceb7ada652bc5ba37132564aa3be9942956383fce84e063

Observation be61820e-baf0-4138-88c9-9fe5c3de2be3 · outbound

This paper cites A diagram is worth a dozen images.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A diagram is worth a dozen images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.666783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.666783Z digest=sha256:416f6af192a2938cda2f1c386436a5945ae90b9f44f7dcbe4fac62c7f620edfe

Observation 5f422613-5648-437a-a5b1-21c108970889 · outbound

This paper cites Segment any- thing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Segment any- thing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.679727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.679727Z digest=sha256:b340259ff2a7098222e2f072bbde67e159548a15e0278b690c123041497cd8e9

Observation f9151211-5ed7-40c3-a995-6a6ec82e37ca · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Lisa: Reasoning segmentation via large language model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.296163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.686421Z digest=sha256:3f8791b029d0d406e0abcb59948284cd0c1aba378d2a1eb013a5124323861153

Observation 9912157e-0ba8-4078-98f6-3d9d0e1c2f97 · outbound

This paper cites Text4Seg: Reimagining Image Segmentation as Text Generation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Text4Seg: Reimagining Image Segmentation as Text Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.692538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.692538Z digest=sha256:d9ce39c4f8ac7e83d3ea93b158bf3600494d748a4e5bf894ff6ee14b613fcc1d

Observation 438b3181-0746-4800-8651-c1a3abe12590 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.699518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.699518Z digest=sha256:ba624edc02d332d23d1e1fce1cf7be8843b4e8f7b226c5fe78904cf43708c8d8

Observation e96330e5-cbf0-4c07-b9e9-957de476c346 · outbound

This paper cites OMG-Seg: Is One Model Good Enough For All Segmentation?.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance OMG-Seg: Is One Model Good Enough For All Segmentation?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.706354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.706354Z digest=sha256:1bd100ee28a2bc529caaef9b273ee289ec2ca1dc7175eae18bb28b628a075c93

Observation 6c3b72e5-80cd-47c4-ade6-52e4b22c0a6a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Evaluating object hallucination in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.267371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.713875Z digest=sha256:263a1c9c4c1a0bd0e572f79128b2a12204552c8c76a0fe6735e4c5fa558b019a

Observation c1effccf-7352-46f9-ac3a-7c10cf449f54 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.719308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.719308Z digest=sha256:ee2b414ccecc3270983fe4ed9953a4567b33e51f2428e4c5f97dbec028f42911

Observation 776f9d73-49bf-4390-8d13-8be844bf18ac · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.248510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.725997Z digest=sha256:b3a06bc4b2e0482f346b063e5f1fd6b3b87a7c9ae22e561f0c5eb164a7dd31bc

Observation 66c9b9c1-9e14-4c63-98cd-3117bb196e78 · outbound

This paper cites Gres: Gener- alized referring expression segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gener- alized referring expression segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.231802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.731847Z digest=sha256:01d55c5b2857c841c51bd45ef7558d4fb7ef328dc69850a85c6b26b6f1d17140

Observation 1309b116-be41-4492-8150-93af04f8a6ef · outbound

This paper cites Gres: Gen- eralized referring expression segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gres: Gen- eralized referring expression segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.214525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.737729Z digest=sha256:9453c955612d854c36880fdd85596596a7ac54f633170995b227e6547a7e4aa3

Observation 25cbb403-d745-40c7-9d46-33681c38d58f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.197835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.743293Z digest=sha256:3874cacec5c738688ae08a8f28117ae8c50ccc3e14f6a1d9069590458daa3399

Observation 3a6f45e1-600a-4186-b0b4-b66d8c1959aa · outbound

This paper cites Visual instruction tuning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.181474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.750364Z digest=sha256:b6157f32f9c6c8be56bdcf1425966a558ce68aa87831005823f2037499397198

Observation b6798a7d-cf83-4917-bc88-4894cf994b78 · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.756176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.756176Z digest=sha256:46e9258a09c93e5d642f9d9aebd916bdef68d6062b8d1417a3d4bc1b26f0d721

Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.762351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.762351Z digest=sha256:385b03219b31d043253dc6ef1c290517b0ff578abd588c80dacf870eb9a97ea3

Observation 608ab5c5-c239-4760-b351-dfcd833a97c9 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.768004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.768004Z digest=sha256:79977bdaf67d229e2203d36f2d5605eea0bb14bb512787db2e8edbfdf3d1f3fa

Observation 4b577851-b868-4a87-a58d-e891afe63e52 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.773169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.773169Z digest=sha256:c7b0baea25c40b3561bbcc41b4d4b135a07b84c87b565d69c273bb612619836a

Observation 2e4a3485-1b56-4e99-860b-5402c90fa88f · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.779348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.779348Z digest=sha256:584e5a514840aeaf43a301ff74a76599d53fa396cfb84098771bdd6a007c67bd

Observation be5cb8d7-a020-43c8-a85c-1bb846236cd4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.152584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.786553Z digest=sha256:d225038f5214c422c96eddd94cb6bd0bed2009f9144a45837ca6d087a211ba48

Observation c64fb980-aa3b-45df-be78-a21e7f0a2fe3 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.794185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.794185Z digest=sha256:328362c4a0ce5c3f48274d3a23ce6d492b836da3a8d2c43791310d3e2fc5bdff

Observation 327d3d31-edab-4095-808e-0eb026ee171d · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.135326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.800568Z digest=sha256:15a0b423041983c9f945fb1acf38f284bc6ddb2dab81844130890b429ec474de

Observation 5625cbc4-79a1-42c8-8255-c8f792690acb · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.806358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.806358Z digest=sha256:2b6724f8011ee91966b2fdc376a1ed8b6be072c684d586073712ebba5974f4f1

Observation 446b98f5-a107-4b23-8d14-f98ce24d6a7b · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Mod- eling context between objects for referring expression under- standing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.811525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.811525Z digest=sha256:b8c73b4dc7da650b75ceec1e159553b2809a96229c3343d87c66ffdead33270a

Observation a79209ef-18ee-4598-bffc-0ed393c22227 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.816755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.816755Z digest=sha256:770dfd1d82940eb3db0d5dac7c389f8c0fca5bd21dee16d5226f41ef88aa3d15

Observation 087fa912-cbf8-437d-bf28-1d9225376729 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.822075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.822075Z digest=sha256:e43e05636127fb5d5ec7fc7dc031c989abb8bea8a8b249c286aba73831a790d2

Observation 3ce61a4c-e2d4-4cfb-9076-cd599d3f71df · outbound

This paper cites Reasoning to attend: Try to understand how¡ seg¿ token works.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Reasoning to attend: Try to understand how¡ seg¿ token works

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.108283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.828356Z digest=sha256:08789cdb1b7f024e994a5b407b3c057fe828db083245d01e0e0704c6c15e0543

Observation f2eca709-119d-497e-88a1-f6efe28401e5 · outbound

This paper cites 10 Glamm: Pixel grounding large multimodal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance 10 Glamm: Pixel grounding large multimodal model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.091454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.834105Z digest=sha256:fd8a0d5bd54d7c44c55393ad66874a312fc9e5a20f29566b6fa09ca3f587ea69

Observation 321b42bd-a7ad-405b-bdc7-4d7fd7a93ac3 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Pixellm: Pixel reasoning with large multimodal model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.075155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.839396Z digest=sha256:7bc99bf08033d4a92d7a302f5592e663340237202b19ae049113b58a66a2f94b

Observation ef42190d-5df3-4e58-8f53-d637060f1479 · outbound

This paper cites Object hallucination in image cap- tioning.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Object hallucination in image cap- tioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.846541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.846541Z digest=sha256:2a63323805935720a6e8f7072fe75f303dcbb7d32161dbd1a91c6ea5af609d17

Observation 08f1824d-f5a8-4416-970e-15639a0b1d05 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.047364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.851753Z digest=sha256:5022b3ad3121741e77a73dbaffc2b201d690d9a4827a45cb1bb3b8f40b783745

Observation 70e555ba-6fe2-4a70-b233-858ccd208a78 · outbound

This paper cites Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Tinylvlm-ehub: Towards com- prehensive and efficient evaluation for large vision-language models, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:04.027874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.856963Z digest=sha256:ad6b42b82dca6ec1863b961bcb0a3c0731cca0393e361c410372cb92abe85ab2

Observation a1553ffa-e220-4930-84d1-3e7668aa78e4 · outbound

This paper cites Towards vqa models that can read.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.862897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.862897Z digest=sha256:c8f4da1813ef0bdae94f1281ba1ceb8f2a8486b74a6a5559e03d10137c456d40

Observation 951a132c-acf1-4cbd-b158-19561d0a73d8 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.997513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.869346Z digest=sha256:4ba0f94d7494b0a5eed257affb6bdbe2a980946c5d611a39e4b4bc41833c955c

Observation 34dc3d2c-51ae-41a9-a577-67bf4810d1a9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.874944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.874944Z digest=sha256:928fa732917b05898c98b687bc043bd9bbfcd40a35fb0b3bf685b3279bd98740

Observation 1834aff1-5fff-49d6-8c83-2fbb9966a293 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.881306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.881306Z digest=sha256:9af2191708ac4c3ea1faf55bf79f7cbecf830e27fe59400fa995b4ff41b3ac5b

Observation 6664a6e4-dea9-464e-bce3-8631bbe9691b · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Visionllm: Large language model is also an open- ended decoder for vision-centric tasks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.980701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.892908Z digest=sha256:49654f8cea7054811ad49e76b2b6dd03e999434f463a0d91b6557d90e28d293a

Observation 7727ef6d-8e2a-4681-add6-7b3405fa1997 · outbound

This paper cites Hierar- chical open-vocabulary universal image segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Hierar- chical open-vocabulary universal image segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.963856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.899849Z digest=sha256:af167f82943c2ba61c1f3c03426e663cacd2ceab70834b3d77647e8d7cbc2ed4

Observation a118ef6e-a4e0-4a5f-a839-c70b4d9bbcac · outbound

This paper cites SegLLM: Multi-round Reasoning Segmentation.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance SegLLM: Multi-round Reasoning Segmentation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.905453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.905453Z digest=sha256:65e65c2034a33e7bf992feb25ece7b4209a345c230eec8cca696a9421b844cb7

Observation 079b2c5b-371f-4f8f-8e52-c5cff871c6e6 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.914864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.914864Z digest=sha256:d55125e95aad476c8339d2966e4111e4caa135455cd97df7069ee43870f6ace1

Observation 0b659962-2df3-4eea-b750-fe0c9decdd8a · outbound

This paper cites General object foundation model for images and videos at scale.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance General object foundation model for images and videos at scale

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.946172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.920393Z digest=sha256:c64fe3000e6e40f4f5b8d9641c5340d35c331a08962e0be5f913d5905c2b713d

Observation a8036437-e011-48ec-9078-abfa965eb13e · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.926461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.926461Z digest=sha256:03d7e0d3464a59c410bdbc928ba9c4a661cc3e42820d1560298c3cdbdc9a467d

Observation 59987610-2c72-4322-9dd0-1d553565e551 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance V*: Guided visual search as a core mechanism in multimodal llms

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.929391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.934981Z digest=sha256:25c3c08ae1df5359921e1141de7f1728bce8556e4a9c2207b3a91a9a25417080

Observation 30aae116-914d-4b0e-8ae2-252453f7df71 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Gsva: Generalized segmentation via multimodal large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.911800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.940703Z digest=sha256:356165bafff15898acd1cb534482352e774df2898419f95dcf5f1a19b6ae119a

Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.946036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.946036Z digest=sha256:b0b7c9e33dfbdf834cee2a57bd74dea898a16ca302739ea50391600fc04e5b5c

Observation 991fa207-564b-42eb-b34a-12de5323ffce · outbound

This paper cites Qwen2 technical report, 2024.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Qwen2 technical report, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.894309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.952327Z digest=sha256:93c6fd7f903fc532019578ef9782419529e187b9aadb484d74ef247e1c9461d7

Observation ae5f731d-8a5b-4e55-8359-535d567b62d7 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.959690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.959690Z digest=sha256:9bf7d2a1e0a70d6873ecefb46f308b8184426ccf038d081f9e0ea1c8af3127c4

Observation 6b3102a1-ab8f-4168-8805-3f6862cb943c · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.876247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.966226Z digest=sha256:f730ce09d1aeece42e0ce6c48562ec123fff7174bc76d946dd7ef7b5e3d801c1

Observation f156745b-ab03-44ca-becf-967ee861c55b · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.971505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.971505Z digest=sha256:435574cdcf0212f7d14db0ad1accde3443aecefb24cba80757b1af1f7a598f66

Observation 2182922a-1b65-4d5c-8bf2-36da9a7155b1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.977524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.977524Z digest=sha256:b7cebf2e3d8bd57110c89ea31075e2db9c7804dfc128ba65c103dfd255f1feae

Observation 836ac2fe-de2b-4583-aa04-dd43e5df3ee7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.857363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.984296Z digest=sha256:b721494b1cdbe9a8c3242fdc965cdbc2a73ec5b73714da73a6c8528df0e4d0e4

Observation fe05fdce-4e5e-4b39-bd41-193b88ab0b47 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Ferret: Refer and ground anything anywhere at any granularity

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.839191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.991047Z digest=sha256:efcf7aa35a9c07135e69ea08fe50f463e37197271fb62f19d09029cbf41b1922

Observation 19788345-0a4a-453f-8243-d4300b29a1ec · outbound

This paper cites Modeling context in referring expres- sions.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Modeling context in referring expres- sions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.822751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.998247Z digest=sha256:0b5343a534096a8c0ed2ca8f6b4955762ab0c96ae672e3c923292eced6ba843b

Observation b9a6494e-4048-411a-9dcd-9946a68d8d5b · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.005676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.005676Z digest=sha256:ff25006bfc405e926491d2f24a3af01d5b2a43c58c61a3902b8eb7f3328dbf54

Observation fa0bda23-d1d4-4f18-ab4c-7c1957fad63b · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.012476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.012476Z digest=sha256:fb98e8778e6e7a7a92dc4739dc607a5613ed01c5faa56ce65501c411d9397e73

Observation 5faef72c-b81c-4aea-8c07-04ecd20f71e1 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.805984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:03.020693Z digest=sha256:f09ca2953017a6f782560355c46b1ce588ffb31a27841df0680c581fb143b80c

Observation d06d3946-b1e1-46c6-82e1-dcc182433d74 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Psalm: Pixelwise segmentation with large multi-modal model

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.787981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:03.026335Z digest=sha256:b15037d6c16790359f6c211d16b58a8346ba39a474cc785119966925fece1391

Observation 068d6238-60ab-4185-b6b0-29f180f24a04 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:03.032130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:03.032130Z digest=sha256:abc7253e4afcb911b07f6465b8c816ea1773487d57f31de3139a90c96c3743a1

Observation ffabad69-66d6-401c-9b0b-cfdb05f446f3 · outbound

This paper cites When the provided information is insufficient, respond with ‘Unanswerable,’.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance When the provided information is insufficient, respond with ‘Unanswerable,’

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:25:03.769446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:03.038918Z digest=sha256:83eeed9b300d0eee7361e2d3a51166c0972dfbd1bf628f77de11b6f289d2a170

Observation fd18e0b8-357d-4092-a28a-536a65f6a9a0 · outbound

This paper cites an unresolved cited work.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T19:25:04.323807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:25:02.672598Z digest=sha256:fc1aa4048e1757fe28ec778d8a26843ccef7080236c5d8f9956acb77f90d5945

Pith citing papers

Observation fa7ad16a-b5f6-4ed2-bd67-b6d1ce5d30c6 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.111055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:4f6a9266736961f16fad42d3fa3c3b24217c7c4d92d0082a4ea08ff9c9577d84