Pith. sign in

Paper Citation Record · LEDGER

LISA: Reasoning Segmentation via Large Language Model

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2308.00692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.00692 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:29:35.371817Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:44:27.641390Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e7bc3476-42d1-48eb-a25f-60a498b89f61 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:41.946901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:73be33d31598892ea9ca62144df28124df538df58ae856b0629d5465bc792db1

Observation e800d743-7f54-4c14-9bca-fbd2d8db1d48 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning LISA: Reasoning Segmentation via Large Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.878106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:bea94ddaca928b5f33481d357255f52121935fe506ae4a67759893845146d645

Observation 440c846a-a9c6-47c4-a20a-2a8032f7999a · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks LISA: Reasoning Segmentation via Large Language Model

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:09.950334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:16447472383a0652bd89a287b3a80c1904febefc4e9dada2d359d50861416ebd

Observation 74b58cc9-8d61-4a47-9fe9-799368bf3e75 · inbound

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels cites this paper.

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels LISA: Reasoning Segmentation via Large Language Model

Reference 220

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:35:48.066345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-15T16:35:47.826165Z digest=sha256:d03c22118c0d40f8be2920903c86163153f71b167a83fc61965da4b66afb429f

Observation 99565bf8-9a88-4cd7-a99c-c54b89af5064 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.320280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:e9880b5c4d5edd76ecfa69307520119bb561fd6215d56f0d2466d79d24464c6c

Observation f2f92ab5-4202-4cf0-a0bd-328334a5928e · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LISA: Reasoning Segmentation via Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.005910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:cb54ab06d821ee813824f3eba9ae4a86b35ea4deb877f736895630506f114c5b

Observation 3b836275-5181-4428-b431-2b6e90226b10 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LISA: Reasoning Segmentation via Large Language Model

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.331924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:1febd468bd458e2612ad36994f5448ef774b3451e288e91136305803b40f085e

Observation e25276ac-475d-41b5-9f0b-b2043af2d0a1 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.401414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:919a0f622a59d1be93c3c94fa21ee6b64ad4e42b403329d59a5abeb905afc02b

Observation 90548e61-3ef4-40fe-99a3-20453b93b910 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites LISA: Reasoning Segmentation via Large Language Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.004457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:5284d4c5b283a0df5912ca7518decac150cf846b440a2cbdd1fbc1081b75db9e

Observation d9973341-a54d-4963-89ee-fdaeb533ec6d · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey LISA: Reasoning Segmentation via Large Language Model

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.904702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:e3bee35e4003fea48144191138b67a53e2320ffe938b5564028fa6f88294634b

Observation 16432e8c-11ae-4110-8cbf-15c8e11bc4ab · inbound

Retrieval Augmented Recipe Generation cites this paper.

Retrieval Augmented Recipe Generation LISA: Reasoning Segmentation via Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T21:29:35.371817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:29:35.371817Z digest=sha256:ca99a4f1dbf4d360e4fe8cd578766a6d94b8b13344f6fb63c5b4166221ea2831

Observation 57d4be7d-a371-4c96-8e2e-ccfd64ff52f3 · inbound

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level cites this paper.

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level LISA: Reasoning Segmentation via Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:46.439797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:46.439797Z digest=sha256:ee543c74c2d9858290fddb9cc9277fa532da43cd906188044f6f590d93ed306e

Observation 4001b619-8fb9-4b15-9588-efb73c8522e5 · inbound

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era cites this paper.

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era LISA: Reasoning Segmentation via Large Language Model

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:24.652762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:24.652762Z digest=sha256:065d87c2e9eb1f939c74570189ae064e29fbb80422669389c5d07cb0da2d2670

Observation 2f37db29-59a6-42d9-bc99-caed9a1506e3 · inbound

InsightEdit: Towards Better Instruction Following for Image Editing cites this paper.

InsightEdit: Towards Better Instruction Following for Image Editing LISA: Reasoning Segmentation via Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:19:22.472727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:19:22.472727Z digest=sha256:5bf8fc7f68d5179890e1074c506cf575a07e9eb5727269593bfbd90fa4a6b1e8

Observation 2e9ac932-b0fd-48a0-875c-929d4c08e1b0 · inbound

HyperSeg: Towards Universal Visual Segmentation with Large Language Model cites this paper.

HyperSeg: Towards Universal Visual Segmentation with Large Language Model LISA: Reasoning Segmentation via Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:03:54.124205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:03:54.124205Z digest=sha256:c56ea2b65c455c4ae690635b254fa0c3fdd7fdafc11724e4566edd78171cfca2

Observation abd9a563-b0ef-4a01-992e-f118bbf77b27 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LISA: Reasoning Segmentation via Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.570163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.570163Z digest=sha256:bc8aac9cb87550a613089439dc2fc6706b78f5c90f131920087267635177eafe

Observation c081d1cf-2291-40a6-add3-b52aaad0501a · inbound

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases cites this paper.

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases LISA: Reasoning Segmentation via Large Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:37.036030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:37.036030Z digest=sha256:46d5ead6af77d4214791c274e810288b32393b7c75f051f79ce773049dde2d37

Observation 948e5ad0-0b29-4ebc-b27c-8d4548f9c9af · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM LISA: Reasoning Segmentation via Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.053665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.053665Z digest=sha256:6ff028e48d4f8b35670807573a0ece6c92746fee448551a0c196c58404ed45cd

Observation c37f5ad3-671c-4868-afa1-58cd9c28b6cf · inbound

AIpparel: A Multimodal Foundation Model for Digital Garments cites this paper.

AIpparel: A Multimodal Foundation Model for Digital Garments LISA: Reasoning Segmentation via Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:02:03.848908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:02:03.848908Z digest=sha256:11f91ce337a2e4887c424504058942fb81700e47aa041a18ff8284f8b2e9cbf1

Observation c35744ab-b3ba-4957-b82b-3cb9040a231d · inbound

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models cites this paper.

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:59.529983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:37:59.529983Z digest=sha256:5bd74331929fd34707b4ce5a9e0c00e2f2c3820bc5b30a0b29ed5b5af2aafb57

Observation ab6b5c80-ea47-4e38-ab44-fadae128fc91 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LISA: Reasoning Segmentation via Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.951403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.951403Z digest=sha256:3db5c2386f2b0664fc31b97c097cf6c2b30c67c0199724382bc368a3f007efc4

Observation b1f5dab7-0915-4a14-8676-0d8951c142fe · inbound

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing cites this paper.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LISA: Reasoning Segmentation via Large Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.552690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.552690Z digest=sha256:fdd105444df98486f206d5dec9f851c6b5a54d50bc72905b3bcc2486e7684c80

Observation 57551ff5-db4a-46ff-b974-7b06e4df53af · inbound

Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation cites this paper.

Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation LISA: Reasoning Segmentation via Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:25:41.922273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:25:41.922273Z digest=sha256:1c07dcc5508c27006a08de2ebc2c42c320991f837fd8c8dfdcc0ff253ca16680

Observation 2ba89b2b-c607-46b3-9c58-1a1a3c8095b5 · inbound

Pixel-Level Reasoning Segmentation via Multi-turn Conversations cites this paper.

Pixel-Level Reasoning Segmentation via Multi-turn Conversations LISA: Reasoning Segmentation via Large Language Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T21:32:35.621068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:32:35.621068Z digest=sha256:1668c338e2904468b04bbbddd06b3f1ba269212af1fa7ce2b6751a09387bb2e0

Observation 69848a13-51bb-4a10-a771-a3e89c5162d9 · inbound

On the robustness of multimodal language model towards distractions cites this paper.

On the robustness of multimodal language model towards distractions LISA: Reasoning Segmentation via Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.469758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.469758Z digest=sha256:00dd80c69533bf9935aef54adf8c81399bd68a000f15b5b052c0620864189523

Observation 22f767ef-56a2-4155-a22e-5bbff6497ade · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LISA: Reasoning Segmentation via Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.756541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.756541Z digest=sha256:233c925ae321ee406b2c98bd225a8d073ffaf223c6baa07104d9e2adc1e07268

Observation 11113140-16c8-4c23-abf7-a76ebf07b908 · inbound

STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset cites this paper.

STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset LISA: Reasoning Segmentation via Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:41:45.382441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:41:45.382441Z digest=sha256:af48769b23f51889e2a05f5d9656e7c3ed3d31a4d5d52ae5c050e145f4d07535

Observation e6abe27b-963b-47ef-8e03-5daf892aa2b1 · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing LISA: Reasoning Segmentation via Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:27.498137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:27.498137Z digest=sha256:0b01c77222388246fe8610c09788f1c9e27951792b9049aa870bd9d0b982eea5

Observation 3e5e4640-7317-478a-b48b-99a5e8a3afb8 · inbound

MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models cites this paper.

MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:29:18.093734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:29:18.093734Z digest=sha256:b23396ca12393d85c848783959fce445478d6657de87ca4d811e7b17b8b4d963

Observation 273f83b8-2f55-4422-b866-b56ff5e11ddb · inbound

ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation cites this paper.

ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation LISA: Reasoning Segmentation via Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:41.170936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:41.170936Z digest=sha256:c261fbb4708c4cbc937645b380f38cf77dd5a736bde50b1b0a167485efa2f320

Observation e1982561-f511-4d16-a33c-2e135b68235f · inbound

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive cites this paper.

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive LISA: Reasoning Segmentation via Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:42.435568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:42.435568Z digest=sha256:1860307fc939eb05f09be14f063a9e8e23c051b14da16d9afb33036c68def2c7

Observation 096df928-9178-48e5-8a28-324a82b4db17 · inbound

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model cites this paper.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LISA: Reasoning Segmentation via Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.963924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.963924Z digest=sha256:7e5d6ca9d371c91bdf614028fbabee50fb8141340e53c4208ad778dfe579f3c9

Observation 781612d5-9bc4-4af2-812d-dee76c2e90ef · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning LISA: Reasoning Segmentation via Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:56.916800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:56.916800Z digest=sha256:108b97580be03b273699527a5db76973cf52e8ab913e8c46ffa5469930788c2a

Observation 78686686-890a-4340-87c9-c72b035b5a40 · inbound

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding cites this paper.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding LISA: Reasoning Segmentation via Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.413798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.413798Z digest=sha256:6473ea5fe2d27412a67c0104e9a576bdd843aa71964f2ec3481b103300dab949

Observation 8230abef-3c1f-45fa-a6dd-28c842cf370b · inbound

Advancing Visual Large Language Model for Multi-granular Versatile Perception cites this paper.

Advancing Visual Large Language Model for Multi-granular Versatile Perception LISA: Reasoning Segmentation via Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.031233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.031233Z digest=sha256:f7cd837bb1ba5e43506c851b04bfdd4a3fc450e826fc30bdf77e65ebc4fad4f3

Observation 9fca09ee-667d-40b0-9ce7-6108f3bd8d1f · inbound

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver cites this paper.

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver LISA: Reasoning Segmentation via Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:32:44.554728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:32:44.554728Z digest=sha256:8bebbd02d43fc79ebce50237adc332d844b21d90f49182174d63dec79c7585cd

Observation 942099c8-ae51-4134-9c6f-d153f2cce095 · inbound

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes cites this paper.

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes LISA: Reasoning Segmentation via Large Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:36:34.014108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T15:35:30.549656Z digest=sha256:4d74fadf0e1594c0123170e00826790c8d06c0d305a92d687c1a2c02736cef59

Observation e27b91bd-f254-44cf-a1a9-9939fb1c3a6e · inbound

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models cites this paper.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.580737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.580737Z digest=sha256:86432235629dc0821104ff0478b607114d4a569082ed078124fc80b1d3ea3656

Observation 67864beb-28a3-4eae-bb71-be687ea84ac1 · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding LISA: Reasoning Segmentation via Large Language Model

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.760001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:d15f46102249d8f51d7a4799e0eb091766dbc988ef8327ff9216d11a00717f12

Observation 708a6f30-9b32-4808-ba10-c61796725e5d · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM LISA: Reasoning Segmentation via Large Language Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.377974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:f1693cd994fffb423c1a34e77455de29134902b095553b902359abd2c1a1501e

Observation 996455f2-3463-4e25-982e-58266447e76f · inbound

Moondream Segmentation: From Words to Masks cites this paper.

Moondream Segmentation: From Words to Masks LISA: Reasoning Segmentation via Large Language Model

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:13.708703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T20:14:45.629804Z digest=sha256:98a3e22736d38b9bb862fa37c1dc5692da9a16d3fb34b95b7f53985592d325e4

Observation f26d572c-4106-4450-a5e5-e066623f7380 · inbound

WildDet3D: Scaling Promptable 3D Detection in the Wild cites this paper.

WildDet3D: Scaling Promptable 3D Detection in the Wild LISA: Reasoning Segmentation via Large Language Model

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:26:00.672225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:38:13.336003Z digest=sha256:f1e4abc0f523f00d3e15934c76299c9ae08fdfd012b647b9b0a1dd30ce4e6aa8

Observation c053b920-dbe8-4a68-a714-ab6cf6b669dc · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation LISA: Reasoning Segmentation via Large Language Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.407245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:ce13a98b11673513cc69f71b8b74fc30ae8b354e19e1fe604fbe05955ee67c7e

Observation 563f29a5-3d2d-4865-9859-216e933c1bf7 · inbound

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs cites this paper.

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs LISA: Reasoning Segmentation via Large Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:34.444982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T19:19:09.199731Z digest=sha256:14375e82b7faeace0b08b15524eaf27d57fd634f244d55b52d6c937e9e22f0d6

Observation f38ad578-1b2e-4d30-9651-b8bca572e351 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models LISA: Reasoning Segmentation via Large Language Model

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.643086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-08T02:44:00.608590Z digest=sha256:d99fad9081193f95c3d6f34e453fd9b23c7e4689095a40911f321f7edc45a4bd

Observation 92a261bf-a0aa-4bd0-af57-212de3b6fb8d · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models LISA: Reasoning Segmentation via Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:910068ef6ed28b66f2c539fbc0ad579601ef8a6a3db65a00378dec6b9a094f0f

Observation 4cdf2020-110e-4e22-9c09-901df2ceb4c8 · inbound

Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation cites this paper.

Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation LISA: Reasoning Segmentation via Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:48:46.261946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:48:46.261946Z digest=sha256:1ccbc0c829017e4fbd780e94b90b495c2d0843959247e483e755f57b887867d5

Observation dcdd51bf-6c3a-4e16-a828-df52d40381e8 · inbound

Vision-Language Grounding as Bidirectional Concept Correspondence cites this paper.

Vision-Language Grounding as Bidirectional Concept Correspondence LISA: Reasoning Segmentation via Large Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.830162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.830162Z digest=sha256:f6199f20c1fe5376e4d267b4b0c33ef63ed394aa0ce8c5776147648e5de230b7