Pith. sign in

Paper Citation Record · LEDGER

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling

As of 15 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 2 inbound Pith citation observations for arXiv:2509.06321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06321 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:22:12.020753Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:00:20.216268Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:04.772242Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 716f03b8-86ab-4e08-8c8d-74761c3720c9 · outbound

This paper cites A survey on multimodal large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling A survey on multimodal large language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.644442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.644442Z digest=sha256:8ee4398223a6b3fb2fa5ead51551a3d185ffb74980906283439ff5cb3726a2d9

Observation 0f62a146-a2f4-4d63-9bc6-f5571e8f2534 · outbound

This paper cites Visual instruction tuning,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Visual instruction tuning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.649338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.649338Z digest=sha256:b938c912ab9d2d8ee5e30cd7ce5649cfa740fa4ff6c38f96e1a92a96e64afc1e

Observation 8dc1dd1d-1d75-4478-99a3-d81a985a0b0c · outbound

This paper cites Deepseek- vl: Towards real-world vision-language understanding,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Deepseek- vl: Towards real-world vision-language understanding,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.653667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.653667Z digest=sha256:5074766d12095c552b0f422e1d0bc3891b6dd06254bd7183c5cd83aad03371b7

Observation b1d28aa1-1afb-47d4-89ab-1b0f6dd3b58c · outbound

This paper cites Improved baselines with visual instruction tuning,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Improved baselines with visual instruction tuning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.657945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.657945Z digest=sha256:80e4e48c8eb3c2a5a3b64b6db078971ebca28bdd6377265affca5e173558da8a

Observation 4d2d77bb-cbc9-4ddb-b413-26d7105d731f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.662006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.662006Z digest=sha256:c82bdacc3b164d345396ce69438229ed013d5800365500fb29d69c76c9f48feb

Observation 5a9bf968-02f8-4a54-bb14-9d970c2c971a · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.666105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.666105Z digest=sha256:543f802b129194318bccb1c123515fd61b8d78d098483c01401e38ce386c97d1

Observation caee47b3-16b0-426b-83da-88d573f67b5c · outbound

This paper cites Moma: Multimodal llm adapter for fast personalized image generation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Moma: Multimodal llm adapter for fast personalized image generation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.670520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.670520Z digest=sha256:4dca6e7469903a54545c3ed2909e7d8e0ee01407936db3a956ef39f8c718b35a

Observation e6fbfd26-846e-440f-b710-2f666d48e2b6 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image generation and editing,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Genartist: Multimodal llm as an agent for unified image generation and editing,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.674849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.674849Z digest=sha256:c9d3b2fb80b5b0df6017695545bd4a308619e902d2420ad3d80d83d3ce59853b

Observation f1d1bcbb-da5a-485d-80d1-2c248c5cca09 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.678251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.678251Z digest=sha256:6a6ba397a8d66470c8ba1e17a369c0739cf324579890d11db8d952938a36f057

Observation b331bf85-980c-48c2-9934-994d1a0e9ce7 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Groma: Localized visual tokenization for grounding multimodal large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.682149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.682149Z digest=sha256:ad9abda39251fe857b729adc60c700c1fb615265f0858490c38ba282170e6f25

Observation 94569e89-143d-482b-b03f-8cb59540ccbe · outbound

This paper cites Towards open vocabulary learning: A survey,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Towards open vocabulary learning: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.686481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.686481Z digest=sha256:9df8a64f57f6002321ac23a61c6d9cb0fd4807f5fbb50cd89fe2f5195cfa8ffe

Observation 8b45ef7e-c52e-4f57-bdae-c0cd4a2de9df · outbound

This paper cites NExT-Chat: An LMM for Chat, Detection and Segmentation.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.690392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.690392Z digest=sha256:dfc2f81c77738c182af23b721869b0d9998fd98864506c2327aeedd3bc982de8

Observation c90b2773-5ec2-4b61-b6bc-3826548b16ba · outbound

This paper cites Transformer-based visual segmentation: A survey,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Transformer-based visual segmentation: A survey,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.694066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.694066Z digest=sha256:1b828f4d366450e6d5826290d82fcf8678a18e4d93b5bb881156803124d67e66

Observation 4cae6a04-e525-4a39-8aec-5334820936e0 · outbound

This paper cites Smooseg: smoothness prior for unsupervised semantic segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Smooseg: smoothness prior for unsupervised semantic segmentation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.698288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.698288Z digest=sha256:208b13cbdb49c9315dfa4d774f0f7bbc63ac9ecd09d45a83cb5d9955756a8b52

Observation 4c81b0c7-3da7-41b1-a837-d777d9576340 · outbound

This paper cites Lisa: Reasoning segmentation via large language model,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Lisa: Reasoning segmentation via large language model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.701900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.701900Z digest=sha256:c03c66895b4cdae1833bca41a139eb7fde1f5ba153e3c918f711b89a928f8682

Observation 39fbcf90-08c0-4bca-a33e-f78a08036d04 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Gsva: Generalized segmentation via multimodal large language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.705717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.705717Z digest=sha256:a5aa650251fee2098a9553db15dbea0dfd3cdb490c2b9332871498bc5703bea3

Observation 53d79d32-df32-4514-b455-bf21be807aa1 · outbound

This paper cites Groundhog: Grounding large language models to holistic segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Groundhog: Grounding large language models to holistic segmentation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.709253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.709253Z digest=sha256:171a11d3b7f59908818743a3a5c5377f7c11bfc38cd20e22f669c6360ed7f5b3

Observation d4bfe972-1a09-4597-adb3-ad423ed28433 · outbound

This paper cites Multi-modal instruction tuned llms with fine-grained visual perception,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Multi-modal instruction tuned llms with fine-grained visual perception,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:13.092047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.712961Z digest=sha256:94bf7e0f4a3a5fb392eb18f900a2bc86bd4f30ed921516b5b27dde4f7c0574af

Observation c682b7c6-6b0b-431b-b2b6-d712e43d3a44 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Pixellm: Pixel reasoning with large multimodal model,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.716782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.716782Z digest=sha256:95e7bf7ac854d80220d3143b32a68eeb3888e47962129d519cd789234791cd96

Observation 8d824f5d-66f3-4eb5-920f-0e1ba053b51e · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Glamm: Pixel grounding large multimodal model,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.720335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.720335Z digest=sha256:ba19d77503e459b16c747e993585dbdc8b6c30b3c12b5c516f2de984ee6853c7

Observation 4b0231c8-78cc-47bb-bd3a-4d8fb6362b3f · outbound

This paper cites Segment anything,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Segment anything,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:13.061678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.723550Z digest=sha256:0ebcdb08b113f659e59f54ebaab5d4aa62a72779db6b91ae1468ffcc0ce178e5

Observation 2757456d-95a0-4a9d-8364-b9ccd471cf09 · outbound

This paper cites Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:13.049975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.727281Z digest=sha256:5ce86be26fc864d15599f3bc9e82ac97073b3eac8d654b6d5aa32f0a62ea0b65

Observation 8ce7c468-7c2d-4911-b0df-892351472ea5 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Florence-2: Advancing a unified representation for a variety of vision tasks,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.730552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.730552Z digest=sha256:d8f2839ef2727862c339e4400e7d5c799a53da9f77ec6e1e0307cef42cdca13e

Observation 5f83132b-68e7-42eb-a17b-282989b0cc3d · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:13.028829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.734262Z digest=sha256:cd6f122dfdfb7d2c2f9d234152e124794b3efd9ee4db293a9b2b85d5d1fb73e4

Observation 2c00a354-3c52-442e-95d3-202f7297586f · outbound

This paper cites Text4seg: Reimagining image segmentation as text generation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Text4seg: Reimagining image segmentation as text generation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:13.016913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.737875Z digest=sha256:28f47b393fcf9d97e3e96d69d582ac9280de000a2041ffb098ec24d4348467e8

Observation 3340fec4-8112-4d05-b5d3-b798fd76bfaf · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.741490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.741490Z digest=sha256:3b00f6f01f591ecbf87c189c8830ff97acd091860e6ba61eb6a56a16b400563d

Observation 89d2949c-35d2-43a8-b989-53a6f514d1df · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.745868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.745868Z digest=sha256:cbf2ecfdc9a2e3483c61e7586fb0efcc12fb0cc76b95b919998b29500e67cff9

Observation a42fe197-fe8a-4913-9466-e31f781d7447 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.749956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.749956Z digest=sha256:7b79dd5bbe88bc3ca6c6da4ee5d35d9ef773b791d4b10b1fe79a764f61d352a5

Observation 82cff4a6-f430-488c-b07e-fa39a01ab7d9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.754625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.754625Z digest=sha256:793feafb75232a3fcf4cef3f68d4d45fd063e080dc42f07ce24c4b4f9f1060c2

Observation edecca24-37f4-4c2e-9ca0-c4797a88165d · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Flamingo: a visual language model for few-shot learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.763615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.763615Z digest=sha256:3f5b41c4880f5a60e41bb5df9a75c70679e5dfeb45a2adccf01116748e8a3d5e

Observation 21d16457-5f00-4542-a3f3-967ad991e07b · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.768709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.768709Z digest=sha256:56938c956654834622b5b832d7e5ff0ff3bb0c92bd255097bd2350fb9d3274a2

Observation 446592e4-4679-473f-b55a-3ee3f890c9fb · outbound

This paper cites Otter: A multi-modal model with in-context instruction tun- ing,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Otter: A multi-modal model with in-context instruction tun- ing,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.772986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.772986Z digest=sha256:2b6fe9cad85c66c1efee18cbd438378c6697efb69f0bf7a7d5fb7b2139b02bea

Observation b276f24d-9047-4ef1-8600-15cbdc992481 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.777847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.777847Z digest=sha256:3baca4f57740afbc3cae4b6fe9cad3dd7af8cb77799d42702241fda960050c67

Observation 128d0749-aefc-4acd-ab02-6d2d38d36b82 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling InstructBLIP: Towards general-purpose vision-language models with instruction tuning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.972530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.781394Z digest=sha256:203ed135bc8bdd356fc0d434068493c9bd43d00b55864ab2b7dde33f550da9bd

Observation dc0a4841-5d66-46b6-b9dd-a889b2aad7f5 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.785177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.785177Z digest=sha256:3ec66ab4a830b33691b9a983d6b5a1cb03d4ed9d15c416e8f26db6c04a71f1cd

Observation 05423cf7-ef37-4460-866a-4b857ceccefc · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.789970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.789970Z digest=sha256:5affc07ebcafda720ff0e1cd48444a93920d10adde7bf470aad1e31d33b94b52

Observation e7ebdc94-4c9d-4f1e-aaac-b5d4b732a407 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.936360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.793527Z digest=sha256:4c011f06f9c3f3369da1d4aecdc0010e3a80f7e59bba3bc748934b5378309bb7

Observation 3bd5248c-a49a-44c3-b423-d01028ee0af6 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling LLaV A-onevision: Easy visual task transfer,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.797042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.797042Z digest=sha256:0c811f82666206b85cb6f68349f9bd24756cb2d1604a00d71cf49b6d64a15a19

Observation 0524b036-e703-4ede-8226-be2ca8ba1850 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.801178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.801178Z digest=sha256:dcfe98e3ee31cc6d418180fcb3a1fc055fb432682f520d02979c4c1639c6931b

Observation f96255a2-8309-447a-b909-56648d063aca · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Monkey: Image resolution and text label are important things for large multi-modal models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.914788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.805233Z digest=sha256:89b91e02497a955da3fbe5efff0801deac826b0297de9776dbf410fadb5bffc9

Observation b9539a27-86bc-45bb-b126-6a766f17464a · outbound

This paper cites Sphinx: A mixer of weights, visual embeddings and image scales for multi-modal large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Sphinx: A mixer of weights, visual embeddings and image scales for multi-modal large language models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.903353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.808587Z digest=sha256:7615d9a0d8570914568a414eee742f8868af89d7e9fda338bc7d72e0977c73d3

Observation f5fd926b-d934-461f-92c3-477900688a3a · outbound

This paper cites InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.811999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.811999Z digest=sha256:05b07aa6ae83534986b87477eb55ddfb6f8317bf31ed25f4942c3c4fca8d6aa2

Observation 45383fa1-5c15-439e-8906-6ef399568545 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.816944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.816944Z digest=sha256:2c5e43be5fe393af2dc28c780a86e1d66d5205ff71438fe753a7e17bf0bffed1

Observation c6acb05d-fd47-4252-bbd5-d27fcc0f1210 · outbound

This paper cites Qwen2.5-VL Technical Report.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Qwen2.5-VL Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.821043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.821043Z digest=sha256:6e89165ea339ddc7db42169d66e05204ab9ac08e33a61b81879a0f1409673848

Observation dc969136-352c-41b6-b41a-a505387dc42b · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.889964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.824965Z digest=sha256:9bbe93196b51053da2507bda42396c3116d15f9ce7fcc44ecfa4fcccc395a3eb

Observation 506b73e9-c295-4ebe-a4b0-bb695d08020b · outbound

This paper cites MMR: A large-scale benchmark dataset for multi-target and multi-granularity reasoning segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling MMR: A large-scale benchmark dataset for multi-target and multi-granularity reasoning segmentation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.877817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.828896Z digest=sha256:c65f32d22b9029297b64b0fb72b3fd0a4ce1873cccfbfd6bdddf515cff32aadd

Observation a4d57b4b-8fd7-4272-8936-30d94812745c · outbound

This paper cites SegLLM: Multi-round reasoning segmentation with large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling SegLLM: Multi-round reasoning segmentation with large language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.865219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.832846Z digest=sha256:d775f7d75c8fd1b3a3185600d42e042a7e371f4c7bfe8bf2a933f70789626fe5

Observation 9e40fef4-cb57-4c04-9145-434de89ae1fb · outbound

This paper cites SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.836256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.836256Z digest=sha256:157f6b6f9ebdbc1830a607d9b3decc00b921a3acb79d65aa04b2bdbdd642caa7

Observation 2d391006-c002-41de-887e-13e3b3027ba3 · outbound

This paper cites Geopix: A multimodal large language model for pixel-level image understanding in remote sensing,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Geopix: A multimodal large language model for pixel-level image understanding in remote sensing,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.853105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.840953Z digest=sha256:de3e30e9eb40adaed2cf6a68be486383a26745025889ad69738e2388269c6b5d

Observation b1685c66-070a-4431-85ea-9f62e4d035c8 · outbound

This paper cites Ufo: A unified approach to fine-grained visual perception via open-ended language interface,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Ufo: A unified approach to fine-grained visual perception via open-ended language interface,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.844728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.844728Z digest=sha256:0eb8ef7743c4bf4cb746ae7802a349c77c54c5266928141075c4decd8fa5c47f

Observation a11488e5-ef7a-44eb-8211-4e66ff77f7f4 · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.849167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.849167Z digest=sha256:8802fcdefe7b110981130fade79857e4eb20cef0512ad22f467ab7dc0175a843

Observation 84c93dfc-4775-4d2c-845e-0d00910c8ed7 · outbound

This paper cites HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.853172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.853172Z digest=sha256:e5565cf7c9ca9c049ea6a0d8c2755b45624c4cd34316d0d057f3cea2abb240e1

Observation c98a7b17-c7fc-43fc-a776-a22a0b3c0603 · outbound

This paper cites Alto: Adaptive-length tokenizer for autoregressive mask generation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Alto: Adaptive-length tokenizer for autoregressive mask generation,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.857799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.857799Z digest=sha256:44691bbf076b413de0e077e248fd29bf709b2cb22ca1d5fdecb627daef3f1dcb

Observation d15d53f2-8816-4c93-b271-92667ee7490e · outbound

This paper cites Pix2seq: A language modeling framework for object detection,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Pix2seq: A language modeling framework for object detection,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.839915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.862017Z digest=sha256:bc55d3d993c02a82cc1c16c7d960dc9f7c17782305ca467fc26eed79b6b5de4a

Observation fe671d6a-c133-46d4-b53b-5648672120c8 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.865970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.865970Z digest=sha256:d7468460e3579395b7c4e75cb4a65a82cb8ace2e88c8f13782dd20bc367186ae

Observation 306794c4-6ea0-41ed-8e0e-06b6f31e3e57 · outbound

This paper cites Universal instance perception as object discovery and retrieval,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Universal instance perception as object discovery and retrieval,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.818231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.869533Z digest=sha256:a954c2ce9c12b32a564632ba6d9409ffde439ccfe83c148e14356ada6fd6fa64

Observation 0d5e6583-3433-414d-bfc1-87373ee463fe · outbound

This paper cites Grounding multimodal large language models to the world,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Grounding multimodal large language models to the world,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.803605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.873438Z digest=sha256:5c751ecffc0b739adb2576dc0a9012960702a3f10818bb53e0693fc5eb2ac065

Observation 320582a0-6bac-4264-a043-f1fcc1cac410 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.882102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.882102Z digest=sha256:38bf91e3bfef116fbc215fa52cfd66dd1dcc0dc1876ea1432e0b401ddf1e8ee2

Observation e05b92de-6617-498c-99a7-d8d3fe58a783 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, edit- ing,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, edit- ing,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.777035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.886029Z digest=sha256:5cfa9a9d487e47e23c69a561f0ae10fba0db15f18da43bf74d6f076d19c35a61

Observation 4b781d89-4d0c-4727-89d3-7ea2c2337776 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Learning transferable visual models from natural language supervision,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.889729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.889729Z digest=sha256:7590da62b4d1c1e7a2a147563c2bd23c5fa15f8cbd349c18121d76eb051687e3

Observation 230337c0-0811-4a7f-8cf2-040084193265 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Sigmoid loss for language image pre-training,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.893149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.893149Z digest=sha256:c93d579adb2fbb4ee5847bbc350c3d7d1602c38cb47109965b63cd967a7a9081

Observation 0bfcbda7-bbfb-401c-9a30-7e642fb9b842 · outbound

This paper cites Qwen3 Technical Report.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Qwen3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.896690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.896690Z digest=sha256:9e1d4986fff4139c2f5c88f162212060ea4bc8b12062509919abf5578096062a

Observation 84748385-428c-4c6c-a591-2c664d0d6fd0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.900181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.900181Z digest=sha256:16d667431438dd3e768e45e879500395b0f391de91c63b0b6083ddc387094619

Observation 52b70454-8e07-4d83-b61c-ef088c3260e4 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Referitgame: Referring to objects in photographs of natural scenes,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.904054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.904054Z digest=sha256:bf3e04a8865e72231d9a103948f38a27b60522d3cb0763c372778a4870bbc9a7

Observation ef7d8813-6477-4dd3-88cf-18cd7bda100c · outbound

This paper cites Run-length encodings (corresp.),.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Run-length encodings (corresp.),

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.736085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.907469Z digest=sha256:4549c9ec5d9268d587fc30732baf0418d4df7a503a4e16cbd9b2ef2f036a7c8d

Observation f785469b-2cee-4162-ad49-b1de33aa02b5 · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling LoRA: Low-rank adaptation of large language models,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.911086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.911086Z digest=sha256:87162a271594fdf9ce60d19504209d98cdba76e61c008acce5167cf9e65ff52b

Observation f5456534-717c-4bc7-bc20-d58eaba8d97c · outbound

This paper cites SAMRefiner: Taming segment anything model for universal mask refinement,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling SAMRefiner: Taming segment anything model for universal mask refinement,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.714414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.915243Z digest=sha256:a30846a27a1e55cb2bd02f5f9713b3a42250c5f88ebbc61316a4e353ef269eb0

Observation f3d74f23-94dc-4fbe-b79d-835510d30056 · outbound

This paper cites Swift: A scalable lightweight infrastructure for fine-tuning,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Swift: A scalable lightweight infrastructure for fine-tuning,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.918817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.918817Z digest=sha256:35aee1341e7ff443ba2eaadd7bf1194dbd13e6acf9163743af5a8094b07e8c02

Observation 64db9197-ac38-4595-b229-491b711ab833 · outbound

This paper cites Decoupled weight decay regularization,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Decoupled weight decay regularization,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.922230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.922230Z digest=sha256:e8166840925e74a4c5ff593b9cfb03b9cf73ea932e31964e6683e7b760f378e6

Observation d8d3f090-35ac-47ac-baae-11165d5b497a · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Zero: Memory optimizations toward training trillion parameter models,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.925703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.925703Z digest=sha256:b34087cbf08f94d8f0c28efc069e155e5ce8bf35f293a05c3ba54c7df0e1423e

Observation 8c2f4997-aefa-4d50-9056-382aceb3729f · outbound

This paper cites Phrasecut: Language- based image segmentation in the wild,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Phrasecut: Language- based image segmentation in the wild,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.686003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.929150Z digest=sha256:830af68cfa1db6691d349389d1a607c4e092a7dc2193387da0f1026af97a17c5

Observation e0ae129a-3705-4b2c-9108-52d019562233 · outbound

This paper cites Gres: Generalized referring expression segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Gres: Generalized referring expression segmentation,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.933479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.933479Z digest=sha256:18652c40475f6a5ae0a756b3ea3752266508b66f281fb711d4d7ae582d3d7e9e

Observation 3686281b-e532-4e50-ac3c-a56d9b657180 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Generation and comprehension of unambiguous object descriptions,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.664754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.937004Z digest=sha256:5311e147e70f3311dcea5ce504ffa42f472f60cfc776299c6bae2c69c3022189

Observation 60c9acc2-87f3-4fc6-8bff-7b8dca37a979 · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Polyformer: Referring image segmentation as sequential polygon generation,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.649362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.940648Z digest=sha256:bb6d87604fe0502859b102aac76eccf0e7decdd271bc744b65573649c331dd04

Observation 4247aed3-20bc-4b26-8943-9cf3d65b7d7e · outbound

This paper cites Language-aware vision transformer for referring segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Language-aware vision transformer for referring segmentation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.636010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.945194Z digest=sha256:a01d538d772ba521b8b4a1556be95c37c38874b4d38a570852ebe4c2b60a9707

Observation 234ead33-0ceb-4629-a3c0-72245f0cad71 · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring expression segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Sam4mllm: Enhance multi-modal large language model for referring expression segmentation,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.624866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.949061Z digest=sha256:5fc7187fb2b9845b1d92b464c8ed2e9f20cedd5142460d585e1dad77d1b4be21

Observation 2c238a44-1cd0-4a52-8fb2-93a1ec3b2511 · outbound

This paper cites Segagent: Exploring pixel understanding capabilities in mllms by imitating human annotator trajectories,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Segagent: Exploring pixel understanding capabilities in mllms by imitating human annotator trajectories,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.612245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.952590Z digest=sha256:10e89d997f0cc0b3b8224013f70c87615a74d36f9c6e69a69807ba7398d66202

Observation 6a763a57-4dd6-4ce9-84ff-dea2c2b65cc7 · outbound

This paper cites Popen: Preference-based optimization and ensemble for lvlm-based reasoning segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Popen: Preference-based optimization and ensemble for lvlm-based reasoning segmentation,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.600552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.955957Z digest=sha256:dddcc9ca960c30eaaefdda99e82df9de694a28368a775fcce577ea3b0872eb5c

Observation bef2a781-708f-493b-870f-7e3811d59447 · outbound

This paper cites Microsoft coco: Common objects in context,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Microsoft coco: Common objects in context,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.959710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.959710Z digest=sha256:9f64e4da982da8b2b57a06fc4bc65d7b11fca246f4ea4f35572b89d10392e4a6

Observation d3a068ef-ac45-4ca4-8037-8a2be6265799 · outbound

This paper cites Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.963065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.963065Z digest=sha256:6ddf568cd31e80a46fbe98e561801dc261c7cb7f6a97b67ad4f6226f62c42821

Observation 8b1d6538-e8bc-4143-8f47-7199505655a3 · outbound

This paper cites Rotated multi-scale interaction network for referring remote sensing image seg- mentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Rotated multi-scale interaction network for referring remote sensing image seg- mentation,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.967009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.967009Z digest=sha256:d133ba897692cfd03ae47775e26f097f75c1737227750ffcb2f9fd208e05d26c

Observation 17537b21-d030-4ca4-aeea-3d170b5374cb · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.970567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.970567Z digest=sha256:3bb7e24ed8ab7a1152e5be563a7f7053daf0299c5b8ca7af1ca61c8d9974f911

Observation 0d2acdad-31d8-414b-8f75-d6c98ce3ada1 · outbound

This paper cites Clearclip: Decomposing clip representations for dense vision-language inference,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Clearclip: Decomposing clip representations for dense vision-language inference,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.571985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.974102Z digest=sha256:911cfd07bd7c9711236c7556093d1cf7b975af6e8bbb7ff6fd7996b3656d99b3

Observation df2477ff-adf3-4639-ae1c-b2ed00f958c9 · outbound

This paper cites Proxyclip: Proxy attention improves clip for open-vocabulary segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Proxyclip: Proxy attention improves clip for open-vocabulary segmentation,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.559234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.977841Z digest=sha256:bdef47fbf714fb7ff3ea6f7fc4d3dbc56b34dacbfd5af62220711fdc9cf74eb0

Observation 9b5a401e-d20c-4b66-90d3-87f9f5ec3a74 · outbound

This paper cites Open-vocabulary universal image seg- mentation with maskclip,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Open-vocabulary universal image seg- mentation with maskclip,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.546367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.981115Z digest=sha256:2169c3bfbb82597eca43112a1105e1211a7605148b36c2370097601d8e28ad61

Observation 978982d6-d188-4c57-8389-88d97b2b0a29 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Groupvit: Semantic segmentation emerges from text supervision,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.985304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.985304Z digest=sha256:dd1c57dded4c48781e17adf2273c9a632fd755b51723c4e926a850ca77f5a74f

Observation 66daabaf-eb5a-4188-8002-d8627b4eda68 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Open-vocabulary semantic segmentation with mask-adapted clip,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.526750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.990287Z digest=sha256:b1a2b660843f1beae3a2ddc9ec545d5653d3446593ff4043bb8d1b5284ee9a31

Observation ed057f24-fdde-4764-90b4-5af354059273 · outbound

This paper cites San: side adapter network for open-vocabulary semantic segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling San: side adapter network for open-vocabulary semantic segmentation,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.513496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.994740Z digest=sha256:f93a4cf5e2581c4573c185d6b63f58dd05e716eabf9c1b9322c5e71e25d538ed

Observation 1865d8d8-7f24-4bf2-bcee-fc1d0f0a6a95 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.998453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.998453Z digest=sha256:e91752bf9c10a074d2fcfd10076db1402b52236ca17d8b744b2cd7bbf1e595ac

Observation 087076ab-fc36-4a0d-8bcb-19179f2882ed · outbound

This paper cites GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:12.003667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:12.003667Z digest=sha256:f58655b746f7d254672b88bd15669e08245da6145dbd51bccd609f584794dbcd

Observation 7cc02e5b-4d07-447e-b3c4-5230c55aa13b · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Semantic understanding of scenes through the ade20k dataset,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:12.008111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:12.008111Z digest=sha256:044fcb98d9010791618fb0d1c61d375bf5eb9227c6ad68d3d1c724d4ee447870

Observation a824abb5-81a9-416a-8517-7eccd65fa637 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling The role of context for object detection and semantic segmentation in the wild,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:12.012802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:12.012802Z digest=sha256:2cc184ff800bd07febec9aa1f4ac9bc2ebd25b13bdcf7547f04d750d417784cc

Observation 9d36d4c7-c8c8-4889-bc75-68c29c1516b4 · outbound

This paper cites The pascal visual object classes challenge 2007,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling The pascal visual object classes challenge 2007,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.485552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:12.016591Z digest=sha256:ea17185879337bbb71375698b40ec7813c8cae10585bb9469adde71b5efd58ce

Observation 01d8ea53-ab5f-4988-bb0a-26d37fc44b82 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation,.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Side adapter network for open-vocabulary semantic segmentation,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:12.020753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:12.020753Z digest=sha256:10ad41880b63b86e38fe4153811364e8ba9c8f9273c51645438cf8875e33fdec

Observation 7ece75ed-0efc-4d9b-82aa-09cc53b2e085 · outbound

This paper cites Available: https://openreview.net/forum?id=lLmqxkfSIw.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Available: https://openreview.net/forum?id=lLmqxkfSIw

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:12.790745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:22:11.877495Z digest=sha256:c067ccd8b86a157a79f7cc51d7a0979f010ad5c88767b0126190db9bbba61858

Pith citing papers

Observation 76f9a169-93e0-42b8-a8fe-e1941f75b22b · inbound

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs cites this paper.

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs Text4Seg++: Advancing Image Segmentation via Generative Language Modeling

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:59.696291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:00:20.216268Z digest=sha256:697d16e05ad624e80aac26ed71ca5fcee73515776d6f6b90871790a69cbac70f

Observation 68a75981-2033-4edd-8651-8fe608d4a57d · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Text4Seg++: Advancing Image Segmentation via Generative Language Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:04.784266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:b705e4d099d20226a2b676396c45a634911e69292db7b92179d546f47cffba49