Pith. sign in

Paper Citation Record · LEDGER

PixelThink: Towards Efficient Chain-of-Pixel Reasoning

As of 19 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 5 inbound Pith citation observations for arXiv:2505.23727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23727 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:49.325182Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T13:38:03.546834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.964053Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9c530e1-301e-452a-aa53-0f4f72eded95 · outbound

This paper cites Modeling context in referring expressions.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Modeling context in referring expressions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:40.528887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:40.528887Z digest=sha256:0fb4444cc90ab5c7b08550acf7ad7b049232011e2f0d36512e12630e01afca02

Observation c85d2a63-db7d-4920-a030-357df615d65e · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Lisa: Reasoning segmentation via large language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:56.662142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:40.599512Z digest=sha256:8236b82d2f71c32bd32e208c266e54642690b4928c4ffb08ca91ec3adea828b0

Observation 3b7ac6dd-0d72-41c7-ad02-498abbe7684f · outbound

This paper cites POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:51.290601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:40.721525Z digest=sha256:eef84ad7eed50579572188194b62300006c397535a69918a294690481f74068c

Observation abded3a6-ab1e-4a87-badf-df1d9f0098f5 · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:56.424010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:40.821394Z digest=sha256:9ab398d97d3c5de10657cdc24754550f4ae121498ac15f2f45ac24944159e07c

Observation aadcc3f0-481f-4d1a-b49f-36a5aec754c1 · outbound

This paper cites Masked- attention mask transformer for universal image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Masked- attention mask transformer for universal image segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:56.179089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:40.960493Z digest=sha256:122dc4496cfbe9381d96fc415cd99f9c8bdfaf7b4c7b9e609e33861d13e51d34

Observation e5c4088e-6bba-49d7-86a4-ff1ea31018a1 · outbound

This paper cites Mask r-cnn.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Mask r-cnn

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.940441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:41.066913Z digest=sha256:227c84760e24bc53ed98fb196337abc4ed7f9673f37a3bd5b0845a27c6a64627

Observation f1fd865f-308e-44e0-8842-1b526f9c4c76 · outbound

This paper cites Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.Advances in Neural Information Processing Systems, 36:26650–26685, 2023.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.Advances in Neural Information Processing Systems, 36:26650–26685, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.682226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:41.202004Z digest=sha256:7fc0199ee8f6da3bea85419303e6131544a3e7fa8f86ece1616c882f8cb6283c

Observation 4f7a7ac9-3408-41b2-a512-b91ee6d9304d · outbound

This paper cites Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.330241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.330241Z digest=sha256:3a16b1fb9d92bc1dbfd7e0a28595970511f778320387e8d966647014db6e70bb

Observation 5bcc5003-2b08-4172-b713-edda78ccc6d8 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.425108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.425108Z digest=sha256:c2a41e49c465c6fb8318fc108c60db73bd30ec9d592b4d22e8df64aa91223f58

Observation 39bd3ae6-0ea5-48c6-a8d7-f6e4be4c5309 · outbound

This paper cites Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.564697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.564697Z digest=sha256:cab471682fbf2fa880fddc0669bd040f87c09a57cfcd28d271dd911303e9f831

Observation 29c39849-422d-4183-ab12-bc75d4212f85 · outbound

This paper cites Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.678292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.678292Z digest=sha256:050ca8c4dea5a043c6ab57aaae1311cfcb01af01a43ed67e529a2bef584298f6

Observation c14fc619-572c-43c3-9df3-d068ec168ecb · outbound

This paper cites Improved baselines with visual instruction tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.433582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:41.807075Z digest=sha256:bea12e3413d84aad0a83c7fc94aa7371c459af9932e9b685c416c8377821bee4

Observation 25340ec9-6150-4d27-bc9a-75ff2dc3bd1a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.894174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.894174Z digest=sha256:b29eba0decdf19606d5ac930f29813e63b25c162d73ee1563b1c1df8e8280f29

Observation 172337ab-ef93-4e18-aa86-a59beb93d021 · outbound

This paper cites Qwen2.5-VL Technical Report.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.033959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.033959Z digest=sha256:aa5263a9edd63281e1a29c7364fa28108cc5a6da15cc39cf3d929b1e114643bc

Observation 9e789f4d-afc6-42ea-8178-4bad8686c8a5 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Pixellm: Pixel reasoning with large multimodal model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:55.209420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:42.147050Z digest=sha256:f568b37d590c26d82e7f3228c7df0b664f77cea47579e8b939728b44267b9141

Observation 5809c1e1-7142-4f80-bbad-4ed6e7bb8d8a · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.893435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:42.202679Z digest=sha256:8ece6843060bad6ef94f229d30d83d184dcf6f1f52b8a9dfab49eebba4e4bcd2

Observation 93fe76aa-8ec5-43d7-bc3c-054e0241a974 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning One token to seg them all: Language instructed reasoning segmentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.328505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.328505Z digest=sha256:3fd2b34a9b8a22c4eda2786e0cae883edd4c2b1189c0bd9954e97514ed82627f

Observation a7f9f543-9be1-45d9-9a85-494486069e7f · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.478046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.478046Z digest=sha256:aa29eea0cab6e7ebcaf4b2c8b5da9dc53cda98e750182264f83b4dc9646dc363

Observation ad735f36-99db-4e88-9149-6f6265565e40 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.624891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.624891Z digest=sha256:7f6e6208f030ee2e45ef58446c5aea4d1c143752f3d9a1336fd60e7b321fde57

Observation 9babf3c8-dab2-4c2a-8a6f-d0b944961067 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.775228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.775228Z digest=sha256:4bf8bd72493f805a661de1c7d760d16775fbf60fd0382b64fcd4ea32adb790ab

Observation 758ab3e0-ee7c-4a81-a153-d223a9a14567 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:42.916533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:42.916533Z digest=sha256:47960e7cdc5c73102425d97d0f74f639c009de236c0d35bb1b4114281b81e07b

Observation 509240d8-4f57-4570-83e5-c62b8bc25153 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.033198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.033198Z digest=sha256:9f0697f4648e59530b7ee04e52574088e1dec00acd73dd40c364136e68ca7acf

Observation 5b48a37c-a04e-4115-b913-a8236ed149d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.144525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.144525Z digest=sha256:ede28654a89fbbd1074999874d378a684ba2194c716e9a69ca52dfe46536ccb5

Observation ea4145ac-7bb6-4db7-a498-3cd8d0a6ce32 · outbound

This paper cites R1-V.https://github.com/Deep-Agent/R1-V?tab=readme-ov-file, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning R1-V.https://github.com/Deep-Agent/R1-V?tab=readme-ov-file, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.647463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:43.256104Z digest=sha256:5c4144b0bfaade25ae77555bcfb98305e13981fc8220e57ab36d7af1c27a882e

Observation 9498b81c-33b4-406e-8272-b15b92f26e6f · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.373987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.373987Z digest=sha256:0c467fd8945127d02b74d9309b64898bd683151b7636e3c97b0182ea7361274c

Observation a51758b7-d85c-4ca0-9bb7-f7403779b27d · outbound

This paper cites Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.494739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.494739Z digest=sha256:669a4cbf9ef54bf669db0512bfb84a9342d4d799fab2fa6c1b8efbf551592cac

Observation 4b7ce38a-bfce-4751-a9ed-51581b000b01 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.636519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.636519Z digest=sha256:5234213716dece6ded3cc0340d809197c9d06c82a24d9b8c9a21c827ac9b95e0

Observation f5d7b6ea-9da5-4967-b817-2dc30cd3c00a · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring expression segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Sam4mllm: Enhance multi-modal large language model for referring expression segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.763499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.763499Z digest=sha256:c82f7bc6808d61e837f31e8deae5f69a44e23a8569d4a48ccca10fbf21abd6c2

Observation f10a5cc4-8fc3-457d-bfd3-bd3ff5b4bb19 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.891818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.891818Z digest=sha256:30e0b84e63ec22316ec6a284c0ad5b49a61cec1d25ca3315212d83c2994744f5

Observation aeb88d7b-eadd-4863-a39f-ccce7300b521 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Referitgame: Referring to objects in photographs of natural scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.420700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.013359Z digest=sha256:40fa9d1d99ef1708c2b3e5a76f8daa38d7fd3c72f8fbb3a0064e8bbe624d3f57

Observation def141f6-31f6-4bd8-9d3f-90333b08f475 · outbound

This paper cites Towards robust referring image segmentation.IEEE Transactions on Image Processing, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Towards robust referring image segmentation.IEEE Transactions on Image Processing, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:54.170448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.109823Z digest=sha256:2f2651cfac11becb0171e350e73f759102ff4a847aa599471dbbe5978298bc99

Observation dcab0df8-cc3b-46ea-8408-b9dc3b5e632a · outbound

This paper cites Remamber: Referring image segmentation with mamba twister.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Remamber: Referring image segmentation with mamba twister

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.980694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.249632Z digest=sha256:69af921ecfa6ee38022cb6e1bc77ec7e191bd206506c5fc09db86205db01296d

Observation 7c9a332f-a64a-45dc-82a1-15e3cf5f4a24 · outbound

This paper cites Mask grounding for referring image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Mask grounding for referring image segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.369152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:44.369152Z digest=sha256:1cf31f86210d5114b39e54bc0700b1cddce69ab3f61237dffc91e91142201048

Observation 120cba28-27d0-4f29-8a95-a4fba548ade5 · outbound

This paper cites A mutual supervision framework for referring expression segmentation and generation.International Journal of Computer Vision, pages 1–16, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning A mutual supervision framework for referring expression segmentation and generation.International Journal of Computer Vision, pages 1–16, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.751979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.489010Z digest=sha256:598f1bc000e73dacab2fd257feebe20e7526602c6e9fac375fbbff26ddc64a21

Observation a12b4204-9fa7-4b6d-888b-6f2e11a24395 · outbound

This paper cites Pixel-sail: Single transformer for pixel-grounded understanding.arXiv, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Pixel-sail: Single transformer for pixel-grounded understanding.arXiv, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.553534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.593422Z digest=sha256:bbe543e9d9a121ac4e60c53ac7e85ac3d24285c2160829f1d5e09dddef411c11

Observation 4f3ef293-2857-499a-8ead-9cf4ea864dc7 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning U-net: Convolutional networks for biomedical image segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.343753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.733818Z digest=sha256:601a0777e8e724be4a3469af1e7aac564833369a32ef758e0a5351cab2f3269b

Observation db27352c-e588-46ed-b42e-3b1aab24532b · outbound

This paper cites Segment anything.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Segment anything

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.055280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:44.872837Z digest=sha256:d98ca7414025d0a4a7ff90c7e16393793b8b621ff149595721a3f40d2f4719dd

Observation 223c34e4-a933-4b81-a9e9-357fc0e6451b · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.966934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:44.966934Z digest=sha256:9ca594a70206519933f4239385139469b6e59aaef6e3834b2bcb2f8e5c9c8c7e

Observation 0a77e824-65fd-4963-b8e0-73070732e6bf · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos.arXiv, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos.arXiv, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.800808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:45.065669Z digest=sha256:cb21a1692eb574664e76ebe25688ebde15c800bafb21473ca923b70258a4c65d

Observation ce371911-473f-4dc6-9c50-f66fe4f1888f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 35:24824–24837, 2022.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 35:24824–24837, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.183443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.183443Z digest=sha256:5610434c43f1459a8151e19775a3e9d3f2b2ec5ca58c4d6ad114e98672d75518

Observation c06fddbc-7fd3-4fb6-bc7f-1c763812e03a · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.327386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.327386Z digest=sha256:c58613207e5f88796be434be5e8543b258c37807375ddd003ad0e472907d6fec

Observation 953f057e-1778-438b-9544-48944fbe92ca · outbound

This paper cites Towards Better Chain-of-Thought Prompting Strategies: A Survey.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Towards Better Chain-of-Thought Prompting Strategies: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.422348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.422348Z digest=sha256:47ef9071523a87e9a0b71bb719ad24ba932d143fb035eff8cf58bb2b6e31b952

Observation 18fc7829-96fa-47f1-9703-a59bde8fb76f · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.543162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.543162Z digest=sha256:23e514c70b726fabd73cf863fde074ac9a0de18c4074e34b3f4657262c421057

Observation 98bcf2ab-b1af-405d-b0dd-0ffed01820e9 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.679803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.679803Z digest=sha256:4581c46b4cee639646b8ba110882ac62f5b7dfd02f4763979a5d6193a6479fc5

Observation 495b41d2-2427-436a-9c92-bc82e99c77d2 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.805388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.805388Z digest=sha256:c99c39e244010d65e2366e799ad844a04ebd45e4bc88e4b9a90ce93618f7ca07

Observation 07be56fd-eeef-47ac-88c1-94ce2181fd92 · outbound

This paper cites OpenAI o1.https://openai.com/o1/, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning OpenAI o1.https://openai.com/o1/, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.891852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.891852Z digest=sha256:ea6f28b598294971e89dce6747b9d9fe85fa78f06a82534f04333d0b093e35e3

Observation 27c9ec78-1897-4946-8a51-8a205da80bc8 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.977413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.977413Z digest=sha256:1f73a9f4a37ab76bbe4393cc0dcf4aed871b3de5247294df48796face4ca16ca

Observation 07d63e1e-308a-4252-b7d1-74c29b4ca781 · outbound

This paper cites s1: Simple test-time scaling.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning s1: Simple test-time scaling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.077612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.077612Z digest=sha256:f83b3122b248b60eed3455f354d1689015372a095f6b7db30c850622d539d330

Observation e8b51122-ee9a-4f4c-a762-8a2f07fc51e9 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.144862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.144862Z digest=sha256:6100b9f4c06971d2bc6842de3d8c700dd885cc840c00b62d1c38a951658b3d25

Observation 28802efd-f994-49e1-ae31-14b1dbdd4de2 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TTRL: Test-Time Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.198402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.198402Z digest=sha256:a57641641ae9303e8d73dd436b769861ec2e1e0d6b0d8376b46b17e700268de3

Observation 504f63f2-81b5-433b-b5cc-aebbe956ac3b · outbound

This paper cites Open R1 Multimodal.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Open R1 Multimodal

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.576238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:46.260444Z digest=sha256:0f75c6f39e1deca4b56abfa838c6290e0b407de131d82ebd40fe91022606de31

Observation ba38a7ee-5ee6-472f-b365-1379f93f506a · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.355639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.355639Z digest=sha256:2f2e8cd36867e356a7fcf1605e3d4d1cfee25e941c6a2ccff1d8d887b814dc54

Observation c03c36a2-7818-47bc-9124-33934af2f74a · outbound

This paper cites Efficient Inference for Large Reasoning Models: A Survey.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Efficient Inference for Large Reasoning Models: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.479922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.479922Z digest=sha256:4313714c0437096640d5fb494753b2e756f28f04c8b4849b308b612b9c8bc8d2

Observation 1d42fa26-364f-44fe-b663-9eff29dccd5e · outbound

This paper cites A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.575425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.575425Z digest=sha256:07f7fe531b1dc7f9ed39462f08d30ac1fd13252f0803b29dc7600077e2ee73ca

Observation e7fa8cc1-75ad-43ec-82b9-63bea8c08157 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Token-Budget-Aware LLM Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.718839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.718839Z digest=sha256:096695e6709f2d78484fa9e54ffd48b1254c39e91490b5c8c583be9202758969

Observation 14b0b5c8-56d1-4a7d-9f8f-3a8a885a9fb9 · outbound

This paper cites CoT-Valve: Length-Compressible Chain-of-Thought Tuning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.813417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.813417Z digest=sha256:7a8ea6075712bda53e35de5c1605dbf969d44a685803aacca8afb3b6b96babdc

Observation b9e5165b-591a-431a-a5ed-f3d76ab2ad52 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.890731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.890731Z digest=sha256:bf4d27c96b583aef28aebc5d7ad5abfe0599ba0816f79c6cd208206ec744dfa7

Observation 2568004f-7a97-4103-a80f-b990e3e95664 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Chain of Draft: Thinking Faster by Writing Less

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.989840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.989840Z digest=sha256:e1d703370c0b9200073c77092d99abfbb75cc89a94c03613ce8185e62bb6d587

Observation 42aededb-4303-48bd-8e9c-30030cd75dd7 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.103317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.103317Z digest=sha256:54bc438b44d6fc99317acb9f2b3f94b7e98270e0e9f252f6d8013aae6bad2cd4

Observation bc08dba0-d75b-4fc8-90db-9494c4c596c0 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.190857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.190857Z digest=sha256:5b8a92f153c6110771cf1c604751aa1649f23775799e469e5a20b74a382aec80

Observation d88c928a-1794-4b88-86e6-323e3c1d21b5 · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Self-Training Elicits Concise Reasoning in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.270161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.270161Z digest=sha256:ee509b7e386d05918eb5864757c5afdf6786e5e263ee67d7dbb777312bf1745a

Observation d5f147fd-2b85-48dd-9d0f-711bc0cbc900 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.369962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.369962Z digest=sha256:28b977b29c4eb9577a85e6826c5948515d827fd6ea41b4b9d466dc88dda67974

Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.521580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.521580Z digest=sha256:9975a395ffab56c37f86a2c05cb787c24c3eb1cca185b0aeeec7aa8905e8cf09

Observation ee0d314a-5c6d-4fb8-8488-8e40fee127b2 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.643037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.643037Z digest=sha256:8fa28f3336f8827b96622039b72037feaeba58f954c13aa5a1d69e7b94da25f9

Observation a8b5ad51-a5eb-4969-9f15-cc7d751afe62 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.arXiv preprint arXiv:2412.04467, 2024.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Visionzip: Longer is better but not necessary in vision language models.arXiv preprint arXiv:2412.04467, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.770684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.770684Z digest=sha256:eb14faa4af5fdcca27038acf169c6896f58dec1f61bb2970bd1075446420a3af

Observation fc00c54a-baa8-4102-a00c-f34c8b0ac5ab · outbound

This paper cites Learning to Inference Adaptively for Multimodal Large Language Models.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Learning to Inference Adaptively for Multimodal Large Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:50.149770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:47.885550Z digest=sha256:b34d5c70cc3b253efa2a0107bd660f6c2922026e934707ade9aba6f2580178eb

Observation 05fd2193-04e6-4edf-87ed-77bfb4816150 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning SAM 2: Segment Anything in Images and Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.987803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.987803Z digest=sha256:c41b14adb146430fe1848875f9d86c6a168b540b1ad5cd9ec63684e6ccb4a015

Observation e9080b37-d52e-497f-a384-d23c3b099d9b · outbound

This paper cites Minimum-Margin Active Learning.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Minimum-Margin Active Learning

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:49.840522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:48.090813Z digest=sha256:37fce20b4e5e81297aecb2c0f6389298b91fe16e1c996fc4626c027e24c0b20e

Observation 55a21d8a-69df-46c6-849d-d76362d7ac9e · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Chain-of-Thought Reasoning Without Prompting

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.207589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:48.207589Z digest=sha256:87cb757848eebb2fc3131cabca408734a4779ba3c0c721dade963c5e1260de5f

Observation f1e504b5-8e26-417b-986c-6ce42b072558 · outbound

This paper cites Qwen2.5 Technical Report.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Qwen2.5 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.336602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:48.336602Z digest=sha256:7c42bde7e3d86faf75342b1dcee3fd1ae7d63708f402f684a5af341c72bdfdf8

Observation 0bd56ddb-6426-4bf2-a465-792667874fef · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.361036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:48.453340Z digest=sha256:7b45d248235183ddfb9406cbf7a2b299014466e52716fbf627f84f2a8752152f

Observation 3e48d4c8-304e-443f-a12c-9202b5c50058 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Open-vocabulary semantic segmentation with mask-adapted clip

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:52.155138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:48.580296Z digest=sha256:3cf93758e2213b37d07bccf7cb7f57b41f6d272e465516ac216caf7741386f5a

Observation cd92b246-cc7e-46fa-96a7-300299d3d122 · outbound

This paper cites Gres: Generalized referring expression segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Gres: Generalized referring expression segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:51.959281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:48.721713Z digest=sha256:b5d43578ed9b21823e12dc0518a95ac80d45c820b37016f9191d780609653393

Observation c89fc2ea-7041-4cd5-b4f4-0305f4e92a61 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.814103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:48.814103Z digest=sha256:9122f8517178bc56b8cce0524785e0c0c20b5bad481a213d558f455c5db1044c

Observation 92fdd583-08e4-482b-ba24-7947dd364e63 · outbound

This paper cites Lavt: Language- aware vision transformer for referring image segmentation.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Lavt: Language- aware vision transformer for referring image segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:51.710334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:48.921057Z digest=sha256:5b496ba93fdbd8a16dadff9175e332e5be06e27ddf76f89cb55e53251eac79a7

Observation ce38aa49-aec0-4121-b092-3fff4802a0c3 · outbound

This paper cites Perceptiongpt: Effectively fusing visual perception into llm.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Perceptiongpt: Effectively fusing visual perception into llm

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:51.493970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:45:48.978262Z digest=sha256:c166fcdbeccd00d6e329c30301878cd6ee246de4ae45995c5ca4f55513ad68ac

Observation 15688ab7-2a7c-4dd4-b0ac-c2039a0c824a · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.arXiv preprint arXiv:2503.16188, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.arXiv preprint arXiv:2503.16188, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.027105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.027105Z digest=sha256:f531a9bf7894210ddc559291036cfc429aafc70ff2efc12a4dbf398cb157c858

Observation d9f8daec-e419-4683-a770-7aee54f155c4 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.097448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.097448Z digest=sha256:fe2a82e88b8ea3ecd84a4c8bcace3712c875249b32b7c45e425e482a46611de3

Observation 3146d7b9-8f90-4442-8c82-8e85a243b33e · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.152191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.152191Z digest=sha256:83b0172040416988aef9be02ca9644fe5f1a2325cd1252b9cf27534f7a21553e

Observation 8d47115b-baf7-4cf6-a4dc-632836c4a70c · outbound

This paper cites Microsoft coco: Common objects in context.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Microsoft coco: Common objects in context

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.226688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.226688Z digest=sha256:8250329f8950a86f29953333a0465b6bb49831f7f4d4782b2f33a7dc81ae0ea4

Observation a459d7db-a80b-4667-86b6-87f82ee0a03a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.325182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:49.325182Z digest=sha256:f20140a3c217ce5282e0601cb0d06314033e0a3cd3cd45d6e05c834a6c86af0a

Pith citing papers

Observation 65e3b23c-75ec-4b0c-b5cb-825c7e0ef0d9 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:06.680384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:34fdcb90d0da15dbb22e4dda69bb223bdbcc935437899513f3703b1196830293

Observation 2aff3429-372d-4715-8e88-22264670b1a0 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.842628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:84dda3b9d7ac8208e359a0c68c63943afed03af062161d15193047838657f91a

Observation eef658b0-ee47-4ffd-a49d-4c206216bc64 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.966001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:ddab1d9071ebfdd80fe8e970c72d3422e932398341909fd5a9f4b2c5a0f5f329

Observation a391c2ba-5640-49db-8efe-ca035b958005 · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.636192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:850e8cbd8909a2087f248aa4a893f128f9c8b5e22ed813a3c0a7ed67215343ff

Observation 93f3c467-b83d-4ba6-b091-288392f93fe6 · inbound

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation cites this paper.

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T13:38:03.546834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T13:38:03.546834Z digest=sha256:f92fab8478e01fcb09647f06c2bd875a0180ca2b1d80eee038ca2a64a5e14d8f