Pith. sign in

Paper Citation Record · LEDGER

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2504.10465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10465 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:28.136196Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:18.024069Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d36d20c-b0ee-41cd-bbc2-03c20d0ca365 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:39:22.544748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:fb0d65e322cecfd5eeb424f704135ef53d783bd20fe34880057dc31b3a0f9753

Observation 443d7519-58aa-4878-8b9a-b7cf2ec0812a · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.136196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.136196Z digest=sha256:aff4c31d0ecb7627036664a6f81a5adc16528fa6c9e65c6fb7f855cec9b09223

Observation d43770c4-9261-42b9-901c-4092cd57a8e4 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.749504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:21b6acda3b82d2c21b821982281339a361fff4b7397009e657ae6a5a64f79875

Observation 5984fb9c-150c-4bbd-9f4c-cb0289374dcf · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.273762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.273762Z digest=sha256:49256b7df12fe0646cffb4a3104d0d989c146f4768e5ce224184c61ac77ddb8b

Observation b399ce0c-bec0-4a01-8791-97bb3e3cc2dc · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:55.688403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:55.688403Z digest=sha256:54ca54ed515479905fb40dc999a8c4d39659cd85712c9be4f76c62e66817101c

Observation 773c2860-b542-4f6e-a3d2-2c71c3acce28 · inbound

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning cites this paper.

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:05:07.431761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:05:07.431761Z digest=sha256:3d9677fd4e9e741c1c090f61dbe11fff8eb624b138a163d0d6de355bdf9adb80

Observation 925c7ed3-2ea6-488a-9936-4a04bdc85585 · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.751792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.751792Z digest=sha256:440c863c042f6ec10876035fe92fb705d630dd35cea0bb77b56af66dff9a7ecb

Observation 5d347ad1-9a51-40ec-851f-1a87105d9ce2 · inbound

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs cites this paper.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.464562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.464562Z digest=sha256:a55602a608c977f2fb53f5c0508a124752390ee3a8d26fc47afb4b56070845e8

Observation a11488e5-ef7a-44eb-8211-4e66ff77f7f4 · inbound

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling cites this paper.

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:22:11.849167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:22:11.849167Z digest=sha256:753e0a9459c46cabdbae1f6dc781b8dd84bbe5e0b02288b281911413afd0805f

Observation 926876d0-961d-4cb2-a56b-95a4b776275c · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:13:16.235635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:c12189e5b419497d4e469d075f8bb1edad57a86f6c7dc24452b720c2daee6e4f

Observation cf4aeb30-1799-4ff3-a2a5-7e0bbab87631 · inbound

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models cites this paper.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.026551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:0f9561b8c9d6f51a45534e3d7c731971ba9404dbef2e5901605bd640b80b7256

Observation d5f801c7-3793-4c59-998c-4ebfce3039e8 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:aae7d83598e672aa26a0b4179f2c5dda854cdcbfa28c6cda132385f44412969d