Pith. sign in

Paper Citation Record · LEDGER

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2504.10465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10465 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.273762Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:18.024069Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d36d20c-b0ee-41cd-bbc2-03c20d0ca365 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:39:22.544748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:cd9be7ccc9f9928dc2fee622efb772abe56df538410d8eeb6ac152a9c7a9e6a5

Observation d43770c4-9261-42b9-901c-4092cd57a8e4 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.749504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:830ed0b1afb2b9f1cc262e1f822be9561484c605f465d18649a71e65172d1455

Observation 5984fb9c-150c-4bbd-9f4c-cb0289374dcf · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.273762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.273762Z digest=sha256:81a0ff66dddc3cfbd757cef1444101a72f8f2d13b483b8c0c92fcf082b673828

Observation b399ce0c-bec0-4a01-8791-97bb3e3cc2dc · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:55.688403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:55.688403Z digest=sha256:ceeaa4480354305c8a4890a2a44ffdae891dc5df08b4158ff9879321454a3778

Observation 773c2860-b542-4f6e-a3d2-2c71c3acce28 · inbound

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning cites this paper.

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:05:07.431761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:05:07.431761Z digest=sha256:162cd0126367c60eec1e148984d0d285bc6b736463615f7a4309d7ab519b27fb

Observation 925c7ed3-2ea6-488a-9936-4a04bdc85585 · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.751792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.751792Z digest=sha256:7f960bc42cc565dc8ec98f9fdc0b872e23f55674cf7bb6c52f067d91df2cd791

Observation 5d347ad1-9a51-40ec-851f-1a87105d9ce2 · inbound

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs cites this paper.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.464562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.464562Z digest=sha256:a4d86ecfbe730fcd3f3f82e596cbddfc5690335933173b27daa8c69acbcfe415

Observation 926876d0-961d-4cb2-a56b-95a4b776275c · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:13:16.235635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:feca949aee5def7aa3b17f8816dba5404eb7bae2b8d06f9fcb2d87f3cc34d7c9

Observation cf4aeb30-1799-4ff3-a2a5-7e0bbab87631 · inbound

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models cites this paper.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.026551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:6dd4040db32de3f3eff613298145bade8f785e1d22c472454ab838d1ee59fbc7

Observation d5f801c7-3793-4c59-998c-4ebfce3039e8 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:3798bbc60e7cef3940daf6fc904ea0ece9af9f25144b49207f274f885337c77d