Pith. sign in

Paper Citation Record · LEDGER

PixelLM: Pixel Reasoning with Large Multimodal Model

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2312.02228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02228 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:33.993033Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:03.024533Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26fbb55f-4266-4bfc-bb44-8f5d254b1555 · inbound

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level cites this paper.

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:46.653302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:46.653302Z digest=sha256:e7bae5aaa97914f6d66a45f0cb4d1800556177773ef6acc41bd28fd8bf980932

Observation 8f174c74-546b-4ec3-9044-5f2bc7b5e25e · inbound

HyperSeg: Towards Universal Visual Segmentation with Large Language Model cites this paper.

HyperSeg: Towards Universal Visual Segmentation with Large Language Model PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:03:54.184388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:03:54.184388Z digest=sha256:109c8a757b96bd8c1b10321c50bd8cbf4f05208b871831fb7c147c8ceb16872f

Observation 078c33a5-c77f-4ce9-a57b-a11819d7d899 · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.077636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.077636Z digest=sha256:14b9840504fda9540c1367fd9b577f3e22eceacadbabf8bcf57a6d9e970e15c3

Observation d6c41679-8549-44bd-86de-4ff915752f29 · inbound

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models cites this paper.

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:59.602850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:37:59.602850Z digest=sha256:1339bfb4b5bd450252bbc605d0e3c1484d8d12264e72831e82c99e097da6a960

Observation 86681f54-61b8-4a88-a6f5-273f4ef7fc9f · inbound

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing cites this paper.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.653877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.653877Z digest=sha256:61c609b275b80f0afec8ef8cbf10c3119a8ac0a81075bb76ec4bcb17cabe4f12

Observation 5cf9db23-a1b3-4486-9bd3-6d6313aaeca9 · inbound

Advancing Visual Large Language Model for Multi-granular Versatile Perception cites this paper.

Advancing Visual Large Language Model for Multi-granular Versatile Perception PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.107930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.107930Z digest=sha256:010d7fa163c7876ab99bedfc18010069a04299431a80456ba36c24c5bde37e96

Observation c34e7079-30e6-4004-92c1-01f0c58eccf7 · inbound

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System cites this paper.

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:33.993033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:29:33.993033Z digest=sha256:18aa2a9edcbf5084a85586c89a6c3f0c7aab5236e93b4bfbc6e2a018e6cbeebe

Observation 1a35217e-96be-4a4a-bc27-61c00f6786e8 · inbound

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes cites this paper.

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:36:34.027052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T15:35:30.549656Z digest=sha256:5af19d0ef01cd4a2b109ddc0bca1812a8771267344ddf3b5d1decb3c1b49915a

Observation cd626809-10df-47f8-9bf5-15e7b2ba0485 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.475649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:fda90e4674196d1ea3837acc248a9e11bea458a73e35afba959eb856064f1f43

Observation 592de972-91a7-4d8e-b634-46c85a1f6496 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.025941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:d90f90aa4b3019bf2a0159f37bebc54c6a547d00aaa1a5b0d327a00dd6d7e91f