Pith. sign in

Paper Citation Record · LEDGER

TULIP: Towards Unified Language-Image Pretraining

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2503.15485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15485 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:57.321888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.381543Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4293610-9839-41e7-889f-03aa336b86c7 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence TULIP: Towards Unified Language-Image Pretraining

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.879831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:56298eea9b7e90a4495c2a9c5d0551a4edd2e4fff603ed9d3b63a76605b50af9

Observation b78fa165-3361-4744-9e81-cf90a451b496 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence TULIP: Towards Unified Language-Image Pretraining

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.385092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:efc7575bbb9e528f64cdb68fc9d5ad070fcff726e3602a2778de00fcccdc6d63

Observation 1ca84f57-c161-401b-939f-f32b0ecf4a29 · inbound

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets cites this paper.

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets TULIP: Towards Unified Language-Image Pretraining

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:57.321888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:57.321888Z digest=sha256:19fda7fd20f0d21eef119a54079557952e0cff6b20f63a9d99d3ac1b97d55d06

Observation 8ffa87e2-96fa-4345-8c8b-49c5e676b941 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning TULIP: Towards Unified Language-Image Pretraining

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.682149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:86e111f0857852bdccdc82c98763a34bf4c2eab42cbd341928f1ca8b64b03959

Observation 0519322e-3314-4085-ba16-d45541650657 · inbound

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics cites this paper.

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics TULIP: Towards Unified Language-Image Pretraining

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:40.606092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T04:25:58.275705Z digest=sha256:92da2099d83675eb4fa3090ae2972b171fedf930316ba1043d22482d405dc39f

Observation 13e07a65-2d5a-469d-bf35-4e99ee468c0e · inbound

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models cites this paper.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models TULIP: Towards Unified Language-Image Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.402513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.402513Z digest=sha256:fbe8016f01ef3b9aa9fe1c2a54078eae55a886198bcc5f20b2e4b8339f5dae46