Pith. sign in

Paper Citation Record · LEDGER

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

As of 6 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 8 inbound Pith citation observations for arXiv:2604.12012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12012 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:00:12.466258Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:43.056117Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact11
  • verified fuzzy29
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e82e152e-3443-434e-98f4-278ec7102e27 · outbound

This paper cites Alabdulmohsin, X.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Alabdulmohsin, X

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.343823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:fadbc2454c30220393939b647a1a5cd031aeb047dff493dd70945da5a0d0dcae

Observation 79dd9586-3c98-4b69-865e-2fc480337926 · outbound

This paper cites Assran, Q.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Assran, Q

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.346511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:f26ae723373791180c669288266ff0e74a6ea6f6bc3de84d8e09966a82c67c5b

Observation 5b53f90c-5cd1-4b42-a2c3-dfa9d80eca4f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment PaliGemma: A versatile 3B VLM for transfer

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:74f5052b7c384ceaa0a4303fd934511d956bc1428e35fc40245ab7bfddab1ee8

Observation 41690dcd-4721-4494-ab71-858385071944 · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Perception Encoder: The best visual embeddings are not at the output of the network

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:16.084670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:6b4fb671a6c7835bdd5249ce683ab11621cb6586d5b56faae13d93746884a52b

Observation a141721a-5b17-4af4-81e0-49e1cc5a9faf · outbound

This paper cites Caron, H.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Caron, H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.336117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:c5b81419af93d555583721c6244bba09b1c6cd6405f36836e5c00885c99f2783

Observation 9edf73e5-e280-4f36-819a-549fdee5dd8f · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.279447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:9dc883d5b75cd9dc72923bf8c0af8f4f57fc70379a40ad463176e3ced8eabb55

Observation 6363c580-a099-41df-8417-195c842d9a58 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.327880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:e7d612574724c2b0ecec7d60e0ae5327e861ad8d01e6ac8fb8852285e37b527c

Observation 6eb3bc23-8ff8-4672-b700-2f559dca0f0d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:bceeecb894901f20d4f8a5e6d2e09b3e08b41d78b0144f5cfd0e2a32698b1c2e

Observation e23681a5-ecae-4fe2-89bc-2682dd498f5d · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.414621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:a888e4601424e308243687422784908901ff17170c9aa42a955ea69f7a727aeb

Observation 9ed60abe-4795-4411-9846-7189b3961bc7 · outbound

This paper cites Cherti, R.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Cherti, R

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.324882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:7e020d7c9d3f7747e995d7270d2a2738e8e6d47b336be534d068c68d88395628

Observation da1d535e-a3a3-44f9-9f56-a5e90fdbcc0b · outbound

This paper cites Chuang, Y.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Chuang, Y

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.282075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:80224894b43d065f7b358920c2f22001684b0c7e84675d0f83231544f1948ea4

Observation b0dbe556-f6f8-459a-8de3-5ef898166eba · outbound

This paper cites Darcet, M.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Darcet, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.322086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:a257bae987f22db2fe4b985705594c77e0115db516444750905874c469e6f326

Observation 10c52f89-4891-4304-a475-4c1326fdb408 · outbound

This paper cites Dosovitskiy, L.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Dosovitskiy, L

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.379921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:c74f547be244bd0abc21ce6fd8543c46a79182eca85fe809baa4327f1f1e039d

Observation 964d3db5-51e4-477f-af25-0b8f71f6f82b · outbound

This paper cites Everingham and J.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Everingham and J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.385270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:bae9cb20addd8877c817063f13859488991d25cd7b24ee4bc4a0e05585678252

Observation e296c886-c839-42b0-b1aa-0708702ba190 · outbound

This paper cites Everingham, L.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Everingham, L

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.319566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:25266061b9f8b130f842d8de5e1ee9680675ccf1a287d223ae645648f73effdb

Observation 96813c63-c740-421e-a79a-aafe56e4e5a7 · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:01.874816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:82bbabaa9eaacd5127a5c9051f6b7cd8043c9ad43a73494f0b9b5ded40c4cb5a

Observation a9f71bcb-4a7a-43df-bbb0-3f8d6410bf79 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.382458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:28fc1801096b010505e71643ec8580bbd8e806ef0844c5934d720bd1050fd8e8

Observation 5e93eaae-4378-447d-8796-5bede516bd49 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.330718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:9312f667fa2bfab07db4f3178f2e324ff3cdbf144e2d35c28b20842d730fbf05

Observation 6cbbf9bf-1f72-42d4-9342-9fcd29f7bcf0 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.371586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:cac71adaf151955d32d1030cbd82ade529d5b43f8caddabd31622793aa5ccb53

Observation b7e1a9bf-9e27-4933-b3c3-8ffda34e3b0e · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.357604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:f19b6d3160083b20f0b1936679cfe1c189e76d86e59178a4867528f5e7597f77

Observation 9bf9c295-3100-4b22-bcc8-2daf794b338f · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.352329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:4fc6c2868d96e3b3a6aec2c02a076cce408f2f222ae2d378fd877d5b4c26bcc7

Observation 6323718e-be41-407e-ad5d-400726e3fabe · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:31:01.908565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:a4a1c8121199bb1c252ca76b434b111ff3e27e5fe87444a81e81f63f64cb91c9

Observation fe65e40c-4e41-4536-ad36-03f146071b76 · outbound

This paper cites Grill, F.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Grill, F

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.290050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:6ea55026e680d5b0e463678c006087eabdba65f35de2efb45413316bd947dc10

Observation 6d67c1ce-cd5a-4268-b0c7-4bb69fe1980b · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.284706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:1821b2f21e5981863c838db23167bd96eae022b36cd2de02a331d871a6037085

Observation dcb37d35-930b-4dbe-98de-c2e280496883 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.292817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:dd0c7e9a63a30f14b373d3bc922ded2a2828eb3429034d82fda7f384c9affe16

Observation e5052216-6309-4739-8ab5-0b44d655ce8e · outbound

This paper cites Hinton, O.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Hinton, O

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.341144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:59eb24f188674dbc2e401788a511b17981bea139c2d90cf931b6e5dcd40f4e87

Observation cc55c400-198a-4610-913e-3e73de8832e3 · outbound

This paper cites Jampani, K.-K.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Jampani, K.-K

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.368888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:1a0b5d0f2753679b9aa8c0fe46e24cb2214279eebde6db6d3f4a369564391ebf

Observation 645d39e2-0104-425a-ae4c-42b2ab3cf69c · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.374550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:a14bbbe3d3f7ee8d139f5fe549e12681b4b3b3fc04900215997f8edf75565dd6

Observation 1e6d563c-6c83-4ac3-9973-60dd020fbd92 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.287422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:148aeefedb5d1cb51f9f020b32dc4013f09f631defe0d175304ff59362ae979c

Observation e9953794-1853-4b28-9d8b-9fae203e4fa4 · outbound

This paper cites Krause, M.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Krause, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.411642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:58c05684f463acacdadfd9e2403608d796ad572a6b800d69282c14d6aa587dd0

Observation 9fd68227-7d58-46de-9276-d1a99aa48cc4 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.417616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:3dbf4f4a135bb8278af3342227a2e9f074035c8f4196cae3e56ccfd57d107c4d

Observation 3fc269b1-fff3-4476-a022-c966b0259b8a · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.423422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:c216dfa23f2ac2667677cca5020700c9e377816b12282f286b07e2a55043be45

Observation d6573569-bf40-46d0-acad-f0207e381a36 · outbound

This paper cites Maninis, K.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Maninis, K

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.408426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:12a2318cf07c27d0b044351c5d6705c8d4bcaba2fc4267e3ef857c1263edf4a3

Observation 5cef3dcb-c7c6-40ac-9838-75e0f2aa6fcc · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.396889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:d03bf344cea1a9fb821404b77f6623e4a1a555bd364b02b638d621624b3a4944

Observation ddd2ec76-0513-49c2-9eb0-3c5c708968f3 · outbound

This paper cites Mottaghi, X.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Mottaghi, X

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.429902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:e89efc1a33f650a49c9d14bc3df924eda7ca00bc842af892ccceb93ea3b0269c

Observation e27d8f1f-28d1-416c-b9fb-e79222611833 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.359978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:9004c8b5e728cc2cf988dda9bdc8397236d6dfdb9f01a777bf6fe909cce9155c

Observation 4cff6c34-b52d-4a45-ad53-108bfc8b5afe · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.391019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:7150e1065fea86e29bac500d4b68ccdbf50715529fdb68fdab4b929d5652c1e3

Observation 70dde919-7acf-49a7-8fb7-cee055fd6e98 · outbound

This paper cites Oquab, T.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Oquab, T

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.316944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:66d4a4579b625e6494de1e6ab9a919a55eef1a66ddb9e192f81d251d4bb6ca8d

Observation 0ada5ce0-2a64-4a83-ba59-206b92fde74e · outbound

This paper cites RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:01.842217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:b34f474e0024263b5a83240e0c8ded639b29012966191ade699b13f6c603c281

Observation 0a57ae36-25eb-4b29-81fa-aff34995568f · outbound

This paper cites Radford, J.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Radford, J

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.295456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:362866efa0e652639d19a667df13a8f5cf875a4f87b57a1ae82dfe35bc891e7d

Observation b05c0a01-6b6f-4b30-a28e-08f622711759 · outbound

This paper cites Ranftl, A.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Ranftl, A

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.276712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:60d5ef86988f97205646392984af44516e38f4c9b7c2463bb54cecab90f0077c

Observation 038d09b1-1cde-4dab-be26-8b335aa78ec3 · outbound

This paper cites Russakovsky, J.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Russakovsky, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.402500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:c08c0c392607b1be33ca272c62d599c257f7d542492136d67e2dce7766f47c9f

Observation d99a704b-8c36-409d-b595-250a9bbd119f · outbound

This paper cites Shazeer and M.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Shazeer and M

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.311216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:b1998233add5495987a182a99018e4e549af462a7cafece98abf77a11bd985d5

Observation eae4c2e3-b60c-482f-b7b2-ae78991f1a04 · outbound

This paper cites Silberman, D.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Silberman, D

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.305592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:619a5367fc1b0089e4b6ca4b20a244dc99029b327d239e917bcff6b28da53f31

Observation f06c4df7-ca5d-4caa-b92b-7476671ae4af · outbound

This paper cites DINOv3.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment DINOv3

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:31:01.854255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:d54892c8a0d41ec63478a1f9744a1dadc6edb66e9bf60b1eaec20e7b34a1c7cd

Observation 648a415e-8c3a-4c43-9d4e-3298c9635949 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.338749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:2c2837dcbea1f9205970ad69407a7228a07161fc2a5ba8123607c5139599d195

Observation b5e025aa-e39d-4f11-8b3f-05b21f4cdd43 · outbound

This paper cites Stone, H.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Stone, H

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.405278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:b8037a16f060b276e9f2e4f9d374da832304d839a376414e494997e51b6a77a3

Observation f2ca36ba-bf71-4b66-99b2-967c260953f0 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:f0be34324842db67ab7affc2d1cdb014cebd5559b45f14b423a69d607f7893e6

Observation de1ef774-5ba3-49ae-9a21-61a55ee5f3b4 · outbound

This paper cites Tschannen, M.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Tschannen, M

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.399794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:1d1cd682be853d1cdb8a2fdc9a3f7ebdc1364ba5557b8201bdff75681649796b

Observation 167a5b7f-e425-426b-b5ac-94a09fae16f5 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:31:01.933355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:0ae48999aa7b6513b1673e280628119cdce30a07af6ea31db9a18b7265f67a0b

Observation 0985afab-5336-467d-b415-8f33bb1a71cf · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Representation Learning with Contrastive Predictive Coding

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:31:01.946402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:25ebb35b2cc5f51d6aae45fec66d64d2efb40b155f92dface2790c7ffa21b9e7

Observation 572e44a2-5d81-4632-8f90-1d8d796b1186 · outbound

This paper cites Van Horn, O.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Van Horn, O

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.388035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:b1e1cf7f82fecdead484afd055272e098a76c19819281b89fd41e85b3219e38c

Observation a2def039-d0fc-4929-94a3-b6ce16831e01 · outbound

This paper cites Vaswani, N.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Vaswani, N

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.420456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:672702339bd04c66ce200d68d5ccdf8eae550f82cababcf5ffd3a0eb1a601580

Observation 640b57de-2b43-4f8f-9f77-d0d276849686 · outbound

This paper cites Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:31:01.917283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:9bf077cef3111a424679915a9bb837d38320ce2ac47212fe4d9b3b47103c7676

Observation e07374f3-584b-4554-84a8-f22b4d12a38a · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.393814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:6a95290c04d7e70f48db212193385ae74437e8f97dfbe776678f19c063467ad3

Observation d8481022-3bb4-4673-9103-e216b13344eb · outbound

This paper cites Weyand, A.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Weyand, A

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.365284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:720872daed0baf45f07b1dd66787e2bd65634ad4d003e3360264faf6ceb6f10d

Observation cfca63ea-637d-4322-aec9-827bb74e7d28 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.298022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:879ca857320572b181a9768d4a9c5c8d45cdf3190c404b43fe4ca829c4868030

Observation b187b90f-7b56-4b58-914a-2e6fa0b40cbd · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.300560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:155872ebda8f5cac35813bb47df584327e338143371ec6011f6be2320e606548

Observation de571cbc-35f6-4680-b8dc-5e4edcc7248e · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.314077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:bbd806e06b23a74550cce0570759adcf1a9747fb7e0b67a8e955fbe756776037

Observation ccc206b0-b278-4455-9a43-12e4fa3ea331 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.308306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:c737422c0ab3dd599a7a03043f487015b1177740113bcf70283c3bd2a44f1e23

Observation 252acf1c-3cf9-46eb-afa5-7b29109e5ff2 · outbound

This paper cites Young, A.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Young, A

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.354805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:cc10d724f5936f1ebf79b0f1ba660e83ce67658a9b82dac9f47bae35f1e54d90

Observation 95392362-76fb-4f61-a4b7-abf8d142ee33 · outbound

This paper cites Ypsilantis, N.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Ypsilantis, N

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.349636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:18dc9ee603f653c245a6a5f643a69ce859845b75df87a27d1cf5bfb3f1cd1217

Observation 772431e3-6a08-4288-a73a-19be9b542269 · outbound

This paper cites Ypsilantis, K.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Ypsilantis, K

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T17:30:01.377332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:d2de2909b5a653e4c68f2798f485e13e3fecd8cfefdfc9d30c006fcd1fe9ce4b

Observation 7339b91e-48e2-4dcd-aa37-80e4bc7015f8 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.426627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:2c8f223bae51dac6a6a086c1c5a6addbd546f92edc52cc45c05d9640582335dc

Observation 7af2400c-2dbe-4fdf-a6fe-1fb210ea3e4d · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.273974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:e35fd059d8a67fd278b7d14d83eb9b3f19e376aa68d876f4fe81a11f30275d4b

Observation 6bc6cfb5-29ff-4d80-a008-3aea236cfd1c · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.333418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:52022804939401e530790216addfe24fd72d14507ab2110dd2640e40aa697309

Observation fd5b29c7-623c-4dd9-a777-ba7b49569ef6 · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.362658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:2c69b69d251c16338b0956b7834b5a4052480ffe28997976ded92ceb7b69d47b

Observation 0642d745-9913-44dd-bde1-1c436fcedadb · outbound

This paper cites an unresolved cited work.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-17T17:30:01.303042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:1bd7c62b39895a07efed6fddf0e178e2a547be0c52308b21685b304c4e8e8938

Pith citing papers

Observation 85e2c683-0ee5-43af-a5e5-a03131446bf8 · inbound

Image Generators are Generalist Vision Learners cites this paper.

Image Generators are Generalist Vision Learners TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:41:06.045062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:14:05.034951Z digest=sha256:685e5c8fbac7f72c31cd64ba979319bb2e41afb0d0f59ac3ff6025b095cfe627

Observation e64482e8-370d-4625-9114-fec3b11aa0d3 · inbound

Image Generators are Generalist Vision Learners cites this paper.

Image Generators are Generalist Vision Learners TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:45:14.646817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:40:46.090808Z digest=sha256:9de64117400146c69c063f2e0e43237fceee169d434925132da3ba38a12f21ee

Observation 153ea559-a258-4bed-8a87-1ee1dc28e482 · inbound

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute cites this paper.

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:40:52.145228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:36:50.406353Z digest=sha256:c4685f9cf38b0d7dc3f692cf779c5868f4480bbdea0b1294ce1b5ce5fd832288

Observation 5c6d84be-dede-4489-8e39-2c5ad9f0611c · inbound

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image cites this paper.

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:51:24.724170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:53:40.505699Z digest=sha256:ccae8b89bbadbc121c40156818de3696db995dd3f921cfb479ecc5835680728d

Observation 2eee143d-e8cc-47de-9714-d5fe4de6612c · inbound

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers cites this paper.

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:27:47.864848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T21:26:36.707554Z digest=sha256:ff1bdd7f0ff9f6e18e55c0d21b4924d9773b742a9b58d97c36665e46312a6acb

Observation accff3d6-4c9b-4513-9ea7-7a1daea97668 · inbound

ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks cites this paper.

ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-26T20:59:58.035717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T20:53:01.922037Z digest=sha256:06a80e26cd06b1f4f32b13fa327f5d70f2da71eae291859b2f2ed73416e4640e

Observation a0ab3fbe-5941-493f-8f03-52b1b2ebbd3e · inbound

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models cites this paper.

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:09:37.353134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T14:50:10.041000Z digest=sha256:6512d4eb7283bd627e6fc0511aa1b6d88a983654b87e2e23536d2625f18bc8a7

Observation de75f00c-62f0-4962-9043-dd50da99aee3 · inbound

Self-Supervised Learning of Structured Dynamics from Videos cites this paper.

Self-Supervised Learning of Structured Dynamics from Videos TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.056117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.056117Z digest=sha256:13b0b1cebdfed2ec076f28f05e76e7937b40be22b6fb3bbc4b457bdfaf085ad9