Pith. sign in

Paper Citation Record · LEDGER

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2310.09199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.09199 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:57.383936Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:56:35.167882Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 197ddb6a-225a-41f0-a0b0-87643d848315 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:46:36.300773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:146131ecb319fd69fcaffbabd07d124dd766ea91e7888e99183927dbb21d9fd3

Observation 0123981b-f859-4870-8551-d2729f5609ca · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.162525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:3a77862b20e9f3323221deb1c734b2581f8a5ec783127411d229acfd056471d3

Observation c9085aad-b652-4417-917f-ead2dd30e829 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.585432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:3d00437be2863efe65b1fb29615c98fdb041dd9249c89f9f0b807a260f839cc6

Observation 4fba7893-3477-4fa5-9ad8-fb06bb1015ce · inbound

Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics cites this paper.

Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:16.340460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:16.340460Z digest=sha256:eb7c22940eb4aafc15acd2c3d1d6dbd0dab73c77b92512c2a506fb2ffe997ef6

Observation 02ae1770-1abe-4fc2-8916-527203cf8e6e · inbound

PaliGemma 2: A Family of Versatile VLMs for Transfer cites this paper.

PaliGemma 2: A Family of Versatile VLMs for Transfer PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:07.620366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T09:15:07.523565Z digest=sha256:38a069ac5621fc54a70881298ffbe21b5ecbd63560b274d36ca5477b84ea899d

Observation 4dfe4c90-b52b-453b-8d31-5c68720cd02e · inbound

AnyBimanual: Transferring Unimanual Policy for General Bimanual Manipulation cites this paper.

AnyBimanual: Transferring Unimanual Policy for General Bimanual Manipulation PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:26:41.737682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:26:41.737682Z digest=sha256:2072d0c8cc851e0823d70d323b39749d95f6af88444b5e2b94672ca3de60b912

Observation 36c54bed-25d8-4b12-8d24-137be9caaf0a · inbound

7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement cites this paper.

7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-11T20:27:18.498337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:27:18.498337Z digest=sha256:7c3c7a37ed18447cac0829efdacb2318230f75a2a04a1ba9e413fdbf434bd72b

Observation 0db02ade-4d5c-4fc9-80c0-dd8a8cc24eb9 · inbound

DocVLM: Make Your VLM an Efficient Reader cites this paper.

DocVLM: Make Your VLM an Efficient Reader PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.716518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.716518Z digest=sha256:41df98086b204a43446d8bfb4a22a40d09eae4b376feb4520fb75c7ef86ad65d

Observation 00211690-4b76-43c9-af30-58d775ede52b · inbound

Neptune: The Long Orbit to Benchmarking Long Video Understanding cites this paper.

Neptune: The Long Orbit to Benchmarking Long Video Understanding PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:28.603722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:58:28.603722Z digest=sha256:71dc65a6609ff5d94350741a3081fe2f2fc8d7afbc541447625c8593a02a4ad0

Observation 77c68856-9872-4508-9a4a-6ba6bc4232ca · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:27:22.842553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:0b9b6013c61c65031c35aedcd1e6250294e736850128972d0c5b808e7bd4b4a6

Observation ef648ba7-9177-4563-a7ce-4c4bc552e631 · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.076893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.076893Z digest=sha256:7612aa9cc3cd0c6ce0f2ddb6595157b704a840d15f1c2b7d373e1cc0da2a93d9

Observation 41fc879e-a193-4ee4-a796-5aae5b8012bf · inbound

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images cites this paper.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.479694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.479694Z digest=sha256:5c77e7ed3140a3af1a7ec75058402913da8554014479ca8a250a96f9b866d27f

Observation 9cf15b16-9dd9-4d6f-b357-9ebd2a45c8fd · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 214

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.948642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.948642Z digest=sha256:aa876a8896ca08cb23f461a58d4d17ba15f3df920aa1409421a61764da17f9ba

Observation 1fd1c37b-bf4b-4494-80a9-446b801a26ed · inbound

BiFold: Bimanual Cloth Folding with Language Guidance cites this paper.

BiFold: Bimanual Cloth Folding with Language Guidance PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T13:12:46.368082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:12:46.368082Z digest=sha256:2d7a96eedf923ae520b1e4a958b35d638fc9b8da45497e2ac1a526ff08356bed

Observation 62146370-3bbb-4834-85ed-c1be60865c89 · inbound

From Image to Video: An Empirical Study of Diffusion Representations cites this paper.

From Image to Video: An Empirical Study of Diffusion Representations PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.182071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.182071Z digest=sha256:380572e7151686153c1c61f6d9e7e94a9f8dbc0b800e8937fe0cdb0566fe413d

Observation 1cb02f62-f99a-4e2e-9ece-1e73d1980bd6 · inbound

Scaling Pre-training to One Hundred Billion Data for Vision Language Models cites this paper.

Scaling Pre-training to One Hundred Billion Data for Vision Language Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:12:30.662861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:12:30.662861Z digest=sha256:ad1350eec900fa41552e9f4dbfbf4ffa98e4883a4191d49745dfd6a44cec34a1

Observation 1369fc8e-ca41-41cd-ab5d-3da27cc877e4 · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.853944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.853944Z digest=sha256:1aa1153e88cda65b45a144baf66a047e6aa97bb39e00034164e9268f255fa9b4

Observation e559251c-59cf-49e4-a982-993066342d07 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:21:45.212918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:0a820d761e8c7e66ea1a1581abf427952d96257c69b88339501f6e3931fa939f

Observation ca28d16d-5870-4756-92d0-de895c98bbbc · inbound

Early Accessibility: Automating Alt-Text Generation for UI Icons During App Development cites this paper.

Early Accessibility: Automating Alt-Text Generation for UI Icons During App Development PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:57.383936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:57.383936Z digest=sha256:adcd04acf02b3800bc5a843fe60c045ee21e2a45ce1cdb29b3a3d6ffb869c26a

Observation 21c85dcd-64af-4bee-9355-dd6fd3ba3694 · inbound

Context and Pixel Aware Large Language Model for Video Quality Assessment cites this paper.

Context and Pixel Aware Large Language Model for Video Quality Assessment PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:21:35.454864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T13:20:59.840347Z digest=sha256:5830fa7326c295ecd9dfbe6fde71d70ec64973db5b615738d444b113c944f4b6

Observation 7ef6ce2e-1f01-4d44-aac5-6835a73a4daa · inbound

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation cites this paper.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786935Z digest=sha256:e3f2f25cdb314a9689a4d2ebf821fb0435dd721bbaba0b70fbfd384aec13d18e

Observation e027e3ac-455f-4719-bcb7-ec0c4b1b4b45 · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:51.104029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:51.104029Z digest=sha256:091b8f7c24fe1f9e657cdcf703a3e1edc91a641c412a79fda1f5c948fdb33144

Observation f73bcae5-be14-4df6-9a04-4a9e7ad0463b · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.260193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:85fadca0182d785f6ea571933b7eeb01614e2f324a8a02296561b0b24fe59e98

Observation 62903617-4d48-4494-b541-78cd6317ea22 · inbound

Generalizing vision-language models to novel domains: A comprehensive survey cites this paper.

Generalizing vision-language models to novel domains: A comprehensive survey PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 290

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:06.067923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:06.067923Z digest=sha256:c36c263d1db0f2a46c7db8ef16c5ef9d2d4d39c97180742fa88aec55167f85c4

Observation 579e83de-05a9-4ea4-bd85-97c9bf19cca7 · inbound

From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs cites this paper.

From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:59:54.338478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:59:54.338478Z digest=sha256:99a5893cf261102111f49b4910a4b79444a0746d6fc930158d0f02ed4ddd4761

Observation 946f5be3-94c9-4f85-bc6b-7c95eeeabba7 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.592591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.592591Z digest=sha256:ea024d6c6a78f2003f3e9362008000e2ced8cff86132b20090a4888d16be2725

Observation e906fc2b-413a-4ce2-bdcc-f17a0a8e76e9 · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:29:09.886292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:19a32a89091765532833c5ec88958a5448974af48c67ab5c20065bfa7b796843

Observation 15147e5a-4388-4fdd-916c-7690b0fc2495 · inbound

CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion cites this paper.

CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:54:14.512459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T15:53:44.767068Z digest=sha256:ad35a1809b20a74935dcc32a796200e9263b869602ea29eb5665cc3c3be0cc12

Observation 01a365a8-5847-4109-a518-e62a857f273d · inbound

On The Application of Linear Attention in Multimodal Transformers cites this paper.

On The Application of Linear Attention in Multimodal Transformers PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:59.530104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:12:49.695614Z digest=sha256:455f0eb7d2fc6b2a115e80f945faa9ce730a596c64faea1a678383775083b98a

Observation 9805146f-5aef-47fd-b028-9beb53e77963 · inbound

Weak-to-Strong Knowledge Distillation Accelerates Visual Learning cites this paper.

Weak-to-Strong Knowledge Distillation Accelerates Visual Learning PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:15:10.317380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:13:31.161933Z digest=sha256:ab7c084d44994a23e915af2f0404a03b6e5efa5a48ea431097cd38cb0d799181

Observation 52b97257-e3af-4ce2-b985-68a1af823548 · inbound

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models cites this paper.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.871267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:90cc0751b72133c70b0cc616c04f7ceb1eba8efcfe76cefcb538debe5640ebe2

Observation 88cbd8bc-e83e-4754-a297-3b68df6a8f80 · inbound

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models cites this paper.

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.230558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T08:07:51.697353Z digest=sha256:c2c13845cb3bea919485a958257ec8b75c4487eecc6c7d297c64e67e6b1b6be9

Observation cd14ee36-d963-4581-8dd7-b9403db91f00 · inbound

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models cites this paper.

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.000288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T17:48:02.228941Z digest=sha256:a215e24bf9a0329b5790a08f5931531d3f224db509af951bcb72111b8149fd9b

Observation 36590b86-5a2b-4ea6-abae-3dbb20de2961 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.899095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:22d4b719005b1434c80e56ec3b96bb49bdfca038249efae4a4a461ba8706fce5

Observation b69f6fe3-46b2-4bf2-9025-187b05735030 · inbound

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders cites this paper.

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.244445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T19:23:08.100056Z digest=sha256:6d679211a4eb8a138a9e61ab3b3a6d48db571648f4cc859a1dfea4b2ca2c0a94

Observation db1537cb-77a4-478e-851c-ae1c3c7dcd5e · inbound

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning cites this paper.

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.169406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T09:31:56.712360Z digest=sha256:f3af53aa5aa76e872904462cf42c89989d7805e4bd32dc2a4f86b513aab8edc1