Pith. sign in

Paper Citation Record · LEDGER

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2202.03052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.03052 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:08:42.037290Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T18:15:14.590918Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2984da1e-7e1d-4986-91ec-658eb3347da7 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.292817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:026bc8d3414d2a9bba55f9d34af0ceafc61f5400771663125fd3d98fe12ccffd

Observation cb6d162d-7fc2-436d-bea7-e30364cc991f · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.391341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:265b01cfc45611334b0a43ad9c6715124a87883ea8860af80d14c7f572232f02

Observation a9cbffaa-0912-4a70-9df9-fab7bd93e46e · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.723201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:5a129830de5dd9c3bcadc0ef0bb4c12579450f0ae402bfed880a7d545510171b

Observation 75c31969-c348-48af-9347-9793304c8979 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.240416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:6e6ff1b1770cc324999ea56cd9246c8dc7c2bad0dd4f6c1b87d1cadcc31155a7

Observation 3d71af40-d070-4805-a9de-c6fc64f9c9ec · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:15:14.595172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:0e3de0a39d746cdc5368e6704ca8f458a1cd092712cfbbff0af250b27b02821c

Observation fd024c18-b775-46ce-af79-d6fafcb83187 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.175626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:3cc7bfc11ad0c9620d5962d50077f9d0ddbef769e2b91b2a7ccc042bc3caf471

Observation b0e88f0c-1d92-4bd9-9257-708632af24b9 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.037290Z digest=sha256:13e6087dc86414f613a16b021265b3a2472c09b9390001ff95d6bbaeafeecb13