Pith. sign in

Paper Citation Record · LEDGER

Unifying Vision-and-Language Tasks via Text Generation

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2102.02779.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.02779 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:39:57.120440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T02:13:52.177920Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b755161-29c7-41e2-9d9d-841baa9d06fa · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Unifying Vision-and-Language Tasks via Text Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:10:49.032939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:85998823d42eeab56ae25647e84c6f2bba51b669beead600835cb76d340023ae

Observation 02a3adda-1471-46e1-b6c2-21b3311a8c23 · inbound

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning cites this paper.

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning Unifying Vision-and-Language Tasks via Text Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:13:52.181936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T02:13:52.097263Z digest=sha256:56a6cc6dd170ee6dbdd37386ff22a0c13c70d04d9edcd1544310c023f563ea81

Observation 6b33a499-2173-4ca7-89a1-638fce6d92f6 · inbound

FLIP Reasoning Challenge cites this paper.

FLIP Reasoning Challenge Unifying Vision-and-Language Tasks via Text Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:39:57.120440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:39:57.120440Z digest=sha256:4f34bff2a637d70a53622384113664270a620708b4bde6eb12863c8e13fc4538

Observation 5824ec6d-c163-485b-85a7-6c84408a0665 · inbound

Multi-modal single-cell foundation models via dynamic token adaptation cites this paper.

Multi-modal single-cell foundation models via dynamic token adaptation Unifying Vision-and-Language Tasks via Text Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:57.017796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:57.017796Z digest=sha256:5b69edcbabac50de57f7f2483d55c6820ac19b5809610468ffdb603ac24b07a8

Observation 4701650d-827b-4e98-8b9d-2bb1acff2070 · inbound

Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems cites this paper.

Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems Unifying Vision-and-Language Tasks via Text Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:58:00.486496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:58:00.486496Z digest=sha256:9189e4e6cef9ab93c0f8a6956213457f2d8fead95ab426b7fb253ee24ff1bea6

Observation 95752b53-1ad2-42a0-bedd-60e424aee94d · inbound

AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning cites this paper.

AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning Unifying Vision-and-Language Tasks via Text Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.224300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:13:52.705187Z digest=sha256:da799b24c0524729d0f7ab0efb0fc9c3d5189b005854d905b71d316ad8f614d5