Pith. sign in

Paper Citation Record · LEDGER

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2107.07651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.07651 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:30:42.957175Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T23:34:26.612743Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation acdf0181-1d6e-48cb-8011-b3a27f2ffeee · inbound

Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model cites this paper.

Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T20:20:41.193473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:20:41.193473Z digest=sha256:58956d6eeeea33033c5576fdd331b4c4833d5351a09c637e5fc689423b212f7d

Observation bb4f4c0e-5a67-4ac4-9cd6-ce8f9e74fbe6 · inbound

Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation cites this paper.

Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T10:30:42.957175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:30:42.957175Z digest=sha256:a233b264a805732c9f104776c15077aebd11248ca279ee107872a7f32f4595db

Observation e12c1a3e-1799-4555-b5cb-49fdce6416fe · inbound

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? cites this paper.

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:34:26.614390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T23:31:40.691896Z digest=sha256:e02534a4acea753161878043066f0ac9870368ca2155f1a5e905612682138e9b

Observation b5622a6b-20b4-49cd-aa99-715188ae19a9 · inbound

Multi-modal encoder-decoder neural network for forecasting solar wind speed at L1 cites this paper.

Multi-modal encoder-decoder neural network for forecasting solar wind speed at L1 Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:56:09.816132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:56:09.816132Z digest=sha256:46bbe9df5c00470bcb7df04622090ed749d1d24cf1a8eb02302041c922ced124

Observation d1ddf609-711b-453e-8456-6e6b0383d7c5 · inbound

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis cites this paper.

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:22.016905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:22.016905Z digest=sha256:3d0ca1a57d49bf6f1ac640ee25e25640f2d68226da5f4151223d113927ccb577

Observation 724739de-1ce5-4e61-83a1-496263fd3782 · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:24:21.338167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:df4f20c7c89b699f515f14cfa4ebc334b8369e8f89a8622322178c4a40075d56

Observation b3fefd16-2022-4c83-adeb-e0478648b496 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:33.131873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:33.131873Z digest=sha256:f3af4a325c71e7d9ec94906c6584712f249ba5b1cbd53ad76ab53fdd2786c98d

Observation 78a038ea-4f53-4a2e-a438-ecd7daacff3c · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T17:19:37.624395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:37.624395Z digest=sha256:6d9c77bcd82fcb085410c4a297b69fb2bdc24b0f705f3e5cb20d6f4e4c3b2143