Pith. sign in

Paper Citation Record · LEDGER

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2107.07651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.07651 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:30:42.957175Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T23:34:26.612743Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation acdf0181-1d6e-48cb-8011-b3a27f2ffeee · inbound

Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model cites this paper.

Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T20:20:41.193473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:20:41.193473Z digest=sha256:c7e7ae9342c927f1539c5dcb1d1d8b4709b51a3b923cb0576bfe26c720d75042

Observation bb4f4c0e-5a67-4ac4-9cd6-ce8f9e74fbe6 · inbound

Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation cites this paper.

Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T10:30:42.957175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:30:42.957175Z digest=sha256:bbee2176191dd677c83cc4e59b1fb56c4ff7ee9476a67d7a0ca1a25eb3fb3e9d

Observation e12c1a3e-1799-4555-b5cb-49fdce6416fe · inbound

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? cites this paper.

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:34:26.614390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T23:31:40.691896Z digest=sha256:cd62f53dd75df0abab6a6062ebde994e5e858cb356e7250ff4e7810ceeb65145

Observation b5622a6b-20b4-49cd-aa99-715188ae19a9 · inbound

Multi-modal encoder-decoder neural network for forecasting solar wind speed at L1 cites this paper.

Multi-modal encoder-decoder neural network for forecasting solar wind speed at L1 Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:56:09.816132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:56:09.816132Z digest=sha256:46bbe9df5c00470bcb7df04622090ed749d1d24cf1a8eb02302041c922ced124

Observation d1ddf609-711b-453e-8456-6e6b0383d7c5 · inbound

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis cites this paper.

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:22.016905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:22.016905Z digest=sha256:3c27f9f7916e5a43dc8b950e37c4d8546bfcb7542f7792d45a8664d5116b5466

Observation 724739de-1ce5-4e61-83a1-496263fd3782 · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:24:21.338167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:6e815fe9692d4c9322e392c0b4c5dfaab1c2d5da55ff660331384f29045db518

Observation b3fefd16-2022-4c83-adeb-e0478648b496 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:33.131873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:33.131873Z digest=sha256:f3af4a325c71e7d9ec94906c6584712f249ba5b1cbd53ad76ab53fdd2786c98d

Observation 78a038ea-4f53-4a2e-a438-ecd7daacff3c · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T17:19:37.624395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:37.624395Z digest=sha256:6d9c77bcd82fcb085410c4a297b69fb2bdc24b0f705f3e5cb20d6f4e4c3b2143