Pith. sign in

Paper Citation Record · LEDGER

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2502.06355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06355 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:45:46.182206Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:41:37.241236Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 047863be-79c3-4071-8719-f096f8fb55be · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.511457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.074894Z digest=sha256:a5a19e04504a63b8258fd5fe0e1faa2d2107cf4d76e9dcabc396e84adc5ffb73

Observation b481ac32-0527-466c-bfb5-1ec656a21d18 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Imagebind: One embedding space to bind them all,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.458426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.099243Z digest=sha256:9c8ee2e779c419203c2432a3658f20152d3a23ea385037b4ceed5e9e863a9bad

Observation e69c759e-9c45-43be-b0f2-2f2826b2543d · outbound

This paper cites Dis- tributed learning of deep neural network over multiple agents,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Dis- tributed learning of deep neural network over multiple agents,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.447966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.106043Z digest=sha256:7e663752b5e281d31d0813d5966a30877dcddc4a645a4d4255fd460abbc7698c

Observation afc58a40-9189-4937-859e-37974fad4339 · outbound

This paper cites Parameter- efficient transfer learning for nlp.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Parameter- efficient transfer learning for nlp

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.425705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.112255Z digest=sha256:6550af71f9340743878401836b18b58dfefdc1f64f813c29af50afe904a7c454

Observation d384fa3a-feb2-4ae9-b880-3e2ffc2774d7 · outbound

This paper cites ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:45:46.270073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.119825Z digest=sha256:c5ee3ecd342ecb00dd7e2d17adaecab0dd85772b21b87fec4378ba04852ab653

Observation c5538d8b-5472-42ba-9f67-dc161476ce5d · outbound

This paper cites Federated learning on non-iid data silos: An experimental study,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Federated learning on non-iid data silos: An experimental study,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.123879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.123879Z digest=sha256:c29308da1624f22f4d4b6c262ab692392d2539a7a11b175341a92923a3fbe343

Observation 88d47e6a-9630-4a15-87ed-9c1a425d8649 · outbound

This paper cites Microsoft coco: Common objects in con- text.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Microsoft coco: Common objects in con- text

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.396507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.127552Z digest=sha256:2972ea78d5342d781c851ce2eab5077d2c82f866ade04c5815f2be14fddce4ab

Observation 6e99a235-226d-40a1-9f01-331a6e6b765e · outbound

This paper cites FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.135137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.135137Z digest=sha256:50262f8b75956239b815d231f365ad80de848580077e8de13f0bac48bc31381b

Observation ff2e8f44-388b-4317-aab3-d1c0534f33e9 · outbound

This paper cites Scalable aggregated split learning for data-driven edge intelligence on internet-of-things.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Scalable aggregated split learning for data-driven edge intelligence on internet-of-things

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.374599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.139246Z digest=sha256:e352fac7fcfbccf8332d7af649759f4ffef0f32d45856723df07e98dfec22e60

Observation 76c39fa3-52d1-4317-b6f7-97b9a3ded190 · outbound

This paper cites Communication-Efficient Learning of Deep Networks from De- centralized Data.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Communication-Efficient Learning of Deep Networks from De- centralized Data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.364187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.143029Z digest=sha256:3a7ec35ce649c544fa2fc1e385fbcbb47c232836d1f5003a3ed78605e1f1f69b

Observation 88096161-94a3-4b34-b601-c554372c80c6 · outbound

This paper cites Mix2sfl: Two-way mixup for scalable, accurate, and communication-efficient split federated learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Mix2sfl: Two-way mixup for scalable, accurate, and communication-efficient split federated learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.353629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.146614Z digest=sha256:330c8c941812d2a7f95569f8a00236ebb7ff0ce865ea60559e0e53d282ae21db

Observation bc431f59-4408-4cbf-afad-156d1ee9d040 · outbound

This paper cites Server-side local gradient averaging and learning rate acceleration for scalable split learning,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Server-side local gradient averaging and learning rate acceleration for scalable split learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.341856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.150833Z digest=sha256:ad37a0d2f0653d2a9a57781a9a4930d23ea4ad55d1ee7f2a22092e6c862877d0

Observation 867ab09c-342e-4aff-97cd-37ca8eb9c915 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.331970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.154589Z digest=sha256:168b16d7a892b6d4e6794a914e32b38b3bda0d78fe4b186cfe6d40b581612df8

Observation e9e6aa05-c7d5-4b4c-b220-c5a596f52fd3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Learning transferable visual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.322614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.162140Z digest=sha256:11ad72591a7d9a51509622b3696c0ce9027798e2e960c1d2a96948f3b8e56fce

Observation c6e27adf-2f4b-4aed-96c9-94f094e9df24 · outbound

This paper cites Exploring models and data for image question answering.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Exploring models and data for image question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.312961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.166079Z digest=sha256:f9e69e827b6af2734e4d66d7bb87637e83913caa069ffbf843ae3a6aefe8002d

Observation 068d5ab2-9a8a-42b7-95a8-03906bba3f95 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.169912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.169912Z digest=sha256:b61f5a4c98980b49f52bdafc27c5e7a9826d9b4f7eed580400f6d44a42837eb7

Observation 7ed6be78-5418-4f27-896b-ac766681927e · outbound

This paper cites Permutation Equivariance of Transformers and Its Applications.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Permutation Equivariance of Transformers and Its Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.177880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.177880Z digest=sha256:d98d35d25f42a631cc496ab410eb2b031ffaf15c3cb8246b67bfba84baa12cb8

Observation b28a1464-fd51-485c-b3ce-5993f5cd603c · outbound

This paper cites Meta-Transformer: A Unified Framework for Multimodal Learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.182206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.182206Z digest=sha256:dcc98049b7e8eb70abf148a674f09a1e3a8fdab55647f366cda93d24d80c8869

Observation b2faa4c5-5bf5-4e57-9a74-ab440ff1a163 · outbound

This paper cites Cross-media learning for image sentiment analysis in the wild.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Cross-media learning for image sentiment analysis in the wild

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.302732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.174092Z digest=sha256:32c572eb63c95a88d3e2e6413d56fde08ebdc05207c7c93b8e4f618394b39c8c

Observation 203b684d-ddf3-4051-9698-06c6b2d55a2d · outbound

This paper cites Efficient par- allel split learning over resource-constrained wireless edge net- works.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Efficient par- allel split learning over resource-constrained wireless edge net- works

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.385583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.131345Z digest=sha256:6ed98885151a482ce3cdf07d68ce825f65325e168e7e0c059f75969fd83b6ff1

Observation 476a4860-d038-4dbb-988b-e16c8978d0a9 · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.158254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.158254Z digest=sha256:1e8f39b697ed297303209454327380220cf9661b500d3d10a002d718332bbe49

Observation 1919314f-7d34-4b0e-90b4-2d8e4a4a7660 · outbound

This paper cites The iot breaches your household again.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach The iot breaches your household again

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.488910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.083782Z digest=sha256:36c10fc4576446145b7453f4b93245110972eed622cdbec9118f6154b59c0508

Observation fb285ac0-c69a-48ed-af88-c8e18d841bb3 · outbound

This paper cites Accelerating federated learning with split learning on locally generated losses.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Accelerating federated learning with split learning on locally generated losses

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.436925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.109116Z digest=sha256:44c0a0c23e4a9d5cfbd7ddf9e186861cdaa57ed480f302933159d391ed5e510c

Observation b61cd8e7-21e4-460c-8769-d1045fc486f5 · outbound

This paper cites Privacy-sensitive parallel split learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Privacy-sensitive parallel split learning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.414118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.115887Z digest=sha256:bdfec41331f56133c7266b234441be95fe02056dbfff53c64892a26d6189b76f

Observation a2c380c3-b741-4d10-88a1-fe7d9271c522 · outbound

This paper cites Multi-modal align- ment using representation codebook,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Multi-modal align- ment using representation codebook,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.478560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.091790Z digest=sha256:cba2fae793e9f820053bd66823011b19cd4d4fa8afa407acb7893f9695b02584

Observation b483d407-0919-4b42-840a-94bbdca68a30 · outbound

This paper cites Look, listen and learn.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Look, listen and learn

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.499614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.079585Z digest=sha256:2281e1a9ce31d662c24246dd6ba1f39e8a1803e1cc17670f17e5faee369e399e

Observation 46967e75-311e-4564-ac23-cebf4f614e40 · outbound

This paper cites Split learn- ing of multi-modal medical image classification.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Split learn- ing of multi-modal medical image classification

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.468989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:45:46.095406Z digest=sha256:cddd0e5c4446752d19977b6104b41847135e0dd5e339023226cf284bf85891d4

Observation 01a0685b-d2c9-4e6d-8743-21b5b06fdc70 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach AST: Audio Spectrogram Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.102491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.102491Z digest=sha256:75e7f50d4f8f3b46da72f1fb90afdb2980c99138feb1e0d7d1abbc07a699ee11

Observation 633f2018-36aa-45ce-90cd-b1a626ebee3a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.087545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.087545Z digest=sha256:c9035a75426f42fec69f131803c7ff97e9a246ade0fc68f76856f9eb60abd4a1

Pith citing papers

Observation aebefd0f-aadf-4d82-8629-9886ad8242a6 · inbound

AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning cites this paper.

AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:41:37.241236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:41:37.241236Z digest=sha256:d8ffb8e5de34718c125905bc5262719a9e23f841beaf6a936bfe0865ebdc9e28