Pith. sign in

Paper Citation Record · LEDGER

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets

As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2501.03332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03332 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:41.937067Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cefb8f7e-1b1c-4885-895d-066ea60b3fd7 · outbound

This paper cites Multimodal personality recognition using cross-attention transformer and behaviour encoding.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multimodal personality recognition using cross-attention transformer and behaviour encoding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.746006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.747478Z digest=sha256:279da3d85c86f5d8d1d9300c6fe7aa5968372121558ec202b9e7e992839466e3

Observation 9a3d884a-bb0a-4e77-989b-edcb82da6259 · outbound

This paper cites Multimodal vision transformers with forced attention for behavior analysis.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multimodal vision transformers with forced attention for behavior analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.730891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.752893Z digest=sha256:955a298ba7a8372a1d260290ad7f538a31eeb959bb776e0c31bc80609bace9e9

Observation 1eaf3175-b942-4fb1-8eb4-6ab022416907 · outbound

This paper cites VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.757556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.757556Z digest=sha256:c7a2caeb70279f6578b45cdecb3bc24afb21d4561b1bb21d70a547c4cdfbd0fa

Observation 1c420e10-9362-479c-a3a8-236326b87c4a · outbound

This paper cites Vivit: A video vi- sion transformer.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Vivit: A video vi- sion transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.716215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.762729Z digest=sha256:d2fcf380e1d2b45cb1bb40c8b67477c2773f8144576610b27b99de4b058ee846

Observation 3e79fad3-0e5a-4983-a89b-43454524ad2d · outbound

This paper cites Bodily be- haviors in social interaction: Novel annotations and state-of- the-art evaluation.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Bodily be- haviors in social interaction: Novel annotations and state-of- the-art evaluation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.700369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.767362Z digest=sha256:1ac9faf6cd12b358096225e4ef268e9807fc14e0e81ff2306ab3dbebb11181f1

Observation be8de12e-1c86-499f-b53b-51028f9afdec · outbound

This paper cites AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.772310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.772310Z digest=sha256:db40bfd3fee01b64904b0ff5ffec41ef0d459e553466d747bb9a9f6b15e1696f

Observation 29b47da0-5ead-46ef-98c8-e46eb20e0c02 · outbound

This paper cites Rescaling Egocentric Vision.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Rescaling Egocentric Vision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.777674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.777674Z digest=sha256:5eabe67271ca98b3c5ef61a5a741978ae9828d1e589391155b7a04e820408e15

Observation 2d7d4960-2027-4a56-9d79-a5fc39f7bf2c · outbound

This paper cites A transformer-based joint-encoding for emotion recognition and sentiment analysis.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets A transformer-based joint-encoding for emotion recognition and sentiment analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.685448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.782611Z digest=sha256:c8f3684dd1fb709a1eb93080250678a6cfb900d871f75cd78322b0c9f4198f12

Observation 517d6aa6-979a-4842-a0e5-e3e409206b80 · outbound

This paper cites PPT: Pre-trained Prompt Tuning for Few-shot Learning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets PPT: Pre-trained Prompt Tuning for Few-shot Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.786939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.786939Z digest=sha256:919fe988161f41ceb0cca2f50240c845c93ec60ecc21a6541c7ba4f7d9a5bf80

Observation 57d12086-2d8c-486e-96e5-af7792988656 · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.791474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.791474Z digest=sha256:d4663d4de304a6d20605ea1081c9a60c9bc90c60ae52d66a7ce7d94b5b5e5d66

Observation 17187cc0-9144-4617-b319-154d44b933fc · outbound

This paper cites Parameter-efficient trans- fer learning for NLP.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Parameter-efficient trans- fer learning for NLP

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.669455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.796366Z digest=sha256:f90220cc2fd4149340710597cbdb7480ca0f9f8fd72749745a3f3e390616fdc4

Observation d50bb71c-9b26-4cca-9ca9-3d8dff2a5375 · outbound

This paper cites Lora: Low-rank adaptation of large language models, 2021.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Lora: Low-rank adaptation of large language models, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.652332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.801052Z digest=sha256:9908668b40b7a499182b239f751ae494589ee1d8848838227ba34c8e9f937a44

Observation 5f3ad6b7-82a6-4838-98de-3a7c28a17858 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.805566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.805566Z digest=sha256:45c6819c5afca6a5b198c6b9d22e488f8556af82b8d77f95da30242c269ff129

Observation 47638112-c3bb-4db1-b42e-0a4f6c7fee1f · outbound

This paper cites Mumu: Cooperative mul- titask learning-based guided multimodal fusion.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Mumu: Cooperative mul- titask learning-based guided multimodal fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.637267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.810254Z digest=sha256:be450add4879786002c67899f5bba7700ce5010bfb6c587a3236ea3380f7bde7

Observation 1e937211-8683-4eed-a255-fd36306eb630 · outbound

This paper cites Vi- sual prompt tuning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Vi- sual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.622072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.814539Z digest=sha256:45068049eb67ab151f20aced66672d9d80dc343aa160c85a96d55353442ae4bd

Observation 062b865c-3421-4328-b567-87405a285e13 · outbound

This paper cites Compacter: Efficient low-rank hypercomplex adapter layers.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Compacter: Efficient low-rank hypercomplex adapter layers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.607216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.819154Z digest=sha256:7f394c6e1e83d214d9cb0df173575d156cf563d05b5bb1e93010a10588f8fc01

Observation 37329977-4498-4e23-aac8-b181ca94c9e5 · outbound

This paper cites Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.591462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.823645Z digest=sha256:f3ace5dd652710321cbc20d4c32399044b0f3fc96b48ba29dafa4dd736469293

Observation 039df51b-eee5-410e-b860-2f411990162c · outbound

This paper cites Transformers in vision: A survey.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Transformers in vision: A survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.828211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.828211Z digest=sha256:3bfbb5f77443af5117c92c1fe87285bf67ea852f5b7374152d353fd3ff9c37e6

Observation ecc15948-d0c0-46eb-b0a4-d2f7e61e1c0c · outbound

This paper cites Gated Mechanism for Attention Based Multimodal Sentiment Analysis.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Gated Mechanism for Attention Based Multimodal Sentiment Analysis

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:57:42.254175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.832911Z digest=sha256:ad54b87731dad01f9eed61bc0604ef458f92f4fb09c899635bc8723cc2049993

Observation f7680172-9a00-48f8-b83c-08944022ab9a · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets The power of scale for parameter-efficient prompt tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.566531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.837764Z digest=sha256:d634210fca0a443ef0875928c58cb4d38d4df32f4c8f2e5f3d2d2d8634a00cba

Observation a2452fab-058f-49ef-a192-af15732c1147 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.842107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.842107Z digest=sha256:c0d0db738726985d2501448f0a937691572c146c5ea0fa09b94234964ee1487a

Observation 0ec927f3-ed76-41bd-aca2-7d3101e3d2f0 · outbound

This paper cites Video swin transformer.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Video swin transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.551485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.847102Z digest=sha256:63443ec0feb92b6fba46adf1efdebbd153f38c387e750556f63cf2969abf69bd

Observation c4eee4e8-5c4c-4682-9226-73692234e2f5 · outbound

This paper cites Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.851696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.851696Z digest=sha256:0913b8324c64241f63fc3e4a0962b6e4aea0017024ee8af26190114a5647c37d

Observation f125b378-f6e3-48ad-9f6f-80238fe6635f · outbound

This paper cites UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.856385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.856385Z digest=sha256:d5ec010ae5f3eb8fdc13ef05628959995239947e07deed1d9a85aaf0d1ae7035

Observation 791ab3e0-3809-4a6c-a617-1c3875df456e · outbound

This paper cites Tiny adapters for vision transformers, 2023.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Tiny adapters for vision transformers, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.536574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.861544Z digest=sha256:ebb2e6a37c52a9d598676293d754254b9fb446176845621ed8a01edb56e7a350

Observation 3ab3513c-5b1e-444a-84ab-a79ca64d5114 · outbound

This paper cites Context- aware personality inference in dyadic scenarios: Introducing the udiva dataset.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Context- aware personality inference in dyadic scenarios: Introducing the udiva dataset

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.521422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.865979Z digest=sha256:e57438632648175b0ef6532728c4773b697d5a42366e29c8b2dbd739ed7fe7cc

Observation f5f113a2-2ebc-4f1c-9de3-62d70bf10728 · outbound

This paper cites ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.870525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.870525Z digest=sha256:3b0ed4f7c4fb432bc5f48f8f73a9956fc104af06dafe82625c08f4b0fb11953d

Observation 70d7c14b-1d06-43ed-a5ee-6204b3fdae4c · outbound

This paper cites Dual-path adaptation from image to video transformers.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Dual-path adaptation from image to video transformers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.505307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.875584Z digest=sha256:96bb9608bd475eaef01d2dc956381d318288dff248a425399017922ad987ff00

Observation c53a2f50-64df-4e3c-9c6f-09538f2edcfb · outbound

This paper cites Learning transferable visual models from natural language supervision.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Learning transferable visual models from natural language supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.490340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.880284Z digest=sha256:731b3b91106f8c8ebffae14e9eaf4da3aa9c57d340fe6f772fd65af038e96537

Observation 7dcef711-74ce-4ef1-a1d0-c7859b704e57 · outbound

This paper cites Learning multiple visual domains with residual adapters.Ad- vances in neural information processing systems , 30, 2017.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Learning multiple visual domains with residual adapters.Ad- vances in neural information processing systems , 30, 2017

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.475392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.884972Z digest=sha256:666533d7d8752dd5a675360628becc5580349026adb8596e4c835a88d0e6dad7

Observation b3cddac2-21f3-4624-99d5-391385aecde6 · outbound

This paper cites Efficient parametrization of multi-domain deep neural net- works.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Efficient parametrization of multi-domain deep neural net- works

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.459033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.889613Z digest=sha256:19923aeffa10f5e84f579c7600c932ea4f09b74bb665f8464e52f966121ef50f

Observation 05eb2b85-d972-4d7e-bba8-745c2bd5e3f7 · outbound

This paper cites Multilingual Detection of Check-Worthy Claims using World Languages and Adapter Fusion.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multilingual Detection of Check-Worthy Claims using World Languages and Adapter Fusion

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:57:42.167868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.894087Z digest=sha256:15d8768c721afa439eff8e5fc4734aab8d81b465309234f0adf106328ae18780

Observation 332d9f91-1186-490b-a9da-593c90a511ec · outbound

This paper cites an unresolved cited work.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:42.442110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.899007Z digest=sha256:b8f9cfbd06a2c7cd9970e86e592a022a120e7b8d34ab15df7a524a2d96acd5ba

Observation 811a7111-1a81-440f-b2b4-57d8cfb7adc8 · outbound

This paper cites Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.425548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.903694Z digest=sha256:343caf3e4aed4c2f4ec8c5338629890003d4567fc17aa2c7786788b0d70ef16d

Observation df61c7e2-bb0c-4dfa-929a-4697adc19e80 · outbound

This paper cites Training neu- ral networks with fixed sparse masks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Training neu- ral networks with fixed sparse masks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.408691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.909217Z digest=sha256:7e23321d4407d03d4d5aece2e766e6a683ea5a29bf95ce18018bd961fc6a5f64

Observation 1051fd7a-9059-410d-b500-fc5f29f641da · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.393480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.913810Z digest=sha256:ee2d8793bcef703f504cb18c9ddb1627f362b556c11bffa468102cfe904227f2

Observation 3e61394d-7238-44fd-9e1c-9f45e904f98e · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.918538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.918538Z digest=sha256:aa2310e15561a769986ea21e379037f1c2708003dd9b3db67bda6d93d4a08316

Observation a9217669-4064-436c-9976-7f7f8fe6acda · outbound

This paper cites M&M Mix: A Multimodal Multiview Transformer Ensemble.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets M&M Mix: A Multimodal Multiview Transformer Ensemble

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.923255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.923255Z digest=sha256:289a3db2e89f34ce58dcba335534737b272c21eb49d2dcfb1b152638deaeaf70

Observation 31ca9c65-cca4-4d16-a0fd-a23491cd5239 · outbound

This paper cites Multiview transformers for video recognition.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.376564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.928345Z digest=sha256:a922f82353d2df172558707e43a5c5d8ca49fd0f1434c0506f0a5bd7b516e163

Observation 80143f1f-e6e2-45f7-9206-914917274a3d · outbound

This paper cites Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.932616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.932616Z digest=sha256:ecba184b2d7f68b6e492f1faea47cbae8f1599c7962808c1053fdb1e5b17fa37

Observation fab0796f-1759-49e8-a8e4-9bd83f972ab1 · outbound

This paper cites an unresolved cited work.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:42.361307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:57:41.937067Z digest=sha256:02e770f6d46e4e34b916ac0df5ba1872a904d27fb1716b6b67803cc258259710

Pith citing papers

No inbound Pith citation observations are available.