Pith. sign in

Paper Citation Record · LEDGER

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.11465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11465 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:09:43.057849Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50dfc4cd-ded7-4194-bd00-11fd38177eae · outbound

This paper cites write newline.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.068767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.068767Z digest=sha256:0ead75033ec575c05dbc6e15e4b437d0c41843c338695bb40ab54488fbfe71bc

Observation d76e886a-75f2-4af7-a52a-86e4cffbefcb · outbound

This paper cites and Zisserman, A.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer and Zisserman, A

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.559290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.226847Z digest=sha256:b421c6b6dc7bea4a9156916313fe7274186a24d87b366815f9ce778d661989cf

Observation 09d7dbc0-1d54-49b5-a6f2-ecb3c3651dfa · outbound

This paper cites Singular value decomposition tutorial.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Singular value decomposition tutorial

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.544180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.350243Z digest=sha256:5c012a2a9a369440dad48585ad4680ef5d8a31a6e51f4dedc197d559102a407b

Observation 82bf73e3-e43b-46b8-8094-7ab82b49670d · outbound

This paper cites Multimodal machine learning: A survey and taxonomy.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal machine learning: A survey and taxonomy

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.530010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.427179Z digest=sha256:8dde3fc97fdede0472ae80f845aff36c6727d7354e596e959a8029ee486ed4ef

Observation 2cafc4cc-ed47-42e3-880b-9fcd5e97f458 · outbound

This paper cites Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.515557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.498054Z digest=sha256:5efb5ff8f3af4132edd18ec65967c6c89ceb95cab75f512686e646a555736b4d

Observation ec1d95fd-8ee2-48a8-accb-0c867feded1f · outbound

This paper cites G., Keutmann, M.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer G., Keutmann, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.500337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.594262Z digest=sha256:bc01bef055450892675b63c423ba72e712de0d65a3161f1023d3e6e1069151cf

Observation e20a43c3-110e-47fb-afea-7212f0bc8b75 · outbound

This paper cites End-to-end object detection with transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer End-to-end object detection with transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.679875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.679875Z digest=sha256:8605d26efafa839cb5d5a68ad69f26a9d31edca72630d0c382883b0c6a46d702

Observation 7885befd-1227-4d1b-8bef-696afd9da023 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Vggsound: A large-scale audio-visual dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.477032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.753191Z digest=sha256:5a1c58df90f0696cbb2fa9e601f84c48535d88e001b8f7750725bfa13022eadd

Observation e4c3c1a8-38f2-4f74-8bdd-d088994c2cca · outbound

This paper cites T., Rubanova, Y., Bettencourt, J., and Duvenaud, D.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer T., Rubanova, Y., Bettencourt, J., and Duvenaud, D

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.823412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.823412Z digest=sha256:89f9ce6d18cae4f0702fb192e6fde60dbb2a3e522921837be0a60f8f0043cc4b

Observation 4786108c-a60a-452b-8d9a-2da6020d58dc · outbound

This paper cites Self-attention fusion for audiovisual emotion recognition with incomplete data.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Self-attention fusion for audiovisual emotion recognition with incomplete data

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.453909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.893997Z digest=sha256:7d2330db2aa4f8dd3d22397f2647efe900c6395dac28ae463329d91656162aa3

Observation cf10ac09-1ebc-4fd9-a007-290cdf6daac3 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer What Does BERT Look At? An Analysis of BERT's Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.980072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.980072Z digest=sha256:1220541efac3929551c2e2bd496fa2bf6acdd15199b21a313a62b89024627792

Observation 0551a8ae-6918-4c74-8dfb-8f47cf9b286c · outbound

This paper cites Addressing failure prediction by learning model confidence.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Addressing failure prediction by learning model confidence

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.438443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.064181Z digest=sha256:c05f72b83f2eb93bad69ba87fdf3b51dbda5d29ac951a81cde8cf2d233f3b542

Observation 55f3978c-c448-4684-ae99-900215ee25cf · outbound

This paper cites MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:09:43.650849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.130768Z digest=sha256:d24e2c8dfa7b3780d2da1574ce8625f8442539c7cf513cde201793d324f69e5a

Observation b7bb59bc-732c-44d1-b58a-d335e21ac48a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.218605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.218605Z digest=sha256:429b295d38e7b6b5d7d8ca6a833b49d820c5c68096504ca2158ac941aa06a6f1

Observation f19b3643-e38e-41ca-813b-463a025a7aaf · outbound

This paper cites Pmr: Prototypical modal rebalance for multimodal learning.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Pmr: Prototypical modal rebalance for multimodal learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.422317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.294459Z digest=sha256:2e6d7a8b1f60d5ede7a4c9a8a3b08037a6a56aa50d14d55680510171c5057dbd

Observation 545a75d6-ae72-469f-a773-b66ee6fd6e5f · outbound

This paper cites A survey on deep learning for multimodal data fusion.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer A survey on deep learning for multimodal data fusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.341220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.356065Z digest=sha256:53e4e38167198d41c2442ff59cd02410be13834646c44c7d1c593cf232a396dc

Observation a4fd3df9-1b5d-4eb6-8613-e142ab914055 · outbound

This paper cites C., Wang, X., and Li, H.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer C., Wang, X., and Li, H

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.199632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.433527Z digest=sha256:66244698453dd9216beca596da8c95d40e4950cb45171943186de5c679f6ebcf

Observation 8820c915-2613-45bc-ac88-5c32f18a1484 · outbound

This paper cites A mathematical perspective on Transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer A mathematical perspective on Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.501652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.501652Z digest=sha256:5e15e5fc267083d548d786257696b16391a447a817e46bd3bba67e55a03146a2

Observation 03b11aec-86b4-4fd2-9e88-275a6673e170 · outbound

This paper cites Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.139355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.571302Z digest=sha256:8b45555b8e37c46c0cba3566c754bca4b81d2e9b9cf0fbf3e09c8e6aeef0c071

Observation 0074a38c-b600-4ccb-8213-a9b30af6ad1f · outbound

This paper cites Deep residual learning for image recognition.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.650765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.650765Z digest=sha256:9650ca6bee031489984cd620f4ca423fe38733abb8c84266520c84aa05ca4cdc

Observation ea0983aa-2509-4c86-9810-d1956fd73305 · outbound

This paper cites Masked autoencoders are scalable vision learners.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Masked autoencoders are scalable vision learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.718540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.718540Z digest=sha256:0c07814c64adfb2cef0a423918e77ef37622025be499231121a38ebdd80876f8

Observation 6b0631c4-cabe-4c4a-baa7-ce27b0913306 · outbound

This paper cites ReconBoost: Boosting Can Achieve Modality Reconcilement.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer ReconBoost: Boosting Can Achieve Modality Reconcilement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.795317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.795317Z digest=sha256:44c62f888c29fee553a75e32e6e3cefd09ac4eef6a9b8bc93d3911ebebfea8cf

Observation 7aaf9aba-4f25-42eb-8a9f-d8ec915f37d2 · outbound

This paper cites Adaptive unimodal regulation for balanced multimodal information acquisition.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Adaptive unimodal regulation for balanced multimodal information acquisition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.090763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.862671Z digest=sha256:22f059a3954deeb2ee5b63e38d1a6a29518a949121719ed178fa64cef63dcc0e

Observation e306a950-b2aa-40f3-8948-a3b310b95012 · outbound

This paper cites Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably).

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.016974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.922194Z digest=sha256:b23aeebde28624a31dd60e9b817140170fda6bcbd49bf4d35cd5e6461fffbebc

Observation 23ea6d0e-235a-495f-bd8c-67dcabef7be7 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.997161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.997161Z digest=sha256:236ff7765215cbcea31c495e0868361345533be6ace01972ce50d95618c71283

Observation 25da82be-d28c-4481-bd7e-c6c802b68b09 · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:48.867236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.072016Z digest=sha256:ed97868301fd71f45d8e44826ca7c8864b68b0288376001c4a3691d4260d7dcf

Observation 4320cb08-331b-4022-b18d-cbd218295782 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Vilt: Vision-and-language transformer without convolution or region supervision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.736558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.154442Z digest=sha256:7f4a333e6e0b91882929b283a1026ff13a50c1054f41dca7ef14e41ed7e197db

Observation dd8abe0d-cd06-417d-a8f4-04b7f5f7f8d8 · outbound

This paper cites Revealing the Dark Secrets of BERT.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Revealing the Dark Secrets of BERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.244763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.244763Z digest=sha256:641fcb6c1728e72745dff5a3a7bdafc3cfbed8fb6e5cd6768cb867b848900838

Observation e4c048b0-6165-4175-9943-dff252466a8f · outbound

This paper cites Hmdb: a large video database for human motion recognition.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Hmdb: a large video database for human motion recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.699781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.308098Z digest=sha256:05cff3753f5efadc6b67d39f79acd5444aec1b2749cec098e94f5b8dba6b1407

Observation d5a6dd0e-544b-4395-a0b8-db9e4b7cb64e · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.379262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.379262Z digest=sha256:d665ffc98caa548f4267630eee146f8a95f1f83ad2305ae70eceb92a785eedb7

Observation 9ff55cff-2ff8-4236-b7ff-edc47bc4d7aa · outbound

This paper cites P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.581228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.437438Z digest=sha256:c7e575b7016acc9c3c4570b5c9852df1e151225501bf682dfe2b601541a8b505

Observation a18a3558-b526-4404-a362-657ea7b69717 · outbound

This paper cites Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.521348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.521348Z digest=sha256:a6ad5aa99161e7203eeedebc185f3f2b5fc8edad98c36623d4becfaa843c0a54

Observation d7a3ff90-7b6b-4e7e-b8a0-5b3b1f6a6866 · outbound

This paper cites Efficient Low-rank Multimodal Fusion with Modality-Specific Factors.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Efficient Low-rank Multimodal Fusion with Modality-Specific Factors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.621399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.621399Z digest=sha256:abf83f5206423f8323f1d1e96da7ccf4190fe7fc0b561da4c3f11e60827505f7

Observation e0e52db5-9606-435e-a9c9-4036ca3f6cdb · outbound

This paper cites Attention bottlenecks for multimodal fusion.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Attention bottlenecks for multimodal fusion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.420532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.697091Z digest=sha256:cc98ba1f7c2b96120b5a94ba953db4029532893afd5016f56bf691ce7f7b1e9c

Observation 239cea9e-0360-41c7-bea2-1287ac56cd59 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Balanced multimodal learning via on-the-fly gradient modulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.259625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.821931Z digest=sha256:d0e8a3a342007f3cf598cad1e4fec98a3530b99267487bdbd730b0b177e0949a

Observation 547cea5c-7798-47df-9452-779a9b14c033 · outbound

This paper cites Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.892599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.892599Z digest=sha256:22b4d467b2c5cde08a20e19d962437986fede5e6c4d241d4058cf5295b785f75

Observation 16ce0f8b-0daf-4bc3-967d-d938b0dc3d2b · outbound

This paper cites Zorro: the masked multimodal transformer.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Zorro: the masked multimodal transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.977183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.977183Z digest=sha256:cfb948915ab1dede836d7a60b46775df31db1d24029687e834ec979b5a2d7fb2

Observation d46087ef-f70c-4d11-83e1-910194c6e1e2 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer ImageNet-21K Pretraining for the Masses

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.086921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.086921Z digest=sha256:3b0d5d0dd2269ef78885b679e58d08b8a599a0d68c544f42da94f38347a2399d

Observation ea171b79-85fd-47e4-ae91-6a6c7b111834 · outbound

This paper cites and Monro, S.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer and Monro, S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.138516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.216679Z digest=sha256:319bf470a3d3ec70a733b7f18d5e4e68ee0afc2eef9a464edc15a66ccf2c74a8

Observation ebcad2ad-3b0e-4e4e-9aba-961f1d4e01b2 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.333805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.333805Z digest=sha256:3166ce74436f9fe25ea0f9c286bb19a586967b3318d84878e8427014e1a49aaf

Observation 5ffad669-1854-4da8-9703-6fcf17fd33b9 · outbound

This paper cites L., Tickoo, O., and Huang, J.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer L., Tickoo, O., and Huang, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.883478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.440088Z digest=sha256:116b889d2e41c1ebaf559c3edc3ba852f8c68c095938fb5044e873531c1df8c4

Observation d0858fa7-edf5-4ef4-812c-4c2c3767b01d · outbound

This paper cites H., Bai, S., Liang, P.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer H., Bai, S., Liang, P

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.606166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.552413Z digest=sha256:45839aeba53f759e9f50249b05a16db8d76e172cd62d54402f6c1319ae041ae0

Observation 2fe37389-a3c9-4df0-afb3-bb45496bb3d9 · outbound

This paper cites Attention is all you need.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Attention is all you need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.645870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.645870Z digest=sha256:2f1421cb75a7ed0391db157c6e6b755f0d5dc375d6aee973c59ff0de620f86e5

Observation ca52aa6e-4920-443d-9036-e6abdfdecfe2 · outbound

This paper cites H., Zeeshan, M.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer H., Zeeshan, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.235792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.762413Z digest=sha256:9db837803ad088163d8f28dbb866fbaea1ec2b0e82b14811debb9b4ad51b5a8e

Observation a82048ad-0376-4017-b5ad-0b4f1e841597 · outbound

This paper cites What makes training multi-modal classification networks hard? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 12695--12705, 2020 a.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer What makes training multi-modal classification networks hard? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 12695--12705, 2020 a

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.029696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.879818Z digest=sha256:92e044239a141e9dbd3863f406b3d733ca114c1d892f4f78e580b159860e7f48

Observation a99a2d56-7b1f-4a7f-a537-54692516e151 · outbound

This paper cites Deep multimodal fusion by channel exchanging.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Deep multimodal fusion by channel exchanging

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:46.828432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.996523Z digest=sha256:e4c6c17abd5ff805c5bbd1770bd62f7783b4d771d12d3a215aa7c5f14cafde8e

Observation f21ef0bb-bf4f-461c-800f-d0a1362c8203 · outbound

This paper cites Multimodal token fusion for vision transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal token fusion for vision transformers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:46.490537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.101140Z digest=sha256:ca7a125763c9932f8268f622c1e1eb3340ebb9097103d21f2e945ab239c2e6a1

Observation 2d4e5a44-063e-4e35-8056-c28ae24e966e · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:46.245419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.170844Z digest=sha256:66835c4a635a47af5f08cf38d65794138a420d4923544a9445d4619771954da5

Observation 7323a5ea-4aaa-4e4c-aa74-bd1a0b5851e2 · outbound

This paper cites Enhancing multimodal cooperation via sample-level modality valuation.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Enhancing multimodal cooperation via sample-level modality valuation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:46.014133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.227002Z digest=sha256:5b810837001226fba2ebf88bda6f1bed3aad23376255e458339f8741d6470204

Observation a1a04a73-0842-4a2e-beba-ec1c69b862ca · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:45.793727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.283576Z digest=sha256:145a25014cc24f7b3b46530af17a5c4de2ba2138412ce3dc166fcab12a0f23fe

Observation 9441ca14-8e29-4259-abb4-a2903efd00ea · outbound

This paper cites Multimodal fusion with co-attention networks for fake news detection.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal fusion with co-attention networks for fake news detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:45.527009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.335445Z digest=sha256:c51053acd6a449abeaaf331ec1d9da5cf815f38e324de9dbeb277742589b47e9

Observation 049e80b9-65b8-495a-8450-e1a6d4325ab9 · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal multi-loss fusion network for sentiment analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:45.332401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.388586Z digest=sha256:8bc9f466bd7d27af8f6e1f18905ccf5e6d9654111dd603a99c9ac5a96b699b18

Observation 2b312de6-de2a-406f-800e-4eb2aee2f3c7 · outbound

This paper cites Balanced Audiovisual Dataset for Imbalance Analysis.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Balanced Audiovisual Dataset for Imbalance Analysis

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:09:43.246686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.449291Z digest=sha256:ddc97d60a2a992d6e501505e1927050c845aedff1cb3eb1b4ade8f1abf11e592

Observation f9ef6733-d0c6-460b-9bd2-99b47c4c3490 · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:45.069587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.505015Z digest=sha256:20b0cc177d8a0ac3d7f4109ec57704b70213895185c80c991c50d9dd173c4877

Observation 66ed191c-5d3d-4682-beeb-b3e45e79ae60 · outbound

This paper cites Facilitating multimodal classification via dynamically learning modality gap.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Facilitating multimodal classification via dynamically learning modality gap

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.948613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.553697Z digest=sha256:499355433debc2782c56e1362c330f6bb5825fe353b69c5afd50a5df8e998ae6

Observation f45c8981-637e-44af-9e1a-b8e6becbbf24 · outbound

This paper cites Learning to rebalance multi-modal optimization by adaptively masking subnetworks.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Learning to rebalance multi-modal optimization by adaptively masking subnetworks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.656662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.619478Z digest=sha256:ab0fc71d2bf747f7fd501221cc0fb7d41490c0633ec9e3090e3d870e541cae2a

Observation e4aab786-902f-4950-b114-ce9358f957a4 · outbound

This paper cites Deep modular co-attention networks for visual question answering.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Deep modular co-attention networks for visual question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.419762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.689333Z digest=sha256:71f2842e9a8cba315b93b5540c9634fa109df4bfced8e313396afcccd06461f6

Observation 1b29b99a-b9c2-4b6a-87ff-895202aa7c49 · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:42.779895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:42.779895Z digest=sha256:ebe194b64e37f0a9c371a7441fbf546fb284de87e01a2e1a997d0389bf7d51d9

Observation 33865e64-506f-4578-9ab0-f0365b44c665 · outbound

This paper cites B., Liang, P.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer B., Liang, P

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.221027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.869769Z digest=sha256:20ac8ff5869fc5364509f9935ef46b7e1ddd72d8c717172498a568da1c2001e3

Observation d43883cf-fe5d-4810-887c-33cfd95e3fd6 · outbound

This paper cites T., and Peng, X.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer T., and Peng, X

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:43.959333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.959924Z digest=sha256:7a1b3dfc018bff6feca25472dce941133bebde45c5ca9de97e2f79bc8a0b8ac9

Observation c2704c55-8763-4037-aeaa-3d9c784a8eef · outbound

This paper cites Multimodal Fusion on Low-quality Data: A Comprehensive Survey.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:43.057849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:43.057849Z digest=sha256:f0d01d670a7601b95ab401bcf5bae710c55b92c18241855abb8c81c90dfeb9dc

Pith citing papers

No inbound Pith citation observations are available.