Pith. sign in

Paper Citation Record · LEDGER

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2606.06249.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06249 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T02:22:03.908592Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b9a41a9-0836-4205-b67f-dc54e8dbb939 · outbound

This paper cites Attention is all you need,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:56bb9b3263faf1eefdded9e63756b978e97279877b85cf11433fc9c88d8505dc

Observation 03688220-1d4b-454f-be79-71d144cc6f30 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:e663454571c17d4e91fa5d7692cb455fe5f44b5c30c0fa3b36cce1ad80db1460

Observation cae0a385-ff23-4dcd-8674-276b0e2170f2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Flamingo: a visual language model for few-shot learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:537e52eb27ec6f220400eb75b696d58919ce869ea166acf657b963b5c8a7f1f9

Observation ef9746f6-d3bc-4e35-95bb-9400911d1ddd · outbound

This paper cites mPLUG-Owl3: Towards long image-sequence understanding in multi-modal large language models,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention mPLUG-Owl3: Towards long image-sequence understanding in multi-modal large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:e9ae80a2651e445350366c1ab934e53385dae2da776d6ca37cf4b8319402c08e

Observation dd8144cc-360c-493f-82f3-ab2414b5bd6f · outbound

This paper cites an unresolved cited work.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:ad6cebda9b6a76064669bdfb174e872a84231ee8bc6c8346b7cbf627fef12c23

Observation 73f13c84-d371-4f92-b3c5-c08213adbe79 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal transformer for unaligned multimodal language sequences,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:e94f22baeec7699a23512383479edb9b160d3b253aef9aefbef53f3575444108

Observation 2390d52d-04ee-4ef0-8654-15683e390acb · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Scaling rectified flow transformers for high-resolution image synthesis,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:97a1451907ece163bff34fe22a54da3072d24dcfe4cfa416b8d574fc5363f1fa

Observation 83cb7406-5622-4893-817b-f2fb0bc2730f · outbound

This paper cites TACA: Rethinking cross-modal interaction in multimodal diffusion transformers,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention TACA: Rethinking cross-modal interaction in multimodal diffusion transformers,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:939fbeffb16ed46f61888e6a40029fe4c2f1cf52d154b5c6441e68b46ea726f2

Observation 1d74a2d7-d699-44d9-8a59-d325be4c69cd · outbound

This paper cites Videobert: A joint model for video and language representation learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Videobert: A joint model for video and language representation learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:987ed616880fc3785566d5857db72ebcdb6140ecdf6597b24875d4c3d9e9dc47

Observation f6614d7f-bf8e-4446-aeb1-00290c02c046 · outbound

This paper cites Unveiling the power of audio-visual early fusion transformers with dense interactions through masked modeling,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Unveiling the power of audio-visual early fusion transformers with dense interactions through masked modeling,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:6f62bed7b98f4dee9f21756123b3bae6b27a3f91991c90ecfcccec724929d483

Observation 580c99a0-9cec-4b12-98ee-b0200493e02b · outbound

This paper cites Language is not all you need: Aligning perception with language models,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Language is not all you need: Aligning perception with language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:c3277413be7570c7ac0520cf5f2e09f15af35c4e272dd0fe772bd62201b9ccbc

Observation 04c8752d-ea5c-43b3-82cd-390b4a45e3fc · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:5fcce1b748bda240b0ea258967f9025279c99fc56485084a8ff3a1b28cd1d6d2

Observation e5fea75d-0e50-4926-b31d-33b9fd434dd7 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:0af02c749dfdb2172f56c905ff31b4b2ef297be84dde344bd5ad46d51ae0ec04

Observation c4deb054-34da-4569-83b5-4b7ab83153df · outbound

This paper cites Cross-modal gated feature enhancement for multimodal emotion recognition in conversations,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Cross-modal gated feature enhancement for multimodal emotion recognition in conversations,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:f835226f496d4e3306f8a680ac7540f4bdc16256f1ac4e8562845ba70ee7ceab

Observation 910210a4-58ef-47b1-8794-4a6c063d0ac4 · outbound

This paper cites Triplet attention: Rethinking the similarity in transformers,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Triplet attention: Rethinking the similarity in transformers,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:65d4c1d0fae88fc29d9ba457c399a936c6e8de424b72239de2b1d724a479c724

Observation 063a8e05-95f8-403a-a9e6-8178634ac98b · outbound

This paper cites MMT: Multi-way multi-modal transformer for multimodal learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MMT: Multi-way multi-modal transformer for multimodal learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:1c497dc6fdb1fc4ce055dcce2ccf83cb43958e881b95eda1267d2503cd9b2d31

Observation 638a7d16-465c-4543-8acf-f9745b22c83c · outbound

This paper cites Triplet attention transformer for spatiotemporal predictive learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Triplet attention transformer for spatiotemporal predictive learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:4690fb043f7d0c70b9553e88f658ed85496a271ce03d33ca3cc07ed6ce54c223

Observation 8cb5f894-953e-48b7-8a97-c41b3b9971c3 · outbound

This paper cites Gramian multimodal representation learning and alignment,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Gramian multimodal representation learning and alignment,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:a78e0bf482b6c081158732d5866a3fdc8df5145a60b9595a576c2932095f1057

Observation a9bcb709-4416-4fd0-ba6d-fea1b758b80e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Learning transferable visual models from natural language supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:55dbb304a002c0e46d7fe0f834c46ea7ea652ae73c9f4f080c182dd2f12a1dce

Observation 7f45dfb5-4203-447d-9706-46dc4938c6f1 · outbound

This paper cites CLAP: learning audio concepts from natural language supervision,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention CLAP: learning audio concepts from natural language supervision,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:75ccc1e2e5754a0e43c550d227980be3891dd49ef28e79bbd608b180f447e03a

Observation c1ebe48f-4b03-4f94-b42b-b70c7557fb33 · outbound

This paper cites CLIP4Clip: An empirical study of clip for end to end video clip retrieval,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention CLIP4Clip: An empirical study of clip for end to end video clip retrieval,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:c13612643faf671e48ea73ad279d0343dc01095161c12347715273a9e246d502

Observation 2918a388-03ab-462a-bd6d-660b9a1ed674 · outbound

This paper cites ImageBind one embedding space to bind them all,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention ImageBind one embedding space to bind them all,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:86e5ed3d5a898e9261be234e115c28b85fdd217e6301b744ec15100fe5a11dd5

Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.379175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:a75168ede8973b1acbb661a7fe9e414d0e6e6b95c89bbf1976e2b9c5d93258bd

Observation 0ccdc2c4-e73d-4cc0-818e-a976da149cd0 · outbound

This paper cites Flowing from words to pixels: A noise-free framework for cross-modality evolution,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Flowing from words to pixels: A noise-free framework for cross-modality evolution,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:d7568a18bc2ad8c0b7406b3ad4021e04b8bd7a418f27a16039daeba879897be0

Observation edb9787c-3a08-4d06-997f-9f17fdf66d5b · outbound

This paper cites A triangle enables multimodal alignment beyond cosine similarity,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention A triangle enables multimodal alignment beyond cosine similarity,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:c1e0e7bcbb0ec9ff654caee403c5b8ccedf36120ef331211d549da60814213c4

Observation fa27101e-a849-4530-b102-b047e14380c1 · outbound

This paper cites Contrasting with symile: Simple model-agnostic representation learning for unlimited modalities,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Contrasting with symile: Simple model-agnostic representation learning for unlimited modalities,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:992401a044133e15fc8d572e4f6b5ee0f1922879fd2c81a0f0f1ce83af25c9fd

Observation 7701ea74-9db6-4215-9f74-17232d2291a5 · outbound

This paper cites Principled multimodal representation learning.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Principled multimodal representation learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.380593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:42673069bf4cb9507d9a3bd286c8e281d93839d1fd9a5e15ea230e838fdee486

Observation d1bb5936-1a6f-41d2-8381-898fe2dbdf85 · outbound

This paper cites Quadruple attention in many-body systems for accurate molecular property predictions,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Quadruple attention in many-body systems for accurate molecular property predictions,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:60b7af91e331dcd1baa446c00174d83982bb3d8611b43979b9be09d851814d69

Observation 36711569-1dd1-4d0b-a6af-cbc6ecb06317 · outbound

This paper cites Matrix theory,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Matrix theory,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:f9ddd6a1950331e08e1da8e3718f548029cade1b49f8aef7d49bc5c8f7af3aa2

Observation 4b184902-bcf0-48c5-980f-f924bdf5803b · outbound

This paper cites M-SENA: An integrated platform for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention M-SENA: An integrated platform for multimodal sentiment analysis,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:67258f82192c92105e40e059c5f346257373d787f234d67642d96178b109532c

Observation 37a79ce9-3580-4980-92f8-2a720fa67469 · outbound

This paper cites Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:1baa5352f9311c0a67cfa4978856a4d3b72355fba8be872d512ddc3473dd7e07

Observation 3c355997-5adf-4ee1-b462-0d8b52d58547 · outbound

This paper cites LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:671abb703c35d98956a0a8f4a0d5c304c3737b0c731d23de71910dbc62ac7045

Observation 58d5e2cf-fbd0-4ced-833d-e00ca698ed4c · outbound

This paper cites What to align in multimodal contrastive learning?,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention What to align in multimodal contrastive learning?,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:23f7181e985f49fc782a5c7d2f9b11c0f7ea93d7d971fe6f4e522a873b165d51

Observation 90e39488-69d7-4225-9858-9dd613b5a8c9 · outbound

This paper cites Multimodal phased transformer for sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal phased transformer for sentiment analysis,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:e4974cf17c01427c4e16452de27db5a2f9f9c316fb192a6b940a0b5df040abe9

Observation 09ca85b9-217f-4e36-924a-c89dc7e750f0 · outbound

This paper cites Joint fine-grained disentanglement and modal-agnostic fusion multi-task framework for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Joint fine-grained disentanglement and modal-agnostic fusion multi-task framework for multimodal sentiment analysis,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:2fb9d788a70117d80961611e158452c2c08a134540032fb38336aa028c9aaa05

Observation 22878b57-d92f-46ae-9a44-137ddc62e175 · outbound

This paper cites MultiBench: Multiscale benchmarks for multimodal representation learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MultiBench: Multiscale benchmarks for multimodal representation learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:f2354b6540ef7fc869cb7d6d97fc78b7861656d6e6902f58694f28cde3b37c3a

Observation 200fc8ee-eef4-4e2d-bdd3-b050a8985449 · outbound

This paper cites Tensor fusion network for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Tensor fusion network for multimodal sentiment analysis,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:bf490aa1d009b0db8ef3625284219ed4913fd28e2cdb68de7bd5f97af1250964

Observation e296aa89-b7a0-457d-92ef-15e04b6b0182 · outbound

This paper cites Efficient low-rank multimodal fusion with modality-specific factors,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Efficient low-rank multimodal fusion with modality-specific factors,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:8c6c885f4f02e6a212cee542fb4fb03c78c4de5b7dee8ad17d0110cdd1de6d12

Observation 82a988a3-8f6d-4230-871e-fa7ef9e28b1e · outbound

This paper cites MISA: Modality-invariant and -specific representations for multi- modal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MISA: Modality-invariant and -specific representations for multi- modal sentiment analysis,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:7ceae9d5a5666902b6f6ed5410cef92eb8441a99bb10e3d2ecb43bae7ebfc8b6

Observation 4916ba40-bf98-4a59-bc91-7fbe239cc37a · outbound

This paper cites Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:afdcab5634825eb0d52f3eb288447503df4dfa06aadfe6cf0d9bc233b0738a13

Observation a105655c-2091-40cc-8050-337c953bca97 · outbound

This paper cites Integrating multimodal information in large pretrained transformers,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Integrating multimodal information in large pretrained transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:b582e8f752902f9c95b47cdf5dbbb147cef0938e0e61c700909cadf9d48248a0

Observation 5a9deffe-7ea9-44c1-9fdc-467333b4cf0c · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal multi-loss fusion network for sentiment analysis,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:1c6dc3f3d1b0e31fa84368984fa79912920994ba64a648035e943e32e433ee12

Observation 31b21a35-39f2-4f90-8438-3dcafe1ef15f · outbound

This paper cites Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:d5f64543759210dcde20fd605ef511dce423f6ec996656c7d5a3664f713c0331

Pith citing papers

No inbound Pith citation observations are available.