Pith. sign in

Paper Citation Record · LEDGER

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2606.06249.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06249 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T02:22:03.908592Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b9a41a9-0836-4205-b67f-dc54e8dbb939 · outbound

This paper cites Attention is all you need,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:b697cac1503344fdaf6cd849e77e278974f5e509b80e7bfe88e5917155376329

Observation 03688220-1d4b-454f-be79-71d144cc6f30 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:d90f7f2cacc3088b31f53f9d0ad40e7b7f2d1331f1fbf739f305615b1e943e3f

Observation cae0a385-ff23-4dcd-8674-276b0e2170f2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Flamingo: a visual language model for few-shot learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:cdf64f82a8d13db3e8eff98bc24e4821436fa1bc32dc83cd324eecb4c463d609

Observation ef9746f6-d3bc-4e35-95bb-9400911d1ddd · outbound

This paper cites mPLUG-Owl3: Towards long image-sequence understanding in multi-modal large language models,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention mPLUG-Owl3: Towards long image-sequence understanding in multi-modal large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:c68f6b0a2e88fd9ac320ca936dfbb844d05a5b944932be7d3df553ca516f416a

Observation dd8144cc-360c-493f-82f3-ab2414b5bd6f · outbound

This paper cites an unresolved cited work.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:6d17c2685eada2772064abb778f4437af0477d20fb0ea8aeae9502f305037619

Observation 73f13c84-d371-4f92-b3c5-c08213adbe79 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal transformer for unaligned multimodal language sequences,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:538cef44993b2d4cc3c137328fffd3846e162801ab38fb1ceb6f6ec59365fb5a

Observation 2390d52d-04ee-4ef0-8654-15683e390acb · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Scaling rectified flow transformers for high-resolution image synthesis,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:d7c58d5951c4d787f99e685873bb3e1fd491c1e078d223c4fa92cd4adc80f22a

Observation 83cb7406-5622-4893-817b-f2fb0bc2730f · outbound

This paper cites TACA: Rethinking cross-modal interaction in multimodal diffusion transformers,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention TACA: Rethinking cross-modal interaction in multimodal diffusion transformers,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:4fae8188c227447dafcc20c540c3bbd5681fb63185058cf15638ac3c87b56886

Observation 1d74a2d7-d699-44d9-8a59-d325be4c69cd · outbound

This paper cites Videobert: A joint model for video and language representation learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Videobert: A joint model for video and language representation learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:a0bb3c955fa9ebe308259a8e84879ac0b6447f76b56df6ae1bf59656f594e703

Observation f6614d7f-bf8e-4446-aeb1-00290c02c046 · outbound

This paper cites Unveiling the power of audio-visual early fusion transformers with dense interactions through masked modeling,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Unveiling the power of audio-visual early fusion transformers with dense interactions through masked modeling,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:9e840011ea4896a319f12a3445e930a81bc998684168328e86f7399d50458fa9

Observation 580c99a0-9cec-4b12-98ee-b0200493e02b · outbound

This paper cites Language is not all you need: Aligning perception with language models,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Language is not all you need: Aligning perception with language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:bd787c3c2ea402854d58cfb8090a8cee3d359886dbbcc8bb6370653422f922af

Observation 04c8752d-ea5c-43b3-82cd-390b4a45e3fc · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:0a0f42065092d9a1f6bcdc25f3ccfb030852850cc26e4fbb73dced76a0a03470

Observation e5fea75d-0e50-4926-b31d-33b9fd434dd7 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:badb598039853b7bd1c12bd86671fa505261a5f6599914c0ae9f4cf155829bc9

Observation c4deb054-34da-4569-83b5-4b7ab83153df · outbound

This paper cites Cross-modal gated feature enhancement for multimodal emotion recognition in conversations,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Cross-modal gated feature enhancement for multimodal emotion recognition in conversations,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:639a7a3f74d8f667ac418b31698b50ad25a64fe1b6c8dce022ec3aca557d50f1

Observation 910210a4-58ef-47b1-8794-4a6c063d0ac4 · outbound

This paper cites Triplet attention: Rethinking the similarity in transformers,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Triplet attention: Rethinking the similarity in transformers,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:9ef83f1a4b7105b135700c0051aca41524e82e9b2ac02a621ce49eb1171dfbc8

Observation 063a8e05-95f8-403a-a9e6-8178634ac98b · outbound

This paper cites MMT: Multi-way multi-modal transformer for multimodal learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MMT: Multi-way multi-modal transformer for multimodal learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:055ad7c3282ddefa50fd460617ff97553f52698c72aca56b7e13a107f2c03ae5

Observation 638a7d16-465c-4543-8acf-f9745b22c83c · outbound

This paper cites Triplet attention transformer for spatiotemporal predictive learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Triplet attention transformer for spatiotemporal predictive learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:03819f7a8fbedf7969da2606dd0bfc8c6d8661a3af9f2f30cc51fbea082adbfa

Observation 8cb5f894-953e-48b7-8a97-c41b3b9971c3 · outbound

This paper cites Gramian multimodal representation learning and alignment,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Gramian multimodal representation learning and alignment,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:feace7b2217b650d67d28250663e44051b6a317000b45ff1a0cc6f2424364d08

Observation a9bcb709-4416-4fd0-ba6d-fea1b758b80e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Learning transferable visual models from natural language supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:92eccad2e064f1ab3b542ef2149667b44283822056555822609eff3498107504

Observation 7f45dfb5-4203-447d-9706-46dc4938c6f1 · outbound

This paper cites CLAP: learning audio concepts from natural language supervision,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention CLAP: learning audio concepts from natural language supervision,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:46787bfb72d0a18f7046ae16329aa8ddd6964ff94482a85fdb5ee4a3bd83d18b

Observation c1ebe48f-4b03-4f94-b42b-b70c7557fb33 · outbound

This paper cites CLIP4Clip: An empirical study of clip for end to end video clip retrieval,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention CLIP4Clip: An empirical study of clip for end to end video clip retrieval,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:3bc9044455c8299647b35a997d546301e7e0b33468dda72dd23366bd26ce355c

Observation 2918a388-03ab-462a-bd6d-660b9a1ed674 · outbound

This paper cites ImageBind one embedding space to bind them all,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention ImageBind one embedding space to bind them all,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:8bc5d0885e7806bef08583a0226d92193a2ae9f21bc5cc36bfcb55aa57e78c8d

Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.379175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:ba9d9071ef00deb5e6c7d5a49578bb07602c3a5e19f2d753f4e54c73ea9d4f94

Observation 0ccdc2c4-e73d-4cc0-818e-a976da149cd0 · outbound

This paper cites Flowing from words to pixels: A noise-free framework for cross-modality evolution,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Flowing from words to pixels: A noise-free framework for cross-modality evolution,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:7967b2404b4878787133f7a2d4d8d4d866bc0641f7cde87bb1466dd765bcde2c

Observation edb9787c-3a08-4d06-997f-9f17fdf66d5b · outbound

This paper cites A triangle enables multimodal alignment beyond cosine similarity,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention A triangle enables multimodal alignment beyond cosine similarity,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:ce53778e28154b82a66bcc8912e8599b336d94b49372e81c8611b99c0771d4fc

Observation fa27101e-a849-4530-b102-b047e14380c1 · outbound

This paper cites Contrasting with symile: Simple model-agnostic representation learning for unlimited modalities,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Contrasting with symile: Simple model-agnostic representation learning for unlimited modalities,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:ca4f643d97c3069d4c181de25a30933d5a4dd078bdd8bfae6cd585ce169371ea

Observation 7701ea74-9db6-4215-9f74-17232d2291a5 · outbound

This paper cites Principled multimodal representation learning.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Principled multimodal representation learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.380593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:284147764a74c91208b9b3a98daabf6fbf6f54d926f299d5e266b372256fbae4

Observation d1bb5936-1a6f-41d2-8381-898fe2dbdf85 · outbound

This paper cites Quadruple attention in many-body systems for accurate molecular property predictions,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Quadruple attention in many-body systems for accurate molecular property predictions,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:15832d058bb892576fb45882ede6c668acf7d5e1cb511daeac0f64f1b44e825f

Observation 36711569-1dd1-4d0b-a6af-cbc6ecb06317 · outbound

This paper cites Matrix theory,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Matrix theory,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:521c33a9042715a9f2683fa67ff98eb6e7eb0e68f78ff88c2ec2ebd1f1f88b74

Observation 4b184902-bcf0-48c5-980f-f924bdf5803b · outbound

This paper cites M-SENA: An integrated platform for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention M-SENA: An integrated platform for multimodal sentiment analysis,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:82c94ca4760b2a093c0887c662f711bd2e457d3503aa33755f1996663c7ba0f7

Observation 37a79ce9-3580-4980-92f8-2a720fa67469 · outbound

This paper cites Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:3f32cb415a3d8fd1a91bbd4f4cd4cfa57fda712cb0c8e6ec2d53b675e4e01d8a

Observation 3c355997-5adf-4ee1-b462-0d8b52d58547 · outbound

This paper cites LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:988ec44463dfdfe086dd823e3cb6ba88c303ba6a8f9275bff62c45aa0161b79f

Observation 58d5e2cf-fbd0-4ced-833d-e00ca698ed4c · outbound

This paper cites What to align in multimodal contrastive learning?,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention What to align in multimodal contrastive learning?,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:f484a21db2e5b1dea279aa477d603fa11231a2b3bb211a5ced30c64f87f6ee01

Observation 90e39488-69d7-4225-9858-9dd613b5a8c9 · outbound

This paper cites Multimodal phased transformer for sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal phased transformer for sentiment analysis,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:9cc3e82b55adc00b651864acfc34e3e5434b7312fba13d3ab9f410292cca26fd

Observation 09ca85b9-217f-4e36-924a-c89dc7e750f0 · outbound

This paper cites Joint fine-grained disentanglement and modal-agnostic fusion multi-task framework for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Joint fine-grained disentanglement and modal-agnostic fusion multi-task framework for multimodal sentiment analysis,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:ed4310eea9274b29866d1df7fd9738ecd7e2fdcc0f59b465720e5e929ba55c1d

Observation 22878b57-d92f-46ae-9a44-137ddc62e175 · outbound

This paper cites MultiBench: Multiscale benchmarks for multimodal representation learning,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MultiBench: Multiscale benchmarks for multimodal representation learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:224a06bf1cf62c78df5ab1b1beb0d8d092b7cb898ec5c0731a565cf9d91d938f

Observation 200fc8ee-eef4-4e2d-bdd3-b050a8985449 · outbound

This paper cites Tensor fusion network for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Tensor fusion network for multimodal sentiment analysis,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:ae44520695e381f09759b17bbde9692aa13e8c8f114b0ca0e60753eeeb2d6726

Observation e296aa89-b7a0-457d-92ef-15e04b6b0182 · outbound

This paper cites Efficient low-rank multimodal fusion with modality-specific factors,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Efficient low-rank multimodal fusion with modality-specific factors,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:1f58ccca5a00d17b7937d322222ee32d8a6f220b90daa44245acef552688600e

Observation 82a988a3-8f6d-4230-871e-fa7ef9e28b1e · outbound

This paper cites MISA: Modality-invariant and -specific representations for multi- modal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MISA: Modality-invariant and -specific representations for multi- modal sentiment analysis,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:e6fe746643909766f32ec0f8c810492164f4bd1613189e40c19a24dec56d9238

Observation 4916ba40-bf98-4a59-bc91-7fbe239cc37a · outbound

This paper cites Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:07b3455d21b855ec75c7bae116dffee446bec37ae77c64d246fd65631bf02a14

Observation a105655c-2091-40cc-8050-337c953bca97 · outbound

This paper cites Integrating multimodal information in large pretrained transformers,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Integrating multimodal information in large pretrained transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:f6fdede70f24f31f4e98ab35d431839c9c115a84e07adc5f927ca09b877b9493

Observation 5a9deffe-7ea9-44c1-9fdc-467333b4cf0c · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal multi-loss fusion network for sentiment analysis,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:d854b9f0aac16bf1648fa321ee0156daf47dff649d0227b31cf029a5b89f33b2

Observation 31b21a35-39f2-4f90-8438-3dcafe1ef15f · outbound

This paper cites Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content,.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T02:22:03.908592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:d767c761362df0c291b4f8b5bb3a78790fd4d6893ec4acefd43da66cff08e754

Pith citing papers

No inbound Pith citation observations are available.