Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T02:22:03.908592Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2606.06249.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T02:22:03.908592Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1b9a41a9-0836-4205-b67f-dc54e8dbb939 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Attention is all you need,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03688220-1d4b-454f-be79-71d144cc6f30 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae0a385-ff23-4dcd-8674-276b0e2170f2 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Flamingo: a visual language model for few-shot learning,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef9746f6-d3bc-4e35-95bb-9400911d1ddd · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention mPLUG-Owl3: Towards long image-sequence understanding in multi-modal large language models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8144cc-360c-493f-82f3-ab2414b5bd6f · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f13c84-d371-4f92-b3c5-c08213adbe79 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal transformer for unaligned multimodal language sequences,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2390d52d-04ee-4ef0-8654-15683e390acb · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Scaling rectified flow transformers for high-resolution image synthesis,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cb7406-5622-4893-817b-f2fb0bc2730f · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention TACA: Rethinking cross-modal interaction in multimodal diffusion transformers,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d74a2d7-d699-44d9-8a59-d325be4c69cd · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Videobert: A joint model for video and language representation learning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6614d7f-bf8e-4446-aeb1-00290c02c046 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Unveiling the power of audio-visual early fusion transformers with dense interactions through masked modeling,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580c99a0-9cec-4b12-98ee-b0200493e02b · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Language is not all you need: Aligning perception with language models,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c8752d-ea5c-43b3-82cd-390b4a45e3fc · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5fea75d-0e50-4926-b31d-33b9fd434dd7 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4deb054-34da-4569-83b5-4b7ab83153df · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Cross-modal gated feature enhancement for multimodal emotion recognition in conversations,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910210a4-58ef-47b1-8794-4a6c063d0ac4 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Triplet attention: Rethinking the similarity in transformers,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 063a8e05-95f8-403a-a9e6-8178634ac98b · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MMT: Multi-way multi-modal transformer for multimodal learning,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638a7d16-465c-4543-8acf-f9745b22c83c · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Triplet attention transformer for spatiotemporal predictive learning,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb5f894-953e-48b7-8a97-c41b3b9971c3 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Gramian multimodal representation learning and alignment,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9bcb709-4416-4fd0-ba6d-fea1b758b80e · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Learning transferable visual models from natural language supervision,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f45dfb5-4203-447d-9706-46dc4938c6f1 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention CLAP: learning audio concepts from natural language supervision,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ebe48f-4b03-4f94-b42b-b70c7557fb33 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention CLIP4Clip: An empirical study of clip for end to end video clip retrieval,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2918a388-03ab-462a-bd6d-660b9a1ed674 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention ImageBind one embedding space to bind them all,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ccdc2c4-e73d-4cc0-818e-a976da149cd0 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Flowing from words to pixels: A noise-free framework for cross-modality evolution,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb9787c-3a08-4d06-997f-9f17fdf66d5b · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention A triangle enables multimodal alignment beyond cosine similarity,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa27101e-a849-4530-b102-b047e14380c1 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Contrasting with symile: Simple model-agnostic representation learning for unlimited modalities,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7701ea74-9db6-4215-9f74-17232d2291a5 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Principled multimodal representation learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1bb5936-1a6f-41d2-8381-898fe2dbdf85 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Quadruple attention in many-body systems for accurate molecular property predictions,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36711569-1dd1-4d0b-a6af-cbc6ecb06317 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Matrix theory,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b184902-bcf0-48c5-980f-f924bdf5803b · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention M-SENA: An integrated platform for multimodal sentiment analysis,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a79ce9-3580-4980-92f8-2a720fa67469 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c355997-5adf-4ee1-b462-0d8b52d58547 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58d5e2cf-fbd0-4ced-833d-e00ca698ed4c · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention What to align in multimodal contrastive learning?,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e39488-69d7-4225-9858-9dd613b5a8c9 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal phased transformer for sentiment analysis,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ca85b9-217f-4e36-924a-c89dc7e750f0 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Joint fine-grained disentanglement and modal-agnostic fusion multi-task framework for multimodal sentiment analysis,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22878b57-d92f-46ae-9a44-137ddc62e175 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MultiBench: Multiscale benchmarks for multimodal representation learning,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 200fc8ee-eef4-4e2d-bdd3-b050a8985449 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Tensor fusion network for multimodal sentiment analysis,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e296aa89-b7a0-457d-92ef-15e04b6b0182 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Efficient low-rank multimodal fusion with modality-specific factors,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a988a3-8f6d-4230-871e-fa7ef9e28b1e · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention MISA: Modality-invariant and -specific representations for multi- modal sentiment analysis,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4916ba40-bf98-4a59-bc91-7fbe239cc37a · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a105655c-2091-40cc-8050-337c953bca97 · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Integrating multimodal information in large pretrained transformers,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9deffe-7ea9-44c1-9fdc-467333b4cf0c · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Multimodal multi-loss fusion network for sentiment analysis,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b21a35-39f2-4f90-8438-3dcafe1ef15f · outbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.