Pith. sign in

Paper Citation Record · LEDGER

Efficient Quantification of Multimodal Interaction at Sample Level

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.17248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17248 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:32.281803Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e8a7eaa-f421-42c4-8d4e-9feb6071be88 · outbound

This paper cites Baseline We adopt three primary types of multimodal learning paradigms: Feature-level fusion: Integration of multiple modalities at the feature level.

Efficient Quantification of Multimodal Interaction at Sample Level Baseline We adopt three primary types of multimodal learning paradigms: Feature-level fusion: Integration of multiple modalities at the feature level

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.409939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:52:32.278424Z digest=sha256:4d4d490220a7dfe3221adf082bacf970e401c5203eb93e9b9891584a35c06308

Observation 06b52f1a-5fa7-44e0-a9eb-3f9877db56c4 · outbound

This paper cites MultiBench: Multiscale Benchmarks for Multimodal Representation Learning.

Efficient Quantification of Multimodal Interaction at Sample Level MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.252230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.252230Z digest=sha256:88519798d3dd825288bfb28a337c2ea190a127840f27e11a3460d684d8dfd913

Observation d08a618e-4352-4159-a23b-d7c550e86355 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Quantification of Multimodal Interaction at Sample Level UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.260373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.260373Z digest=sha256:df055e924ccd2b73050add07a6d0181c265a263ea406f0f80ea6565d3f7bc3b2

Observation 97c00d96-9983-42c8-8246-12d007582e89 · outbound

This paper cites Understanding Unimodal Bias in Multimodal Deep Linear Networks.

Efficient Quantification of Multimodal Interaction at Sample Level Understanding Unimodal Bias in Multimodal Deep Linear Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.274858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.274858Z digest=sha256:cdb8c14ac1bb6cdc202e5583fd7d0d167090f264d66fd7d0177227ea668809d3

Observation c61b8c16-0215-4891-bcc4-63fa9c6a188b · outbound

This paper cites The specific model variant employed in our experiments features unimodal branches, each consisting of four Transformer layers.

Efficient Quantification of Multimodal Interaction at Sample Level The specific model variant employed in our experiments features unimodal branches, each consisting of four Transformer layers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.397280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:52:32.281803Z digest=sha256:36eaeb4b39060c85a7ab9a2921ecfaec3c59da1b9233cf9062dbf96190fbbda4

Observation c7b47a57-c7c6-4ebb-8f22-277125aad31b · outbound

This paper cites Learning Factorized Multimodal Representations.

Efficient Quantification of Multimodal Interaction at Sample Level Learning Factorized Multimodal Representations

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.264142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.264142Z digest=sha256:73686986bc067027640431e56b4c2212cc273ba529a868ade3b468e2ed2d0736

Observation a73c0a1f-38a5-47c7-a49b-2c28abca96b8 · outbound

This paper cites Food-101– mining discriminative components with random forests.

Efficient Quantification of Multimodal Interaction at Sample Level Food-101– mining discriminative components with random forests

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.433225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:52:32.235428Z digest=sha256:520e23c38c1d1afb398bd22eeef723046528c488bab1c41fb8ad061c79a5b79d

Observation b134af1a-5520-438f-af05-d5e204067563 · outbound

This paper cites Nonnegative Decomposition of Multivariate Information.

Efficient Quantification of Multimodal Interaction at Sample Level Nonnegative Decomposition of Multivariate Information

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.270929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.270929Z digest=sha256:1fd6bafc861a1781e56cf512f98435957e699892c16146e325f5e0d91f059e83

Observation f76dcdcc-cd3a-4c6d-b06b-7760570f1c0f · outbound

This paper cites Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection.

Efficient Quantification of Multimodal Interaction at Sample Level Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:52:32.330835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:52:32.267621Z digest=sha256:f10a3e9b927a026662d71045b3f3617e05f05c44c23c07cf019ff803ffac9625

Observation 5f5153b9-6e44-4bdb-ae19-890ad324f59d · outbound

This paper cites Modality dropout for improved performance-driven talking faces.

Efficient Quantification of Multimodal Interaction at Sample Level Modality dropout for improved performance-driven talking faces

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.421813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:52:32.248738Z digest=sha256:769b429b741babd0cb41943dc839613d1b042678ad8751416c7c3320802db0fd

Observation 76fc0a54-36e4-440e-bddf-b6fdb68f5579 · outbound

This paper cites Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications.

Efficient Quantification of Multimodal Interaction at Sample Level Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.256147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.256147Z digest=sha256:4a2e7d98756af0f17bb83d072576ff978d00941622e916879ec359aeada487ee

Observation c7ac0c3c-ab6a-4d69-8447-200c95f06dec · outbound

This paper cites UR-FUNNY: A Multimodal Language Dataset for Understanding Humor.

Efficient Quantification of Multimodal Interaction at Sample Level UR-FUNNY: A Multimodal Language Dataset for Understanding Humor

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.244311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.244311Z digest=sha256:003dbd361e04e02f61b0ff794f994f1d8105579af3f14945b218d12123250501

Observation 20d53c37-4dac-4566-8139-7ece982f0470 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Efficient Quantification of Multimodal Interaction at Sample Level Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.240171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.240171Z digest=sha256:8261abacc456ac24292e356a72c39b5998f17902711a51bb5ab4c85899725157

Pith citing papers

No inbound Pith citation observations are available.