Pith. sign in

Paper Citation Record · LEDGER

EPIC: Efficient Prompt Interaction for Text-Image Classification

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2507.07415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07415 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:46:00.186219Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e14bf6c4-7ead-4b3e-97fe-5da5a5c3fc0c · outbound

This paper cites Cma-clip: Cross- modality attention clip for text-image classification,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Cma-clip: Cross- modality attention clip for text-image classification,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.354240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.137419Z digest=sha256:0b7e4e5c4edbe92d36065afff392b00348d3fdfc00e0b2deeacb6cd0f2ffe6c8

Observation 33fe7a97-2abf-4ce2-99e4-f5d01c59564e · outbound

This paper cites Future-aware diverse trends framework for recommendation,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Future-aware diverse trends framework for recommendation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.348130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.140117Z digest=sha256:ab1ad1901b0023b2b6951e3a017d70725b2ce17a506e7c79fb6d92ee90dab0f4

Observation 089606f7-3044-4917-b0a3-f1146e4545a7 · outbound

This paper cites Dense fusion network with multimodal residual for sentiment classification,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Dense fusion network with multimodal residual for sentiment classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.341633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.142283Z digest=sha256:67b33563cc404a85ed4e76f1a38c5abdaf4047bdac68dd71329390f38da0f947

Observation 34939d49-d3ad-44b8-923c-2fc5cf032d3a · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

EPIC: Efficient Prompt Interaction for Text-Image Classification Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.144295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.144295Z digest=sha256:4bea22d92b5c75521a3b19dbc1bd3c516066eabf253d442417c1f31dc2be26da

Observation 56960134-b3a1-4bc5-b46a-bfe3b619497b · outbound

This paper cites Memory fusion network for multi-view sequential learning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Memory fusion network for multi-view sequential learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.335901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.146948Z digest=sha256:cadcbf7d2d339a5aaabf7e1a943c6973c5caacc122574bf87d919a784a2b8fc7

Observation 2e0162a0-262e-4d25-897c-33e4adf5e804 · outbound

This paper cites Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.330209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.149114Z digest=sha256:0b879784bb546b585a300c2eb004f1b3ac9b0d2e693e618b9974614b67533a7d

Observation c7805295-3929-48e2-84b1-a43373a9f34d · outbound

This paper cites Centralnet: a multilayer approach for multimodal fusion,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Centralnet: a multilayer approach for multimodal fusion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.324314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.151389Z digest=sha256:80d3fbf826419eac758640182707b29cf5668133ab75f6fd59bda29d13f7f8a0

Observation 2830e517-deba-4ae4-b7cf-f4017f2b3250 · outbound

This paper cites Modular and Parameter-Efficient Multimodal Fusion with Prompting.

EPIC: Efficient Prompt Interaction for Text-Image Classification Modular and Parameter-Efficient Multimodal Fusion with Prompting

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:46:00.247864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.153348Z digest=sha256:bc7c11b298feb7f8ba95d7a85bfee1dd2299db565327c89ca1c8223ac5c3c05a

Observation eeb99690-bec2-4b27-98f8-cf993201bc0d · outbound

This paper cites Efficient multimodal fusion via interactive prompting,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Efficient multimodal fusion via interactive prompting,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.317889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.155562Z digest=sha256:1e7f35558d4a3eda6971a645a0d0951965f2a1657a5f89945b4c8ddf58b2036b

Observation ac88227b-92e0-408c-b950-fae682fc0372 · outbound

This paper cites Supervised Multimodal Bitransformers for Classifying Images and Text.

EPIC: Efficient Prompt Interaction for Text-Image Classification Supervised Multimodal Bitransformers for Classifying Images and Text

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.157567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.157567Z digest=sha256:a97b1b89499f46143c130ec2561ed446a3d24f78da2eb4cd272ff4c7ef2f8e0b

Observation 227179d9-8fa5-4808-8fd0-2c14be0f48ce · outbound

This paper cites Visual prompt tuning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Visual prompt tuning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.311547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.159794Z digest=sha256:b28e1a8f19001619d2c4d14db0b4767c02745f993ec2fee706b118c4f59de247

Observation 98149875-ad99-444a-ade8-d39581cb2e4f · outbound

This paper cites Learning to prompt for vision-language models,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Learning to prompt for vision-language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.305386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.162217Z digest=sha256:934b694035ba7de7373ac020907b6c8eb2ca8f18efeceec464e93812794e769a

Observation 14a89053-f6d6-4a61-b20a-b8f1654001b3 · outbound

This paper cites Con- ditional prompt learning for vision-language models,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Con- ditional prompt learning for vision-language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.298928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.164134Z digest=sha256:b7045bb815069693ce7ec7a6a218c7ee6dda0a6df517ef732665c3ddfc95d765

Observation ccb5fc53-3815-46af-a945-cf2dcb5f7c32 · outbound

This paper cites Maple: Multi-modal prompt learning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Maple: Multi-modal prompt learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.292220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.166005Z digest=sha256:916e5d786c9df1c6f34e619b2e711f8f208c4273e0574cdacbba0cc476b0175b

Observation 50a382c6-55d6-4255-94f5-63a344d52c98 · outbound

This paper cites HUSE: Hierarchical Universal Semantic Embeddings.

EPIC: Efficient Prompt Interaction for Text-Image Classification HUSE: Hierarchical Universal Semantic Embeddings

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:46:00.232501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.167906Z digest=sha256:d96e2d122b0e5e3bce338f443cd7557014a21260212b40f04431f3f726114957

Observation 431cd7f3-ef63-4d69-895f-ff67452f0759 · outbound

This paper cites MultiBench: Multiscale Benchmarks for Multimodal Representation Learning.

EPIC: Efficient Prompt Interaction for Text-Image Classification MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.169992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.169992Z digest=sha256:084c8a8f371bb01274bfd69d46d5b1f169eea9d5d6976cd5c754e60382fdecdc

Observation 38b259c1-d15f-4dd8-94bc-3bc43a8cf9d4 · outbound

This paper cites Dynamic multimodal fusion,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Dynamic multimodal fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.285857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.172055Z digest=sha256:475b8bb5143b0641456ed21abcd63c8d67ec463f61c1204c094746fe72267a5d

Observation dd0298d8-4358-4f2b-a285-0749ee0c0956 · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Unit: Multimodal multitask learning with a unified transformer,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.280004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.173800Z digest=sha256:b3c6f5cc70c61196ac1cf0c17fab65c39c39437cefeb9b551cd8f78edf1180c0

Observation 52dfd6b7-6988-475a-9321-9a35ca8260a2 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.273527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.175634Z digest=sha256:75a5edbc401157993269d3d023a8e23f6cffa022ec86f7f4169d8ed2ee1e20cc

Observation f0d32a1e-ec7f-44e2-9a7f-2c7b5cf7eef9 · outbound

This paper cites Recipe recognition with large multimodal food dataset,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Recipe recognition with large multimodal food dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.267354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.177617Z digest=sha256:404d9c6091546cbca2e05cc9a27bc18d3b24ac84f51f5af4d1fc24e6ecc855a1

Observation 3b6e58a0-f577-4bed-b1c0-e0d3316eaeaf · outbound

This paper cites Gated Multimodal Units for Information Fusion.

EPIC: Efficient Prompt Interaction for Text-Image Classification Gated Multimodal Units for Information Fusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.179567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.179567Z digest=sha256:5c720b72e15449b90bc1a4bbc1faba033646a6465a1f4b76db7b01a0ca1fb9a1

Observation e6d8ccb1-0998-40c7-befa-b26073a20ddd · outbound

This paper cites Visual Entailment Task for Visually-Grounded Language Learning.

EPIC: Efficient Prompt Interaction for Text-Image Classification Visual Entailment Task for Visually-Grounded Language Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.181739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.181739Z digest=sha256:7ed7062f8edd06efccc8514a64b8e34b42e3c456a67cd9bb5d85e88eb0611ec6

Observation 4d0e251c-11f0-4d21-9a81-699fac385c14 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.260276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:46:00.183804Z digest=sha256:27d8fa47cf532ffbc2ffaaa5de99415b7fb1bbd9ae573f6123b7c0a07a9dacb6

Observation fe1d239f-81dd-4190-a66d-455ec0b9a538 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

EPIC: Efficient Prompt Interaction for Text-Image Classification VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.186219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.186219Z digest=sha256:e53776a69bec9a6e3f27bfba67c5dcbcedc81c6592791f1e59b4341d5e311e5d

Pith citing papers

No inbound Pith citation observations are available.