Pith. sign in

Paper Citation Record · LEDGER

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.03531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03531 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:10:40.870961Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f64a285-5486-4481-a848-41d37450c95b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.273653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.500401Z digest=sha256:8c440ba3a47a2a938531313f30273926b871e99f1cbc50dd20ce5e3bc7c9f9b5

Observation 08f5321e-940a-492b-930e-a8d2926d4753 · outbound

This paper cites Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon L.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon L

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.264057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.607543Z digest=sha256:28ec6a81ec3aaf91bc00b2e5860810d8cbd58bdc0d5b8561df3d21838fe2e1ef

Observation a08ef3e8-8b7d-4929-bdad-2564cdb6b6d9 · outbound

This paper cites Advancements in affective and behavior analysis: The 8th abaw workshop and competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Advancements in affective and behavior analysis: The 8th abaw workshop and competition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.254562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.645816Z digest=sha256:a106411e149899bbb27fd2c3462f2e624c49c9f51bb6aaefa17f425da2e020d6

Observation f1e90cbc-c112-4c54-b824-e6368022d9fd · outbound

This paper cites 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:10:41.059074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.759840Z digest=sha256:5af5a340b3139f79dda0a79e3eb46f3d19456ee70109cd4bc73845e00a6d36f9

Observation 35fda226-7fbe-475b-bd42-1b768c83a350 · outbound

This paper cites The 6th affective behavior analysis in-the-wild (abaw) competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding The 6th affective behavior analysis in-the-wild (abaw) competition

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.244620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.781120Z digest=sha256:b801e2ca269fed46c27ebcc4ecacf258f19e3980cf9e8fa269235ce353ab60fd

Observation 7445a316-cce0-437c-9b9a-b346f8832216 · outbound

This paper cites Distribution matching for multi-task learning of classification tasks: A large-scale study on faces & beyond.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Distribution matching for multi-task learning of classification tasks: A large-scale study on faces & beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.235719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.784942Z digest=sha256:d50c56f4777c8fc0ff8ba997c955bc51d6912b1cd0cba1d90af5614888aacf44

Observation f58196a3-b9e7-48ef-aabb-b4550e736c0a · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.226399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.788478Z digest=sha256:7cd4d58cb7768e01a35618b9090f5db7d66c82d555567a97140c03eb71985ed2

Observation 03e8bd93-488b-4c2c-aa4e-c3e8d09a1303 · outbound

This paper cites Multi-label compound expression recognition: C-expr database & network.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Multi-label compound expression recognition: C-expr database & network

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.215669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.791965Z digest=sha256:d60386224601c69d2f6bea8fc94d29e68688abc1055c40170802c5ca8f27c59d

Observation 072d26f9-c235-417b-97cc-5b06ba8674f6 · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.206809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.795345Z digest=sha256:48b3db3804711f1e65fb9dbd514f8412dd351b6d1b1d7709013670cc94e6e998

Observation f63ccf43-2d0e-4eeb-829b-c0f7dcb30879 · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.197017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.799224Z digest=sha256:93661db7b103f6d38238c9d6469f67c25725df74a4a48bdd9c275c5da8a29e6e

Observation a085e703-44d0-4474-9ae1-31069f6f3644 · outbound

This paper cites Analysing affective behavior in the second abaw2 competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Analysing affective behavior in the second abaw2 competition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.187225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.803144Z digest=sha256:7505bdaa59357fcc9dd467063f0e3e26d65d156a7126fbab9a371a8c2f25df39

Observation 87117d46-f234-483e-a142-2a2469971925 · outbound

This paper cites Analysing affective behavior in the first abaw 2020 competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Analysing affective behavior in the first abaw 2020 competition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.177571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.806712Z digest=sha256:44a09b81ade49cfee214c724d3427afa3db8ad178255ef824d0718da9b4f9a57

Observation 1689b15b-0d71-49b2-a3cb-f3e06dc87cfa · outbound

This paper cites Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.810429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.810429Z digest=sha256:895dd297199eba063aad7b5d7027648fe9ed3f5851bf84c4fdd30ae0e0bd2a40

Observation 4b24b0fc-aa14-4630-9a09-26b0a906c5be · outbound

This paper cites Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.814343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.814343Z digest=sha256:06a976f478b6be5c51ce39042200c16777055da9e64111fd26713c69c8ccec43

Observation e6a9dfae-7c77-4cc3-9499-58c56f06e7a9 · outbound

This paper cites Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.817947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.817947Z digest=sha256:606bcd01ae8e4fce3af06cb2adfd9da994b0f99cacef393b45a65019f64d6a5b

Observation 0045ef01-93c9-4bd8-9c95-7beb30d01b9c · outbound

This paper cites Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.821888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.821888Z digest=sha256:ce832bb301801346b0e56832391c65d9213c502790c475af7d082200a6bbb83a

Observation 1d5046d1-8763-454e-828a-e274d0327289 · outbound

This paper cites Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.167438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.825643Z digest=sha256:ed11a2975af1e9e761cf80001b3919e5ba9dd8cafd59c4698c529ac9f8e49346

Observation 2df671bc-8334-47fb-89ab-627009e25e1e · outbound

This paper cites Dvd: A comprehensive dataset for advancing violence detection in real-world scenarios.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Dvd: A comprehensive dataset for advancing violence detection in real-world scenarios

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:10:41.005456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.828773Z digest=sha256:b16177eea47f72aac5554cd832f5fd751e9ca2550489b09fffd298075501661d

Observation 736a5e90-dd88-4472-88b7-674ad358d992 · outbound

This paper cites Deep residual learning for image recognition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Deep residual learning for image recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.157434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.831875Z digest=sha256:fb488062dad8e11626d627b4570ae173cf378c67cc6afa9dd6905987343109f5

Observation 7691a0c4-a1e8-4862-9972-962336ec3e0e · outbound

This paper cites Efficientnetv2: Smaller models and faster training.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Efficientnetv2: Smaller models and faster training

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.148032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.835222Z digest=sha256:2d0da919ef08e7ba6b775c76c53604cabb1778e9c818b26ff646c556d5d297a8

Observation 9fd62586-fcd0-458d-a8b2-606daee5c610 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Masked autoencoders are scalable vision learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.137646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.838238Z digest=sha256:88a774ac60bacb2c5430c7a8f39540d30f2b420d57311463e1818a866eaccef0

Observation c0f5d483-74e6-4ac4-8d94-069b27293e7c · outbound

This paper cites Cnn architectures for large-scale audio classification.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Cnn architectures for large-scale audio classification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.127063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.841654Z digest=sha256:1e9a0243dd562b6115f5a5ce65c3ddfdef75fc7d709b6034dda720797fe0cc7d

Observation 635b295a-8207-4728-a989-d738270442ec · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding wav2vec 2.0: A framework for self- supervised learning of speech representations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.116831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.844737Z digest=sha256:d74e359dbf6661b07c403c03162ef55d2d7bdc1fad604a611c25ec83db5978b2

Observation bbd5d9cb-16bf-45c9-aef4-bd7838eb0c62 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.106839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.847598Z digest=sha256:f796c4d7d706ca872efaffc92a11737df8f820737038356d650f24fcfa540b43

Observation f6bb8e1d-d169-4914-941a-52d1d0f315db · outbound

This paper cites Attention is all you need.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Attention is all you need

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.096282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.850744Z digest=sha256:815b8f84ff655247ab8cb113427da7e379a419127d66446922b40b16f16fe105

Observation 38f6ce61-6586-45ce-9f5f-36e4e8be9a1c · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.085288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.853807Z digest=sha256:3c6b8c895cb602fb980ef81ef546ea8d7098f6156da2a6e1af368e60b44510a4

Observation fae51b66-8d97-4a4e-a4c1-ff63160464e0 · outbound

This paper cites An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.856769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.856769Z digest=sha256:6fd48fdfa1a012837db02e0fa8ed32ac1471b54abf09ef0a678e287b192bcc66

Observation 8ca3f941-0204-4fcd-8856-70ab839f6d5b · outbound

This paper cites Contrastive Training of Complex-Valued Autoencoders for Object Discovery.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Contrastive Training of Complex-Valued Autoencoders for Object Discovery

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:10:40.917485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.861047Z digest=sha256:fbf57db7ab3da056318382f6fc2a3e521c2c1bd4fe52b5541e5b50da7032ddd1

Observation e2935fdb-3a17-4d9c-8fe2-1c96ee5ed57b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.864539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.864539Z digest=sha256:f321fa3e00605a0a933b35d296cfe3e5622bb21e64adca2ad31fed9af6562499

Observation a8410bae-100f-4ba2-b17a-20ffe6b0573e · outbound

This paper cites Focal loss for dense object detection.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Focal loss for dense object detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.867742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.867742Z digest=sha256:310c7ca5fbc9417ab1ed437542655a7c463d7ce71aff47fed221923ebdcedcf3

Observation cbdc3a54-adea-4007-a5d8-8526b33355a3 · outbound

This paper cites Decoupled weight decay regularization.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Decoupled weight decay regularization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.069656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:10:40.870961Z digest=sha256:f7b9d3fcdb6944c5cca91b317de567d591845e3f68246e1f5daec9cbcd464fb6

Pith citing papers

No inbound Pith citation observations are available.