Pith. sign in

Paper Citation Record · LEDGER

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

As of 7 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 2 inbound Pith citation observations for arXiv:2508.02429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02429 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:02:24.778729Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:19:34.140857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T13:55:53.243562Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97ea4dbf-5925-42e3-87e0-8cbc9aab1064 · outbound

This paper cites A review of affective computing: From unimodal analysis to multimodal fusion,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A review of affective computing: From unimodal analysis to multimodal fusion,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.340081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.340081Z digest=sha256:76656fc6056807b68a36346633ed17487d1e59ecc8495698d037e76dc5d8fabf

Observation 11db27f3-e04a-4364-a9a2-a32fe4fb8e4a · outbound

This paper cites A systematic review on affective computing: Emotion models, databases, and recent advances,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A systematic review on affective computing: Emotion models, databases, and recent advances,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:31.520898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.345319Z digest=sha256:d1e7e9ef7a8b06aea1f255a26e18135c5891b52edec768ed1a2a8f6db3bf8355

Observation de78fbb7-3a1d-4f76-84a8-a9aca737d6bd · outbound

This paper cites An effective data fusion methodology for multi-modal emotion recognition: A survey,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting An effective data fusion methodology for multi-modal emotion recognition: A survey,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:31.232058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.350453Z digest=sha256:e360c62a1938aa0b7fe0132692687b82fe411ed57a9504c91642e7c4f8c0f935

Observation b7525118-6794-41e5-a078-930cd45e178b · outbound

This paper cites Surveying the mllm landscape: A meta-review of current surveys,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Surveying the mllm landscape: A meta-review of current surveys,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.355118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.355118Z digest=sha256:39845c733f4beedb38be66764eaddd0449e43bb35747dcb37708eb1f68e3d151

Observation f1b077f7-53f5-49e0-b917-d81568e64559 · outbound

This paper cites Learning by comparing: Boosting multi- modal affective computing through ordinal learning,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Learning by comparing: Boosting multi- modal affective computing through ordinal learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:31.087023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.359998Z digest=sha256:f50af94ded63aac27714f5acdfdb6f29eb5012395b9acc4a1c59b287d690ee3a

Observation e75fac92-5937-4d9f-99cf-2ab6992642da · outbound

This paper cites Llm-based nlg evaluation: Current status and challenges,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Llm-based nlg evaluation: Current status and challenges,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.871599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.365403Z digest=sha256:38dc82093a2858f7fe0063d429d24ecd24c1714a0517f316d5348878447b30f6

Observation f3f1bbd2-9bca-4223-b3f0-c63677aba923 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.371526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.371526Z digest=sha256:bc00fdce5ff1a94cc927438519ed50c9bee284a2ca8735abd0f4d6dd0c9a10bd

Observation 81d76dc5-c191-4d76-9687-2d6e99d50007 · outbound

This paper cites Visual instruction tuning,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Visual instruction tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.718937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.378442Z digest=sha256:e75c63f45511c8cd03f846f1793f32a064982fbc5f464501d1e6939b6b1f461c

Observation 0fe48994-3171-4398-b5a2-43388ce28bca · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.382903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.382903Z digest=sha256:bc2b451ff95423334d414c4116df60b29f1d2b5126b02905663059fb881a2a72

Observation bd611aaf-b1b6-4e25-96de-8140eb598eae · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen-vl: A versatile vision-language model for understanding, localization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.529258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.388809Z digest=sha256:781813c0b7173e0ba806bd377327d59eaadbc11fa25c0639beb7a9965b5907c8

Observation df5a41c4-2648-4435-a407-e74517b4513f · outbound

This paper cites A survey on multimodal large language models,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A survey on multimodal large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.385999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.394458Z digest=sha256:919562d01c7c26e9722dd0f835eaaeeea2de8fc1baae5b9530ad4664e799c943

Observation b0d994dc-4d41-410e-91a1-54be9d9d8e9a · outbound

This paper cites EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.399808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.399808Z digest=sha256:3347f2a9f6ee28ea4ed61a0ff1f9eed816b45f1b2d5d92aee29b038132f1aa31

Observation 54470714-89de-4455-97f4-f8f45267dceb · outbound

This paper cites Eemo-bench: A benchmark for multi-modal large lan- guage models on image evoked emotion assessment,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Eemo-bench: A benchmark for multi-modal large lan- guage models on image evoked emotion assessment,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.405487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.405487Z digest=sha256:ec88768739570bde4b0f8f6eb70f6cbc8a95d68ce66766dcb3306c67d792fb46

Observation d1c5d6d0-bdc4-4b2a-8647-0b82c892956b · outbound

This paper cites Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.412048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.412048Z digest=sha256:02c32e6ffd6b9c1c57d52f2bc729037e207f0c5e557cf5edab8baf73235f250b

Observation 42d0d87b-261b-41e3-8ed8-a2c60f5f6119 · outbound

This paper cites The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.417545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.417545Z digest=sha256:fb05956e1856c200e835cd4b06cadb6353bc0c283f602f4d16862cb91962c224

Observation b1b62d41-cbfa-452e-a7a0-e9afabb2e9b6 · outbound

This paper cites MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.423638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.423638Z digest=sha256:03d82d9be130f068986b6d4277fb2048bbe6af77c0123a2d11fa408ccddf9154

Observation 21db1851-7af0-41ee-ba55-5898ff420f84 · outbound

This paper cites Ch- sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Ch- sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.238274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.429162Z digest=sha256:0b8733c0dc38a0fb081791830061af39d1e3ce95c0cd3c53ed3ea1284bab7059

Observation 3309977f-1b35-4894-a074-94440234caa7 · outbound

This paper cites Make acoustic and visual cues matter: Ch-sims v2. 0 dataset and av-mixup consistent module,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Make acoustic and visual cues matter: Ch-sims v2. 0 dataset and av-mixup consistent module,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.092305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.434158Z digest=sha256:33c84431035c1bbd57ac2cbf6ef90c6dae6513ca7bf470e4eb55983a7c4e7432

Observation d915f76c-e091-4c4e-8496-012f69e9870b · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.438791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.438791Z digest=sha256:af3b405b7acea1d4f902ea594534dc49977cfee914102b168ed5470794b9c01c

Observation 0855267a-b265-4c1c-838b-74924086cff2 · outbound

This paper cites UR-FUNNY: A Multimodal Language Dataset for Understanding Humor.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting UR-FUNNY: A Multimodal Language Dataset for Understanding Humor

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.443408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.443408Z digest=sha256:96eea671024437bdf7385a95526a101f87dc89b451d2b75badae9d8f91a0353e

Observation fd73413d-c3a3-48cf-b9f0-4314da5146bc · outbound

This paper cites Generated Knowledge Prompting for Commonsense Reasoning.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Generated Knowledge Prompting for Commonsense Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.447997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.447997Z digest=sha256:1985e7d200fb322bbd0cf2ebff4d0e8d0f0a0fa657b5693b1a5049a57b385165

Observation 8fc8751e-6fb9-4930-8744-859f0a979492 · outbound

This paper cites Self-attentive feature-level fusion for multimodal emotion detection,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Self-attentive feature-level fusion for multimodal emotion detection,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.971592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.452615Z digest=sha256:e10c4616dd887537cf374472ef7f0c9f8cc9e4e06393f17ff2e0816295b46099

Observation adfaf75b-99c2-44bb-a848-01b9d6828b6c · outbound

This paper cites Emotion recognition using feature-level fusion of facial expressions and body gestures,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Emotion recognition using feature-level fusion of facial expressions and body gestures,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.832233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.457912Z digest=sha256:19a858b582d46f9de7f6719539db9e6c143f607ba6db936f795415feedb9704d

Observation 090ae2fc-4055-4976-84ac-48158ff0fd74 · outbound

This paper cites Decision-level fusion method for emotion recognition using multimodal emotion recog- nition information,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Decision-level fusion method for emotion recognition using multimodal emotion recog- nition information,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.705955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.463005Z digest=sha256:62973a1491578b7bbc270c662acf00ba11f3571628d5a467e62b37339325507f

Observation 1e10bca3-bf67-4ee0-9c05-2e71ca3b791f · outbound

This paper cites Deep learning-based late fusion of multi- modal information for emotion classification of music video,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Deep learning-based late fusion of multi- modal information for emotion classification of music video,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.514991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.467810Z digest=sha256:2b175f52d6313504bace6ac6c1421d55bafa52b3293d84f2d64a9e0b7267b678

Observation 368a5029-19d1-4bcd-9ea0-746739f59219 · outbound

This paper cites A joint cross-attention model for audio-visual fusion in dimensional emotion recognition,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A joint cross-attention model for audio-visual fusion in dimensional emotion recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.311160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.472652Z digest=sha256:5f444eb2314fc08cd1c48b70cbd28f12cd64e373a3c50bf99b8d994c92cc93d4

Observation 404f1533-fea9-4bfe-a18d-5545045a8f36 · outbound

This paper cites Speech emotion recognition with co-attention based multi-level acoustic information,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Speech emotion recognition with co-attention based multi-level acoustic information,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.122837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.476988Z digest=sha256:e30a538d22b91ff4588539bca9da1a5020242f953468b731b0b943c5110ae357

Observation b284676b-38d0-44fd-978f-2acca5ce5b4a · outbound

This paper cites Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.481438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.481438Z digest=sha256:497a836f4a4264f28c15fd4dadf89f5e6f39478c21f12dcea756021ed74c3a13

Observation 74964bcf-796f-4350-a856-84deb9c68674 · outbound

This paper cites OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.486614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.486614Z digest=sha256:9d353e9296148d7c4f57481739c97bbe89eaa1b1889a3d969b83af75cf290f4e

Observation a0fec81d-4e6d-40dc-9b2f-8d553f64eefc · outbound

This paper cites Mellm: Exploring llm-powered micro-expression understanding enhanced by subtle motion perception,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Mellm: Exploring llm-powered micro-expression understanding enhanced by subtle motion perception,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.491839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.491839Z digest=sha256:380ac6b0c3857c1f58536ecd5eef357f9043c8883e7f6c8b736f561ba8b39153

Observation a3f3bde7-2787-49c7-9fca-fcaec3ec39ca · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Learning transferable visual models from natural language supervision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.497289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.497289Z digest=sha256:6fc59a46cedf16374b190d97bab39024035abca2b30b44755f8324944e07f643

Observation c35e7560-2238-4b01-b812-b28bcbea21f0 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.930253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.504123Z digest=sha256:4900e131e8bbb51b05dd315f0653e31a7dbd6c6b68bee0b35f85f9c22c469824

Observation d32b479c-3c7f-4f03-9026-72d988877b7b · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.509860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.509860Z digest=sha256:c1096ac136570d539b05bee7b3494de536ed78fc270c5ddb50f59af4fe686b3c

Observation c7731c3b-89bc-49ca-a089-6d1ee0720fdf · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.516633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.516633Z digest=sha256:86535399d80c6344c2a0c2f1e1eef891792972edb977d84ff492594e44fafae5

Observation 20b4fd2d-8357-4fa8-905d-821aa6e42340 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.523176Z digest=sha256:bd0217c7b88ce9caa5336f276406ea9c158c39d4d7f12d1193faa97273197005

Observation 5af11d9b-7bd0-4ea7-ac40-7c30cab0911d · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.529880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.529880Z digest=sha256:84be078c2ac85d1cb4bb5bc7b1ba7f8ce3a446cf20f12bd056f4613faeef62b5

Observation 11f2de67-3088-49d6-8256-7b9f687d8eef · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.535612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.535612Z digest=sha256:727e98e56abc0f86b614949ae6d0867bfa8f68f99725e51bb78e1f9fe21f0226

Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · outbound

This paper cites HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.540778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.540778Z digest=sha256:57f572166d56775e0370f5ecec78f733a00b7fa0f43fcf21ea412579417ce125

Observation a779eae7-fa98-4831-9eb1-a1dd5b7829b9 · outbound

This paper cites Ola: Pushing the frontiers of omni-modal language model with progressive modality alignment,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Ola: Pushing the frontiers of omni-modal language model with progressive modality alignment,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.768680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.548636Z digest=sha256:9f4de7c4e4c4a3c7f6889da99855c26f349d4ef7489b24c06d8a7d5da7fd2e5d

Observation ce2dff1b-3623-49ca-8b58-96642e13d7f8 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2.5-Omni Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.555547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.555547Z digest=sha256:8ddab45e6a092bd93a6e11410ddc80aff4cbf3fe10eddaf41615ba0965205fd6

Observation 23714ce4-ca41-49be-a199-fdb7fe05d0c0 · outbound

This paper cites Learning emotional prompt features with multiple views for visual emotion analysis,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Learning emotional prompt features with multiple views for visual emotion analysis,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.612232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.563409Z digest=sha256:c107dec5de4a1dc4345c3a27b7a7db81d4d7930a6417943ba86155b1fbb1e4b9

Observation fd4ab5fb-87cc-402c-988d-3457955a13c0 · outbound

This paper cites Visual and textual prompts in vllms for enhancing emotion recognition,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Visual and textual prompts in vllms for enhancing emotion recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.478909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.570432Z digest=sha256:14119a7e260a4664fca4ea0989c1f05e87f7a14f0eb5795707ad12f62d1d9c07

Observation 86ed16c1-d44c-4cd9-80cf-0b421f060ff1 · outbound

This paper cites Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.328838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.578864Z digest=sha256:f9c38c8a9f3e2749d45103d5f94099ad1d2f1747356c89403e6f925a7d3ba786

Observation b4ec7c38-d45f-40db-9361-da55708317ab · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.586681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.586681Z digest=sha256:e0fc5d939005d54f6e7b3df4bcd8b4777ae48695401ad3028f6f563efd1d5b3e

Observation 0dd638df-6b46-4131-9a35-4098b317b9dd · outbound

This paper cites Emotion-llama: Multimodal emotion recognition and reasoning with instruction tuning,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Emotion-llama: Multimodal emotion recognition and reasoning with instruction tuning,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.130166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.593485Z digest=sha256:a310277e2cce7ec51538791ff636bafeafafd599723121007ad4ee224973abb0

Observation a138cafe-fb39-4808-a7e6-ba7c95b95a7d · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting PandaGPT: One Model To Instruction-Follow Them All

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.597936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.597936Z digest=sha256:1f75e0feedd64d4308832086a28406299a14193b31ea6e0cfb6d6ef87655d6ea

Observation 42933506-5da1-4470-bf52-7a1a9093cb09 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Lora: Low-rank adaptation of large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.979157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.602861Z digest=sha256:67c6752d94d069ab22873cffe43bdb7b7f133bd728c120109b4bfcc9c14e9e75

Observation 23935ab7-5e3e-4cd0-bf5e-dfdfc236d63f · outbound

This paper cites Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.838128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.608419Z digest=sha256:b0cb36f9594b03e94bfceb4ba471c3a08e012ac28f49aea2eeff27d0ce495e18

Observation f4ade031-da1b-491c-a62b-695a1d283a13 · outbound

This paper cites Injecting multimodal informa- tion into pre-trained language model for multimodal sentiment analysis,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Injecting multimodal informa- tion into pre-trained language model for multimodal sentiment analysis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.684303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.614031Z digest=sha256:3e37750d72baa2f5eccb2b7d545d06a1bc59c1d74f6d22355f1d1e6be9de8af5

Observation 9364ce03-6343-4e38-8aba-e1f5a2243737 · outbound

This paper cites Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:02:25.031605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.619395Z digest=sha256:37bf0be0c54c6eae64c07e209b08e5358a23c532432dae0368de945b5a509848

Observation 07d8b9bb-8567-4c82-b9d9-4716b492c995 · outbound

This paper cites Hgtfm: Hierarchical gating-driven transformer fusion model for robust multimodal sentiment analysis,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Hgtfm: Hierarchical gating-driven transformer fusion model for robust multimodal sentiment analysis,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.533321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.624420Z digest=sha256:83d518d30630d9e67c92dfa3c88e888b7b2a4271377486fd1b7d1123fe4061a8

Observation 19298fd8-686c-4506-ab67-daa8ca63f8cf · outbound

This paper cites End-to-end Semantic-centric Video-based Multimodal Affective Computing.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting End-to-end Semantic-centric Video-based Multimodal Affective Computing

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:02:24.999298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.630238Z digest=sha256:1f0ed68747b94cee348217fa80182ef93aafe4fdf5644f022b323f3f818e7cba

Observation 721e03b3-f413-49f6-b0d5-41838fea01f9 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.635326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.635326Z digest=sha256:99a8127e02064b7048808ddc31974b01a08c1ac345276323359e7ea78a823040

Observation 50dd4fc4-8c65-4259-999c-7920a4cd450d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.641304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.641304Z digest=sha256:136c150c9acdc580608cf2bf028ef8353f2a4bdc39ed6103c9949d03569e19e5

Observation ef9dcab4-4f63-40a7-a256-d0d59273805d · outbound

This paper cites Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.338026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.647280Z digest=sha256:667fc39923ceb6ba983c890218d4a88d99ae391140392d98587fc1a567d794d6

Observation e1fd8b1b-21fa-48b1-9fa3-0f4edd6ef6ae · outbound

This paper cites Qwen2.5-Coder Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2.5-Coder Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.652010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.652010Z digest=sha256:76549e98471ebf1f4010f7be150c90594b82c07f31a22bd5dee50ae9a7cdf9c4

Observation 85140fb3-e5f7-4992-910b-cca5bdd125be · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Robust speech recognition via large-scale weak supervi- sion,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.657476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.657476Z digest=sha256:28bebefdffc498958407c1a91f5a4dcb6f828ffde3d961110baef7888699baf6

Observation 8c2e31f1-ad87-48ef-89f8-083ee16eb6b6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2.5-VL Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.671751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.671751Z digest=sha256:9817d90957c91c8f11c2074850671ffa28e2d10557dd07fd61bd679a64fcd8b6

Observation 4f549136-c179-4339-8972-b9a90ce551b8 · outbound

This paper cites A survey on vision transformer,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A survey on vision transformer,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.174956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.677944Z digest=sha256:d98aabc65f7c16dd605eab2d2aa5723547e658f247dbf30836c77c3a172e90d7

Observation 4e5f5b46-67cf-4e38-9266-19f98ad5916a · outbound

This paper cites Sigmoid loss for language image pre-training,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Sigmoid loss for language image pre-training,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.008775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.683813Z digest=sha256:b3284a67e5a92cc3eac9bac10bbd140160a502e3885a61dcb8d490c73ef8359c

Observation 8b1132f2-adf1-4797-9d46-f0ba22dcad40 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting LLaVA-OneVision: Easy Visual Task Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.688796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.688796Z digest=sha256:1f650d32c410c287e864d7bb6ccdb2acf75b1b1d2e27e2518323f28449bd0c6f

Observation b95fbbae-b9c3-43f3-8cf9-7f2c29168dbd · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.696608Z digest=sha256:5e5d8d375b48078a6151b723bace8b8cc39e663bc7d7400a7d48d41086b27e1a

Observation 39714fe6-3995-4275-b61f-95aee8211602 · outbound

This paper cites Qwen2 Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.703321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.703321Z digest=sha256:825905def80088b89617dcd17c5e463871bbea3c273027f047347e32b577d902

Observation 5548e265-db02-49f1-82f6-2416ffd59237 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Imagebind: One embedding space to bind them all,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:26.848544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.710416Z digest=sha256:6101bb881b0aad44809b2e7094b1b68158170f0b377b2020127e5ecfa75fa825

Observation 3fa0f8c0-2f9e-4206-a2a3-56bda65f93b1 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.717254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.717254Z digest=sha256:afe004f5dacb8bbda1d02c7852e44ba831859986c8b0b466965047a183cc602c

Observation 263337ce-aea8-41fb-9da2-b727ba36fc0d · outbound

This paper cites Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:26.617560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.723159Z digest=sha256:2a7f390f50414fb65aa7c6849ac3ea1a2a83b5c5995cb59cfe22d7ac6a36f4b0

Observation 0b18373b-846d-452e-9c72-040a0914caa9 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Eva: Exploring the limits of masked visual representation learning at scale,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:26.349226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.730210Z digest=sha256:31be22d8f62537f2559f814f7007929d4a435dadecfe287211a93a9d82c1bb02

Observation 74da8b45-edc1-4c74-897f-a14243688431 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting LLaMA: Open and Efficient Foundation Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.737622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.737622Z digest=sha256:4fb1be4b25d811544b9b6f0af37668c6a370f02389fd4b9ad00bddd80d766c7e

Observation cde85087-202c-4242-87b8-d29ff4cbc18a · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:26.217538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.743337Z digest=sha256:a147656ee9b2e9702c2593e867683ed8e5fc93ae028ba26516b630c884be6355

Observation 917be28d-0b19-4dab-b624-cd745e496966 · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:26.031375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.755334Z digest=sha256:ee0ac7efaae1a2feecc43a1efedd2e7f6e4b3faf8313a7c9c94a52dcf034e989

Observation 311dbcf4-fc63-4d0f-8ad0-e80cfbc0cb9b · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:25.873698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.762492Z digest=sha256:97b06b269dea5a8c53c33c9438ebf720c3805a397ac51ff2e39cb4636a588d74

Observation abe15f6e-c07c-4049-80bb-ed134b901b3a · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:25.826168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.768191Z digest=sha256:fad5d33c3d2f2fe5d7f7187d128165e1ac6043af36d202e9f06b965de2fd8c40

Observation 99597bf0-f020-44c7-84c0-680b7b5a62ee · outbound

This paper cites Its core architecture follows the Thinker-Talker design.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Its core architecture follows the Thinker-Talker design

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:25.786279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.773095Z digest=sha256:7297483ff06ccda80fcfe198184f9e026a4e260568d1612afeb0864cdc3334c9

Observation 27a231c1-e83b-44f8-b2fd-748b9bd889d0 · outbound

This paper cites Its key innovation lies in the ability to simultaneously process visual and speech information in human-centric scenes.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Its key innovation lies in the ability to simultaneously process visual and speech information in human-centric scenes

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:25.746097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:02:24.778729Z digest=sha256:22051d12053fe28ef94827126f87f800699a1886cd107bd243ad9932b10ec123

Pith citing papers

Observation 8b65755c-ea84-4a75-8437-4557e50e4303 · inbound

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis cites this paper.

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:19:34.140857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:19:34.140857Z digest=sha256:fb5f8c97a866f5af1687ce96ac6b383017b845c2402fb0e07de7b1789bca9f4a

Observation 73165972-423e-47ac-9fdc-f2fb77c25237 · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.245819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:5f11018a7ceaad154de22476e2680bc68b6abdf4f47eec62ee165e1272394e44