Pith. sign in

Paper Citation Record · LEDGER

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2506.13971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13971 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:34.050123Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:31.400913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:02:34.268817Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14f71907-93c9-4273-8391-f6dec1e4d4e9 · outbound

This paper cites Although it is an es- sential medium for communication, it has not been sufficiently studied.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Although it is an es- sential medium for communication, it has not been sufficiently studied

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.913476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.108389Z digest=sha256:1b1132b531292ec387842a459465e9f88ebe5565baccd574fa15c60bc6a094ba

Observation 5beace89-127d-4171-87da-9ce7ded35caf · outbound

This paper cites While audio-based SSL has been widely applied in speech emotion recognition, multimodal approaches have only recently emerged [5].

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience While audio-based SSL has been widely applied in speech emotion recognition, multimodal approaches have only recently emerged [5]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.906318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.318650Z digest=sha256:04cad5eb049d065ec2641bee07a637ac053bdb66f10cdcba5b13d637b892dc25

Observation 250380de-4773-4642-becc-ff26f686d791 · outbound

This paper cites Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:34.330862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.400913Z digest=sha256:bdcef6db3156636d896f2169f53f8171a38f266c8891fe02787f85b4775c9bd3

Observation 9f3f8ed2-9ff6-4fb8-954c-66f9a1ebedc8 · outbound

This paper cites Data split To analyze the impact of the labeled data ratio on SSL model performance, we partitioned the targeted clips data into 10 folds.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Data split To analyze the impact of the labeled data ratio on SSL model performance, we partitioned the targeted clips data into 10 folds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.899495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.500877Z digest=sha256:936c32fb70bf90885abbbdf0dbbea5213b88a0ae0e3ef06877caca7d85f59da7

Observation 3721b4df-c3fd-44c2-b52a-b7f439701f61 · outbound

This paper cites They both outperformed SL counterparts at nearly every levels of labeled data in both ROC-AUC and macro F1 score by 1-4% for predicting Enjoy- ment or Fluidity.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience They both outperformed SL counterparts at nearly every levels of labeled data in both ROC-AUC and macro F1 score by 1-4% for predicting Enjoy- ment or Fluidity

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.892482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.640676Z digest=sha256:befddc58c6cc86fa29fd26e714f94165f177aee987e528ca98d8a8ed07ce4582

Observation 4b42839b-4222-4946-852f-59f472b2a021 · outbound

This paper cites We also demonstrated this approach gener- alizes to new sessions with different participants.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience We also demonstrated this approach gener- alizes to new sessions with different participants

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.883948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.734534Z digest=sha256:a39bdc5358752c1a2c4e403609b8baaa800a7ea21493d126f26872350803b06c

Observation 0f6851fd-ef72-4fd4-9693-06c0175db9c6 · outbound

This paper cites are supported by NYU Discovery Research Fund for Human Health.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience are supported by NYU Discovery Research Fund for Human Health

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.876763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.860956Z digest=sha256:86de1d1863d5c6a67d516df6d047b3c12ab7e924c741ca746cc19984a4f12460

Observation 0216becc-1531-4bb6-a341-b865ff7dd3d3 · outbound

This paper cites Sepa- rable processes for live “in-person.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Sepa- rable processes for live “in-person

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.869522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.969102Z digest=sha256:0d525260ec44dad25e1893e35373054b47a87ff96cedf5b88c2403bb6538e561

Observation d6a1ce21-fb37-41f4-b618-189606ad5a20 · outbound

This paper cites Perceiving others through a screen: Are first im- pressions of personality accurate and normative via videocon- ferencing?.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Perceiving others through a screen: Are first im- pressions of personality accurate and normative via videocon- ferencing?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.862748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.089559Z digest=sha256:87ac215bde7bfe8c7e82e80d21e6f0f475e32e4db1f818f4a2597149172a9a71

Observation 9fb2e6b3-aa66-43f7-9561-daa65b84c27b · outbound

This paper cites Virtual (zoom) interactions alter conversational behavior and in- terbrain coherence,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Virtual (zoom) interactions alter conversational behavior and in- terbrain coherence,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.855048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.228978Z digest=sha256:24affd4462028280d2727999627713b08f38da94b231ea5dddfe7c2b9e6b3407

Observation 7c8af8da-14d1-4bb1-8472-e1264d55b435 · outbound

This paper cites The effect of video feedback delay on frustration and emo- tion communication accuracy,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience The effect of video feedback delay on frustration and emo- tion communication accuracy,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.847338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.313479Z digest=sha256:a6a9aa96a925f383190f402ad888684171f87497fc5eab351d0cbc548e1a387e

Observation 4d884b55-e1ab-4fe5-b2cf-73be7dc8d62f · outbound

This paper cites A survey on the semi supervised learning paradigm in the context of speech emotion recognition,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience A survey on the semi supervised learning paradigm in the context of speech emotion recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.839435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.435257Z digest=sha256:1fe320bd338011780482e23f9d9824fee833ec30be691bfdf74cc6ae5e6389ef

Observation 0d0b9bfc-2c06-456a-9dc8-ab550a3f792f · outbound

This paper cites Combining cross-modal knowledge transfer and semi-supervised learning for speech emotion recognition,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Combining cross-modal knowledge transfer and semi-supervised learning for speech emotion recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.832068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.527665Z digest=sha256:7dc7e96eaa503a401de137ed4dec92730a314928f432efc3b7ed91e67a1a2ef4

Observation 712e1a3e-f380-40fa-8c8c-a310fe7cc07e · outbound

This paper cites SMIN: Semi-supervised multi- modal interaction network for conversational emotion recogni- tion,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience SMIN: Semi-supervised multi- modal interaction network for conversational emotion recogni- tion,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.824409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.647703Z digest=sha256:604d045f38c4af318c702878888e9572c12bc8a08a1cebe812612448ab5ebd48

Observation 24621e28-0fef-4fdc-8502-08ac0926270d · outbound

This paper cites Multimodal emotion recognition with vision-language prompting and modality dropout,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Multimodal emotion recognition with vision-language prompting and modality dropout,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.815933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.775464Z digest=sha256:b02d713c9ae1b3a91c4e71275fca0489af6574e16b80874ac1c57a1e9f7fa9b2

Observation a8a013e0-c416-412d-be93-08bf6fd8f456 · outbound

This paper cites Focused or stuck together: multimodal patterns reveal triads’ performance in collaborative problem solving,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Focused or stuck together: multimodal patterns reveal triads’ performance in collaborative problem solving,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.808608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.850168Z digest=sha256:670c0fffdd260e5db85ccc4a876e54e7e0173da6a12ea97537462d4d32fe2e90

Observation fb84d95b-bce3-4858-8a4c-18a3ddeda976 · outbound

This paper cites QoE estimation of webRTC-based audio-visual conversations from facial and speech features,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience QoE estimation of webRTC-based audio-visual conversations from facial and speech features,

Reference 17

Resolution
verified exact
doi, observed 2026-08-07T12:02:34.193424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.906151Z digest=sha256:5bf6982a420c0b11f7cfe18c7182a6a1682dd6f75590ecfe9e415c5aefa267bd

Observation a4041642-c90e-4244-a9e2-07fbc52f954b · outbound

This paper cites Multimodal machine learning can predict videoconference fluidity and enjoyment,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Multimodal machine learning can predict videoconference fluidity and enjoyment,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.800732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:32.960804Z digest=sha256:bdffbe4b5545f6de37b58f4bd3c756b82819c1545c5cd6b13310bae0dfaf1103

Observation bced800f-af80-4f3d-98e8-698201cade0c · outbound

This paper cites A survey on semi-supervised learning,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience A survey on semi-supervised learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.793668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.011780Z digest=sha256:ced19d9f2b15173c0ca6c8f62f8576bb2296bd7353340ef93916dd781a006119

Observation 864c3047-534a-4915-a15b-57963d720772 · outbound

This paper cites Self-training: A survey,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Self-training: A survey,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.786431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.085789Z digest=sha256:e0b4f114160af5d8f8594d17bfebfa2a69c474e71f20dedd638300e41968e6db

Observation 8222a713-f43e-4f7e-b14a-22b717b03ba1 · outbound

This paper cites RoomReader: A multimodal corpus of online multiparty conversational interactions,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience RoomReader: A multimodal corpus of online multiparty conversational interactions,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.780298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.146074Z digest=sha256:07a361f9509f4c74e94d43a2ef5955afddcc87c665e12c0dfe7fa198fda565be

Observation 13a6602f-8406-481c-89c1-fdc64f4a56af · outbound

This paper cites Zoom disrupts the rhythm of conversation.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Zoom disrupts the rhythm of conversation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.774060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.217943Z digest=sha256:2e4c66e89d0179ed89e0d117f466404224285c0f15669757e2cbf45c71586d1c

Observation 8c5eab4d-20f3-457d-84f6-10ab2b8fe3f9 · outbound

This paper cites CNN architectures for large-scale audio classification,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience CNN architectures for large-scale audio classification,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.767947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.290408Z digest=sha256:4fef803866b446fc772bb4aa4902738747f37053b3c043c2542eb0198dc5846b

Observation f6c6d7f2-b2e0-4ff8-92cc-1a616f275b80 · outbound

This paper cites Openface 2.0: Facial behavior analysis toolkit,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Openface 2.0: Facial behavior analysis toolkit,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.760837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.344916Z digest=sha256:fe115e0e8c615aecdd07dff55307f0d37b8bf3536f8588a30125c3f5722e9726

Observation aa76d053-f642-4314-a29a-1d5bb0f78f73 · outbound

This paper cites Sentence-BERT: Sentence embeddings using siamese BERT-networks,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Sentence-BERT: Sentence embeddings using siamese BERT-networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.754529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.397811Z digest=sha256:994cb1fd5b50445600f0f3f4c530743f3fbe54b0f3c7d628e50309dd3caf78e4

Observation 5207f3bf-45a3-4738-850c-1ca9d3ce2efb · outbound

This paper cites Dawn of the trans- former era in speech emotion recognition: closing the valence gap,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Dawn of the trans- former era in speech emotion recognition: closing the valence gap,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.561797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:34.014509Z digest=sha256:d2552c13d4bc41f6a55288fffd59ea2f916a51a013d20a6743b36b8c4532a73f

Observation 98963298-8a30-4102-856f-d280508ca078 · outbound

This paper cites Unsupervised word sense disambiguation rivaling supervised methods,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Unsupervised word sense disambiguation rivaling supervised methods,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.748580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.516715Z digest=sha256:29bc47ef20520c99454493f70b6ba060b8931270631bb698e293bf2601bf2f2a

Observation 03005f86-1240-4a87-b6d5-980d033a85a2 · outbound

This paper cites A new analysis of co-training.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience A new analysis of co-training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.742683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.572295Z digest=sha256:96b94ef9ef9284c86c77ddc6bc13f4c5d7aed54b1e15a10aa96ad926ed8b9f64

Observation 6f9e520e-f388-4c9d-8b4d-c8f99b3b2357 · outbound

This paper cites SSLearn: A semi-supervised learning li- brary for python,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience SSLearn: A semi-supervised learning li- brary for python,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.736454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.658316Z digest=sha256:a4539da04e6413d162d22091abade39a2efe5e0fbb6c86073364d5b1c768e500

Observation d77f066a-8e63-44e7-a549-ec43c0453d09 · outbound

This paper cites Optuna: A next-generation hyperparameter optimization framework,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Optuna: A next-generation hyperparameter optimization framework,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.729208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.725720Z digest=sha256:2443cb6c5f73d11332a367f595703e87c0aac832f3e6a22f791246ac704309cf

Observation c661bc39-f3d8-407d-8117-f1ed423aa00d · outbound

This paper cites 3m- transformer: A multi-stage multi-stream multimodal transformer for embodied turn-taking prediction,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience 3m- transformer: A multi-stage multi-stream multimodal transformer for embodied turn-taking prediction,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.722266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.776863Z digest=sha256:97bca9af7b2a4ccaa1253a867b52efefafb6aed4fe5cc9c96fe644992c435ba3

Observation 6cb51c3a-63e5-4a15-b7cb-7dc0c3f0187f · outbound

This paper cites Pre- dicting conversation outcomes using multimodal transformer,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Pre- dicting conversation outcomes using multimodal transformer,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.715066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.858907Z digest=sha256:b42758cead08e5ea73ef33fa06ae1689e214a456ff1bb87e9797c3bc9ba0df99

Observation fe291e78-ff26-45ad-a4fc-7b6ce2ea9c1e · outbound

This paper cites Dyadformer: A multi-modal transformer for long-range model- ing of dyadic interactions,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Dyadformer: A multi-modal transformer for long-range model- ing of dyadic interactions,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.708267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:33.946759Z digest=sha256:0783a85586e2e8dc55f931f2c48f06ad393588459c87d9d8e84129914a012309

Observation 18b0d01b-1508-4a5d-b42d-74d9253f400c · outbound

This paper cites CTNet: Conversational transformer network for emotion recognition,.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience CTNet: Conversational transformer network for emotion recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:34.477575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:34.050123Z digest=sha256:dcb41f7c06e5748a8a478b5eebad79d12332fe17f812a03abb2ff5ad5e240565

Observation 70a47fcc-fc55-42ad-a552-f1b3cfd54226 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.445605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.445605Z digest=sha256:07e7edd355be3bc7112ef8b6a14b83b3192a143811f121bbfc5a8dfc8bc5a67b

Pith citing papers

Observation 250380de-4773-4642-becc-ff26f686d791 · inbound

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience cites this paper.

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:34.330862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:02:31.400913Z digest=sha256:bdcef6db3156636d896f2169f53f8171a38f266c8891fe02787f85b4775c9bd3