Pith. sign in

Paper Citation Record · LEDGER

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

As of 5 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.07294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07294 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T15:07:52.715880Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact13
  • verified fuzzy30
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0282f24d-3551-4817-bef1-7b320ee434d1 · outbound

This paper cites A Simplest Systematics for the Organization of Turn-Taking for Conversation,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders A Simplest Systematics for the Organization of Turn-Taking for Conversation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.352789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:a1633623f03e18072590cb8692dd379d9fe2ac2de36e68ecc6ed888d411d5179

Observation f9129915-712f-4024-8910-f406a887bc7c · outbound

This paper cites Palinko, L.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Palinko, L

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.358292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b5360d8f793253e5b1f0dcb597d39c2de4eed87da8b46131ee6a1ace801aea88

Observation 52b7dab9-f39a-45be-b973-b72343cca828 · outbound

This paper cites Towards improving turn-taking in social robots using Visual-Only V oice Activity Detection in multimodal dialogue systems,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Towards improving turn-taking in social robots using Visual-Only V oice Activity Detection in multimodal dialogue systems,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.394045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:d60ef28b7f630516e9481ebcaa0df5c0a962c667a80f5415c8ba23220b7cee34

Observation b62a8dfb-9685-4f45-81b0-e8c88c2bdb24 · outbound

This paper cites Design of Social Features for Robot-mediated Cross-cultural Interaction,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Design of Social Features for Robot-mediated Cross-cultural Interaction,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.367807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b8534acb9bbce8939ed6ed6089dd560f32f21ba5d20f7dd0f59987f978eb90ac

Observation 87a5e57a-9755-4805-aaad-905ec97166d4 · outbound

This paper cites Haru in the Care Network: Stakeholder Perspec- tives on Privacy with Social Robots in Pediatrics,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Haru in the Care Network: Stakeholder Perspec- tives on Privacy with Social Robots in Pediatrics,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.376031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5fa2e222dc7bb12b7deeca5c1510e3293e419bc063088138b5348f27a3111194

Observation ea58cff9-99f7-4cdd-bba2-b13bef42da12 · outbound

This paper cites Building Friendships Across Borders: The Role of Social Robot Haru in Children Group Communication and Con- nection Development,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Building Friendships Across Borders: The Role of Social Robot Haru in Children Group Communication and Con- nection Development,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.345600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:95160ad1adce8f567e52235be2396601d1dcc88383e3d40e07fcf4f58bda5cb5

Observation db4b203d-461d-42bc-b452-13de60bc5d7d · outbound

This paper cites Multimodal Transformer Models for Turn-Taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction During Cooperative Gameplay,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Transformer Models for Turn-Taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction During Cooperative Gameplay,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.362684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:78395470720972dde40f34e3bd457222da6d136050274e8553567fee44a8f234

Observation 68bb76cc-eb5f-459f-88f5-3ad4ac8d6b12 · outbound

This paper cites When and How to Express Empathy in Human-Robot Interaction Scenarios,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders When and How to Express Empathy in Human-Robot Interaction Scenarios,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.360215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:6fb5d017d3fbe2f49bc6c23c6dbdbaa782b80c5ff06587720086bc431c452fa3

Observation 73e92f35-0e57-4eba-8dd0-72374184aaf6 · outbound

This paper cites Visual Cues Enhance Predic- tive Turn-Taking for Two-Party Human Interaction,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Visual Cues Enhance Predic- tive Turn-Taking for Two-Party Human Interaction,

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:16:18.180515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:68154507c543cf509374c50325fc5dd2bc873cbd4fdbae24d8cc38a16b393024

Observation fa1a4aef-7966-4b6b-a468-8273d5052685 · outbound

This paper cites Turn-taking in Conversational Systems and Human- Robot Interaction: A Review,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Turn-taking in Conversational Systems and Human- Robot Interaction: A Review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.382925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:e8e314f5fa19cf4614bdc4f14245ec1a9cafec5a2c55863389b0b6977a8242ae

Observation 9466d6d4-c7cf-44d1-8baa-558880af1735 · outbound

This paper cites Data-driven models for timing feedback responses in a Map Task dialogue system,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Data-driven models for timing feedback responses in a Map Task dialogue system,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.356589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:aa70b5a34a103d31eb44ec50aa38c460f23cd84b7d1b0e70bbb2006ad99d3769

Observation 147ab622-c0ef-4294-83e1-e929682a9059 · outbound

This paper cites On temporal aspects of turn taking in conversational dialogues,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders On temporal aspects of turn taking in conversational dialogues,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.403311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:78931aaf77f4782c65a9a26d4a2041197a32b9f0e46a642cfb04cd16bb617d1e

Observation 5a3cfe38-9e60-465c-8841-10c1474ca5f6 · outbound

This paper cites Turn-Taking Modelling in Conversational Systems: A Review of Recent Advances,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Turn-Taking Modelling in Conversational Systems: A Review of Recent Advances,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.386459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:8dd45d6694879551c1a18a124fabdc1c05e402bc369a423750f4086629b1ddd1

Observation ad296677-e421-4ec5-a49f-bad88b7a66d7 · outbound

This paper cites Pauses, gaps and overlaps in conversa- tions,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Pauses, gaps and overlaps in conversa- tions,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.388335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:062373f9cbefd538bfe2271b75134c3cba54f759d00fa19738e9def108ab9a5c

Observation 1782f32b-0171-4208-a4ca-efc617d8e4a3 · outbound

This paper cites Timing in turn-taking and its impli- cations for processing models of language,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Timing in turn-taking and its impli- cations for processing models of language,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.400964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:87daf64e6126121ac55f9b0b5128fbefcdd73eb5b3b68ad38ecb11de119743d6

Observation 52335817-0e8d-4365-930c-919e49bd33b4 · outbound

This paper cites Timed picture naming in seven languages,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Timed picture naming in seven languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.381221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:15b40531935df1371d4a34d04fa3c975a89414463a893258d638b9ac8557d480

Observation a5cecb58-38eb-49ad-a568-5d9f1e1f671d · outbound

This paper cites Multimodal Turn Analysis and Prediction for Multi-party Conversations,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Turn Analysis and Prediction for Multi-party Conversations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.377072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:97a5cb2f9a17ad59c7491d9c8fe3b74bc253a30172a62c5b73aba3e346a54b53

Observation 4bc0f3d5-efba-410c-8251-d83d4b2439f7 · outbound

This paper cites Voice Activity Projection: Self-supervised Learning of Turn-taking Events.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection: Self-supervised Learning of Turn-taking Events

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.178117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:9da52faa5280b55abaa61963e96c35602bc1aea467f33ec5fa20f9ec3f04a4cf

Observation 8ccd277d-0bec-4730-98fa-10db1f7dcc45 · outbound

This paper cites Multimodal V oice Activity Projection for Turn-Taking and Effects on Speaker Adaptation,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal V oice Activity Projection for Turn-Taking and Effects on Speaker Adaptation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.375030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:8c1fab9dafeda3dac4e308b14d6f24f9b29b4e10a44e0f11ee6ca687e6af279b

Observation 470d4985-fcdd-40ae-960d-58c68db749cc · outbound

This paper cites Voice Activity Projection Model with Multimodal Encoders.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.175063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:1acb8cec00ae9ec6e0d55b7afb1c7630f0577d32eabed913233bf75ab0ca5183

Observation 4ad90fe7-4bca-4639-8d5a-dc3befd6f0f8 · outbound

This paper cites Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.187022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:2727bc56c4a28f5ed1443796512defcb6c1dbff46bfaa6526fda298265f79fa6

Observation 8e2c3d74-7b8b-46de-84b1-d06f7598c9ca · outbound

This paper cites Behind the scenes: Mechanistic interpretability of lora-adapted whis- per for speech emotion recognition.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Behind the scenes: Mechanistic interpretability of lora-adapted whis- per for speech emotion recognition

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:16:18.189958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:7866a814a5d31b10027ab33b4bfefdf25d0f24d6a5692cafa0a1a8191eee82f3

Observation 81a4e9cf-4fcc-43ee-8751-69b055435dd9 · outbound

This paper cites GRPO- Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multi- modal Emotion Recognition,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders GRPO- Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multi- modal Emotion Recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.381415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:320090b651bebea9bf1e9e875664662feb45ae69c4052ff600a3c047527715a3

Observation 5ffdd3e1-feae-44d4-8b2e-af516f9e1ffa · outbound

This paper cites LoRA- Whisper: Parameter-Efficient and Extensible Multilingual ASR,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders LoRA- Whisper: Parameter-Efficient and Extensible Multilingual ASR,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.396659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:e6ec26d346c3a3d8c640778107bd8a112459ed7020557575994bf652f1a30a4c

Observation 627f9364-7480-480c-a8e3-e381cc071047 · outbound

This paper cites M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.189364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:8f558fa485f55751150ec5cad084b4e44f6512b1b6d9a32b079c827b8c039783

Observation 60165d4b-3bf1-426a-abff-8715198e1dab · outbound

This paper cites Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.371418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5f9351c27d466fb5aec56604772bb0067020c248cf0d70554392c6bab7105656

Observation e56bc495-1146-412f-a0d5-7d6b5ef30de5 · outbound

This paper cites Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.192224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:69f8d7f07f547cb7af25697d9d7abd168309a3c63d9c8c31a0dbeff65d0aa234

Observation 4692176c-5662-45d6-b5ac-c5fa00ca9c1c · outbound

This paper cites Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.373311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:dcf8e313f9ee7b4198f81dd6748c549da279eeac1e1cabdd5908ea49a9b76ca0

Observation a6bd654a-fe3d-425b-9f65-b98a84a59c74 · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.339998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5a8fbfb7bc73c2fd76bc5a9aa3fc64e430befffc1b8aa8bd4d13b3c00a3e7ab1

Observation 8823cca0-f77c-4553-b411-afb1410635fd · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.184153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:78900285aac1fa704f6e07fd8943d0ab01d7b061c9386cd21276b72c2482b8a1

Observation f39153ad-0790-4acc-a295-d007ff5aaa2a · outbound

This paper cites Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.166295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:6c641f865f65f64a9ddb11f86359247a0a486118064da79ca530d00a93715719

Observation 9eef50a1-4531-47c6-904d-92b43f8c6444 · outbound

This paper cites Triadic Multi- party V oice Activity Projection for Turn-taking in Spoken Dia- logue Systems,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Triadic Multi- party V oice Activity Projection for Turn-taking in Spoken Dia- logue Systems,

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:16:18.195382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:e5af3b7dbac76c7dd6b55e1d009e0a588a5f47e80b8f2ae311dddb5f14de61d3

Observation 6364f2d9-d724-4985-9d88-3fbf7c07ef6e · outbound

This paper cites Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of V oice Activity Projection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of V oice Activity Projection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.363896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:97b4719fe19392f280bb78489826d61f87cfc65113bb3c42abf392761287d3bc

Observation 798ee106-eef7-402b-ad6e-c44d0ada81c3 · outbound

This paper cites Predicting End-of- turn and Backchannel Based on Multimodal V oice Activity Prediction Model,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Predicting End-of- turn and Backchannel Based on Multimodal V oice Activity Prediction Model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.365940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:90c00ec8e66a7d74c133ad8f079a88f7c80b4f50bbf77910b213e4d12234b11c

Observation efab9e88-132a-48c6-bc93-38cca0e4d9a0 · outbound

This paper cites Multi- lingual Turn-taking Prediction Using V oice Activity Projection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multi- lingual Turn-taking Prediction Using V oice Activity Projection,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.366437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:adc73c85729d7321cb3907b6832e96c33db0b93ea0223f067428247ec0c582cd

Observation deaf6d46-f308-4e15-b733-d528c8ebf0c3 · outbound

This paper cites Multilingual Turn-taking Prediction Using Voice Activity Projection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multilingual Turn-taking Prediction Using Voice Activity Projection

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.174784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:31f5cd147ff99bd2ce976616c4e4d956fd7251e2da351834067d95c379936b67

Observation 43e6bba6-d75f-40ce-b5e1-dae18d5ae313 · outbound

This paper cites Investigating the Language Independence of V oice Activity Projection Models through Stan- dardization of Speech Segmentation Labels,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Investigating the Language Independence of V oice Activity Projection Models through Stan- dardization of Speech Segmentation Labels,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.353223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:e850bacc6685bbe1bcf80114d237e99bc54a3f0ee5ca177b6ea4baf8323232e7

Observation e2e4601e-d5ba-4c72-b71c-0438836b5697 · outbound

This paper cites Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.172565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:370f2da5eca40538b157cf329b9d62151e5d32a93c944e15847ecd9647de9576

Observation 15f01bfe-0936-4d00-88a9-dc5b1ca7f709 · outbound

This paper cites Applying General Turn-Taking Models to Conversational Human-Robot Interaction,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Applying General Turn-Taking Models to Conversational Human-Robot Interaction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.349501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:97b2169a929ddf3bd92af43e8ccc23c8bb3509c8180662e51524c791dd913d7a

Observation 1a211102-c65a-442b-9e0e-168a5460d3fc · outbound

This paper cites A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.383096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b1ab013fd77525dec681a1334625f7c55bc7994878a473be6972bdd08c0e5880

Observation 7534f11c-3d9f-4270-8692-5ce08e54edea · outbound

This paper cites Multimodal V oice Activity Prediction: Turn-taking Events Detection in Expert-Novice Conver- sation,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal V oice Activity Prediction: Turn-taking Events Detection in Expert-Novice Conver- sation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.392045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:d37e1853094251874f2d2efa4468411028ed345461f89fa1c98639c1fd9c318e

Observation fbb30dbd-9767-496d-9c1c-d43bcc6e0859 · outbound

This paper cites The NoXi database: multimodal recordings of mediated novice-expert interactions,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders The NoXi database: multimodal recordings of mediated novice-expert interactions,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.361970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:f138fea866faf8f007e9acfe1cca5be79cf06a50d7f10b45b813b23c511dfe22

Observation 20d8a772-8675-41ad-9df8-c90507514cfb · outbound

This paper cites Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on Engagement.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on Engagement

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.186867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:ba5b9db0da4482947a7f790c31e63344f588818c4d0bde3e76de2614e2c5883e

Pith citing papers

No inbound Pith citation observations are available.