Pith. sign in

Paper Citation Record · LEDGER

Voice Activity Projection Model with Multimodal Encoders

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.03980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03980 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:34.901221Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:30.784001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.173817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact2
  • verified fuzzy44
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · outbound

This paper cites Voice Activity Projection Model with Multimodal Encoders.

Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:30.784001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:30.784001Z digest=sha256:a689de9f1ac1e0c4877ee7ff059ccad7b80343656e7306ba4232359130beb028

Observation 73b51346-a3ec-4558-a0b3-360959c5f71a · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:43.054931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:30.879347Z digest=sha256:e01a6f42f7f1b284a8f7c3821017afde220085744d44ae77a894c19e27628211

Observation aa716f32-312b-415a-a2b2-ae805ea34828 · outbound

This paper cites Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units.

Voice Activity Projection Model with Multimodal Encoders Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:30.968030Z digest=sha256:da560fc2f0ca4b9dee634663e64519b698e607f2898c2b4488f061031ec0bc91

Observation c6cd6bf3-e16e-41ee-9126-6c918ecce3f2 · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:42.792038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.049890Z digest=sha256:c553dc3c511755f6d5e9cda177c95c50443adb0b8c42dfdf8eb4edf90af3d7b9

Observation 224039b8-c633-4b71-a07f-f0d11163178b · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:42.629601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.162124Z digest=sha256:7c9b843f5a120e9b48e8f8cf311daef59c3a32de5c3c30c3473a8552ecdeeaf9

Observation 4cdeaaba-ed8d-459b-bdc4-8419da09eb7c · outbound

This paper cites Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities.

Voice Activity Projection Model with Multimodal Encoders Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.446772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.240357Z digest=sha256:ec6549835962d3345f232d24a3ac8ac33f6132262e7c101f96bd2e4b25fe5c6f

Observation 3ed24270-7e68-43f5-96e1-a5766682c055 · outbound

This paper cites Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37].

Voice Activity Projection Model with Multimodal Encoders Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37]

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.334883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.302301Z digest=sha256:1a2a67c5048ba01cecb86d024dd4abc43dd0552e0038c562b3a243f8a03d051c

Observation 780aeee2-700b-4651-a456-3398cd8e9303 · outbound

This paper cites For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder.

Voice Activity Projection Model with Multimodal Encoders For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.153637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.386731Z digest=sha256:bddd35b096c8c0353ac7b4d2983770ac2582a3a03cfaa8c7cc9c5901ad233f5d

Observation 60c487df-7565-4fb3-ade1-fefc984fdf05 · outbound

This paper cites By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models.

Voice Activity Projection Model with Multimodal Encoders By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.009648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.455482Z digest=sha256:575518eeb5e0511c2c859f6a365f7847fe9d2f4abf30036ef778f54ad670981d

Observation eb5de9e2-133c-41a7-8d17-4924acea615d · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:41.838076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.536151Z digest=sha256:87985fc2c54d73d9d674df3691b91e7d34c051c0211a3bb2b03ff7ccf3735df7

Observation 62bce116-9f4b-44ee-9713-047f6d3fcf15 · outbound

This paper cites A simplest systemat- ics for the organization of turn-taking for conversation,.

Voice Activity Projection Model with Multimodal Encoders A simplest systemat- ics for the organization of turn-taking for conversation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.696757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.603372Z digest=sha256:ba449bb265c66005abedfc88d0cd3877c798f8fd5e0ddea3dd795b30de34bc04

Observation 5b02f3d6-2f7a-4823-a7dd-4c68fde28c3d · outbound

This paper cites Turn-taking: A critical analysis of the research tradition,.

Voice Activity Projection Model with Multimodal Encoders Turn-taking: A critical analysis of the research tradition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.521561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.672711Z digest=sha256:3991b71f75ab06f0077c456e66d3785740a2d6e45890f1426766254a0ea1f1a5

Observation 591fbe38-ee34-4ebf-9c78-a272121bf8a0 · outbound

This paper cites Some signals and rules for taking speaking turns in conversations,.

Voice Activity Projection Model with Multimodal Encoders Some signals and rules for taking speaking turns in conversations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.371958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.768806Z digest=sha256:c11ab525d43699df575e1b7d890e5c3828a18eacf71439eff5c2251d05bc37bf

Observation aec0d1ca-d0be-4ae3-ad90-f445f9f4a695 · outbound

This paper cites Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,.

Voice Activity Projection Model with Multimodal Encoders Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.256020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.848267Z digest=sha256:a3ac5e8647a3725783677ad9f800b423c4b1a0331df32f169559e567edaeb6a0

Observation e1f26d21-f463-40f4-91fc-2084d25146de · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:41.114513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.931068Z digest=sha256:90e08af1dd715b1b2f1b87e068d7e5515d93da7eed7515dd4bba359774e35320

Observation cd7c379f-dd7d-47c8-a25b-bde9899b98af · outbound

This paper cites Using uh and um in spontaneous speaking,.

Voice Activity Projection Model with Multimodal Encoders Using uh and um in spontaneous speaking,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.996352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:31.990975Z digest=sha256:09f4dc2731baad121aaf5242e00ff732e769f3e60404c5579f14225e232384f5

Observation 09b69a46-189d-4ea2-83fb-766c2db89b85 · outbound

This paper cites Using prosodic clues to decide when to produce back- channel utterances,.

Voice Activity Projection Model with Multimodal Encoders Using prosodic clues to decide when to produce back- channel utterances,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.844132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.071262Z digest=sha256:3ffc884591f32934b3dddbb4ecf279134e9de6f347a425fac65016751efd7499

Observation ccd90798-8833-4eb4-be36-b1610fea83a1 · outbound

This paper cites Nonverbal behaviours improving a simulation of small group discussion,.

Voice Activity Projection Model with Multimodal Encoders Nonverbal behaviours improving a simulation of small group discussion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.712473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.132742Z digest=sha256:8398542ff69a3126a67dfb12858d5b88bb953d7f110ff9ff34d4f7eb819a1dda

Observation 13f5be9a-d41e-4f27-8a59-05e5a9ac65c3 · outbound

This paper cites Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,.

Voice Activity Projection Model with Multimodal Encoders Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.564372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.220226Z digest=sha256:d8414857f38d72de01b15530b31cd5c053b6706f76051a52086b09d0696547fb

Observation 63bf7315-b13f-4021-ac0c-781f5dc7fbcb · outbound

This paper cites The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,.

Voice Activity Projection Model with Multimodal Encoders The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.417748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.302198Z digest=sha256:adca2b3171ef871363f29298a846acd4c4074a947ec420bfe5cac371e33122ae

Observation d6ffe9f2-4f0a-47e6-92d8-2091d2c9dec8 · outbound

This paper cites Now or when? interrup- tion timing prediction in dyadic interaction,.

Voice Activity Projection Model with Multimodal Encoders Now or when? interrup- tion timing prediction in dyadic interaction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.299939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.384924Z digest=sha256:c9ac7cf499f8dd6cb971236bee49906c155996ea10282abd3b2ccbcf48ddb155

Observation f42f06c1-b3a3-4cfd-bf38-f36ac0dae25d · outbound

This paper cites How turn-taking strategies influence users’ impressions of an agent,.

Voice Activity Projection Model with Multimodal Encoders How turn-taking strategies influence users’ impressions of an agent,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.132579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.475213Z digest=sha256:eca455f0f9b42ce4d13fb6017b29d59e17e7fe4c5417f0c2373311888001773f

Observation 62270099-7a5e-4217-afb9-d21fc7f21aa8 · outbound

This paper cites Timing in turn-taking and its im- plications for processing models of language,.

Voice Activity Projection Model with Multimodal Encoders Timing in turn-taking and its im- plications for processing models of language,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.014896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.568382Z digest=sha256:f363c23187edfb6d53df1037b44b625e941b3503bfb062d86a4284a5a2181995

Observation bb06fffc-88c0-48bf-8da3-211339adf5c2 · outbound

This paper cites Timing in conversation,.

Voice Activity Projection Model with Multimodal Encoders Timing in conversation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.863966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.636080Z digest=sha256:4c747b0ab01d45f462b7bbd1837110af11923f872c53613d4a7bbdfea6299854

Observation df526a09-03c7-476f-95ea-b7f86eab929f · outbound

This paper cites Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,.

Voice Activity Projection Model with Multimodal Encoders Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.721820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.724770Z digest=sha256:bfd7c329a6dee6140a6d599c725e8608212f9b90aa7024b5ba45064aead08b36

Observation 4da0c1f6-3c35-4047-bdf6-26c2e15dce5f · outbound

This paper cites Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,.

Voice Activity Projection Model with Multimodal Encoders Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.546358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.843858Z digest=sha256:cab380803386d15bf8c5588f39d07e0f708999f7b7a318dccf500fb8883c8a08

Observation 0c39117f-2ccc-46af-b5d0-f9a8e5f281eb · outbound

This paper cites Attentive listening system with backchanneling, response generation and flexible turn-taking,.

Voice Activity Projection Model with Multimodal Encoders Attentive listening system with backchanneling, response generation and flexible turn-taking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.388384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.909303Z digest=sha256:ce7666f3c893a10d857d1e66b5accb4e6e440017990b72881fc3115a8cf24114

Observation 4397cd86-ad7e-4d7a-98d4-0b499497769d · outbound

This paper cites V oice activity projection: Self- supervised learning of turn-taking events,.

Voice Activity Projection Model with Multimodal Encoders V oice activity projection: Self- supervised learning of turn-taking events,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.187792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:32.990781Z digest=sha256:59646a10f26c5befce1c54501e9eb050d2cd60ad7de02c416a7c85312cce68f3

Observation 88e17781-a09a-4660-8764-15a4c9979cef · outbound

This paper cites Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,.

Voice Activity Projection Model with Multimodal Encoders Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.016839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.117340Z digest=sha256:443dadb53e0166883c7b8886ac6f9e29b63538f052dfbcc164fef2f875ed502a

Observation 72879ccf-5500-4aeb-9098-a3004ce9138c · outbound

This paper cites Predicting turn-taking by compact gazing transition patterns in multiparty conversation,.

Voice Activity Projection Model with Multimodal Encoders Predicting turn-taking by compact gazing transition patterns in multiparty conversation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.872661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.186043Z digest=sha256:1d73eb6a8d02174580d525a41785ed617a9d8ff9335262fe452c3cdfcee2c78f

Observation f9c1df9b-1a7d-4722-a582-d1657e5baa05 · outbound

This paper cites Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,.

Voice Activity Projection Model with Multimodal Encoders Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.700880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.270468Z digest=sha256:d9c8be9c2a2292764a5266e558232b7a3eb65a41ab2004208ba3541522fe043d

Observation cffd3c23-bf87-4e5b-a3c6-4e63e3116661 · outbound

This paper cites Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,.

Voice Activity Projection Model with Multimodal Encoders Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.574119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.345041Z digest=sha256:a56ca79e987264af0095b2cca7086c02371b6d2a1e2d4da3e29bd78a3848f8bf

Observation 27f3612d-ddcc-497f-9b80-59a3d55a0de0 · outbound

This paper cites Turn-taking and backchannel prediction with acoustic and large language model fusion,.

Voice Activity Projection Model with Multimodal Encoders Turn-taking and backchannel prediction with acoustic and large language model fusion,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.412562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.427059Z digest=sha256:dbf84c01e15c9fdb29d5ba9cba50acba81807c87098c482cc8c892e912a002e5

Observation ee2159a5-e41e-429c-b0ab-71b5bfd41980 · outbound

This paper cites Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,.

Voice Activity Projection Model with Multimodal Encoders Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.272095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.524622Z digest=sha256:a3c870a9ad9639ad89dfeff318b075e1daa9b06c6d02e0deacd9c0cd9b07d6c1

Observation 319c46e6-58d3-46f3-b227-99ef7960d4b4 · outbound

This paper cites Mini-omni: Language models can hear, talk while thinking in streaming,.

Voice Activity Projection Model with Multimodal Encoders Mini-omni: Language models can hear, talk while thinking in streaming,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.118055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.597594Z digest=sha256:a838d1312927c35fae9b14b8d5a4d3eeb210ca103f9db4f39cd220bf12f00b22

Observation 38d5a649-549b-4ec6-bb89-2e1661e190ac · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

Voice Activity Projection Model with Multimodal Encoders Moshi: a speech-text foundation model for real-time dialogue,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.955757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.663036Z digest=sha256:e9fc5f8de39212bc86f72cc0bf11f3af0e556528bc9331bea6872a29aae38edf

Observation ca140c6e-e660-48da-9588-d30903dab38a · outbound

This paper cites [Online].

Voice Activity Projection Model with Multimodal Encoders [Online]

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:53:35.452694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.737389Z digest=sha256:bb41c75bcadb477196a5e20e238b1c06000003ca32ccc9a1d66ab0efb73ec4ce

Observation ec01cbc1-9b8f-4645-8879-73d67df253f6 · outbound

This paper cites [Online].

Voice Activity Projection Model with Multimodal Encoders [Online]

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.793002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.808721Z digest=sha256:394f637fbe6519fb6f646857f4e20ad3fef85d5cc6f1f3a68cc836f1f5c5124d

Observation 881446a4-fe91-4cae-9141-96ada48bfbe2 · outbound

This paper cites Real-time and continuous turn-taking prediction using voice ac- tivity projection,.

Voice Activity Projection Model with Multimodal Encoders Real-time and continuous turn-taking prediction using voice ac- tivity projection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.615970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.886558Z digest=sha256:582220ebe453002a97ac989cdce60c2158ca890bbc8e4088aa614592113eda3f

Observation 3fc479d7-856d-45ad-b697-8ddd0a9c8624 · outbound

This paper cites Multilingual turn-taking prediction using voice activity projection,.

Voice Activity Projection Model with Multimodal Encoders Multilingual turn-taking prediction using voice activity projection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.419432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:33.981132Z digest=sha256:11212a716eb0d67a47c7580d79317d88413d3bfc879e83511f6ace1f4eef0e40

Observation e17c12c0-735c-4ac2-b140-e657a73e465a · outbound

This paper cites How much does prosody help turn- taking? investigations using voice activity projection models,.

Voice Activity Projection Model with Multimodal Encoders How much does prosody help turn- taking? investigations using voice activity projection models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.283172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.064405Z digest=sha256:5bb62f4544fe92a036a837dd9d13f7a4202fcb2adf62ba88801e4e340f25edd6

Observation 2e2e21d0-e993-4a8e-a422-81c281be1e99 · outbound

This paper cites OpenFace: An open source facial behavior analysis toolkit,.

Voice Activity Projection Model with Multimodal Encoders OpenFace: An open source facial behavior analysis toolkit,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.094480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.136834Z digest=sha256:17b4575ff2fe6ab5605aee43f2b3ce6d901a8308a8a74c7c88bd99fde212ab98

Observation 0a9cae1f-8d18-410a-888c-8affc0869fb9 · outbound

This paper cites Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,.

Voice Activity Projection Model with Multimodal Encoders Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.970605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.203584Z digest=sha256:f93753ef1eae3facca6fce376898906fdba6914136f71327c5d2d13ba3328c52

Observation e6fd51c1-dc02-4a73-b9ac-86fccdf24ffa · outbound

This paper cites Former-dfer: Dynamic facial expression recognition transformer,.

Voice Activity Projection Model with Multimodal Encoders Former-dfer: Dynamic facial expression recognition transformer,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.817813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.268089Z digest=sha256:ce7f35c1d28ba2d534e0aeeb0691d819739b1bcde736c26baf35adab10fcef57

Observation 4d281b4e-0199-4e08-b775-6c119c49253e · outbound

This paper cites Unsu- pervised pretraining transfers well across languages,.

Voice Activity Projection Model with Multimodal Encoders Unsu- pervised pretraining transfers well across languages,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.605534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.338299Z digest=sha256:6f29f69d572d3adc303f7850117c1df1844e24116fba4269b65a65436d962ebb

Observation 55d37418-757f-4d48-ba23-cd3102fd6f51 · outbound

This paper cites Dlib-ml: A machine learning toolkit,.

Voice Activity Projection Model with Multimodal Encoders Dlib-ml: A machine learning toolkit,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.456539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.416269Z digest=sha256:2463e5070546e0c9cc4e61e8186c80a1aad16eaa9bf9ebee277407a5304e901a

Observation 0f56658d-9e42-4ba9-923b-fac07bf14a51 · outbound

This paper cites The NoXi database: multimodal recordings of mediated novice-expert interactions,.

Voice Activity Projection Model with Multimodal Encoders The NoXi database: multimodal recordings of mediated novice-expert interactions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.282903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.513926Z digest=sha256:23182169641b924c02ab82bbfce0e7dbb8417add380495c4701a2c8b1eb1fb4e

Observation 4cdced7e-395a-4d45-a249-982e7f1df638 · outbound

This paper cites PyTorch Lightning,.

Voice Activity Projection Model with Multimodal Encoders PyTorch Lightning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.137530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.593139Z digest=sha256:0017159b4b5eb350677f6f6d943c839ba56a4083dd90f906e3d938992f707250

Observation 73512c0f-2b7e-4b8b-9385-91b48365d77a · outbound

This paper cites Discourse as an interactional achievement iii: The omnirelevance of action,.

Voice Activity Projection Model with Multimodal Encoders Discourse as an interactional achievement iii: The omnirelevance of action,

Reference 50

Resolution
verified exact
doi, observed 2026-08-07T10:53:35.150121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.751806Z digest=sha256:afc6947b2a3a18eb973e368815a9e10a2ff999aa96cf2edaff0091b85812a7a4

Observation eda6959a-4dde-4ea0-ab68-e3e9ac2710a4 · outbound

This paper cites Between and within: Alternative sequential treat- ments of continuers and assessments,.

Voice Activity Projection Model with Multimodal Encoders Between and within: Alternative sequential treat- ments of continuers and assessments,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.828855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.819433Z digest=sha256:9aca7c70356a35382173fcbd03293b1da42cfd1876c73f9d192a1dfbaa9f62e6

Observation 5ab26f48-cf14-4575-a333-2892af7ef0f5 · outbound

This paper cites Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,.

Voice Activity Projection Model with Multimodal Encoders Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.653190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.901221Z digest=sha256:44a41917972f3a77cc071bc9a628abb78f95ba0a197e8e2ba41dc9fbd79977f0

Observation fba6514b-15e3-4dd3-90ba-c469c37b185a · outbound

This paper cites Available: https://github.com/Lightning-AI/ pytorch-lightning.

Voice Activity Projection Model with Multimodal Encoders Available: https://github.com/Lightning-AI/ pytorch-lightning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.986850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:53:34.665269Z digest=sha256:6f4f8891d83de4aad10822d6cd4337be21e45ab345bafd71229ea9eb2460453e

Pith citing papers

Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · inbound

Voice Activity Projection Model with Multimodal Encoders cites this paper.

Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:30.784001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:30.784001Z digest=sha256:a689de9f1ac1e0c4877ee7ff059ccad7b80343656e7306ba4232359130beb028

Observation 470d4985-fcdd-40ae-960d-58c68db749cc · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.175063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:ece92d8daa1204512c77317bae60c7dfb1adc3801bf92ce619389c8394a98f5c