Pith. sign in

Paper Citation Record · LEDGER

Voice Activity Projection Model with Multimodal Encoders

As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.03980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03980 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:34.901221Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:30.784001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.173817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact2
  • verified fuzzy44
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · outbound

This paper cites Voice Activity Projection Model with Multimodal Encoders.

Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:30.784001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:30.784001Z digest=sha256:ee81424b8a182179f6389d293431df03a4fd86629ff1471dec1ecb608b5d0e8c

Observation 73b51346-a3ec-4558-a0b3-360959c5f71a · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:43.054931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:30.879347Z digest=sha256:79d30214a83db7c77eb6e65206a05548a53bdf3602a7dcac75aa7a604ece01e8

Observation aa716f32-312b-415a-a2b2-ae805ea34828 · outbound

This paper cites Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units.

Voice Activity Projection Model with Multimodal Encoders Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:30.968030Z digest=sha256:2d62c734d6312a712eca16601661ee5d975141d6fefd3fe1d8da8149eda29ecf

Observation c6cd6bf3-e16e-41ee-9126-6c918ecce3f2 · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:42.792038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.049890Z digest=sha256:0888a4e42d45613118e2e25ee724159a1178b0c2cdaaf96de9b2eaf3884193ca

Observation 224039b8-c633-4b71-a07f-f0d11163178b · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:42.629601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.162124Z digest=sha256:c848546a34462d3d6c26c11fac468d0c3a12a0a46f66afa5d9488b5a73991056

Observation 4cdeaaba-ed8d-459b-bdc4-8419da09eb7c · outbound

This paper cites Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities.

Voice Activity Projection Model with Multimodal Encoders Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.446772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.240357Z digest=sha256:b6a6cf4286a025f64ba6d3c5e1e966c530952b3bfe1c19cc59214a60118aeb0f

Observation 3ed24270-7e68-43f5-96e1-a5766682c055 · outbound

This paper cites Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37].

Voice Activity Projection Model with Multimodal Encoders Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37]

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.334883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.302301Z digest=sha256:37f6ff6eff0f19539c341dbb9078125d7ed24bd7b33f89fff785d7114757383e

Observation 780aeee2-700b-4651-a456-3398cd8e9303 · outbound

This paper cites For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder.

Voice Activity Projection Model with Multimodal Encoders For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.153637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.386731Z digest=sha256:61e2f651dc891b336f46e00756ee7b554028bcbd791302e38800a40e23af67db

Observation 60c487df-7565-4fb3-ade1-fefc984fdf05 · outbound

This paper cites By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models.

Voice Activity Projection Model with Multimodal Encoders By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.009648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.455482Z digest=sha256:d81fde6c346eb91637aada00d956e1a728bcb624b58f079aa62cdd9cc374f8c0

Observation eb5de9e2-133c-41a7-8d17-4924acea615d · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:41.838076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.536151Z digest=sha256:5b41f905dff173eb761facefe28c0d9da9faf8611d5db00b1e03fce355650be3

Observation 62bce116-9f4b-44ee-9713-047f6d3fcf15 · outbound

This paper cites A simplest systemat- ics for the organization of turn-taking for conversation,.

Voice Activity Projection Model with Multimodal Encoders A simplest systemat- ics for the organization of turn-taking for conversation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.696757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.603372Z digest=sha256:c6cfaf6f97eadb5df7c9eb6c262f1eda6f0e73b449aae644b62dc480e331c9d0

Observation 5b02f3d6-2f7a-4823-a7dd-4c68fde28c3d · outbound

This paper cites Turn-taking: A critical analysis of the research tradition,.

Voice Activity Projection Model with Multimodal Encoders Turn-taking: A critical analysis of the research tradition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.521561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.672711Z digest=sha256:2d7bd8f96c46d0487c752ed795690e482d0eaf31fb3267077919aff5161feb85

Observation 591fbe38-ee34-4ebf-9c78-a272121bf8a0 · outbound

This paper cites Some signals and rules for taking speaking turns in conversations,.

Voice Activity Projection Model with Multimodal Encoders Some signals and rules for taking speaking turns in conversations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.371958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.768806Z digest=sha256:a6e0d20562fac3699c86a34f2096f14cf4b403dda4421c083c9b7e50525efc06

Observation aec0d1ca-d0be-4ae3-ad90-f445f9f4a695 · outbound

This paper cites Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,.

Voice Activity Projection Model with Multimodal Encoders Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.256020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.848267Z digest=sha256:b4bf2945f93cb8b94981b2cfaa96d6885dcdb0b462a566f52a225614ef32993f

Observation e1f26d21-f463-40f4-91fc-2084d25146de · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:41.114513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.931068Z digest=sha256:0d778d36c2542aaba22c8cb6a38613489a790df06de9e093aff16c320d2be43f

Observation cd7c379f-dd7d-47c8-a25b-bde9899b98af · outbound

This paper cites Using uh and um in spontaneous speaking,.

Voice Activity Projection Model with Multimodal Encoders Using uh and um in spontaneous speaking,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.996352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:31.990975Z digest=sha256:8f426534b8123ae6dfc4579ce1f347e7a316b3cbe3674c215a12bc0ab7bcc7bd

Observation 09b69a46-189d-4ea2-83fb-766c2db89b85 · outbound

This paper cites Using prosodic clues to decide when to produce back- channel utterances,.

Voice Activity Projection Model with Multimodal Encoders Using prosodic clues to decide when to produce back- channel utterances,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.844132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.071262Z digest=sha256:d98c803b046b190bbd83170de9a7775a8d4565a2ed10ede4b234d915520fb6b4

Observation ccd90798-8833-4eb4-be36-b1610fea83a1 · outbound

This paper cites Nonverbal behaviours improving a simulation of small group discussion,.

Voice Activity Projection Model with Multimodal Encoders Nonverbal behaviours improving a simulation of small group discussion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.712473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.132742Z digest=sha256:d4c2d31380b25dee054ee6889afd7520f4812b65d616891dec9a7f337af4d71d

Observation 13f5be9a-d41e-4f27-8a59-05e5a9ac65c3 · outbound

This paper cites Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,.

Voice Activity Projection Model with Multimodal Encoders Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.564372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.220226Z digest=sha256:d43b465aac3fe689172caa56f13a24b6e33f6afc546220a3979f71e5c643d9df

Observation 63bf7315-b13f-4021-ac0c-781f5dc7fbcb · outbound

This paper cites The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,.

Voice Activity Projection Model with Multimodal Encoders The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.417748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.302198Z digest=sha256:c9fb7bf31aa21e4832ad4db83330595bac67e21ffe92998043804c790e818ee1

Observation d6ffe9f2-4f0a-47e6-92d8-2091d2c9dec8 · outbound

This paper cites Now or when? interrup- tion timing prediction in dyadic interaction,.

Voice Activity Projection Model with Multimodal Encoders Now or when? interrup- tion timing prediction in dyadic interaction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.299939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.384924Z digest=sha256:ef888e518be9f36c968a2f229f248cb14acbdceedb7dfe32f3e3275ae70ce7f0

Observation f42f06c1-b3a3-4cfd-bf38-f36ac0dae25d · outbound

This paper cites How turn-taking strategies influence users’ impressions of an agent,.

Voice Activity Projection Model with Multimodal Encoders How turn-taking strategies influence users’ impressions of an agent,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.132579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.475213Z digest=sha256:61607d1278d6a4b257f7d4c5118f3ac1cb7ae2065bcbe019ed444620882d94d5

Observation 62270099-7a5e-4217-afb9-d21fc7f21aa8 · outbound

This paper cites Timing in turn-taking and its im- plications for processing models of language,.

Voice Activity Projection Model with Multimodal Encoders Timing in turn-taking and its im- plications for processing models of language,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.014896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.568382Z digest=sha256:7e8fbe4ac9388d0fd1d7c94cf2556ee8dd3d00727aecea70753e2b8e693ec620

Observation bb06fffc-88c0-48bf-8da3-211339adf5c2 · outbound

This paper cites Timing in conversation,.

Voice Activity Projection Model with Multimodal Encoders Timing in conversation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.863966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.636080Z digest=sha256:49ccaa97eeea73d7f7a69311c2504ef9e6778cae066498bd1563bbbaf708293c

Observation df526a09-03c7-476f-95ea-b7f86eab929f · outbound

This paper cites Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,.

Voice Activity Projection Model with Multimodal Encoders Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.721820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.724770Z digest=sha256:789d926ba44031220cf2eff0a2aff5b719434a2a98eefb742e0fbdb963acc228

Observation 4da0c1f6-3c35-4047-bdf6-26c2e15dce5f · outbound

This paper cites Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,.

Voice Activity Projection Model with Multimodal Encoders Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.546358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.843858Z digest=sha256:735650a80f51bca9f2b726b20294b5697e9aebefe13e116d72d8bee1a51f5029

Observation 0c39117f-2ccc-46af-b5d0-f9a8e5f281eb · outbound

This paper cites Attentive listening system with backchanneling, response generation and flexible turn-taking,.

Voice Activity Projection Model with Multimodal Encoders Attentive listening system with backchanneling, response generation and flexible turn-taking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.388384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.909303Z digest=sha256:c6c024b960331a51d58b451486bd81d682b913170ed749c1d75e4af2a1abfa5a

Observation 4397cd86-ad7e-4d7a-98d4-0b499497769d · outbound

This paper cites V oice activity projection: Self- supervised learning of turn-taking events,.

Voice Activity Projection Model with Multimodal Encoders V oice activity projection: Self- supervised learning of turn-taking events,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.187792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:32.990781Z digest=sha256:d990007e79024b3d777eaf596239514b18c7e26fa46e1e5912b4754aff4924be

Observation 88e17781-a09a-4660-8764-15a4c9979cef · outbound

This paper cites Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,.

Voice Activity Projection Model with Multimodal Encoders Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.016839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.117340Z digest=sha256:01c375afae00e26f61f62fd057fe4c0863e0249e0a524ef308b0f600efb030ec

Observation 72879ccf-5500-4aeb-9098-a3004ce9138c · outbound

This paper cites Predicting turn-taking by compact gazing transition patterns in multiparty conversation,.

Voice Activity Projection Model with Multimodal Encoders Predicting turn-taking by compact gazing transition patterns in multiparty conversation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.872661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.186043Z digest=sha256:887a7ecaa738c4ad60a0757d86442843896f2180524cd0e74b3fcda2c4e876be

Observation f9c1df9b-1a7d-4722-a582-d1657e5baa05 · outbound

This paper cites Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,.

Voice Activity Projection Model with Multimodal Encoders Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.700880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.270468Z digest=sha256:d1d14f2b30ccac4d626dd2a215c4ff087e434b3a92013a20b18bce0d5e380338

Observation cffd3c23-bf87-4e5b-a3c6-4e63e3116661 · outbound

This paper cites Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,.

Voice Activity Projection Model with Multimodal Encoders Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.574119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.345041Z digest=sha256:d7bb9e4379fb37395f09aeaac91952eb9d4a3c09b22154eb7bd2018e07f983a7

Observation 27f3612d-ddcc-497f-9b80-59a3d55a0de0 · outbound

This paper cites Turn-taking and backchannel prediction with acoustic and large language model fusion,.

Voice Activity Projection Model with Multimodal Encoders Turn-taking and backchannel prediction with acoustic and large language model fusion,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.412562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.427059Z digest=sha256:d209a77d613933d795917f43b6e050e647a221fc0148f4861835737cc5dfcc47

Observation ee2159a5-e41e-429c-b0ab-71b5bfd41980 · outbound

This paper cites Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,.

Voice Activity Projection Model with Multimodal Encoders Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.272095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.524622Z digest=sha256:d3b101542ba1b558f9c500a4a44d84662f4694f0b54c1c234a2f80c19ed215c3

Observation 319c46e6-58d3-46f3-b227-99ef7960d4b4 · outbound

This paper cites Mini-omni: Language models can hear, talk while thinking in streaming,.

Voice Activity Projection Model with Multimodal Encoders Mini-omni: Language models can hear, talk while thinking in streaming,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.118055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.597594Z digest=sha256:e017f12ece773b83fc34d217f83bd8643892e9a351b056375e7834997e9b3ed0

Observation 38d5a649-549b-4ec6-bb89-2e1661e190ac · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

Voice Activity Projection Model with Multimodal Encoders Moshi: a speech-text foundation model for real-time dialogue,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.955757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.663036Z digest=sha256:9e0afb499873813b7ea7d2908f1b952e3e588425ce5f87bab411e8b3ccdea4d9

Observation ca140c6e-e660-48da-9588-d30903dab38a · outbound

This paper cites [Online].

Voice Activity Projection Model with Multimodal Encoders [Online]

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:53:35.452694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.737389Z digest=sha256:7d9ce04c96b7af9f6b01e910210c80ad267deb3bc98cf73f3ccdeada8ca64b9a

Observation ec01cbc1-9b8f-4645-8879-73d67df253f6 · outbound

This paper cites [Online].

Voice Activity Projection Model with Multimodal Encoders [Online]

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.793002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.808721Z digest=sha256:2eb7d3348fa239f7cab337b22d0a1d82768869bcda172e49a837204e6e16bc90

Observation 881446a4-fe91-4cae-9141-96ada48bfbe2 · outbound

This paper cites Real-time and continuous turn-taking prediction using voice ac- tivity projection,.

Voice Activity Projection Model with Multimodal Encoders Real-time and continuous turn-taking prediction using voice ac- tivity projection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.615970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.886558Z digest=sha256:87746eb1c1b9dbda66096d7bb98a2df034e9faf2111cfed8e67f760db67aaefe

Observation 3fc479d7-856d-45ad-b697-8ddd0a9c8624 · outbound

This paper cites Multilingual turn-taking prediction using voice activity projection,.

Voice Activity Projection Model with Multimodal Encoders Multilingual turn-taking prediction using voice activity projection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.419432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:33.981132Z digest=sha256:44330cd1d1e00721fe4dba3c83059d8ae5bf7529e9582213b7c38f5033c64278

Observation e17c12c0-735c-4ac2-b140-e657a73e465a · outbound

This paper cites How much does prosody help turn- taking? investigations using voice activity projection models,.

Voice Activity Projection Model with Multimodal Encoders How much does prosody help turn- taking? investigations using voice activity projection models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.283172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.064405Z digest=sha256:dc3bbb220b43d82e8e0a98d177d030a5155854f4ae6589bc17a03e8007b0a3cf

Observation 2e2e21d0-e993-4a8e-a422-81c281be1e99 · outbound

This paper cites OpenFace: An open source facial behavior analysis toolkit,.

Voice Activity Projection Model with Multimodal Encoders OpenFace: An open source facial behavior analysis toolkit,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.094480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.136834Z digest=sha256:23210e02f299175a0ff5a4de3d728fbc63e3534df2421ea0593b2cd26623f3ac

Observation 0a9cae1f-8d18-410a-888c-8affc0869fb9 · outbound

This paper cites Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,.

Voice Activity Projection Model with Multimodal Encoders Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.970605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.203584Z digest=sha256:4c26b31147a1ae720077d274eaba3d8240e20b098a9a1b8a82b01308a746bcde

Observation e6fd51c1-dc02-4a73-b9ac-86fccdf24ffa · outbound

This paper cites Former-dfer: Dynamic facial expression recognition transformer,.

Voice Activity Projection Model with Multimodal Encoders Former-dfer: Dynamic facial expression recognition transformer,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.817813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.268089Z digest=sha256:c9e8d0313f50901f438375b9a7b98734c339fd78ff1ffc34d7c88d668b0e8062

Observation 4d281b4e-0199-4e08-b775-6c119c49253e · outbound

This paper cites Unsu- pervised pretraining transfers well across languages,.

Voice Activity Projection Model with Multimodal Encoders Unsu- pervised pretraining transfers well across languages,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.605534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.338299Z digest=sha256:94fa8a545e12313fa23abab0d331cd44eb5ccd691b7b98b31714d95283d545e2

Observation 55d37418-757f-4d48-ba23-cd3102fd6f51 · outbound

This paper cites Dlib-ml: A machine learning toolkit,.

Voice Activity Projection Model with Multimodal Encoders Dlib-ml: A machine learning toolkit,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.456539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.416269Z digest=sha256:1b334ac76073fa5511e2274e2a36e709c7870642af57e60e313f487b410fb175

Observation 0f56658d-9e42-4ba9-923b-fac07bf14a51 · outbound

This paper cites The NoXi database: multimodal recordings of mediated novice-expert interactions,.

Voice Activity Projection Model with Multimodal Encoders The NoXi database: multimodal recordings of mediated novice-expert interactions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.282903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.513926Z digest=sha256:db7afa6daf57332d685b0778f3ac97ab254b375fcebf5221a31c55f0408789b2

Observation 4cdced7e-395a-4d45-a249-982e7f1df638 · outbound

This paper cites PyTorch Lightning,.

Voice Activity Projection Model with Multimodal Encoders PyTorch Lightning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.137530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.593139Z digest=sha256:49179eafa634aa55f26f0461138ef192b2d004107f748939d6952f961d55a0ed

Observation 73512c0f-2b7e-4b8b-9385-91b48365d77a · outbound

This paper cites Discourse as an interactional achievement iii: The omnirelevance of action,.

Voice Activity Projection Model with Multimodal Encoders Discourse as an interactional achievement iii: The omnirelevance of action,

Reference 50

Resolution
verified exact
doi, observed 2026-08-07T10:53:35.150121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.751806Z digest=sha256:9cae0ff9cfb09bf7969a55c99dfadcf748026bb5eacc6ef15a42b29005c39b2a

Observation eda6959a-4dde-4ea0-ab68-e3e9ac2710a4 · outbound

This paper cites Between and within: Alternative sequential treat- ments of continuers and assessments,.

Voice Activity Projection Model with Multimodal Encoders Between and within: Alternative sequential treat- ments of continuers and assessments,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.828855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.819433Z digest=sha256:e74b6397fa5dcd85f80840df925d631e59efa0d80ad88b2980831e4b7948ad4c

Observation 5ab26f48-cf14-4575-a333-2892af7ef0f5 · outbound

This paper cites Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,.

Voice Activity Projection Model with Multimodal Encoders Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.653190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.901221Z digest=sha256:6c867f1bd3023f703f6b8e4a405dc5f2501d72e34790eab496465f5f317329ec

Observation fba6514b-15e3-4dd3-90ba-c469c37b185a · outbound

This paper cites Available: https://github.com/Lightning-AI/ pytorch-lightning.

Voice Activity Projection Model with Multimodal Encoders Available: https://github.com/Lightning-AI/ pytorch-lightning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.986850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:53:34.665269Z digest=sha256:573c862ec073efe2f10052be747e8d9ebdd1fa49960274e3598fe016d36adb36

Pith citing papers

Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · inbound

Voice Activity Projection Model with Multimodal Encoders cites this paper.

Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:30.784001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:30.784001Z digest=sha256:ee81424b8a182179f6389d293431df03a4fd86629ff1471dec1ecb608b5d0e8c

Observation 470d4985-fcdd-40ae-960d-58c68db749cc · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.175063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:76320121af775be943ec4f08f640eb1af430fa020978d358623292045a5bcfdd