Print

Voice Activity Detector (VAD)

With Voice Activity Detector (VAD), Avaya Experience Portal can detect the presence or absence of human speech. VAD can help detect if the sound is background noise or actual speech (speech energy).

VADs are used to determine when to send the audio stream to cloud speech resources such as Dialogflow, Google Speech, or Nuance Dialog as a Service (NuanceDaaS). When AEP VAD detects what it determines as speech energy, AEP then sends the audio stream to the cloud speech resource.

A properly tuned VAD reduces the incidence of falsely interrupted announcements when barge-in is enabled.

The AEP VAD configuration can enable and disable the local VAD and use the cloud speech resource instead. The AEP VAD configuration supports the following features:

importantImportant:

Disabling AEP local VAD or using the Hybrid option leads to increased network usage as the audio is streamed immediately to the cloud speech resource. It also leads to increased charges from the cloud speech resource provider since the cloud speech resource VAD is used.

AEP implements the following VAD algorithms:

VAD Configuration

VAD configuration consists of VAD selection and VAD tuning. VAD configuration can be applied on a global system-wide basis or a per-application basis.

AEP supports separate VAD configurations for each application. This overrides the global system-wide VAD configuration for that application.

When configuring an application for the cloud speech resource provider (Dialogflow, Google Speech, or NuanceDaaS ASR), you can configure the VAD parameters for that application.

VAD Mode

VAD Mode determines whether local VAD, the VAD of the cloud speech provider, or a combination of both VADs are used.

The VAD mode values are the following:

VAD type

VAD type determines which local VAD algorithm is used.

The VAD type values are the following:

For more information on Avaya Experience Portal Voice Activity Detector (VAD), see Avaya Experience Portal Dialogflow White Paper.