Skip to content

Speech Recognition (STT) ​

Voice input works in one of two modes depending on the auto_send option of init(). Choose which input mode to use first.

Choosing an Input Mode ​

ModeInitOptionCall flowEnd-of-user-speech handling
Manual send (default)auto_send: falsestartListening() → user speech → endListening()Your app confirms the send with endListening()
Auto sendauto_send: trueCall startListening() onlyThe server detects end of speech and requests the response automatically

Manual Send — auto_send: false (default) ​

Start listening with startListening(), and when the user finishes speaking, your app calls endListening() to confirm the send. Suitable for push-to-talk UIs that treat only the audio spoken while a mic button is held down as input.

js
// Start listening — begins the speech recognition segment
SDK.startListening();

// Call when speech is finished — the send is confirmed and the avatar responds
SDK.endListening();

Auto Send — auto_send: true ​

Suitable for avatars that use voice input as the primary interface. Calling startListening() is all you need — the server automatically detects end of speech and requests the response, so there is no need to call endListening(). To start listening automatically when the connection completes, write:

js
SDK.onStatus((status) => {
  if (status === 'CONNECTED_FINISH') {
    SDK.startListening(); // From here on, the server detects end of speech and requests the response
  }
});

SDK.init({
  sdk_key: 'YOUR_SDK_KEY',
  avatar_id: 'YOUR_AVATAR_ID',
  auto_send: true,
}).catch((e) => console.error(e.message));

startListening() ​

Starts a speech recognition segment. The microphone itself is already on from the moment the connection is established when enable_microphone: true (default). startListening() tells the server to treat this segment as voice input, and re-enables the microphone track if it was disabled by a previous endListening()/cancelListening(). In auto-send (auto_send: true) mode, this single call is enough to start the voice conversation.

js
SDK.startListening();

endListening() ​

In manual-send (auto_send: false) mode, ends the speech recognition segment, disables the microphone track, and confirms the send. The recognition result is delivered as an STT_RESULT signal, and when stt_only=false, the avatar responds based on that result. In auto-send (auto_send: true) mode, there is no need to call it.

js
SDK.endListening();

cancelListening() ​

Cancels speech recognition and deactivates the microphone. Discards recognized text and the avatar will not respond. Can be called in both manual-send and auto-send modes.

js
SDK.cancelListening();

The stt_only Option ​

When initialized with stt_only: true, the avatar does not respond and only the STT recognition result (the STT_RESULT signal) is delivered. The default is false, in which case the avatar response follows the STT result. It is independent of auto_send, so it can be combined with either input mode.

Note that when initialized with enable_microphone: false, the SDK does not request microphone permission, so voice input is not available at all.

STT session and speech detection are separate concepts

startListening() / endListening() control the speech recognition (listening) segment. The microphone itself is on from connection time when enable_microphone: true; endListening()/cancelListening() disable the microphone track, and startListening() re-enables it. USER_SPEECH_STARTED / USER_SPEECH_STOPPED are speech segments detected by the server; while the microphone is on, they can also arrive outside the listening segment (e.g. right after connecting, before startListening() is called). Receiving USER_SPEECH_STOPPED does NOT end the listening segment. In manual-send mode you must call endListening() to end it; in auto-send mode the server detects end of speech and requests the response.

js
SDK.onSignal((data) => {
  switch (data.signal) {
    case 'USER_SPEECH_STARTED':
      // User started speaking (microphone was already on)
      console.log('User speech started');
      break;
    case 'USER_SPEECH_STOPPED':
      // User stopped speaking (microphone is still on)
      console.log('User speech stopped');
      break;
    case 'STT_RESULT':
      console.log('Recognition result:', data.payload.text);
      break;
  }
});

Full STT Flow ​

Manual Send (auto_send: false) ​

Assumes stt_only=false. Because the microphone is on from connection time, USER_SPEECH_* signals may also arrive before startListening() is called; the diagram below shows a typical flow.

Auto Send (auto_send: true) ​

Without any endListening() call, the server detects end of speech and carries on through the response. Assumes stt_only=false.