← Blog9 min read#web speech api#voice#react

Web Speech API in React: Voice Input, Text-to-Speech in the Browser

Add voice recognition and text-to-speech to React apps using the native Web Speech API - no SDK required. Full hooks, error handling, and UI patterns.

colorful sound wave visualization on dark background for voice interface

What the Web Speech API Actually Is

The Web Speech API is a browser-native interface split into two independent pieces: SpeechRecognition for converting spoken audio into text, and SpeechSynthesis for reading text aloud. Neither requires a third-party SDK. Neither hits a billing endpoint. They ship in Chrome, Edge, and Safari - and as of 2024, Firefox has partial synthesis support too.

That said, browser support is lopsided in ways that matter. SpeechRecognition is still prefixed in Chrome as webkitSpeechRecognition. Safari supports it unprefixed since Safari 15, but quirks remain. Firefox? Doesn't do recognition at all without a flag. Plan accordingly - you'll want a feature-detect guard before any of this code runs.

In practice, this API is most useful for quick voice-input flows, accessibility toggles, voice-commanded UIs, and TTS for screen-reader-alternative experiences. It's not a replacement for a production transcription pipeline (Whisper, Deepgram, etc.), but for in-browser prototyping and lightweight features, it's surprisingly capable.

Worth noting: Chrome's SpeechRecognition sends audio to Google's servers for processing. That's not obvious from the API surface. If you're building for privacy-sensitive contexts, you need to know that upfront and disclose it to users.

Setting Up SpeechRecognition in React

Start with a custom hook. Wrapping the API in a hook keeps your components clean and gives you a single place to handle the prefixing mess, lifecycle cleanup, and event wiring. Here's a minimal useSpeechRecognition that actually works:

import { useEffect, useRef, useState, useCallback } from 'react';

const SpeechRecognitionAPI =
  (window as any).SpeechRecognition ||
  (window as any).webkitSpeechRecognition;

export function useSpeechRecognition() {
  const [transcript, setTranscript] = useState('');
  const [listening, setListening] = useState(false);
  const [error, setError] = useState<string | null>(null);
  const recognitionRef = useRef<any>(null);

  useEffect(() => {
    if (!SpeechRecognitionAPI) return;

    const recognition = new SpeechRecognitionAPI();
    recognition.continuous = true;
    recognition.interimResults = true;
    recognition.lang = 'en-US';

    recognition.onresult = (e: any) => {
      const text = Array.from(e.results)
        .map((r: any) => r[0].transcript)
        .join('');
      setTranscript(text);
    };

    recognition.onerror = (e: any) => setError(e.error);
    recognition.onend = () => setListening(false);

    recognitionRef.current = recognition;

    return () => recognition.abort();
  }, []);

  const start = useCallback(() => {
    setError(null);
    setTranscript('');
    recognitionRef.current?.start();
    setListening(true);
  }, []);

  const stop = useCallback(() => {
    recognitionRef.current?.stop();
    setListening(false);
  }, []);

  return { transcript, listening, error, start, stop,
    supported: !!SpeechRecognitionAPI };
}

A few things worth unpacking. continuous: true keeps the microphone open after a pause - set it to false if you want it to auto-stop after one utterance. interimResults: true gives you the rolling draft text while the user is still talking, which makes the UI feel alive. Without it, you get nothing until there's a full stop.

The onend handler fires when recognition stops - both from calling .stop() and from errors or silence timeouts. That's why you always reset listening there, not just in your stop callback. Forgetting this is how you get a "listening" indicator stuck on forever.

One more thing - always call recognition.abort() in the cleanup function of your useEffect. If the component unmounts while the mic is active, you'll get an error event on a dead reference otherwise. The cleanup handles it cleanly.

Building a Voice Input Component

With the hook in place, the actual component is straightforward. The interesting design problem is: what do you show the user while they're speaking? Interim results (the rolling transcript) need visual treatment that signals "this isn't final yet." A simple opacity difference works well.

import { useSpeechRecognition } from './useSpeechRecognition';

export function VoiceInput({ onResult }: { onResult: (text: string) => void }) {
  const { transcript, listening, error, start, stop, supported } =
    useSpeechRecognition();

  if (!supported) {
    return <p className="text-sm text-red-400">Voice input not supported in this browser.</p>;
  }

  return (
    <div className="flex flex-col gap-3">
      <div
        className={`min-h-[80px] rounded-xl border p-4 text-sm transition-colors ${
          listening
            ? 'border-violet-500 bg-violet-950/30 text-white'
            : 'border-zinc-700 bg-zinc-900 text-zinc-400'
        }`}
      >
        {transcript || (
          <span className="opacity-40">Start speaking…</span>
        )}
      </div>

      {error && (
        <p className="text-xs text-red-400">Error: {error}</p>
      )}

      <div className="flex gap-2">
        <button
          onClick={listening ? stop : start}
          className={`rounded-lg px-4 py-2 text-sm font-medium ${
            listening
              ? 'bg-red-600 hover:bg-red-700'
              : 'bg-violet-600 hover:bg-violet-700'
          } text-white transition-colors`}
        >
          {listening ? 'Stop' : 'Start Listening'}
        </button>

        {transcript && !listening && (
          <button
            onClick={() => onResult(transcript)}
            className="rounded-lg bg-zinc-700 px-4 py-2 text-sm text-white hover:bg-zinc-600"
          >
            Use This
          </button>
        )}
      </div>
    </div>
  );
}

Honestly, the error codes you'll hit most often are not-allowed (user denied mic permission), no-speech (silence timeout), and network (Chrome's recognition server hiccup). Map those to readable messages - not-allowed especially deserves a helpful prompt since users often forget they blocked the mic.

For visual flair on a listening state, this pairs well with a pulsing ring animation. A 4px ring in violet at 60% opacity, expanding to 8px over 1s in an infinite loop, reads immediately as "active." Something like ring-2 ring-violet-500 ring-offset-2 animate-pulse in Tailwind gets you most of the way there. If you want something more expressive, check out the animated button components in Empire UI - adapting the pulse pattern from those saves you 20 minutes.

Text-to-Speech with SpeechSynthesis

SpeechSynthesis is the flip side. You give it text, pick a voice, dial in rate and pitch, and it speaks. Browser support here is much wider - it works in every major browser including Firefox. The API is synchronous in terms of configuration but async in execution, which trips people up.

import { useCallback, useEffect, useRef, useState } from 'react';

export function useSpeechSynthesis() {
  const [speaking, setSpeaking] = useState(false);
  const [voices, setVoices] = useState<SpeechSynthesisVoice[]>([]);
  const utteranceRef = useRef<SpeechSynthesisUtterance | null>(null);

  useEffect(() => {
    const loadVoices = () => setVoices(window.speechSynthesis.getVoices());
    window.speechSynthesis.onvoiceschanged = loadVoices;
    loadVoices();
    return () => { window.speechSynthesis.cancel(); };
  }, []);

  const speak = useCallback(
    (text: string, voiceIndex = 0, rate = 1, pitch = 1) => {
      window.speechSynthesis.cancel();
      const utterance = new SpeechSynthesisUtterance(text);
      utterance.voice = voices[voiceIndex] ?? null;
      utterance.rate = rate;   // 0.1 – 10
      utterance.pitch = pitch; // 0 – 2
      utterance.onstart = () => setSpeaking(true);
      utterance.onend = () => setSpeaking(false);
      utterance.onerror = () => setSpeaking(false);
      utteranceRef.current = utterance;
      window.speechSynthesis.speak(utterance);
    },
    [voices]
  );

  const cancel = useCallback(() => {
    window.speechSynthesis.cancel();
    setSpeaking(false);
  }, []);

  return { speak, cancel, speaking, voices };
}

The onvoiceschanged event is critical on Chrome. Voices load asynchronously after page load, so calling getVoices() immediately on mount returns an empty array. The event fires once they're ready. Safari, on the other hand, has them available synchronously. The pattern above handles both.

Quick aside: voice names vary wildly by OS. On macOS you'll see "Samantha," "Alex," "Daniel (Enhanced)." On Windows you get "Microsoft David," "Microsoft Zira." On Android it depends on what's installed. Don't hardcode a voice name - let users pick from the available list, or fall back to voices[0].

Rate values between 0.8 and 1.2 sound natural for most voices. Go above 1.5 and you're in chipmunk territory. Pitch defaults to 1 - values below 0.8 give you that deep narrator vibe. Worth experimenting, but for accessibility use cases, leaving both at defaults is usually the right call.

Combining Both: A Voice Command Interface

The real power comes when you chain them. Listen for a command, parse the transcript, take an action, then confirm verbally. This is the loop that makes voice UIs feel coherent rather than bolted on.

import { useEffect } from 'react';
import { useSpeechRecognition } from './useSpeechRecognition';
import { useSpeechSynthesis } from './useSpeechSynthesis';

const COMMANDS: Record<string, () => void> = {
  'scroll down': () => window.scrollBy({ top: 400, behavior: 'smooth' }),
  'scroll up': () => window.scrollBy({ top: -400, behavior: 'smooth' }),
  'go home': () => (window.location.href = '/'),
};

export function VoiceCommandLayer() {
  const { transcript, listening, start } = useSpeechRecognition();
  const { speak } = useSpeechSynthesis();

  useEffect(() => {
    if (!transcript) return;
    const lower = transcript.toLowerCase().trim();
    for (const [command, action] of Object.entries(COMMANDS)) {
      if (lower.includes(command)) {
        action();
        speak(`Done. ${command}.`);
        break;
      }
    }
  }, [transcript, speak]);

  return (
    <button onClick={start} aria-label="Activate voice commands">
      {listening ? '🎙 Listening…' : '🎤 Voice'}
    </button>
  );
}

Look, this is a simple pattern but it scales. Add fuzzy matching with a library like fastest-levenshtein if you want tolerance for near-misses. The key insight is that transcript changes on every interim result, so your command parser runs frequently - make sure the actions are idempotent or add a debounce before the matching loop.

For a polished UI layer on top of something like this, the command palette component pattern works great - same mental model of "type or speak a command, get an action." You can surface both keyboard and voice entry through the same interface.

One thing to watch: don't fire TTS while SpeechRecognition is active if you can avoid it. The synthesis audio bleeds into the mic on non-headphone setups and generates phantom transcripts. Stop recognition first, speak, then restart. The callback chain makes this manageable.

Accessibility and UX Considerations

Voice UI done well is an accessibility win. Voice UI done carelessly is an exclusion mechanism. Users with accents, speech impediments, or non-native pronunciations get worse transcription accuracy. That's not a React problem, it's a Web Speech API problem - but you still own the design decision to use it and how you handle failures.

Always provide a text fallback. Every voice input should sit next to a regular <input>. Never make voice the *only* way to submit something. The button that triggers start() needs an aria-label that says what it does - "Start voice input" is better than just a mic icon with no label.

Consider announcing TTS output via an ARIA live region too, so screen readers don't double-announce: <div aria-live="polite" className="sr-only">{currentSpeakingText}</div>. That way if someone *is* on a screen reader, they don't get the synthesis voice plus the screen reader reading the same content simultaneously.

Worth noting: on mobile, SpeechRecognition requires a user gesture to start - you can't autostart it on page load. That's actually correct behavior from a privacy standpoint. Just don't build a flow that expects ambient listening without a tap.

If you're building a visually rich voice interface and want something that signals "this is a voice-first experience," look at the aurora background or glassmorphism components from Empire UI. A frosted panel with an animated waveform behind it lands way better than a plain white input box with a mic icon.

Browser Support Gaps and Fallbacks

Here's the honest matrix as of mid-2026: Chrome 33+ and Edge 79+ have full SpeechRecognition support. Safari has supported it since version 15 (2021) with a few quirks around continuous mode. Firefox does not support SpeechRecognition at all without media.webspeech.recognition.enable in about:config. On SpeechSynthesis, coverage is broader - Chrome, Firefox, Safari, and Edge all support it.

Your feature detect should gate the whole recognition feature, not just log a warning:

const isSpeechRecognitionSupported =
  typeof window !== 'undefined' &&
  ('SpeechRecognition' in window || 'webkitSpeechRecognition' in window);

const isSpeechSynthesisSupported =
  typeof window !== 'undefined' && 'speechSynthesis' in window;

For environments where recognition isn't available, a clean fallback is to show a standard text input with an upload button for audio files if you want to go further (and process with a backend). But for most use cases, "not supported in this browser, please use Chrome or Edge" with a clear message is honest enough.

In Next.js, both APIs require typeof window !== 'undefined' guards because SSR will blow up trying to access window.SpeechRecognition. Wrap hook initialization in a useEffect or use dynamic(() => import(...), { ssr: false }) for any component that instantiates these APIs. Skipping this step will give you hydration errors that are annoying to trace back to the source.

FAQ

Does the Web Speech API work offline?

No, not in Chrome - recognition requires a network connection because audio gets sent to Google's servers. Safari may have limited on-device support depending on the OS version, but don't count on offline functionality.

Why is my SpeechRecognition stopping after a few seconds of silence?

That's the default timeout behavior. Set recognition.continuous = true to keep it running, and handle the onend event to restart it automatically if you need indefinite listening.

Can I use a custom voice model instead of the browser's built-in voices?

Not with the native Web Speech API - you're limited to what the OS has installed. For custom voices or higher accuracy, you'd need a service like ElevenLabs, Play.ht, or Google Cloud TTS and handle audio playback manually.

Does this work in a Next.js App Router project?

Yes, but all Web Speech API code must run client-side only. Add 'use client' to any component using these hooks, and guard initialization with typeof window !== 'undefined' to prevent SSR crashes.

Free components in 41 styles
React & Tailwind, copy-paste ready.
Browse →

Read next

Advanced CSS & JavaScript Patterns: Production-Grade Techniques 2026 →Parallax Scrolling in React: useScroll, GSAP and Pure CSS →HTML Canvas Animations in React: Particles, Noise Fields, More →Neon Text Effect in React: Animated Glow with CSS and Framer Motion →