Speech Enhancement: Methods, Python Code, and Production Tests

A speech waveform emerging from background noise
A speech waveform emerging from background noise

Speech enhancement can make a noisy recording easier to hear and a live voice system easier to understand. The hard part is choosing a method that removes the right interference without erasing speech, adding delay, or hurting downstream recognition. Here is a practical path from a spectral-gating baseline to production evaluation.

What speech enhancement does

Speech enhancement is signal processing that improves degraded speech for a listener or a downstream system. In practice, it can attenuate background noise, reduce reverberation, and suppress an interfering speaker while preserving the target voice.

It cannot recover the exact speech information that a microphone, channel, or codec never captured. Bandwidth-extension systems can synthesize plausible high-frequency content, but that content is an estimate. Clipped samples and masked phonemes are also irreversible losses.

Quality and intelligibility are separate outcomes. A quieter file can sound cleaner while becoming harder for a person or an automatic speech recognition (ASR) system to understand. The working objective is therefore narrower: reduce a specific degradation without damaging the speech cues your application needs.

Map the degradation before choosing a method

Degradation or taskWhat it changesIs it speech enhancement?
Stationary noiseA stable hum, fan, or electrical floorYes, usually noise suppression
Non-stationary noiseKeyboard taps, traffic, music, or changing machineryYes, usually needs an adaptive or learned method
ReverberationLate room reflections that smear speech over timeYes, through dereverberation
Interfering speechAnother voice overlaps the target speakerYes, through target-speaker extraction or separation
Acoustic echoLoudspeaker playback leaks back into the microphoneAdjacent task; acoustic echo cancellation uses a playback reference
Automatic gain control or normalizationRaises or stabilizes signal levelAdjacent task; it does not separate speech from noise
Resampling or codec conversionChanges sample rate or representationAdjacent task; it cannot restore missing source information
Speaker diarizationLabels who spoke whenDownstream task; it does not clean the waveform

The location of the degradation matters too. Microphone self-noise, room reverberation, telephony compression, packet loss, and downstream playback each call for different fixes. Save representative audio as early in the pipeline as you can. Otherwise, a later recording may hide where the damage entered.

Choose the least complex method that passes your tests

A good first implementation is the smallest one that improves the intended outcome across representative audio.

ApproachBest fitMain advantageMain limitation
Hosted or desktop cleanupOne-off recordings and editorial audioFastest route to a usable fileLimited control, privacy choices, and uncertain live-stream behavior
Spectral gatingStable background noise and a reproducible baselineSimple, local, and inexpensiveCan leave changing noise or create metallic artifacts
Pretrained neural modelVariable noise, full-band audio, or live integrationBetter modeling of speech and changing interferenceMore CPU, memory, state, and deployment work
Custom modelUnusual sensors, acoustic domains, hardware limits, or licensing constraintsCan optimize for a narrow targetRequires data, training, evaluation, and long-term ownership

Custom training should be the exception. Start with a deterministic baseline and at least one relevant pretrained model. Their failures will show whether your problem is model quality, input conditioning, or a unique domain gap.

A complete test has two branches: what a person hears and what the rest of the voice pipeline receives. Improving only one branch can create a regression in the other.

Noisy speech passes through enhancement to listening and downstream system tests

Build a spectral-gating baseline in Python

Noisereduce documentation describes stationary and non-stationary spectral gating. The library estimates a frequency-dependent noise threshold, builds a mask, smooths that mask across time and frequency, and applies it to the short-time Fourier transform (STFT) of the signal.

Install the package and an audio reader:

pip install noisereduce soundfile

The current API uses y for the input waveform and sr for its sample rate. This example reads mono or multichannel audio, uses the first two seconds as a noise-only reference, and preserves the channel layout on output:

import noisereduce as nr import soundfile as sf # soundfile returns shape (frames, channels) when always_2d=True. audio, sample_rate = sf.read("noisy.wav", always_2d=True) y = audio.T # noisereduce expects (channels, frames) # This slice must contain representative noise and no target speech. noise_seconds = 2 y_noise = y[:, : sample_rate * noise_seconds] enhanced = nr.reduce_noise( y=y, sr=sample_rate, y_noise=y_noise, stationary=True, prop_decrease=0.8, n_fft=512, ) sf.write("enhanced.wav", enhanced.T, sample_rate)

Three settings deserve deliberate tests:

  • y_noise should contain the same noise distribution as the noisy sections. Speech in this reference can teach the gate to suppress speech.
  • prop_decrease controls how much of the estimated noise is attenuated. Start below full suppression. Aggressive settings often remove consonant energy or produce an unnatural, watery texture.
  • n_fft controls the analysis window. The package documentation recommends 512 as a practical speech setting. Confirm it at your actual sample rate because the same number of samples represents a different duration at 8, 16, and 48 kilohertz (kHz).

When you do not have a clean noise-only interval, try the non-stationary mode:

enhanced = nr.reduce_noise( y=y, sr=sample_rate, stationary=False, prop_decrease=0.8, n_fft=512, )

This mode estimates a changing noise floor. It is still a spectral gate, so it is a baseline for live, highly variable, or competing-speech conditions rather than a default production answer.

Try a pretrained neural model before training your own

Neural enhancement models learn a mapping from degraded audio to a clean-speech estimate. The model may estimate a time-frequency mask, complex spectrum, filter coefficients, or the waveform itself. Architecture names matter less than fit with your channel, latency budget, and failure cases.

The DeepFilterNet repository provides a concrete modern starting point. It includes pretrained models for full-band 48 kHz audio, a command-line path, Python integration, and a real-time plugin. The Python package can process a file with:

pip install deepfilternet deepFilter noisy.wav --output-dir enhanced

Treat the 48 kHz input requirement as part of the model contract. Resampling an 8 kHz telephone call to 48 kHz does not recreate the frequencies removed by the telephone channel. Test the model on the original channel conditions your application will receive.

Other useful reference points include:

  • The RNNoise repository, a compact hybrid digital signal processing and recurrent-neural-network design for real-time full-band suppression. It is useful when a C library and predictable local execution fit the system.
  • The NSNet2 baseline, an older Microsoft Deep Noise Suppression Challenge model distributed with Open Neural Network Exchange (ONNX) inference scripts. It remains useful for reproducing past work, but it should not be treated as a universal current recommendation.

Compare candidates under one harness. Normalize input format, keep the same clips, record model and configuration versions, and store the enhanced files. A model that wins on a public test set can still fail on your microphones, codecs, languages, overlapping speakers, or noise mixtures.

Train a custom model only for a proven gap

Custom training is justified when repeated evaluation shows that available models miss a valuable, stable domain. Examples include a fixed industrial sensor, a specialized headset, a strict on-device compute budget, or a license that cannot ship with your product.

The training set must represent clean speech, noise, room responses, devices, codecs, and signal levels expected after deployment. Include the hard cases you intend to preserve, such as quiet speakers, fricatives, crosstalk, and rapid turn endings. Hold out devices and environments, since a random split of near-duplicate mixtures can overstate generalization.

Engineer real-time enhancement as a stateful stage

A file-processing demo says little about a streaming deployment. A production suppressor has to preserve model and filter state while audio arrives in short frames.

Plan for these constraints:

  1. Causality and lookahead. Confirm how many future samples the algorithm needs. Model inference may be fast while frame buffering and lookahead still add conversational delay.
  2. Continuous state. Keep recurrent, filter, and noise-estimator state across frames. Reset it at a defined session boundary, not on every network packet.
  3. Window overlap. STFT systems need consistent overlap and synthesis windows. Independent chunks can create clicks, periodic pumping, or audible seams.
  4. Resource headroom. Measure CPU, accelerator use, memory, and runtime per second of audio under concurrent load. Report median, 95th-percentile, and 99th-percentile processing time rather than one average.
  5. Channel fidelity. Test each sample rate, channel layout, codec, and transport path separately. Browser audio and phone audio are different operating conditions.
  6. Clipping control. Detect clipped input before enhancement. A suppressor can reduce surrounding noise, but it cannot reconstruct a flattened waveform reliably.
  7. Safe bypass. If the model misses its frame deadline, loses state, or produces invalid samples, route the original audio and log the event. This keeps an enhancer failure from stalling the audio path.

Enhancement also changes the inputs to voice activity detection (VAD). Over-suppression can erase quiet speech starts or word endings. Residual bursts can look like speech. Measure false starts, missed speech, and endpoint timing after the enhancement stage.

Evaluate speech quality and system behavior together

Subjective listening is the practical reference because speech quality is perceptual. Build a blinded listening set that represents devices, speakers, languages, rooms, codecs, signal-to-noise ratios, and interference types. Include the original, the enhanced output, and a stable baseline. Randomize order and ask focused questions about intelligibility, naturalness, residual noise, and listening effort.

Objective metrics support that review. They do not replace it.

MeasureWhat it tells youConstraint
International Telecommunication Union recommendation P.863Perceptual quality estimate against a referenceRequires suitable reference and test conditions
Deep Noise Suppression Mean Opinion Score (DNSMOS)No-reference estimate aimed at noise-suppressed speechModel estimate that can drift outside its training domain
Non-Intrusive Speech Quality Assessment (NISQA)Overall quality plus noisiness, coloration, discontinuity, and loudnessPretrained weights use a non-commercial license
ASR word or entity error rateWhether downstream recognition improvesDepends on the chosen recognizer and transcript set
VAD and endpoint errorsWhether speech boundaries improve or regressNeeds labeled timing or carefully reviewed events
CPU, memory, and tail latencyWhether the stage meets runtime limitsMust be measured on target hardware at expected concurrency
Task completionWhether the whole voice workflow benefitsNeeds enough production-like conversations to be meaningful

International Telecommunication Union P.863 is the current recommendation for objective perceptual speech-quality prediction. The organization withdrew the older P.862 Perceptual Evaluation of Speech Quality (PESQ) family in 2024. The P.862 recommendation page directs readers to P.863, so PESQ is no longer the current general standard.

For recordings without a clean reference, the DNSMOS documentation and the NISQA repository can help rank variants. Record the metric version and exact input preparation. NISQA's code is MIT-licensed, while its published pretrained weights use the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 license (CC BY-NC-SA 4.0), which matters for commercial deployment.

For a voice agent, also compare transcription accuracy, entity capture, turn detection, interruption behavior, task completion, and end-to-end latency. Measure on both enhanced and original audio. Noise reduction is a regression when it produces a cleaner waveform and a worse conversation.

Diagnose common speech enhancement failures

SymptomLikely causeFirst corrective move
Metallic or musical noiseSparse, rapidly changing spectral maskReduce attenuation and increase mask smoothing
Missing consonants or quiet syllablesSpeech energy classified as noiseLower suppression strength and improve the noise estimate
Pumping backgroundNoise estimate follows short-term eventsTune the estimator time constant or use a better-matched model
Reverberant smear remainsNoise suppressor does not model room decayEvaluate a dereverberation model or improve capture geometry
Background speaker remainsInterfering speech resembles target speechUse target-speaker extraction or separation with identity controls
Clicks at chunk boundariesState resets or incorrect overlapPreserve state and use the required analysis/synthesis overlap
Clean audio but worse ASRModel altered phonetic cuesSelect by transcript and entity errors, not perceptual score alone
Distorted peaks remainInput clipped before enhancementFix gain and capture; retain a bypass for clipped frames

Save failure clips as a regression corpus. Every change to a model, noise estimator, frame size, codec, or resampler should run against the same set before release.

Deployment checklist

  • Define the target degradation and the desired downstream outcome.
  • Preserve original audio for comparison and debugging.
  • Establish an unprocessed baseline and a simple spectral-gating baseline.
  • Test a pretrained model on the real sample rate, codec, and device path.
  • Measure listening quality, ASR accuracy, VAD behavior, CPU, memory, and tail latency.
  • Keep streaming state across frames and verify overlap handling.
  • Add deadline, invalid-output, and model-load fallbacks.
  • Version the model, configuration, resampler, and evaluation corpus together.
  • Run a limited rollout and compare completed conversations before expanding traffic.

Frequently asked questions about speech enhancement

Does speech enhancement always improve intelligibility?

No. Enhancement can improve perceived cleanliness while removing speech cues. Listening tests and downstream recognition results are both required.

Is noise cancellation the same as speech enhancement?

Noise suppression is one speech-enhancement task. Active noise cancellation creates an opposing acoustic signal at playback, while acoustic echo cancellation removes known loudspeaker leakage from microphone input. Those systems solve different problems.

Can speech enhancement restore low-bitrate or narrowband audio?

It can reduce some artifacts or generate plausible missing content, but it cannot recover the exact source information discarded by the original capture or codec.

Where should enhancement sit in a voice AI pipeline?

Place it after the earliest capture stages needed for echo and level control, and before components that benefit from a cleaner microphone signal. Then evaluate the exact ordering with your automatic speech recognition and turn-detection stack because every stage changes the next one's input.

Evaluate enhancement in the completed conversation

A suppressor can sit before a managed conversational runtime, but its value appears in the full interaction. Our Call Inspector exposes completed-call recordings, transcripts, timeline events, large language model (LLM) interactions, tool executions, and latency breakdowns. Run the same representative call set with enhancement enabled and bypassed, then compare speech recognition, turn boundaries, latency, and completion evidence before you ship the new stage.

Related Posts

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.