Technology

WhatsApp calls are getting noise cancellation — how it decides what to delete, and the calls where you should switch it off

The model keeps speech and discards everything else. That is the feature and the failure: it also discards other people talking, and anything musical.

WhatsApp calls are getting noise cancellation — how it decides what to delete, and the calls where you should switch it off

Anyone who has taken a WhatsApp call on a Dhaka bus knows the ritual: "Hello? Hello? Say again?" while a horn, a conductor and forty other conversations compete for the microphone. WhatsApp is adding noise cancellation to voice and video calls, and if it works as described it is one of the more useful things the app has shipped in years for people who do not take calls from quiet offices.

What is known

WABetaInfo reports the option has appeared for users on Android beta 2.26.14.1, inside the calling menu. It switches on automatically when a call starts and can be turned off. A wider rollout is expected within weeks. Gadgets 360 first reported the details.

How it decides what to delete

The processing runs on your phone, on your outgoing audio, before anything is sent. Worth understanding what it is doing, because the limitations follow directly from the method.

The older approach to noise suppression worked by subtraction. The software listened during the gaps between your words, built a profile of the background, and subtracted that profile from everything. It worked well on steady noise — a fan, an air conditioner, engine drone — and failed on anything sudden, because a horn that was not present during the gaps was not in the profile and so survived untouched. It also left a characteristic watery artefact that anyone who used early video calling will recognise.

The current approach does not subtract anything. A model trained on enormous numbers of recordings — clean speech, and the same speech buried in every kind of noise — learns what human speech looks like, and is asked to reconstruct the speech from the mixture. It is closer to redrawing than to filtering.

That is why it handles the Dhaka bus: a horn is not a problem to be subtracted, it is simply not speech, and anything that is not speech does not get redrawn. It is also why the feature arrived on phones only recently — running that kind of model continuously, fast enough that a call does not lag, needs hardware that mid-range phones have only had for a few years.

The catch everyone misses

It improves what the other person hears. It does nothing for what you hear.

If your caller is standing in a market and has not enabled the feature, or is on a version without it, their noise still reaches you exactly as before. A two-way improvement needs both sides to have it on, which is why switching it on by default matters more than it sounds — a feature that only helps when both people found a setting would help almost nobody.

What else it deletes

Here is the consequence of "keep speech, discard the rest" that catches people out, and it is not a bug.

Other people's voices. The model keeps the speech it judges to be the primary speaker and suppresses competing speech as noise. Usually right, and wrong in the specific case where somebody else in the room is meant to be heard — a colleague adding something to a work call, a relative greeting the person abroad. On that kind of call the other side hears you clearly and the second voice faintly or not at all.

Music and singing. Music is not speech, so the model treats it as noise and deletes it. Playing something down the phone, a child singing, a recitation held up to the microphone — all of it will arrive muffled or mangled. This is the standard complaint about noise suppression on every platform that has it, and the fix is the switch.

Occasionally, the beginning of your own sentence. Any system deciding in real time what is speech can take a moment to lock on after silence, which clips the first syllable. Noticeable mainly if you speak immediately after a long pause.

What to expect on your phone

  • Higher battery use during calls. The model runs continuously for the whole call, which is a constant computational load the phone did not previously carry.
  • Older and low-end phones may get it later, or with lighter processing. The quality depends on what the hardware can run in real time, so the same feature will not sound the same on every device.
  • Test it before it matters. Make one call to someone who will tell you honestly how you sound, rather than discovering the behaviour during an interview or a call with a customer.

Late, and worth it here anyway

Google Meet, Zoom and Microsoft Teams have had noise suppression since the pandemic, and phone makers bundle it into their dialers. WhatsApp is years behind.

It matters more here regardless, because of where the calls are made from. WhatsApp is how a great many people speak to family abroad and to customers at home, often from a construction site, a shop on a main road, or a bus. For a migrant worker on a noisy site, this is not a refinement to call quality. It is the difference between a call that works and one that does not.

Source: WABetaInfo, Gadgets 360, speech enhancement literature, on-device inference practice

Written by

Mahbub Rabbani

Mahbub Rabbani covers consumer technology for Tech BD — phones, apps, operating systems and the platforms most of the country uses daily. He writes the site’s how-to guides, and tends to judge a product by what it costs to live with rather than what it costs to buy.