Bunny Honey ClubBunny Honey/blog
Work with us
← back to indexblog / ai / ai-voice-watermarking-elevenlabs-openai-synthid
AI

Your AI voiceover is now watermarked. Here's the catch.

ElevenLabs and OpenAI now bake an inaudible SynthID watermark into AI audio, and the detectors are free. What it changes for your ads and your phone line.

AH
Arthur HofFounder, Bunny Honey Club AI
publishedAug 01, 2026
read6 min
Your AI voiceover is now watermarked. Here's the catch.

Every clip of AI voice you generated through ElevenLabs or OpenAI in the last few weeks has an inaudible watermark baked into it. Not metadata you can strip. A pattern hidden in the sound itself. OpenAI switched it on for audio on July 31.

Every clip of AI voice you generated through ElevenLabs or OpenAI in the last few weeks has an inaudible watermark baked into it. Not metadata you can strip. A pattern hidden in the sound itself.

OpenAI switched it on for audio on July 31. ElevenLabs started in June.

Both of them also run a free web page where anyone can upload your file and get an answer. That last part is the one that actually changes how you should work.

Jul 31, 2026OpenAI audio watermarking + verification API live
Jun 25, 2026ElevenLabs began SynthID on text-to-speech
Aug 2, 2026EU AI Act Article 50 transparency duties apply
FreeCost of both public detectors, to anyone

The two engines behind most business AI voice now mark it by default

ElevenLabs announced in late June that it was partnering with Google DeepMind to embed SynthID directly into audio it generates. It started with text-to-speech on free accounts and said coverage would expand to all ElevenLabs audio generations over the following weeks.

OpenAI followed on July 31, extending the same watermark from images to audio generated through ChatGPT and the OpenAI API. In the same update it opened API access to its verification tool, so any developer can now wire provenance checks into their own systems.

Those two vendors sit behind a large share of the AI voice small businesses actually ship: ad voiceovers, video narration, IVR prompts, phone agents, podcast reads.

If your voice work runs through either, it is marked. You did not opt in and you cannot opt out.

The watermark survives the things you actually do to audio

This is the part that surprises people who assume watermarks are a metadata field.

SynthID is a pattern hidden in the waveform. ElevenLabs says the mark holds through trimming, speed changes, metadata stripping, compression, and conversion to a different file format. It also says the watermark cannot be copied onto audio ElevenLabs did not generate, which matters more than it sounds: it means a false accusation is hard to manufacture.

Their published bar for shipping it was strict. No added time-to-first-byte latency, high detection rate with low false positives, survives cropping, and no audible quality loss.

So the usual workarounds do not apply. Re-encoding your MP3 does nothing. Speeding the read up by four percent does nothing. Dropping it into an edit and exporting as WAV does nothing.

The free detector is what actually changes behaviour

Watermarking on its own is an internal safety project. Nobody outside the lab notices.

The detector is the part with teeth. ElevenLabs shipped a free Audio Detector page that tells you whether a clip came from ElevenLabs. OpenAI's public verification tool now accepts audio alongside images, and as of last Friday there is an API behind it.

Think about who uses that. Not regulators, mostly. Regulators are slow and understaffed. It will be the ad platform trust team reviewing a complaint, the local competitor who thinks your testimonial video sounds off, and the customer who wants to prove the "person" who called them was not a person.

The cost of checking just went to zero. Things that cost zero get done constantly.

Watermarking does not change what is legal. It changes what is provable. Provable is the thing that changes behaviour.

What we tell every client shipping AI voice

Detection only runs one direction, and that is the trap

Here is the limitation almost nobody states plainly, and it is the reason this post exists.

A detector can tell you a file came from a specific vendor. It cannot tell you a file is human.

OpenAI is explicit about this in its own announcement: no detection method is foolproof, the tool is limited to content generated by OpenAI, and when no signal is found it will not conclude the content was not AI-generated. ElevenLabs' detector answers one question too: did this come from ElevenLabs.

So the failure modes are asymmetric, and both hurt:

  • A genuine human recording run through either detector comes back "no signal found." That is not a clean bill of health, and anyone treating it as one is misreading the tool.
  • AI voice from an unwatermarked engine also comes back "no signal found." Open-source models, smaller vendors, and older exports sit entirely outside this system.

Which means the practical effect of watermarking in 2026 is narrow and slightly perverse. It catches the honest majority using the two biggest, most compliant vendors. It does nothing about the person cloning your CEO's voice with a self-hosted model.

That is not an argument against it. It is an argument for not confusing a detector result with the truth.

Your vendor just handled the machine-readable half of Article 50

Timing is doing some work here. The EU AI Act's transparency obligations apply from August 2, tomorrow, and one of them requires AI-generated audio to be marked in a machine-readable format.

That specific duty sits with the provider of the AI system. If you generate through OpenAI or ElevenLabs, they have now done it for you, and ElevenLabs said outright that it was building toward the growing number of jurisdictions demanding machine-readable marking.

Your obligations are the other ones: telling people they are dealing with AI, and labelling synthetic content where the rules require it. We covered what actually lands on August 2 when everyone was still repeating "the AI Act got delayed," the narrower reality of labelling AI-generated ads across New York, the EU, and Google, and the two US state rules that genuinely reach an inbound AI phone line.

Watermarking does not discharge any of those. It just removes one line item you were never responsible for anyway.

What we changed in our own pipeline this week

We run AI voice in two places: ad and video creative, and the AI receptionist lines we operate for clients, including the six physiotherapy practices we published the real numbers on.

Two changes, both small.

We log the engine per asset. Every generated audio file now carries the vendor, model, and date in its filename and in the job record. When a platform or a client asks where a clip came from, "we would have to check" is a bad answer and "we do not remember" is a worse one. This costs nothing to do at generation time and is nearly impossible to reconstruct later.

We stopped treating detection risk as a creative constraint. Our disclosure rule has been the same since we built our first phone line: if the customer could plausibly be surprised later, say it now. That rule was already stricter than the law in most jurisdictions, so the watermark changed nothing about what we ship. It only confirmed the bet.

Where it did change our thinking is on testimonials and anything voiced to sound like a specific real person. That creative was always the risky category. It is now the checkable one, which is a different and worse problem. We do not build it, and we talk clients out of it.

If you would rather not hold this in your head at all, that is roughly the job: we run the AI content and ad creative pipeline end to end, with provenance logged per asset and the disclosure call made before anything goes live.

The honest summary is that this is good news wearing a scary hat. A durable, standard watermark makes AI voice more trustworthy to use in public, not less. The businesses who lose are the ones who were quietly counting on nobody being able to tell.

— share
— keep reading

Three more from the log.