AI Research17 minAugust 2026

I Built a SynthID Watermark DetectorThat Cannot Detect Gemini

A week reverse-engineering Google SynthID-Text watermarking, a detector that scores ROC-AUC 1.0000, and the one-digit experiment that proved no third party can ever detect Gemini or Claude output.

By Sharon Rosario · August 2026

SynthIDAI WatermarkingLLMGeminiClaudeAI DetectionPython
Published
August 2026
Campaign
SynthID-Text watermark detection
Scope
Feasibility research + reference implementation
Outcome
Built, benchmarked, and shelved on purpose
Platform
Python, PyTorch, Hugging Face Transformers

The Question

Can you tell if a piece of text came out of Gemini?

The short answer, which took me a full build and a week of reading to be sure of: no. Not you, not me, not any company outside Google. Here is the whole story, including the parts where I was wrong.

The question landed the way good questions usually do, casually, in a chat thread. AI watermarking was suddenly everywhere. Google had been quietly stamping Gemini output with something called SynthID for a while. Anthropic had just announced it was doing the same for Claude. The EU had started requiring AI-generated content to be marked. And there was an obvious product idea sitting right there: build a free tool where someone pastes text and finds out whether an AI wrote it, verified by the watermark rather than guessed at.

It sounded very buildable. Google had open-sourced a reference implementation. There was a peer-reviewed Nature paper. Hugging Face had shipped a production port into Transformers. All the pieces were public.

So I started reading, then building. What I ended up with was a working detector that scores watermarked text at ROC-AUC 1.0000, a benchmark of nearly 4,000 samples, and a conclusion I did not want: the tool cannot be shipped. Not because my implementation was bad. Because the thing it needs does not exist outside two companies, and never will.

I am writing this up because I could not find a single honest account of it anywhere. What I found instead was a search results page full of tools claiming to detect SynthID that demonstrably do not. If you are evaluating this idea, I would like to save you the week.

One thing worth separating up front, because almost every article I read blurred it. Watermark detection and AI detection are completely different mechanisms. AI detection looks at the writing itself, statistically, and guesses. Watermark detection reads a deliberate signal planted at generation time, and either verifies it or does not. This piece is entirely about the second one.

Key takeaway

"Watermark detection is not a harder version of AI detection. It is a different problem with a hard cryptographic dependency, and that dependency is what decides whether you can build it."

How SynthID Works

The watermark is not in the text, it is in the choices

This is the part most explainers get wrong, and getting it right is what makes everything afterwards obvious. Nothing is added to the text. No invisible characters, no zero-width spaces, no metadata, no changed punctuation.

When a language model writes, it does not pick one inevitable next word. At each step it has a ranked cloud of plausible candidates, and it samples one. Often several candidates are near-equally good. "The results were surprising" and "The results were unexpected" are both fine. That freedom is unavoidable, and it is where SynthID lives.

Here is the mechanism, in the order it happens:

Take the previous four tokens as context. For each candidate next token, hash the pair together with a secret key. That hash gets mapped to a single bit, either 0 or 1, called a g-value. Since the hash is deterministic, the same context plus the same candidate plus the same key always gives the same bit. Now run a small knockout tournament between candidates, preferring those whose g-value is 1. Repeat this thirty times over, once per key, in thirty stacked layers.

The word that wins is still a perfectly good word. The text reads normally, because the model was genuinely indifferent between the options. But the winner is now slightly more likely to be one whose g-value was 1.

Do that for one token and you have learned nothing. Do it for four hundred and you have a measurable statistical bias. On unwatermarked text, the average g-value sits at exactly 0.5, a fair coin. On watermarked text it drifts upward. That drift is the entire watermark.

I find the coin-flipping analogy the clearest one. Imagine a rule that says: whenever you have a free choice, flip a keyed coin and let it break the tie. Anyone holding the same coin can replay every flip and notice you were winning far more often than chance allows. Anyone without the coin sees a person making ordinary, unremarkable choices.

Two consequences fall straight out of this design, and both matter enormously later.

The watermark needs the model to be uncertain. If one word has probability 0.999, every tournament branch picks it anyway, g-value or not. So high-entropy creative prose carries a strong watermark, while factual writing, code, and low-temperature output carry a weak one. Anthropic says this plainly in its own documentation: watermarking is sparser on factual passages where fewer choices can be made without hurting accuracy.

The watermark is a correlation, not a payload. There is no hidden data to extract. There is only a statistical relationship between the text and a specific key. Hold that thought.

Key takeaway

"SynthID biases which word wins among equally good options, using a keyed coin flip. Unwatermarked text averages 0.5. Watermarked text drifts above it. Detection is measuring that drift."

Building The Detector

Reading the source instead of the blog posts

I made one decision early that saved me from most of the mistakes I later watched other people make in public: I refused to work from secondary sources. Google published the reference code and a Nature paper. I read both, line by line, before writing anything.

Phase 1

Start from the primary sources, not the ecosystem

Every third-party explainer I checked had at least one material error.

  • Cloned Google DeepMind reference implementation and read all nine source files end to end before writing a single line of my own.
  • Cross-read the Nature paper for the parts the code does not explain, especially the theoretical ceiling on watermark strength.
  • Read the Hugging Face Transformers port, which Google itself points to as the production-ready implementation.
  • Deliberately ignored blog posts and tutorials until I had my own working mental model. Several turned out to state things the source code contradicts.
Phase 2

Extract the features, do not reinvent the maths

The watermark maths is published and tested. Rewriting it from scratch would only add bugs.

  • Drove the official Transformers implementation for g-value extraction rather than reimplementing the hashing myself.
  • Reimplemented only the scoring layer in NumPy, so the detector does not need to pull in JAX.
  • Ported the DeepMind reference hashing separately, used purely as a cross-check in tests rather than in the detector path.
  • That cross-check turned out to be the most valuable thing I built. More on that in the next chapter.
Phase 3

Handle the two masks correctly, or every score is wrong

The unglamorous detail that decides whether a detector works at all.

  • End-of-sequence mask: everything from the first padding token onward must be discarded, which forces right-padded tokenization. Left padding silently keeps the padding and throws away the text.
  • Repeated-context mask: during generation, watermarking is skipped whenever the current four-token context was already seen. This is what keeps the scheme distortion-free over repetitive text.
  • The detector has to skip exactly the same positions, otherwise it averages in unbiased noise and dilutes its own signal.
  • This is why tokens analysed is always lower than tokens submitted, and why repetitive text is genuinely harder to verify.
Phase 4

Build the statistics properly, with a stated error rate

A detector without a false-positive rate is not a detector, it is a vibe.

  • Derived the closed-form null distribution: under the no-watermark hypothesis, g-values are independent fair coins, so the score has mean 0.5 and a variance that shrinks with token count.
  • That gives a p-value and a length-aware threshold with no training data required, which matters because the reference README recommends computing thresholds empirically per length.
  • Validated it against Monte Carlo: thresholds built for a 1 percent error rate fired on exactly 1.0 percent of 20,000 simulated cases.
  • Also fitted an empirical null on real text as a check. It agreed closely with theory, so the analytic version is trustworthy on its own.

Key takeaway

"Detection needs no model weights and no GPU. It is integer hashing over token IDs plus a small amount of statistics. Scoring 3,924 texts took 53 seconds on a laptop CPU."

The First Surprise

Same config, same keys, different watermark

This is the finding I am most glad I stumbled into, because it would have quietly wrecked the project and I would never have known why.

Both implementations, Google's reference repo and the Hugging Face production port, derive an identical 64-bit hash from the same inputs. Same keys, same context, same candidate token, same linear congruential generator, same constants. Then they turn that hash into a g-value in two completely different ways.

The reference repo re-applies its hash function twelve more times, shifts the result, and takes a single bit out of position 30. The Transformers port takes a pre-generated table of 65,536 random bits and indexes into it using the hash modulo the table size.

Both are perfectly valid ways to produce a fair pseudorandom bit. Neither is wrong. But they are different functions, so they produce different g-values from the same text.

I wrote a test comparing them across thousands of tokens. They agreed 45 to 55 percent of the time, which is to say, exactly as often as two unrelated coins agree. Chance.

The practical implication is sharp: a watermark applied by one implementation is completely invisible to a detector built on the other, even with identical keys, identical n-gram length, identical everything in the config. If I had generated my test data with one library and detected with the other, I would have seen nothing at all, and the obvious conclusion would have been that my code was broken. I might have spent days there.

So a watermarking configuration does not define a watermark. The full specification is the key list, the n-gram length, the tokenizer, the context history size, and the g-value sampler implementation, including the exact random table and the seed that made it. Change any one and detection returns chance.

I turned that into an assertion in my test suite rather than a note in a README, because it is precisely the kind of assumption that rots silently.

There is a second, quieter version of the same trap. The g-values are hashed from token IDs, not characters. Detect with a different tokenizer and you are hashing a different integer sequence entirely, which produces pure noise. There is no partial credit and no warning, just a score near 0.5. If you do not know which model produced a piece of text, you cannot reliably check its watermark even holding the right key.

Key takeaway

"Two official implementations of the same published algorithm produce mutually invisible watermarks. Detection must match generation exactly, down to the random table, or it silently returns nothing."

Proving It Worked

Building a benchmark that can actually be wrong

It is easy to build a detector that looks perfect. Feed it your own watermarked text, watch it light up, publish the number. That proves almost nothing, and I wanted to know whether the thing was real.

The trap with any detector is that your positives and negatives usually differ in more ways than the one you care about. If I generate watermarked text with a model and compare it against Wikipedia articles, a high score might mean I detected the watermark, or it might mean I detected "sounds like a language model". Those are very different products.

So the core of the benchmark is a control that removes that ambiguity. For every watermarked sample I generated a twin: same model, same prompt, same random seed, watermark switched off. Two texts that are stylistically indistinguishable, that started from identical sampling noise, differing in exactly one variable.

Then I added human writing from Wikipedia and Project Gutenberg, plus output from a different model entirely, to catch a detector that had accidentally learned a house style.

Across 3,924 samples, at a threshold set for a 1 percent false-positive rate:

ROC-AUC 1.0000. Detection rate 100 percent. Measured false-positive rate 1.2 percent, against a 1 percent target, which is about as well calibrated as a few hundred negatives can demonstrate.

The number I actually care about is the breakdown by negative type. Human text produced a 1.4 percent false-positive rate. The other model produced 0.0 percent. And the same-model, same-prompt, same-seed twins produced 1.0 percent, with a median score of 0.4999 against a chance value of 0.5.

That last figure is the one that makes the detector real. Those samples are the hardest possible negatives, and they land on chance to four decimal places. The detector is reading the watermark, not the writing. No perplexity-based AI detector can make that claim about its own negatives, because it has no way to construct that control.

A couple of other things I checked, because they are the standard ways this kind of work goes wrong:

The theoretical ceiling from the Nature paper, assuming a maximum-entropy model, is a mean g-value of about 0.7500. My watermarked text scored a median of 0.6252, roughly half the available headroom. A score above the ceiling would have meant I was double-counting something.

I also refused to let the detector guess on short text. Below about 25 usable n-grams, no threshold separates the classes at a useful error rate, so it returns "inconclusive" with a reason instead of a verdict. Those refusals are reported separately and never folded into the accuracy figures, because trading a refusal for a coin flip is how detectors get to claim numbers they have not earned.

And I never reported accuracy. Accuracy depends on how many positives happen to be in your test set, so it describes your dataset rather than your detector. Rates, ranking quality, and detection-at-a-fixed-error-rate are the numbers that transfer.

Key takeaway

"The decisive test is not how well a detector scores positives. It is whether it stays at chance on negatives that differ from the positives in exactly one variable."

Hitting The Wall

The one-digit experiment that ended the project

By this point I had a detector that worked beautifully. Then I ran the test I had been circling for days, and it took about four seconds to answer the actual product question.

Everything above used the research keys that ship with Google's public reference code. Thirty integers, published so people can reproduce the paper. Real Gemini uses different keys, and those are secret.

I had been assuming this would be a matter of degree. Wrong key, weaker signal, maybe still something detectable if the text were long enough. So I took genuinely watermarked text and detected it with the same keys, incremented by one.

| | score | verdict | | correct keys | 0.6278 | detected, p = 3.26e-26 | | keys off by one | 0.5023 | nothing, p = 0.42 |

Same text. One digit different. The watermark did not weaken. It vanished, landing on chance as cleanly as a human-written paragraph.

That is when the mechanism I described in chapter two stopped being an interesting detail and became the whole answer. The watermark is not concealed information sitting inside the text, waiting for someone clever enough to notice it. It is a correlation between word choices and a keyed pseudorandom function, and that correlation only exists relative to the key. With the right key, bias. With any other key, fair coins.

SynthID is specifically engineered to be distortion-free: the statistical distribution of the text is provably unchanged by watermarking. That is a feature, so watermarking does not degrade output quality. It is also a wall. There is no entropy artefact, no perplexity anomaly, no fingerprint of any kind to find, because by construction there is nothing there until the key makes it appear.

I went looking for a way around it anyway. Every route dead-ends on the same property:

Brute-force the keys. Even assuming the small integer range Google uses in its research config, that is on the order of 10^90 combinations. And you would still need their exact tokenizer, depth, n-gram length and sampler variant, none of which are published.

Extract the keys from the model. They live in the provider's server-side sampling loop. They are not in the weights and never reach the client. For a closed model there is nothing local to pull them out of.

Steal the watermark statistically. This is the serious candidate, and researchers at ETH Zurich have done it. Bulk-query the model, learn an approximation. Against SynthID it achieved 4 percent spoofing success at baseline budget, rising to 15 percent at 90,000 queries, far more resistant than other schemes which collapse above 80 percent. But it detects whether a model watermarks, across a corpus, not whether one document does. It also needs an unwatermarked reference from the same model to compare against, which does not exist when every response from that model is watermarked.

Detect side-effects. Ruled out by the distortion-free property. And anything you did find would be generic AI detection wearing a costume.

Read provenance metadata. Works for images and files. Text pasted into a box has none.

So the honest answer to "can we detect it, ignoring which company made it" is no, and dropping the attribution requirement does not help at all. Presence and readability are the same problem here. You cannot tell that an envelope is sealed without the key that opens it.

This is not an obscure limitation, either. It has a name in the literature, public verifiability, and it is an active research problem. The dilemma is stated cleanly in a USENIX Security 2026 paper: the key cannot be public, or an adversary can strip and forge watermarks at will; but it cannot be private, or detection is opaque to everyone else. Proposed solutions involve building zero-knowledge proofs into the watermarking scheme itself. That requires the provider to adopt new machinery. It is structurally not solvable from outside.

Key takeaway

"Watermarked text detected with the wrong key is indistinguishable from human writing. Not degraded, identical. That is a mathematical property of the scheme, not a gap in tooling."

What Survives Editing

Three things about watermark fragility that surprised me

Since I had a watermark I fully controlled, I could measure something nobody can measure with a black-box detector: exactly how much editing it takes to destroy the signal. The results were not what I expected.

I ran every sample through a set of deterministic transformations, applied to positives and negatives alike so that each transformation has its own false-positive rate rather than borrowing one from unedited text. Detection rates below are at a 1 percent false-positive threshold.

Vocabulary does not matter. Word order does.

Replacing every single word that appeared in my synonym table left detection at 100 percent. Untouched. Meanwhile, swapping 20 percent of adjacent word pairs, a much smaller-looking edit, dropped it to 85 percent, and stacking edits together took it to 48 percent.

The reason is structural. Each hash covers a five-token window. Displace one word and you damage up to five overlapping windows, not one. So an edit that preserves sequence is nearly free, while an edit that disturbs adjacency is expensive. Swapping synonyms one-for-one keeps every window's shape intact. Reordering shreds it.

Making text shorter does not help.

I expected truncation to be the easy attack. It is not. Cutting text to a quarter of its length kept detection at 100 percent. The score is a rate, an average over positions, not a total that accumulates. Removing half the text removes half the evidence and half the positions, and the average sits exactly where it was.

What shortening does change is confidence. Evidence grows with the square root of token count, so a short passage becomes genuinely undecidable rather than falsely clean. Which is a much better failure mode for a detector than it sounds.

Low-entropy text is fragile in every direction.

I ran the whole benchmark twice, once at temperature 1.0 and once at 0.5. At the lower temperature the watermark starts weaker, a median of 0.5579 against 0.6252, roughly half the signal, simply because the model had fewer free choices to encode into.

And every attack then became dramatically more effective. Adjacent word swaps went from costing 15 points to costing 70. Heavy rewriting went from 48 percent detection to exactly zero, with ranking quality collapsing to 0.60, barely above a coin toss.

So robustness is not a property of the watermark alone. It is a property of the watermark and the entropy it was hidden in. Factual writing, technical explanation, anything constrained: weaker watermark, easier removal.

This lines up with what everyone else who has measured it reports. ETH Zurich found SynthID easier to scrub than comparable schemes, with ordinary paraphrasers achieving over 90 percent removal. Anthropic states it directly: light editing probably will not remove the watermark completely, but a rewrite where every word is replaced will. Brookings reaches the same conclusion from a policy angle, noting that modest paraphrasing can erase the signal and that watermarking is therefore of limited use in high-stakes settings.

Which raises the uncomfortable practical point. Even if you could detect these watermarks, the texts most likely to be edited before submission are exactly the ones you would most want to catch.

Key takeaway

"Word order is the vulnerability, not vocabulary. Truncation does nothing. And low-entropy text carries a weak watermark that ordinary rewriting removes entirely."

Who Can Actually Detect

Where each provider actually stands, as of August 2026

Once I understood the key dependency, the whole confusing landscape resolved into something simple. Here is the real state of play, with the one development that might genuinely change things.

Phase 1

Gemini: watermarked, and only Google can read it

Deployed at scale since 2024, following the Nature publication.

  • Google has been applying SynthID to Gemini text output for a while, and it is the first publicised large-scale deployment of LLM watermarking.
  • Google runs a SynthID Detector portal, but it is a limited-access pilot rather than a public service.
  • There is no public API for SynthID text detection. The consumer-facing routes that do exist, like the Chrome right-click check, cover images and media rather than arbitrary pasted text.
  • Net: nobody outside Google can verify a piece of Gemini text.
Phase 2

Claude: watermarked recently, and a detection API is promised

This is the one to watch, and it is a genuine change.

  • In August 2026 Anthropic confirmed that its text watermark is a version of the same SynthID-Text technique from the Nature paper.
  • The driver is regulatory: the EU AI Act has required providers serving the European market to mark AI-generated content since 2 August 2026, and Anthropic signed the transparency code in July 2026.
  • Anthropic applies the watermark globally rather than only in the EU, saying it does not yet have a durable way to scope it regionally.
  • Crucially, it states it will soon offer a watermark detection API, plus its own file-upload checker. That is provider-operated detection rather than a key release, and commercial terms are not public yet.
Phase 3

ChatGPT: no text watermark at all

A different situation that gets lumped in with the others.

  • OpenAI has not deployed text watermarking in ChatGPT.
  • So for ChatGPT output there is no watermark to detect, as opposed to a watermark that exists but is locked.
  • Worth being precise about, because "we check for AI watermarks" implies coverage that does not exist for the most widely used assistant.
Phase 4

The tools claiming to detect SynthID: almost all misleading

I checked. This is the part that annoyed me most.

  • One site marketed specifically as a SynthID detector states in its own copy that it does not read or cryptographically verify the watermark, because that requires Google private keys which are not public.
  • What it actually returns, in its own words, is an independent likelihood score. That is conventional AI detection with SynthID branding on top.
  • Credit for the disclaimer, which many do not bother with. The product name and marketing still tell a story the technology does not support.
  • There is no cross-provider watermark detector. Brookings has noted that building one would require the model developers to collaborate on a shared detection standard, which is likely years away.

Key takeaway

"Only the provider that generated the text can verify its watermark. The single realistic route to third-party checking is a provider opening up detection access, and Anthropic is the first to say it will."

What I Kept

A negative result, and a measuring instrument

I did not ship the tool. I want to be straight about that, because the temptation to dress up a dead end as a launch is real, and the industry is full of people who gave in to it.

Here is the thing I had to sit with. A public watermark checker built on this would be technically flawless and completely useless. For ChatGPT text there is no watermark to find. For Gemini and Claude text there is one, but we cannot read it. So the output would be "no watermark detected" for every single real submission, forever. Not inaccurate. Just a constant function wearing a lab coat.

And the failure mode is worse than useless, because people would read "no watermark detected" as "not AI-generated". The tool would be right and the user would be misinformed. That is the exact trap the existing tools fell into, and I did not want to add to the pile.

So what was the point?

First, a defensible negative result. "We cannot build this" is a real deliverable when it is backed by an implementation rather than a hunch. I can now explain precisely why, cite the mechanism, and produce the one-digit key experiment on request. That is a different thing from reading a few blog posts and guessing, and it is the reason I trust it enough to write this. Getting to a confident no required building the thing.

Second, and more usefully, a measuring instrument. Every AI-detector benchmark I have run has the same weakness: you are scoring against someone else's black box. When a number moves you cannot tell whether your text changed or their model did. Ground truth is borrowed.

Here, I control the watermark completely. I know exactly which texts carry it and exactly how strong it is. That makes it possible to ask a question you cannot otherwise ask cleanly: how much statistical signal does a given rewriting process actually remove? Not "did a third-party detector's score go down", but a measured quantity against known ground truth. Feed in watermarked text, run it through any rewriting pipeline, measure what is left. That harness works, and it outlived the product idea that motivated it.

Third, transferable knowledge. The fragility findings do not depend on whose key was used, because the mechanism is identical either way. Word order matters and vocabulary does not. Truncation is not an attack. Low-entropy text is weakly marked. Those hold for Gemini and Claude too, and they are the kind of thing worth knowing before you form an opinion about whether watermarking will work as a policy instrument.

The reproducible parts, for anyone who wants them: 95 tests, the closed-form false-positive threshold validated against Monte Carlo, mask computation verified byte-for-byte against Google's reference code, the two-implementation divergence asserted rather than assumed, and a wrong-keys test that fails loudly if the detector ever stops being key-dependent.

If I am honest about the ledger: a week spent, a working detector, a benchmark I still use, and a product idea correctly killed. I would rather kill an idea in week one with evidence than in month six with a support inbox.

Key takeaway

"The build was worth it for the negative result and the benchmark harness. It was not worth shipping, and no amount of engineering would have changed that."

If You Are Evaluating This

The questions to ask before you build a watermark detector

If someone is proposing a watermark-based AI detection feature, these are the questions that decide the outcome. In my experience the first one ends most conversations.

1

Whose key will you be detecting with?

If the answer is "the public research keys", you can only detect watermarks you generated yourself. If the answer is "the provider keys", you do not have them and will not get them, because releasing them would let anyone forge or strip watermarks. There is no third answer. Ask this first and you can save the rest of the week.

2

Do you know which model produced the text?

The watermark is hashed from token IDs, so detection needs the exact tokenizer of the generating model. On arbitrary pasted text you do not know this. Guessing wrong produces a clean-looking score of 0.5, which reads as a confident negative and is actually a silent failure.

3

Which implementation applied the watermark?

Google reference code and the Hugging Face port compute g-values differently from the same hash, and agree only at chance. Same keys, same config, mutually invisible watermarks. Detection must match generation down to the random table and its seed.

4

What false-positive rate are you quoting, and on how many negatives?

A score without a stated error rate is meaningless, and a threshold that is fixed across text lengths is wrong, because the null distribution narrows with token count. Claiming a 0.1 percent false-positive rate needs at least a thousand negatives to resolve it. Below that, report that you cannot measure it.

5

Have you built a same-model, same-seed control?

Generate paired samples that differ only in whether watermarking was on. If your detector fires on those negatives, it is reading style rather than the watermark, and you have accidentally built a mediocre AI detector. This is the single most informative test in the whole suite.

6

What happens on text too short to decide?

Below roughly 25 usable n-grams no threshold separates the classes usefully. The detector should refuse and say why. If it emits a verdict anyway, your reported accuracy includes coin flips.

7

Have you tested realistic edits, not just mechanical ones?

Synonym swaps and deletions put a floor under robustness loss, not a realistic estimate. Real LLM paraphrasing degrades detection considerably more. If your robustness numbers come only from mechanical transforms, label them as such.

8

What entropy regime will real inputs be in?

Low-temperature, factual, templated and technical writing carries a much weaker watermark. Benchmarking only on high-entropy creative prose gives you a best-case number that real traffic will not reproduce.

Key takeaway

"Question one settles it in almost every case. Everything after it only matters if you are working with a watermark you control, which is a research setup rather than a product."

Quick Answers

Common questions about SynthID and watermark detection

The questions I was asked most while working through this, answered as directly as I can.

Can you detect a SynthID watermark in Gemini or Claude text?
No. Not unless you are Google or Anthropic. Reading a SynthID-Text watermark requires the secret key that was used to generate the text, and those keys are private. Without the correct key, watermarked text is statistically identical to unwatermarked text. I verified this directly: the same watermarked passage scored p = 3.26e-26 with the correct keys and p = 0.42 with keys that were off by one digit. Not weaker. Gone.
What is SynthID-Text and how does it actually work?
SynthID-Text is a generative watermark for language models, published by Google DeepMind in Nature in 2024. It adds nothing to the finished text. Instead, at every step of generation, the model runs a small tournament between candidate next-words, scored by a keyed pseudorandom function called a g-value. Words that score 1 tend to win. Across hundreds of tokens this leaves a measurable statistical bias, so the watermark is the token pattern itself rather than any hidden character or metadata.
Why can a watermark not be detected without the key?
Because the watermark is not concealed information sitting inside the text. It is a correlation between the words chosen and a keyed pseudorandom function, and that correlation only exists relative to the key. SynthID is deliberately distortion-free, meaning the text distribution is unchanged, so there is no anomaly, no entropy artefact and no statistical footprint to find. With the wrong key you measure exactly 50/50 coin flips.
Is there a public SynthID detector or API?
Not for Gemini text. Google DeepMind runs a SynthID Detector portal, but it is a limited-access pilot with no public text API. Anthropic is the exception worth watching: in August 2026 it confirmed Claude uses a SynthID-Text variant and stated it will soon offer a watermark detection API plus a file-upload checker. That would be provider-operated detection, not a key release, and commercial terms are not yet public.
Are the online tools that claim to detect SynthID real?
Mostly no. They are conventional AI detectors with SynthID branding. One site marketed as a SynthID detector concedes in its own fine print that it does not read or cryptographically verify the watermark because that requires Google private keys, and that what it returns is an independent likelihood score. That is perplexity-style AI detection, which is a completely different mechanism.
Does ChatGPT watermark its text?
No. OpenAI has not deployed text watermarking in ChatGPT. So for ChatGPT output there is no watermark to detect at all, which is a different situation from Gemini and Claude, where a watermark exists but is key-gated.
Does editing or paraphrasing remove a SynthID watermark?
Editing that preserves word order barely affects it. Editing that disturbs word adjacency destroys it quickly, because each n-gram spans five tokens, so one displaced word damages up to five of them. In my benchmark, replacing every word in a synonym table cost nothing at all, swapping 20 percent of adjacent word pairs cost 15 points of detection, and stacked heavy rewriting cut detection from 100 percent to 48 percent. On lower-temperature text, heavy rewriting drove detection to zero.
Does making text shorter defeat a watermark?
No, and this surprises people. The detector score is a rate, not a running total, so truncation does not lower it. Cutting text to 25 percent of its length kept detection at 100 percent in my tests. What shortening does reduce is statistical confidence: evidence grows with the square root of the token count, so a very short passage becomes genuinely undecidable rather than falsely clean.
Why does watermark strength depend on temperature?
Tournament sampling can only bias positions where the model was genuinely uncertain. If one word has probability 0.999, every tournament branch picks it regardless of its g-value. So high-entropy text carries a strong watermark and low-entropy text carries a weak one. At temperature 1.0 my watermarked text scored a median of 0.6252 against a 0.5 baseline. At temperature 0.5 the same setup scored 0.5579, roughly half the signal, and every removal attack became far more effective.
Do all providers share one watermarking key?
No. Each provider holds its own key, which is why a cross-provider detector does not exist. Brookings has noted that building one would require the model developers to collaborate on a shared detection standard, which is likely years away. Until then, only the originating provider can verify its own output.
Could you brute-force or steal the watermark key?
Realistically no. The key is a list of 30 integers, and you would also need the exact tokenizer, watermarking depth, n-gram length and sampler variant, none of which are published. Researchers at ETH Zurich have shown a related attack, watermark stealing, which learns an approximation from bulk API queries. Against SynthID it achieved only 4 percent spoofing success at baseline budget and 15 percent at 90,000 queries, and it detects whether a model watermarks rather than whether one specific document does.
Is keyless watermark detection an open research problem?
Yes, and it has a name: public verifiability. The dilemma is that the key cannot be published without letting attackers strip or forge watermarks, but keeping it private makes detection opaque to everyone else. Papers such as PVMark at USENIX Security 2026 propose zero-knowledge-proof schemes to allow third-party verification without revealing the key. Those require the provider to adopt the machinery, so it cannot be solved from outside.

References

Primary sources

Everything in this write-up traces back to one of these. I have deliberately weighted it toward primary sources: the paper, the source code, and the providers own statements. Where I cite a measurement of my own, it came from the implementation described above.

Scalable watermarking for identifying large language model outputs — Nature (2024)

The peer-reviewed SynthID-Text paper by Dathathri et al. The Supplementary Information is where the theoretical ceiling on watermark strength comes from, including the result that a maximum-entropy model tops out around a 0.75 mean g-value.

https://www.nature.com/articles/s41586-024-08025-4

google-deepmind/synthid-text — reference implementation

Google DeepMind open-source reference code, Apache 2.0. The files worth reading are the logits processor, the hashing function, and the two detector implementations. This is where the public research keys and the default configuration live.

https://github.com/google-deepmind/synthid-text

How Claude text watermarking works — Anthropic

Anthropic explanation, published August 2026, confirming Claude uses a SynthID-Text variant. Contains the statement that the watermark is detectable to anyone who has a key that encodes it, the promise of a forthcoming detection API, and a candid section on how a full rewrite removes the watermark.

https://www.anthropic.com/news/claude-text-watermark

Probing Google DeepMind SynthID-Text Watermark — ETH Zurich SRI Lab

The most rigorous public attack analysis of SynthID specifically. Source for both the difficulty of spoofing it, 4 percent at baseline budget rising to 15 percent at 90,000 queries, and the ease of scrubbing it, with over 90 percent removal by ordinary paraphrasers.

https://www.sri.inf.ethz.ch/blog/probingsynthid

PVMark: Enabling Public Verifiability for LLM Watermarking Schemes — USENIX Security 2026

The clearest statement of the core dilemma: the key cannot be public without enabling removal attacks, and cannot be private without making detection opaque to third parties. Proposes zero-knowledge proofs as a route to public verifiability, which requires provider adoption.

https://arxiv.org/abs/2510.26274

Detecting AI fingerprints: A guide to watermarking and beyond — Brookings

The most readable non-technical treatment I found, and useful for policy context. Notes that provider detectors can only find their own watermarks because they depend on that company keys, and that a cross-provider detector would require developers to agree on a shared standard.

https://www.brookings.edu/articles/detecting-ai-fingerprints-a-guide-to-watermarking-and-beyond/

Watermark Stealing in Large Language Models

The paper behind the stealing attack referenced above. Demonstrates that querying a watermarked model enables practical spoofing and boosts scrubbing, with success rates over 80 percent against other state-of-the-art schemes, though SynthID proved substantially more resistant.

https://arxiv.org/pdf/2402.19361

SynthID — Google DeepMind

Google overview of SynthID across text, image, audio and video, and the entry point for the SynthID Detector portal that remains in limited-access pilot.

https://deepmind.google/models/synthid/

Found this valuable?

Share this architectural breakdown with your network.