What they don’t do is filter out every web page that has the canary string. Since people put them on random web pages (like this one), which was not their intended use, they get into the training data.
As others have mentioned, this seems kinda crazy and bad. I was surprised you didn’t think this.
“Unrelated” question, but are you under a non-disparagement agreement with GDM that would prevent you from criticizing things like their data-filtering practices?
I am not under nondisparagement agreements from anyone and feel free to criticize GDM. I do still have friends there, of course. I certainly wouldn’t be correcting misapprehensions about GDM if I didn’t believe what I was saying!
As others have mentioned, this seems kinda crazy and bad. I was surprised you didn’t think this.
“Unrelated” question, but are you under a non-disparagement agreement with GDM that would prevent you from criticizing things like their data-filtering practices?
I am not under nondisparagement agreements from anyone and feel free to criticize GDM. I do still have friends there, of course. I certainly wouldn’t be correcting misapprehensions about GDM if I didn’t believe what I was saying!
Thanks!