The problem is that they trained the models using millions of pirated books in standard english.
AAE is mostly used when spoken: they also pirated also millions of tv series and youtube videos that can contain that, but as of now, it was mostly for training voice recognition models
The problem is that they trained the models using millions of pirated books in standard english.
AAE is mostly used when spoken: they also pirated also millions of tv series and youtube videos that can contain that, but as of now, it was mostly for training voice recognition models
(proof that they pirated television content and youtube videos to train whisper: https://community.openai.com/t/subtitles-created-by-amara-org-qtss-etc/462561 - https://gist.github.com/riotbib/3b3c5f817b55b68801d14b8bdb02df09)