How to train your data | The Vergecast
发布时间 来源
Episode 设置
摘要
Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they'd really prefer you not know what's in it, and whether training data could ever be a fair trade.
00:00 Intro
01:02 90 Seconds on The Verge
03:18 Why Training Data Matters
08:43 Common Crawl and Filtering
11:51 Academia and Data Laundering
15:37 YouTube as Data Mine
20:01 Synthetic Data Myth
21:59 Paying Creators for Data
23:13 Wrap Up and Credits
Subscribe: http://goo.gl/G5RXGs
Like The Verge on Facebook: https://goo.gl/2P1aGc
Follow on Twitter: https://goo.gl/XTWX61
Follow on Instagram: https://goo.gl/7ZeLv
Watch The Vergecast on YouTube: https://bit.ly/40RFRkg
The Vergecast Podcast: https://bit.ly/3WQDexZ
Decoder with Nilay Patel: http://apple.co/3v29nDc
More about our podcasts: https://www.theverge.com/podcasts
Read More: http://www.theverge.com
Community guidelines: http://bit.ly/2D0hlAv
Wallpapers from The Verge: https://bit.ly/2xQXYJr
Shop our Verge merch store here: https://bit.ly/4kPCmEc
Subscribe to The Verge: https://bit.ly/3FT6n5S
If you buy something from a Verge link, Vox Media may receive a commission without exerting any influence on editorial content. For more information about our ethics policy, visit: https://bit.ly/3ZWTlLs
GPT-4正在为你翻译摘要中......
