How to train your data | The Vergecast

发布时间    来源
Episode 设置


登录已过期或未登录,无法修改。请先登录后再试。

摘要

Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they'd really prefer you not know what's in it, and whether training data could ever be a fair trade. 00:00 Intro 01:02 90 Seconds on The Verge 03:18 Why Training Data Matters 08:43 Common Crawl and Filtering 11:51 Academia and Data Laundering 15:37 YouTube as Data Mine 20:01 Synthetic Data Myth 21:59 Paying Creators for Data 23:13 Wrap Up and Credits Subscribe: http://goo.gl/G5RXGs Like The Verge on Facebook: https://goo.gl/2P1aGc Follow on Twitter: https://goo.gl/XTWX61 Follow on Instagram: https://goo.gl/7ZeLv Watch The Vergecast on YouTube: https://bit.ly/40RFRkg The Vergecast Podcast: https://bit.ly/3WQDexZ Decoder with Nilay Patel: http://apple.co/3v29nDc More about our podcasts: https://www.theverge.com/podcasts Read More: http://www.theverge.com Community guidelines: http://bit.ly/2D0hlAv Wallpapers from The Verge: https://bit.ly/2xQXYJr Shop our Verge merch store here: https://bit.ly/4kPCmEc Subscribe to The Verge: https://bit.ly/3FT6n5S If you buy something from a Verge link, Vox Media may receive a commission without exerting any influence on editorial content. For more information about our ethics policy, visit: https://bit.ly/3ZWTlLs

GPT-4正在为你翻译摘要中......

中英文字稿