跳到正文
Hacker News · AI· cdrnsf·· 3 小时前AI 评分12

AI 公司是寄生虫

AI Companies Are Parasites

AI 导读

文章将 AI 公司比作寄生虫,指其抓取全网数据、摧毁书籍、无偿攫取内容并以"合理使用"为名训练模型,再按量售卖模型访问权。作者称流量换索引的旧契约已消失,爬虫压垮基础设施、点击率与社区崩塌,而 Reddit 等数据授权依赖的仍是人类生成的社区内容。

正文

Merriam-Webster:

An organism living in, on, or with another organism in order to obtain nutrients, grow, or multiply often in a state that directly or indirectly harms the host.

That's it, right? The whole thing, their entire business model.

  1. Scrape the whole of the internet, destroy books, siphon everything they can from whoever they can without compensation, and call it fair use.
  2. Train models on said stolen data.
  3. Sell metered access to the models.
  4. Repeat until you've killed the host.
  5. Hope you can head off the effects of AI inbreeding.

The old bargain of traffic for indexing is gone. Scrapers overload infrastructure. Clickthroughs collapse. Communities collapse.1

Yes, licensing deals exist. Reddit is selling access to its data, but that data is human-generated and depends on a flagging community the company seems, at best, indifferent to.

They'll even show up in your community and set up a data center that nobody (well, except for the people profiting from it) wants. Maybe they'll spin up gas generators and send pollution your way. Or they'll drive up your electricity and water rates. But what about jobs? Most of those only exist during construction and, who knows, they might just hire folks from outside your area for that.

They never should've attached themselves to society, and it may be too late to dislodge them. At least we get stilted prose, images that look like Pixar rehashes, and code nobody understands.


  1. How's Stack Overflow doing lately? ↩︎

来源:Hacker News · AI · coryd.dev