跳到正文
Hacker News · AI· danielmorozoff·· 3 小时前AI 评分50

Kevin Roose 公开 Dario Amodei 2017 年内部文档《Big Blob of Compute》

Big Blob of Compute: The Essay That Started the AI Race

AI 导读

记者 Kevin Roose 在其 Substack 上公开了 Dario Amodei 2017 年在 OpenAI 内部撰写的文档《The Big Blob of Compute Hypothesis》前 10 页。

正文
This is what AI thinks a “big blob of compute” looks like

In 2017, a curly-haired OpenAI researcher named Dario Amodei had a hunch.

At the time — and keep in mind, 2017 was a bygone era in AI research, just after AlphaGo but just before GPT-1 and the LLM revolution — most AI researchers still thought of themselves as toolmakers. They’d build specialized models to go after specific, narrow problems: language translation, speech recognition, robotics, games. Each of these domains required different algorithms, different data sets, and different approaches.

But Amodei had a different idea. He thought that a better way to make AI more intelligent was simply to make it bigger. In his own experiments, including a speech-recognition system he’d worked on at Baidu, he observed that a lot of improvements had happened as a result of scaling — adding more data and more compute to a model’s training process, letting it train for longer, and watching it magically improve. And he knew that in the natural world, all else being equal, animals with big brains were usually smarter than animals with small ones. He believed that if AI followed a similar principle, the best way — and maybe the only way — to make real progress was to stop building small, narrow models, and start scaling up general-purpose models instead. (Similar ideas were later popularized by Richard Sutton’s famous essay “The Bitter Lesson,” but Amodei beat Sutton to it by a few years.)

The document Amodei wrote to explain his views on scaling to his colleagues — titled “The Big Blob of Compute Hypothesis” — became a classic among early OpenAI researchers. It described, in florid specifics, why Amodei thought that most AI researchers were being too clever. They would engineer all kinds of complicated solutions to whatever problems their models were having, when a much simpler solution — just make the thing bigger! — would work better.

“Creating any given intelligent behavior is mostly about providing a large, minimally structured mass of computational capacity, and then giving it shape and form via interactions with a rich environment and a training process that drives it towards behavior appropriate for the task at hand,” Amodei wrote.

To illustrate what he meant, Amodei used the analogy of creating snowflakes. If you wanted to make a snowflake, you wouldn’t try to assemble individual ice shards with tweezers. Instead, you would create a container with the right conditions, and let the snowflakes form on their own.

The way to make a snowflake is not to think in terms of its pieces but to know the laws of physics (training target), have enough raw material (compute) and a large enough chamber (parameters), set the temperature, pressure, and humidity correctly (normalization, shape), and wait for long enough (quantity of experience). Furthermore, this is your only way to make snowflakes and your only leverage over their shape; trying to piece together a single one from little bits of ice is basically hopeless.

Amodei also thought that AI safety researchers suffered from micromanagement. Some researchers in his orbit, like the ones at Eliezer Yudkowsky’s MIRI (more recently in the news for other reasons) were arguing for a cybersecurity-like approach to AI safety, engineering complicated safety protocols to prevent models from going off the rails. Amodei disagreed. He thought that a better way to create safe AI was to get the high-level conditions right, and let models figure out on their own what was safe or unsafe.

“Big Blob of Compute” was never published outside of OpenAI — and Amodei didn’t want to comment on it — but the essay has taken on mythical status inside the industry. Several early OpenAI employees told me that it had reshaped their thinking about AI, and convinced them that scale mattered more than they thought. In a very real sense, “Big Blob of Compute” may have started the AI race, since it led OpenAI to start scaling up its language models in order to test Amodei’s hypothesis — the effort that eventually resulted in GPT-2, GPT-3, and everything we see today.

I obtained a copy of “Big Blob of Compute,” and am publishing it here for a few reasons.

  1. It’s an important historic document — something that should end up in the Museum of How We Got Here, when the robots get around to building one — and I think it’s useful to hear how Amodei, now the CEO of Anthropic, developed his conviction in scaling. If AI transforms society the way Amodei thinks it will, documents like these may be the equivalent of Darwin’s notebooks, and I think there’s a strong public-interest argument for letting people read them. (Note: the full document runs to more than 25 pages. I am publishing only the first 10, since the rest is very dense and technical and consists mainly of Amodei arguing with various other AI safety researchers about why their suggested approaches are wrong.)

  2. Reading “Big Blob of Compute” gives you some insight into Amodei’s personality, how he communicates with his fellow researchers, and the intellectual influences that shaped his thinking — all of which are important to understand about someone who runs one of the world’s most important companies. It also proves how ridiculous the whole conspiracy theory that Amodei and other AI leaders recently started talking about exaggerated doomsday scenarios as a tactical smokescreen for regulatory capture is. These guys have been thinking and writing about AI doomsday since before there was anything to regulatory capture!

  3. I am trying to tempt you to buy my book, The AGI Chronicles, by showing you that there are many juicy secrets and never-before-published research documents like “Big Blob of Compute” in it. The book comes out tomorrow, and pre-orders (which are disproportionately helpful to authors, since they count toward first-week sales, which in turn count toward bestseller lists) are especially valuable. I don’t charge for this Substack, but if you’d like to do the equivalent of putting a tip in my jar, you can order The AGI Chronicles at Amazon, Audible, Bookshop, or wherever you buy or listen to books. I worked insanely hard on it, and I’m proud of how it turned out. I think it’s a great book to give to someone who is just getting interested in AI and wants to understand the lay of the land, and it has enough new revelations in it to surprise even the most AGI-pilled insiders.

Now here, for posterity, is “The Big Blob of Compute Hypothesis” by Dario Amodei:

A few other housekeeping notes:

  • I’m leaving today on a ten-day book tour that will take me to Washington, New York, Morristown, Greenwich, Cleveland, and San Francisco, and will involve me chatting on stages with human superintelligences like Thomas Friedman, PJ Vogt, Joanna Stern, and Sebastian Mallaby. If you’re in one of those places, come say hi! Full tour schedule is available here. (Some of the events are sold out, but there are still tickets left for my 92NY event with PJ in New York on Wednesday night, which should be very fun.)

  • When I get back from book tour, we’re starting production on Machine Gods, the new video-first podcast I’m hosting with Casey Newton in partnership with NPR. I’m so excited about this, and we’re lining up some really special stuff for the first few episodes. If you want to hear them, you can subscribe to the new feeds on Spotify, Apple Podcasts, YouTube, or wherever you get your shows.

  • I am now a contributing writer at The Atlantic! (The Atlantic also published the first excerpt from The AGI Chronicles, which is the backstory of the feud between Dario Amodei and Sam Altman, and why the Anthropic crew left OpenAI in 2020 to start their own company.) Co-hosting Machine Gods is my main gig, and will remain my main gig for the foreseeable future. But in the event I decide to start writing words again, it will be fun to have an esteemed home for them.

  • Press for the book is starting to roll in. There’s a SF Chronicle story, a CBS News story, and a few starred reviews. (Kirkus and Publishers Weekly both liked it.) I am also in a Netflix “Instadoc” that apparently premieres this week, on the subject of AI and the HuggingFace attack. If you enjoy watching my hair blowing in a gentle Bay breeze while I’m being filmed outside an abandoned airplane hangar (long story), you’ll want to tune in.

  • Lastly, I’m sorry if you subscribed to this Substack expecting more regular updates. This was supposed to be a kind of reporter’s notebook where I’d drop little nuggets as I wrote. But in the end, the effort required to report, research, and write a 450-page book in a little more than a year, while also hosting a podcast and writing occasional newspaper columns, was a lot more than I bargained for. If and when I get some free time after this tour, I’ll probably use it to catch up on TV shows or hang out with my family rather than writing on Substack. But you never know.

No posts

来源:Hacker News · AI · kevinroose.substack.com