Why Google Still Needs Nvidia; The AWS-OpenAI Marriage That Could Have Been
Google unveiled its next-generation artificial intelligence chip at its annual enterprise software conference, Google Cloud Next, on Tuesday. But it still felt the need to tout a different kind of win: a partnership with Nvidia to offer that company’s state-of-the-art AI chips through Google Cloud alongside its own hardware. Nvidia CEO Jensen Huang even appeared on stage at the conference—wearing his signature leather jacket of course—to field questions from Google Cloud CEO Thomas Kurian about how the company’s chips would benefit Google’s customers.
The dynamic reflects a cold reality for Google. Many AI engineers prefer to use Nvidia’s hardware, which is in short supply. But beggars can’t be choosers. Those same developers will take computing capacity wherever they can get it, given the industry-wide shortages for such specialized hardware, according to several attendees. “There’s never enough compute, and it’s never at a low enough cost,” said Saurabh Baji, senior vice president of engineering at AI startup Cohere, at a panel discussion on Tuesday.
That has Google playing a balancing act. It needs Nvidia’s hardware to draw customers to its platform. But it also wants to drive usage of its Tensor Processing Units, the specialized chips in which it has invested heavily in recent years. Google is taking steps to make it easier for developers to use both chips. Google’s software for building large machine-learning models will now work for Nvidia hardware in addition to Google’s chips, the company announced Tuesday.
But that alone likely won’t be enough to drive meaningful adoption of TPUs. One AI developer told me that many companies want to avoid relying on Google’s software tools, because doing so would make it harder for them to use Nvidia’s chips. Some companies also lack the specialized knowledge needed to get the same performance out of TPUs as they would from Nvidia’s H100s, this person said. That helps explain why Google has seen muted uptake from customers even as it has successfully deployed the chips internally. Still, some companies, like Anthropic, use both TPUs and GPUs.
The holy grail for Google would be a chip so good that companies would happily change their software to use it. But cost and availability are decisive factors as well. If Google can offer a close substitute more cheaply, it would have an edge in signing up startup customers. There are signs Google might be making progress. While earlier generations of TPUs were better for training models than for running them, the more recent designs are useful for both aspects of AI development, Ori Goshen, co-founder and co-CEO of AI startup AI21 Labs, told me. Read our quick-hit summary of the event here.—Jon Victor and Anissa Gardizy
Here’s what else is going on…
How AWS Lost Its Footing
They say that pride often comes before a fall. Amazon may have learned that the hard way.
The cloud provider, which pitched in on the original funding of OpenAI when it was formed as a nonprofit research group in 2015, turned down a later opportunity to invest in the large-language model developer, according to an in-depth story out this morning by my colleagues Anissa Gardizy and Kevin McLaughlin.
OpenAI’s deal would have required AWS to provide it with cloud computing resources, crucial in the training of large-language models, but the startup offered the cloud provider no equity in exchange. It's a pity AWS didn't negotiate harder and reach a deal similar to the one Microsoft landed, which would have allowed it to integrate the startup's technology into its products the way its competitor has.
In many ways, a partnership between AWS and an LLM developer like OpenAI would have made perfect sense: AWS’ SageMaker product, which helps companies build machine learning models, could have been used by OpenAI to develop its models, and OpenAI could have gotten quick access to millions of AWS customers.
On the bright side, AWS dodged a major bullet by choosing not to release its own LLM for cloud customers last November during its annual developer event. ChatGPT launched days later, which would have put AWS’ LLM to shame, and then some.
The question now is how AWS can come back from its blunders. One path is its in-house chips, which the company has been peddling to customers in the face of rising shortages of Nvidia H100’s. Already, about 40,000 AWS customers use its general purpose Graviton chips, and its Trainium and Inferentia chips for training and running models, respectively, could also serve as potential draws for customers. With more startups putting models into production, those Inferentia chips could start to look especially appealing, as we explained here.
And as AI startups look to curb outlandish spending and move toward sustainable business models, they may take steps such as using smaller models that require fewer resources. AWS’ chips could offer a better cost-to-performance ratio in those cases. Chetan Kapoor, director of product management at Amazon EC2, told me recently that the company’s Trainium chips are well-tuned for training small to medium sized models while also offering cost savings.—Stephanie Palazzolo
OpenAI is Probably a Real Business :)
We all know Nvidia is making bank as the hardware provider for the artificial intelligence boom. But how much money can large-language models that run on Nvidia chips make? Put another way, how much value can the models create for everyday customers? It’s a trillion-dollar question that has hung over the industry this year.
One barometer is OpenAI, which stormed out of the gate with a paid version of ChatGPT earlier this year and also, it turns out, has plenty of big customers buying expensive access to a speedy version of GPT-4, the language model that powers the chatbot, according to my colleagues Amir Efrati and Aaron Holmes. If OpenAI is already generating more than $83 million per month in revenue, as the story suggests, then it seems like a matter of time before the company can turn a profit. (If that profit materializes, Microsoft is going to get a lot of it, per Microsoft’s agreement to fund OpenAI to the tune of $10 billion earlier this year.)
Another nugget that stood out to us was a notable hire by Jane Street, the secretive Wall Street firm that OpenAI counts as a major customer (as does Microsoft, we just found out). David So, a key large-language model researcher at Google who worked on its Palm model, now works for Jane Street. That suggests the company will either tweak existing models from OpenAI and other proprietary providers or it wants to build some models of its own as it looks to improve its gigantic market-making business. Either way, for all the talk about LLMs’ downsides, finance folks seem to have figured a way to make money from them.—Amir Efrati