Nvidia’s New Hedge Against Chip Competitors? Partner with Them
The AI chip leader’s new deal with d-Matrix shows how Nvidia CEO Jensen Huang is adapting to a world with more competitors.
The Takeaway
- Nvidia partners with AI chip rivals like d-Matrix to combine hardware.
- Strategy helps Nvidia adapt to competition and generate new revenue.
- Using chips made by different designers to handle the same AI task is a fairly new concept.
Nvidia has a plan to deal with the growing list of AI server chip competitors: partner with them.
In the latest example, Nvidia and AI chip startup d-Matrix are combining their respective hardware in a new system to power AI models, the companies told The Information exclusively.
The move comes a month after another AI server chip designer, SambaNova, said it worked with Nvidia to make its chips for powering AI models work in concert with Nvidia graphics processing units.
Nvidia is now hinting that more such deals could be in the works.
“I won’t pre-announce the others,” said Dion Harris, an Nvidia senior director of high-performance computing.
Companies such as d-Matrix and SambaNova don’t require partnerships for their chips to work with Nvidia’s—they can simply plug directly into an Nvidia GPU using readily available Ethernet cables. But with the recent partnerships, Nvidia engineers work with those other chip companies’ engineers to tweak software controlling the GPUs so they work better with the other chips. And a formal collaboration increases the chances a potential customer will want the combined system.
The strategy shows how CEO Jensen Huang has adapted to a growing array of potential threats to Nvidia’s dominant share of the AI chip market. By working with would-be rivals, Nvidia is putting itself in a position to generate revenue alongside them if the newbie chips succeed.
Nvidia “definitely had a change in strategy and their view of the world,” said Thomas Sohmers, chief technology officer at AI chip startup Positron, which isn’t currently working with Nvidia but is open to it. “Nvidia is playing nice and building out so that they’re actually part of that heterogeneous ecosystem rather than fighting it tooth and nail.”
The partnership strategy may also help it beat back past allegations that Nvidia pressured customers to stick with its hardware. The Department of Justice two years ago initiated an investigation into such allegations, The Information reported at the time, but no case has publicly materialized.
Harris said Nvidia has always wanted to be a broader AI infrastructure company. “We are not just a chip company,” he said.
The effort to broaden Nvidia’s approach accelerated last year when Nvidia said other chip developers would be able to run their chips in conjunction with NVLink, the company’s networking gear that connects AI servers to each other so they operate more efficiently. That could help Nvidia sell more of its networking gear, if not its chips.
“We would always rather sell something than nothing,” Harris said.
In December, Nvidia paid $20 billion to license technology and hire key staff from Groq, an AI inference chip designer that was trying to compete with Nvidia. The move was akin to an acquisition, though Groq has continued to operate as a standalone firm while Nvidia develops a server rack that combines Nvidia GPUs with Groq’s chips. It isn’t clear how much demand Nvidia is fielding for the product.
Then came the SambaNova partnership. Nvidia sees the startup as “more partners than people might expect,” SambaNova CEO Rodrigo Liang said.
To be sure, Nvidia in recent years has actually increased its share of the market for AI inference chips despite rising competition from the likes of Google and Amazon, according to The Information’s estimates. And Nvidia CEO Jensen Huang insists the company’s GPUs can do all inference more effectively than its competitors, though he isn’t leaving it to chance: Nvidia is using its powerful balance sheet to help more companies buy its expensive AI chips, including backstopping young cloud providers that want to rent out the chips.
But numerous other firms are also getting into the inference market or considering how to do so, including Microsoft, Meta, OpenAI and, most recently, Anthropic. That could change Nvidia’s share of the market down the line. (See The Information’s AI Chip Database.)
An OpenAI spokesperson said the company hasn’t yet decided whether it will work with Nvidia to operate its chips with OpenAI’s simultaneously. Nvidia has invested heavily in OpenAI and is in talks to backstop a giant data center project for OpenAI in Ohio.
“I think the world is going to a place where [Nvidia] is perfectly comfortable” with a mix of chips powering different pieces of the same AI task, said d-Matrix CEO Sid Sheth.
Upstart chip designers have historically positioned themselves as Nvidia killers. But given Nvidia’s substantial lead, they’re increasingly seeking to work in conjunction with Nvidia’s GPUs rather than replace them entirely, just as Nvidia is trying harder to work with those newer firms.
In fact, it was Nvidia that approached d-Matrix about partnering, according to a person who spoke to the startup about it. (Both companies are based in Santa Clara, Calif.)
AI developers such as OpenAI have long used different Nvidia chips to power the same AI task, after realizing less powerful GPUs were more efficient at handling certain workloads. But using chips made by different designers is a fairly new concept.
Nvidia isn’t the only chip developer to embrace that concept. Amazon, which develops Trainium chips used by Anthropic and soon OpenAI, said in March it would develop a combined AI inference server system with AI chip designer Cerebras.
Splitting up the process of running AI models between completely different types of chips is sometimes known as disaggregated inference. In the version involving Nvidia and SambaNova or Groq, Nvidia GPUs handle prefill, the most compute-intensive part of running an AI model, while the upstarts’ chip handle the rest, known as decode.
In d-Matrix’s case, its chips as well as Nvidia’s each handle some prefill and decode. D-Matrix focuses on powering a process known as speculative decoding that speeds up an AI model’s performance. When an AI customer asks a model to complete a task, the d-Matrix chip runs a small AI model that guesses what a bigger model’s answer to the customer should be, and the bigger model running on an Nvidia GPU then verifies and accepts the guesses.
‘Perfectly Comfortable’
“I think the world is going to a place where [Nvidia] is perfectly comfortable” with a mix of chips powering different pieces of the same AI task, said d-Matrix CEO Sid Sheth. “It’s not going to be a GPU-only world we live in moving forward.”
D-Matrix, founded in 2019, last raised $275 million at a $2 billion valuation in November but is in talks to fundraise again, according to three people with direct knowledge of the discussions.
Sheth said d-Matrix sets itself apart by putting compute and memory on the same chip and by not using the kind of high-bandwidth memory that Nvidia uses for its chips, and which is in short supply these days.
Taiwan Semiconductor Manufacturing Co., which produces AI chips for Nvidia and many others, started producing d-Matrix chips earlier this summer and plans to produce thousands of them per month by the end of this year, Sheth said.
The startup is currently generating single digit millions in revenue, he said, adding that he hopes the company’s chips will consume 30 to 40 megawatts of power in data centers next year—a relatively small sum—and run AI coding as well as voice and video.
Parasail, a young AI cloud provider based in San Mateo, Calif., will be the first company to purchase the combined Nvidia/D-Matrix server system. Parasail plans to rent out the system to its customers starting later this year. Parasail CEO Mike Henry said the combined server system was attractive because it helps his company avoid being too dependent on buying new Nvidia hardware.