Jensen Huang, co-founder and CEO of Nvidia, is increasingly being viewed as the tech industry's leading visionary. During a press Q&A session held on the sidelines of the GTC 2024 conference in Silicon Valley, he projected an extraordinary surge in data center infrastructure spending driven by generative AI. According to Huang, the market for data center infrastructure, currently estimated at $250 billion per year, will grow to $1 trillion by 2028 as enterprises invest heavily in new AI-ready systems.
Huang's confidence rests on Nvidia's dominant position in accelerated computing. The company's most recent quarterly revenue reached $22.1 billion, a 265% increase compared to the same period last year. That figure nearly matches the $25 billion in quarterly revenue reported by Dell, the world's largest server maker. Intel, the leading supplier of server processors, generated $15.4 billion in revenue during the same quarter, while AMD, Intel's main rival, posted $6.17 billion.
The Data Center Ambition
Huang was quick to clarify that Nvidia's ambition extends far beyond graphics processing units. "We are not the only ones making GPUs, but the GPU market is not the one we are targeting. Our commercial opportunity is the data center," he said. He predicted that the data center market would grow at an annual rate of 20% to 25%, driven by investments in accelerated servers and dedicated AI infrastructure.
Nvidia, Huang explained, is no longer just a chip vendor. The company designs entire data center infrastructures, including compute nodes, networking fabrics, and management software. "We are capable of assembling an entire data center, configuring it for maximum performance, and then decomposing it into functional blocks so that you can build your own data center with the network, storage, and administration console of your choice," he said.
He also emphasized that Nvidia works with traditional server makers rather than competing with them. "We make DGX servers that you can buy, but we also sell the components to Dell so they can build and sell HGX configurations based on their own servers. And we help Dell sell them," Huang stated. The company follows a similar model in the cloud, offering DGX Cloud through hyperscaler platforms while simultaneously helping those providers deploy Nvidia architectures in their own data centers.
Rethinking Compute Architectures
According to Huang, the rise of generative AI is fundamentally changing the technologies used inside data centers. He argued that future servers will increasingly rely on High Bandwidth Memory (HBM) integrated into processors, rather than traditional DDR memory modules on motherboards. This shift, he said, is key to improving energy efficiency and enabling more complex AI workloads.
As an example, Huang pointed to Nvidia's Earth-2 climate prediction software. He claimed that a compact cluster using Nvidia GPUs and HBM can predict weather anywhere on Earth with 3-kilometer accuracy while consuming 3,000 times less energy than a massive cluster of standard servers that would have been required to achieve similar results.
"The data center infrastructure market is wondering how to push the limits. But pushing the limits is no longer just about gaining power. It's about saving energy," Huang said. He explained that training a model like GPT-4 would have taken 90 days on 8,000 H100 GPUs, consuming 15 megawatts of power. With Nvidia's newer B200 GPUs, the same training run could be completed in 90 days using just 2,000 GPUs and only 4 megawatts. The B200 chip includes more than twice as much HBM memory as its predecessor.
Huang also suggested that Nvidia's architecture would eventually extend to personal devices. He painted a future where smartphones, PCs, and tablets generate most of the data they display locally rather than fetching it from remote servers. "What costs time and energy today on smartphones, PCs, and tablets is that everything you ask them to do requires sending a request over the network to download data," he said. "In the future, that will no longer be the case. Your devices will generate the essential data you expect locally, in a way that is relevant to your usage context."
This vision is supported by Nvidia's new Blackwell architecture, which includes dedicated circuits for transformer models. According to Huang, these circuits are particularly well-suited for generating photorealistic video frames from a compact embedded memory, using only a short description downloaded from the internet. He drew a parallel to vectorization, a technique long used in graphics to describe images mathematically rather than storing every pixel.
"Our GPUs started with the generation of synthetic images, evolved into AI learning computers, then into generative AI processors, and now they are going to do what they did from the start: generate pixels."
The idea, he said, is to reduce latency and energy consumption by eliminating the need to constantly exchange data with remote data centers. In this future, devices would be smart enough to generate relevant pixels locally, tailored to the user's context.
Software as the Real Moat
When asked about competition from specialized AI chips, such as those from startup Groq, Huang dismissed them as too narrowly focused. "Token generation must be specially optimized model by model," he said. "With such dedicated chips, you build configurable servers to fine-tune results for one type of model. Our GPUs, because they are massively parallel, are used to build programmable servers that can be optimized for all models."
Huang then pivoted to what he considers Nvidia's most enduring advantage: its software ecosystem. While AMD and other chipmakers rely on outside developers to optimize their hardware, Nvidia has been shipping development kits, libraries, and functional frameworks for its GPUs for over a dozen years. "The power of AI is not so much about chips as it is about software," he said.
He elaborated: "Pre-trained models are not completely usable at the base. You still have to adapt them, tune them, protect them, give them access to proprietary information, and so on. This requires a lot of services and tools. Our job is essentially to simplify the creation of the next ChatGPT. It is also about making sure that you don't have to know how to program in C++ to work with a generative AI and refine the quality of the outcomes."
Among the company's recent software offerings are the NIM microservices, which allow developers to integrate pre-trained models into their applications with minimal effort. Huang mentioned that NIM was initially developed for biological research, helping scientists manage complex AI models for drug discovery. The resulting BioNeMo application uses state-of-the-art biomolecular models to accelerate pharmaceutical research.
Nvidia's broader software stack, called NeMo, includes tools for preparing source data, training or fine-tuning models, and deploying them through inference or retrieval-augmented generation (RAG). The company has also expanded its Omniverse digital twin platform, which allows businesses to simulate and predict industrial operations in a virtual environment.
"We are a complete technology platform. SAP wants its own AI. ServiceNow also wants its own AI. NetApp wants its own AI. Industrial robot manufacturers want their own AI. We have the technology, the expertise, and the tools to help them build these AIs," Huang concluded.
The GTC 2024 conference, according to Huang, was not primarily about selling Nvidia products. Rather, it was a showcase to attract developers to the company's platform. "See us at the same level as x86 processors, RAM modules, or Ethernet networks. We are this essential architecture. Except that those technologies no longer need developers. Ours is just starting its career and needs developers," he said.
With Nvidia's market value soaring and its products embedded in nearly every major AI initiative, Huang's declarations carry significant weight. Whether or not the data center market reaches the $1 trillion mark by 2028, Nvidia has already positioned itself as the indispensable vendor for the next era of computing. The challenge ahead will be sustaining that dominance as rivals like AMD and specialized startups fight for a share of the booming AI infrastructure market.
Source: LeMagIT News