Please ensure Javascript is enabled for purposes of website accessibility

AI Infrastructure Beyond GPUs

Meeting with NextGenInfra.io, Dudy Cohen, VP of Product Marketing at DriveNets, explains why the most expensive resource on Earth is a GPU waiting for some resources and explores the networking innovations required to keep these resources fully utilized.

Major trends reshaping AI infrastructure

Dudy Cohen, VP of Product Marketing, DriveNets reports from the AI Infra Summit in Santa Clara and focuses on two major trends reshaping AI infrastructure: scale-across networking solutions that enable AI clusters to span multiple data centers, and the emergence of multi-vendor GPU ecosystems. Cohen explains how DriveNets addresses the challenge of connecting heterogeneous AI environments where different GPU vendors work together on the same workloads.

Full transcript
Hi everyone. We are here at the AI Infra Summit in Santa Clara. It’s a very busy and exciting couple of days because we talk about AI infrastructure, which is super important. I think everyone understands the importance of AI infrastructure. It’s not only about the GPUs and CPUs, it’s about the infrastructure that feeds them.

The most expensive resource on Earth is a GPU waiting for some resources. Be it the power in the data center or a network to carry the information it needs in order to complete its calculation. And I think that us as a networking company hear a lot about 2 main things this year in AI Infa Summit. One is scale across, the ability to connect 2 data centers and to overcome the limitation in terms of power mainly of a single data center while maintaining the growth of a single AI cluster is mind-blowing. We have it working in the field and basically a high-performance scaled-core network allows any customer to extend its AI cluster size without the limitation of the local data center.

So this is one thing. The other thing we hear a lot about is more and more vendors that are coming with ASICs. We know NVIDIA, but we hear a lot now about other ASICs, other GPUs that are entering the market, both as a GPU and as a rack scale. There’s a lot of talk about the AMD Helios rack that is coming to the market. And we hear and understand that we are going into an environment which is multi-vendor and sometimes heterogeneous in which multiple ASICs are co-working on the same workload.

Mainly in inference in the same cluster. This, of course, requires a special attention to AI infrastructure and especially to networking that needs to accommodate this move between the GPUs and needs to be capable of fulfilling the requirements and— of fulfilling the requirements and special traffic patterns of a heterogeneous environment. So this is what we talk about here in AI and Infra. If you are around Santa Clara, you are more than welcome to come and talk to us about anything that has to do with AI. Thank you very much.

Continue reading

Collateral

DriveNets AI Fabric

Full-stack networking portfolio for AI infrastructure. Cluster builders need a networking solution that delivers high pe ...

Read more

Collateral

AMD and DriveNets Reference Architecture

Building high-performance AI clusters is no longer a matter of adding GPUs.

Read more

Blog

Inside WhiteFiber’s long-distance scale-across deployment

We previously discussed an operational reality for NeoClouds and AI infrastructure builders. AI workloads are increasing ...

Read more