AI networking and full-stack optimization for AMD clusters

AMD and DriveNets are helping cluster builders deploy high-performance AMD-based AI environments faster and with lower risk. We are working closely together to provide a complete AI infrastructure solution that includes AMD Instinct GPU systems and DriveNets networking domains, spanning scale-up, scale-out, scale-across and frontend & storage networking. The joint solution provides a fully integrated software stack that is optimized for performance and functionality, delivering a competitive alternative to a full NVIDIA AI solution.

System-Level Optimization for AMD Instinct GPU Clusters

  • Higher cluster performance: Joint AMD–DriveNets optimization across RCCL, networking, and the full cluster stack improves end-to-end workload performance.
  • Faster, lower-risk deployment: A validated reference architecture, automated provisioning, and built-in testing simplify cluster bring-up and enable repeatable deployments.
  • Proven results at scale:Real-workload testing demonstrates measurable performance gains against Spectrum-X, InfiniBand, and baseline AMD cluster configurations.
  • Continued performance innovation:Ongoing joint R&D collaboration including Helios rack-scale infrastructure, liquid-cooled switching, and automated, adaptive, and agentic optimization to boost performance for future AI models.

AMD Instinct Fabric Reference Architecture

AMD and DriveNets released a validated reference architecture document for clusters built with AMD Instinct MI355X GPUs, AMD Pollara NICs, and DriveNets scale-out and frontend solution.

The reference architecture document provides a comprehensive end-to-end blueprint for building a high-performance, scalable AI GPU cluster, and a repeatable deployment model that reduces integration and configuration risk.

Reference Architecture

A validated reference architecture document for clusters built with AMD Instinct MI355X GPUs, AMD Pollara NICs, and DriveNets AI Fabric solution

Faster LLM Inference on AMD Requires Rethinking All-Reduce

Blog

Faster LLM Inference on AMD Requires Rethinking All-Reduce

AMD Instinct GPUs offer real hardware advantages for large AI clusters, including higher HBM3 memory capacity and compet ...

Read more
How AMD Instinct Shines in Real-World LLM Inference

Blog

How AMD Instinct Shines in Real-World LLM Inference

AI workloads have moved beyond experimentation into deep production environments. Today, performance is no longer about ...

Read more

Blog

Optimizing AMD Instinct AI Clusters with DriveNets’ Lossless Ethernet Fabric

The DriveNets AI Fabric networking fabric solution delivers the highest performance AI connectivity for any GPU, NIC or ...

Read more