Banner Image

AMD Radeon™ AI PRO Solutions with GIGABYTE

Enterprise AI Inference, Right-Sized and Ready to Scale 

Where Powerful Performance Meets Open, Scalable AI

Not every AI deployment needs a flagship training cluster. Most need to serve inference reliably, privately, and at a cost that survives a budget review. The AMD Radeon™ AI PRO R9000 series with AMD ROCm™ pairs 32 GB of memory on every card with open multi-GPU scaling, and we put as many as 16 of them in a single server — a practical foundation for large language models, RAG pipelines, and agentic AI running on infrastructure you own.

Choose the Right GPU to Accelerate AI Inference

With up to 64 compute units, 128 second-generation AI accelerators, and native support for FP8, FP16, and INT8 precision, the AMD Radeon™ AI PRO R9700S and R9600D GPUs bring real versatility to today's most demanding AI workloads. With clean multi-GPU scaling through AMD ROCm™, developers building on open AI ecosystems can push beyond single-card memory limits and deploy cost-effective acceleration for large language models on Linux. Both cards are passively cooled, moving the thermal work to the chassis, where our airflow and liquid-cooling engineering does its job.
RDNA™ 4
Architecture
PCIe Gen5
Bandwidth
2nd Gen
AI Accelerators
32GB
VRAM
DisplayPort™
2.1
Content Image

Specifications

AMD Radeon AI PRO R9700S

AMD Radeon AI PRO R9600D

Compute Units6448
3rd Gen Ray Accelerators6448
2nd Gen AI Accelerators12896
ArchitectureRDNA 4
Raw Peak FP16 Matrix TFLOPS19199
Memory Interface Size256-bit
Memory Bandwidth640 GB/s
Physical Memory32 GB GDDR6
Host InterfacePCIe Gen5
Board Power (TBP)300 W150 W
PCIe Form FactorFull Height, Full Length
Thermal, Form FactorPassive, Dual SlotPassive, Single Slot

Open Software, an Open Ecosystem

Hardware is only half the decision. AMD ROCm™ is an open software stack, so the frameworks and model runtimes your team already uses run without a proprietary lock-in tax, and multi-GPU scaling is built in rather than bolted on. For enterprises standardizing on open models and open tooling, that openness is the point: you keep the freedom to change models, frameworks, and vendors as the field moves.

The Inference Engine for Enterprise Agentic AI

The AMD Radeon™ AI PRO R9000 series is built for where enterprise AI is actually being deployed — on the factory floor, in the hospital, on the trading desk. Across edge AI, healthcare, and financial services, the AMD Radeon™ AI PRO R9700S and R9600D accelerate inference for intelligent video analytics, medical imaging and AI-assisted diagnosis, fraud detection, risk modeling, and autonomous AI agents. Combined with GIGABYTE high-density GPU servers, they handle an enormous number of concurrent AI requests and agent operations at once, keeping system utilization high, scaling out as demand grows, and holding total cost per inference down.
Edge AI

Edge AI

Retail | Smart Cities | Manufacturing
Healthcare

Healthcare

Medical Imaging | Genomics | EMR
Financial Services

Financial Services

Fraud Detection | Algorithmic Trading | Risk Modeling

GIGABYTE Server Solution for AI Inference

G494-ZB0-LAP1 <div>(8 x AMD Radeon™ AI PRO R9700S)</div>

G494-ZB0-LAP1 
(8 x AMD Radeon™ AI PRO R9700S)

Pioneering AMD platform in a 4U DLC GPU server, built for data center deployments where maximum GPU density and power efficiency both matter.
  • CPU + GPU DLC solution with up to 23% power saving
  • 2 x AMD EPYC 9005 Server CPUs
  • 8 x AMD Radeon AI PRO R9700S
  • Compatible with AMD Pensando™ Pollara 400 AI NIC

W793-ZU0-LA71<div>(4 x AMD Radeon™ AI PRO R9700S)</div>

W793-ZU0-LA71
(4 x AMD Radeon™ AI PRO R9700S)

Closed-loop DLC in a workstation form factor, for teams that need serious AI capacity in a quiet office or SMB environment.
  • 1 x AMD EPYC 9005 Server CPU
  • 4 x AMD Radeon AI PRO R9700S
  • High DLC coverage with CPU/RAM/GPU/PSU
  • Low noise, under 50 dB
G294-Z43-AAP2<div>(16 x AMD Radeon™ AI PRO R9600D)</div>

G294-Z43-AAP2
(16 x AMD Radeon™ AI PRO R9600D)

High-density AI inference platform holding 16 x AMD Radeon™ AI PRO R9600D, for serving the most concurrent sessions per rack unit.
  • CPU / GPU independent air-flow channel design
  • 2 x AMD EPYC 9005 Server CPUs
  • 16 x AMD Radeon™ AI PRO R9600D
Content Image

High-Density AI Inference DLC Rack Solution with AMD Radeon AI PRO 

Powered by 32GB AMD Radeon™ AI PRO R9700S GPUs and 8 AMD EPYC™ processors across four G494 4U DLC nodes, this 42U rack-scale platform delivers high-density inference for generative AI, RAG, and agentic AI. An in-rack cooling distribution unit and integrated G-REX DLC power and cooling management keep thermals and power in check. Add GIGABYTE POD Manager (GPM) for cluster-scale management, deep telemetry, and infrastructure provisioning, and scale capacity with validated partner storage or GIGABYTE storage servers to provide a turnkey foundation for secure, efficient, enterprise-ready private AI.

Feature Icon

Enterprise Private AI

Running AI models on infrastructure the organization owns and controls, so proprietary data never leaves the building: internal copilots, document search, and knowledge assistants grounded in company information.
Feature Icon

Agentic AI

AI that plans and acts across multiple steps and tools, rather than one prompt at a time. Every step is another inference call, so sustained throughput matters more than single-query speed.
Feature Icon

Edge AI

Inference performed directly at the edge, where latency, footprint, power draw, and acoustics become major constraints. This is exactly where the single-slot, 150W AMD Radeon™ AI PRO R9600D and GIGABYTE edge platforms fit.
Feature Icon

RAG / MCP

Retrieval-Augmented Generation grounds answers in your own documents and databases, while Model Context Protocol gives models a standard way to reach live tools and data. Both trade raw size for accuracy, favoring GPUs with generous memory and strong inference performance.
Feature Icon

Multi-user AI Services

One AI platform that serves many users concurrently, as a shared internal service. High GPU density per node plus GIGABYTE POD Manager lets IT teams partition capacity and track usage across entire departments.