AMD Radeon™ AI PRO Solutions with GIGABYTE
Enterprise AI Inference, Right-Sized and Ready to Scale
AMD Radeon AI PROSoftwareAI Inference
Where Powerful Performance Meets Open, Scalable AI
Not every AI deployment needs a flagship training cluster. Most need to serve inference reliably, privately, and at a cost that survives a budget review. The AMD Radeon™ AI PRO R9000 series with AMD ROCm™ pairs 32 GB of memory on every card with open multi-GPU scaling, and we put as many as 16 of them in a single server — a practical foundation for large language models, RAG pipelines, and agentic AI running on infrastructure you own.
Choose the Right GPU to Accelerate AI Inference
With up to 64 compute units, 128 second-generation AI accelerators, and native support for FP8, FP16, and INT8 precision, the AMD Radeon™ AI PRO R9700S and R9600D GPUs bring real versatility to today's most demanding AI workloads. With clean multi-GPU scaling through AMD ROCm™, developers building on open AI ecosystems can push beyond single-card memory limits and deploy cost-effective acceleration for large language models on Linux. Both cards are passively cooled, moving the thermal work to the chassis, where our airflow and liquid-cooling engineering does its job.
RDNA™ 4
Architecture
PCIe Gen5
Bandwidth
2nd Gen
AI Accelerators
32GB
VRAM
DisplayPort™
2.1
Specifications | AMD Radeon AI PRO R9700S | AMD Radeon AI PRO R9600D |
| Compute Units | 64 | 48 |
| 3rd Gen Ray Accelerators | 64 | 48 |
| 2nd Gen AI Accelerators | 128 | 96 |
| Architecture | RDNA 4 | |
| Raw Peak FP16 Matrix TFLOPS | 191 | 99 |
| Memory Interface Size | 256-bit | |
| Memory Bandwidth | 640 GB/s | |
| Physical Memory | 32 GB GDDR6 | |
| Host Interface | PCIe Gen5 | |
| Board Power (TBP) | 300 W | 150 W |
| PCIe Form Factor | Full Height, Full Length | |
| Thermal, Form Factor | Passive, Dual Slot | Passive, Single Slot |
Open Software, an Open Ecosystem
Hardware is only half the decision. AMD ROCm™ is an open software stack, so the frameworks and model runtimes your team already uses run without a proprietary lock-in tax, and multi-GPU scaling is built in rather than bolted on. For enterprises standardizing on open models and open tooling, that openness is the point: you keep the freedom to change models, frameworks, and vendors as the field moves.
The Inference Engine for Enterprise Agentic AI
The AMD Radeon™ AI PRO R9000 series is built for where enterprise AI is actually being deployed — on the factory floor, in the hospital, on the trading desk. Across edge AI, healthcare, and financial services, the AMD Radeon™ AI PRO R9700S and R9600D accelerate inference for intelligent video analytics, medical imaging and AI-assisted diagnosis, fraud detection, risk modeling, and autonomous AI agents. Combined with GIGABYTE high-density GPU servers, they handle an enormous number of concurrent AI requests and agent operations at once, keeping system utilization high, scaling out as demand grows, and holding total cost per inference down.
Edge AI
Retail | Smart Cities | Manufacturing
Healthcare
Medical Imaging | Genomics | EMR
Financial Services
Fraud Detection | Algorithmic Trading | Risk Modeling
GIGABYTE Server Solution for AI Inference
G494-ZB0-LAP1 (8 x AMD Radeon™ AI PRO R9700S)
Pioneering AMD platform in a 4U DLC GPU server, built for data center deployments where maximum GPU density and power efficiency both matter.
- CPU + GPU DLC solution with up to 23% power saving
- 2 x AMD EPYC 9005 Server CPUs
- 8 x AMD Radeon AI PRO R9700S
- Compatible with AMD Pensando™ Pollara 400 AI NIC
W793-ZU0-LA71(4 x AMD Radeon™ AI PRO R9700S)
Closed-loop DLC in a workstation form factor, for teams that need serious AI capacity in a quiet office or SMB environment.
- 1 x AMD EPYC 9005 Server CPU
- 4 x AMD Radeon AI PRO R9700S
- High DLC coverage with CPU/RAM/GPU/PSU
- Low noise, under 50 dB
G294-Z43-AAP2(16 x AMD Radeon™ AI PRO R9600D)
High-density AI inference platform holding 16 x AMD Radeon™ AI PRO R9600D, for serving the most concurrent sessions per rack unit.
- CPU / GPU independent air-flow channel design
- 2 x AMD EPYC 9005 Server CPUs
- 16 x AMD Radeon™ AI PRO R9600D
High-Density AI Inference DLC Rack Solution with AMD Radeon AI PRO
Powered by 32GB AMD Radeon™ AI PRO R9700S GPUs and 8 AMD EPYC™ processors across four G494 4U DLC nodes, this 42U rack-scale platform delivers high-density inference for generative AI, RAG, and agentic AI. An in-rack cooling distribution unit and integrated G-REX DLC power and cooling management keep thermals and power in check. Add GIGABYTE POD Manager (GPM) for cluster-scale management, deep telemetry, and infrastructure provisioning, and scale capacity with validated partner storage or GIGABYTE storage servers to provide a turnkey foundation for secure, efficient, enterprise-ready private AI.
Enterprise Private AI
Running AI models on infrastructure the organization owns and controls, so proprietary data never leaves the building: internal copilots, document search, and knowledge assistants grounded in company information.
Agentic AI
AI that plans and acts across multiple steps and tools, rather than one prompt at a time. Every step is another inference call, so sustained throughput matters more than single-query speed.
Edge AI
Inference performed directly at the edge, where latency, footprint, power draw, and acoustics become major constraints. This is exactly where the single-slot, 150W AMD Radeon™ AI PRO R9600D and GIGABYTE edge platforms fit.
RAG / MCP
Retrieval-Augmented Generation grounds answers in your own documents and databases, while Model Context Protocol gives models a standard way to reach live tools and data. Both trade raw size for accuracy, favoring GPUs with generous memory and strong inference performance.
Multi-user AI Services
One AI platform that serves many users concurrently, as a shared internal service. High GPU density per node plus GIGABYTE POD Manager lets IT teams partition capacity and track usage across entire departments.