Specifications
Form Factor
4RU rack server — 800 mm deep, 438 mm wide, 175 mm high, approx. 60 kg
GPU Architecture
NVIDIA MGX modular reference design
Primary GPU Options
2, 4, 6 or 8x NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB GDDR7)
Additional GPU Options
NVIDIA RTX PRO 6000D (84 GB), RTX PRO 4500 (32 GB), H200 NVL, H100 NVL, L40S, AMD Instinct MI210
GPU Memory
Up to 768 GB total with eight RTX PRO 6000 96 GB cards
Processors
Dual AMD EPYC 9005 (e.g. 9575, 64 cores) — 2-way platform
System Memory
24x DDR5-6400 RDIMM, up to 8 TB
PCIe Slots
8x PCIe Gen5 x16 GPU slots plus 5x PCIe Gen5 x16 FHHL
Storage
20x E1.S NVMe Gen5 bays, up to 15.3 TB per drive
Networking
4x single-port 400G NVIDIA NICs (ConnectX-7 / BlueField-3) plus OCP 3.0 slot
NVLink Bridging
2-way and 4-way NVLink bridge kits for H100 NVL / H200 NVL configurations
Cooling
Two hot-swap fan zones — 5x 80x80 mm upper and 5x 80x80 mm lower, separate GPU and CPU chambers
Power
Redundant hot-swap power supplies with separate 12 V CPU and 54 V GPU rails
Management
Cisco Intersight, Redfish API and CIMC out-of-band management
Software Stack
NVIDIA AI Enterprise certified; ready for NVIDIA AI Data Platform deployments
Target Workloads
GenAI inference, fine-tuning, RAG, agentic AI, physical AI, VDI, HPC and analytics
Overview
The Cisco UCS C845A M8 is Cisco's flagship NVIDIA RTX PRO inference platform. It takes the NVIDIA MGX modular reference design and packages it into 4RU, accepting between two and eight accelerators so a single SKU family can cover pilot inference clusters and production serving fleets alike. With RTX PRO 6000 Blackwell Server Edition cards the chassis holds 768 GB of aggregate GPU memory, which keeps large models, vector indexes and KV cache resident on the accelerator instead of spilling across the fabric.
The 4RU design keeps the GPU complexity and the CPU/memory complex in separate thermal zones — five 80 mm fans per chamber — which lets the platform sustain high tokens-per-second under continuous inference load without throttling. Twenty Gen5 E1.S bays give local high-throughput storage for model weights and datasets, while four 400G network interfaces plus an OCP 3.0 slot cover both scale-out and management traffic.
For H100 NVL and H200 NVL builds, 2-way and 4-way NVLink bridge kits are available so paired GPUs communicate over NVLink rather than PCIe. The platform is delivered through Cisco Intersight for lifecycle management and is positioned as the accelerated compute tier of Cisco's Secure AI Factory reference architecture.
Key Benefits
Right-sized scaling: the same 4RU platform accepts 2, 4, 6 or 8 GPUs, so capacity can grow without a re-platform. Memory-resident inference: up to 768 GB of GPU memory plus 8 TB of DDR5-6400 keeps long-context and multimodal models hot. MGX ecosystem: modular bays and a documented reference design shorten qualification and spare-part planning. Dual thermal chambers: independent GPU and CPU cooling zones hold clock targets under sustained serving load. Core-to-edge consistency: Intersight policy and observability follow the same operational model from data centre to edge deployment.
Applications
Enterprise generative AI inference and fine-tuning, retrieval-augmented generation services, agentic AI workload hosting, smart-factory and robotics (physical AI) inference, virtual desktop infrastructure with GPU acceleration, HPC simulation, and analytics/video transcoding clusters.
Request a Quote — CISCO UCS C845A M8 — 4RU NVIDIA MGX SERVER FOR RTX PRO 6000 BLACKWELL
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →