Specifications
Accelerators
8x Positron Archer transformer accelerators
Memory per Accelerator
32 GB HBM
Total Accelerator Memory
256 GB HBM
System Memory
24-channel RDIMM DDR5 — 24 x 16 GB = 384 GB (supports up to 2 TB)
CPU
Dual AMD EPYC Genoa 9374F — 64 cores total, 3.85 GHz base / 4.3 GHz max boost
Boot Storage
1x 1.92 TB NVMe M.2, PCIe Gen3 x4 (OS/boot drive)
Data Storage
1x 1.92 TB U.2 2.5-inch SSD
Drive Bays
Four 2.5-inch SATA/SAS hot-swappable bays
Networking
1x 10 Gb/s LAN (X550 compatible)
Management Network
1x 1 Gb/s LAN (Intel I210-AT)
Expansion Slots
2x FHFL PCIe Gen5 x16 (GPUs, NICs, etc.) + 1x FHHL PCIe Gen4 x16
Power Supplies
2+2 2000 W redundant, Titanium level (96%)
System Dimensions
177 mm (H) x 482.2 mm (W) x 743 mm (L)
System Weight
100 lb / 45.4 kg
Software
Positron Inference Engine on Ubuntu 22.04.4 LTS (Jammy)
Reference Benchmark
Llama 3.1 8B BF16 (no speculation, no paged attention)
Performance per Watt
4.54x vs NVIDIA DGX H200 (vendor measurement)
Performance per Dollar
3.08x vs NVIDIA DGX H200 (vendor measurement)
Support
24-hour SLA from a US-based team
Price
Quote upon request
Availability
Shipping today
Overview
Positron Atlas is a rack-mount inference server built specifically for transformer workloads rather than repurposed for them. It carries eight Positron Archer transformer accelerators with 32 GB of HBM each — 256 GB of accelerator memory in total — behind dual AMD EPYC 9374F Genoa processors with 64 cores, 384 GB of 24-channel DDR5 system memory on 24 DIMMs, and a 2+2 configuration of 2000 W redundant Titanium-level (96%) power supplies.
Positron's published comparison for a Llama 3.1 8B BF16 workload (without speculation or paged attention) reports 280.00 tokens/sec/user against 182.00 tokens/sec/user on an NVIDIA DGX H200, translating to 4.54x performance per watt and 3.08x performance per dollar versus the DGX H200 system. Both systems are cited at roughly 28-30 W per accelerator equivalent, but Atlas draws 2000 W system power against 5900 W for the comparison platform.
The chassis is 7.0 inches high, 19.0 inches wide and 29.25 inches deep, weighs 100 lb, and provides four hot-swappable 2.5-inch SATA/SAS bays plus a 1.92 TB NVMe M.2 boot drive and a 1.92 TB U.2 model-storage drive. Two full-height full-length PCIe Gen5 x16 slots and one FHHL Gen4 x16 slot remain free for NICs or additional accelerators.
Software is the Positron Inference Engine on Ubuntu 22.04.4 LTS, and support is delivered under a 24-hour SLA by a US-based team. For organisations constrained by rack power or by GPU allocation, Atlas offers an alternative path to high-throughput token serving at a fraction of the system power draw.
QS Compute supplies Positron Atlas as a quote-based system order alongside its catalogue of NVIDIA, AMD and AI-accelerator server platforms.
Key Benefits
256 GB of HBM across eight Archer accelerators for large-model inference. 2000 W system power against 5900 W for the comparison DGX platform. 3.08x performance per dollar (vendor measurement vs DGX H200). Shipping today with a 24-hour SLA support commitment. Four hot-swap bays and three free PCIe slots for expansion.
Applications
LLM token serving and inference APIs; enterprise RAG and agentic back-ends; multi-tenant model serving; on-premises generative AI in power-constrained data centres; private-cloud inference replacing GPU-allocated clusters.
Request a Quote — POSITRON ATLAS INFERENCE SERVER — 8X ARCHER ACCELERATORS, 256 GB HBM
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →