Specifications

Accelerators

8x Positron Archer transformer accelerators

Memory per Accelerator

32 GB HBM

Total Accelerator Memory

256 GB HBM

System Memory

24-channel RDIMM DDR5 — 24 x 16 GB = 384 GB (supports up to 2 TB)

CPU

Dual AMD EPYC Genoa 9374F — 64 cores total, 3.85 GHz base / 4.3 GHz max boost

Boot Storage

1x 1.92 TB NVMe M.2, PCIe Gen3 x4 (OS/boot drive)

Data Storage

1x 1.92 TB U.2 2.5-inch SSD

Drive Bays

Four 2.5-inch SATA/SAS hot-swappable bays

Networking

1x 10 Gb/s LAN (X550 compatible)

Management Network

1x 1 Gb/s LAN (Intel I210-AT)

Expansion Slots

2x FHFL PCIe Gen5 x16 (GPUs, NICs, etc.) + 1x FHHL PCIe Gen4 x16

Power Supplies

2+2 2000 W redundant, Titanium level (96%)

System Dimensions

177 mm (H) x 482.2 mm (W) x 743 mm (L)

System Weight

100 lb / 45.4 kg

Software

Positron Inference Engine on Ubuntu 22.04.4 LTS (Jammy)

Reference Benchmark

Llama 3.1 8B BF16 (no speculation, no paged attention)

Performance per Watt

4.54x vs NVIDIA DGX H200 (vendor measurement)

Performance per Dollar

3.08x vs NVIDIA DGX H200 (vendor measurement)

Support

24-hour SLA from a US-based team

Price

Quote upon request

Availability

Shipping today

Overview

Positron Atlas is a rack-mount inference server built specifically for transformer workloads rather than repurposed for them. It carries eight Positron Archer transformer accelerators with 32 GB of HBM each — 256 GB of accelerator memory in total — behind dual AMD EPYC 9374F Genoa processors with 64 cores, 384 GB of 24-channel DDR5 system memory on 24 DIMMs, and a 2+2 configuration of 2000 W redundant Titanium-level (96%) power supplies.

Positron's published comparison for a Llama 3.1 8B BF16 workload (without speculation or paged attention) reports 280.00 tokens/sec/user against 182.00 tokens/sec/user on an NVIDIA DGX H200, translating to 4.54x performance per watt and 3.08x performance per dollar versus the DGX H200 system. Both systems are cited at roughly 28-30 W per accelerator equivalent, but Atlas draws 2000 W system power against 5900 W for the comparison platform.

The chassis is 7.0 inches high, 19.0 inches wide and 29.25 inches deep, weighs 100 lb, and provides four hot-swappable 2.5-inch SATA/SAS bays plus a 1.92 TB NVMe M.2 boot drive and a 1.92 TB U.2 model-storage drive. Two full-height full-length PCIe Gen5 x16 slots and one FHHL Gen4 x16 slot remain free for NICs or additional accelerators.

Software is the Positron Inference Engine on Ubuntu 22.04.4 LTS, and support is delivered under a 24-hour SLA by a US-based team. For organisations constrained by rack power or by GPU allocation, Atlas offers an alternative path to high-throughput token serving at a fraction of the system power draw.

QS Compute supplies Positron Atlas as a quote-based system order alongside its catalogue of NVIDIA, AMD and AI-accelerator server platforms.

Key Benefits

256 GB of HBM across eight Archer accelerators for large-model inference. 2000 W system power against 5900 W for the comparison DGX platform. 3.08x performance per dollar (vendor measurement vs DGX H200). Shipping today with a 24-hour SLA support commitment. Four hot-swap bays and three free PCIe slots for expansion.

Applications

LLM token serving and inference APIs; enterprise RAG and agentic back-ends; multi-tenant model serving; on-premises generative AI in power-constrained data centres; private-cloud inference replacing GPU-allocated clusters.

Request a Quote — POSITRON ATLAS INFERENCE SERVER — 8X ARCHER ACCELERATORS, 256 GB HBM

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

SambaNova SambaRack SN50 — RDU Inference Rack Tenstorrent QuietBox / Galaxy AI Systems GigaIO Gryf — Composable AI Infrastructure