QS Compute
MPN: B300-SXM6
NVIDIA B300 Blackwell Ultra
288GB HBM3e · 15 PFLOPS FP4 · 8 TB/s
Blackwell Architecture · 1,400W TDP · NVLink 5 (1.8 TB/s) · ConnectX-8 (1.6T) · Shipped Jan 2026
Specifications
Model
NVIDIA B300 (Blackwell Ultra)
Architecture
NVIDIA Blackwell
High-bin variant
FP4 Dense
15,000 TFLOPS
15 PFLOPS
FP8 Dense
7,000 TFLOPS
~3.5× H200
TDP
1,400W
Liquid cooling required
Interconnect
NVLink 5
1.8 TB/s GPU-to-GPU
Networking
ConnectX-8
1.6 Tbps per GPU
Form Factor
SXM6
OAM module
Fits 70B Model
Full FP16 without sharding
~100GB KV cache headroom
8-GPU System
2.3 TB total GPU memory
400B+ model (no model parallel)
Software
CUDA 12.x
TensorRT-LLM 0.15+
FP4 quantization ready
System Configurations
DGX B300
8× B300 SXM6
Intel Xeon 6776P
2.1 TB GPU memory
144 PFLOPS FP4 sparse
~14 kW peak
HGX B300
8× B300 baseboard
OEM integration (Supermicro, Dell, Lenovo)
User-selectable CPU/chassis/cooling
Same GPU performance
GB300 NVL72
72× B300 + 36× Grace CPUs
Rack-scale, liquid-cooled
Massive inference for frontier reasoning models
Reservable via Spheron
Cloud Pricing Reference (July 2026)
On-Demand (Spheron)
~$9.16/GPU-hr
Per-minute billing
Marketplaces
$10+/GPU-hr
Variable availability
Premium Cloud
$12–18/GPU-hr
Managed stack
Overview
The NVIDIA B300 Blackwell Ultra is the highest-bin Blackwell GPU, shipping since January 2026. With 288 GB of HBM3e memory at 8 TB/s bandwidth and 15,000 TFLOPS of dense FP4 performance, it represents a 50% memory capacity increase and 67% higher FP4 throughput over the B200. The B300 is the first GPU where FP4 inference is truly first-class — a single card holds an entire 70B-parameter model in FP16 without tensor parallelism, with ~100 GB of headroom remaining for KV cache and batching.
Infrastructure requirements are significant: 1,400W TDP per GPU mandates liquid cooling, and an 8-GPU DGX B300 draws approximately 14 kW peak. However, the density payoff is unmatched — 2.3 TB of total GPU memory in an 8-GPU configuration enables 400B+ parameter models with zero model parallelism, dramatically simplifying deployment architecture. NVLink 5 provides 1.8 TB/s of GPU-to-GPU bandwidth, and the integrated ConnectX-8 SmartNIC delivers 1.6 Tbps of networking, doubling inter-node bandwidth compared to B200.
QS Compute supplies NVIDIA B300 across SXM6 (DGX B300, HGX B300) and GB300 NVL72 rack-scale configurations, with liquid-cooled infrastructure support. Pricing is available on request.
Need NVIDIA B300 Blackwell Ultra?
Contact QS Compute for DGX B300, HGX B300, and GB300 NVL72 configurations.
Request Quote