MPN: B300-SXM6
NVIDIA B300 Blackwell Ultra
288GB HBM3e · 15 PFLOPS FP4 · 8 TB/s

Blackwell Architecture · 1,400W TDP · NVLink 5 (1.8 TB/s) · ConnectX-8 (1.6T) · Shipped Jan 2026

Specifications

Model

NVIDIA B300 (Blackwell Ultra)

Architecture

NVIDIA Blackwell
High-bin variant

VRAM

288 GB HBM3e

Memory Bandwidth

8 TB/s

FP4 Dense

15,000 TFLOPS
15 PFLOPS

FP8 Dense

7,000 TFLOPS
~3.5× H200

FP16 Dense

3,500 TFLOPS

TDP

1,400W
Liquid cooling required

Interconnect

NVLink 5
1.8 TB/s GPU-to-GPU

Networking

ConnectX-8
1.6 Tbps per GPU

Form Factor

SXM6
OAM module

Release Date

January 2026

Fits 70B Model

Full FP16 without sharding
~100GB KV cache headroom

8-GPU System

2.3 TB total GPU memory
400B+ model (no model parallel)

Software

CUDA 12.x
TensorRT-LLM 0.15+
FP4 quantization ready

System Configurations

DGX B300

8× B300 SXM6
Intel Xeon 6776P
2.1 TB GPU memory
144 PFLOPS FP4 sparse
~14 kW peak

HGX B300

8× B300 baseboard
OEM integration (Supermicro, Dell, Lenovo)
User-selectable CPU/chassis/cooling
Same GPU performance

GB300 NVL72

72× B300 + 36× Grace CPUs
Rack-scale, liquid-cooled
Massive inference for frontier reasoning models
Reservable via Spheron

Cloud Pricing Reference (July 2026)

On-Demand (Spheron)

~$9.16/GPU-hr
Per-minute billing

Marketplaces

$10+/GPU-hr
Variable availability

Premium Cloud

$12–18/GPU-hr
Managed stack

Overview

The NVIDIA B300 Blackwell Ultra is the highest-bin Blackwell GPU, shipping since January 2026. With 288 GB of HBM3e memory at 8 TB/s bandwidth and 15,000 TFLOPS of dense FP4 performance, it represents a 50% memory capacity increase and 67% higher FP4 throughput over the B200. The B300 is the first GPU where FP4 inference is truly first-class — a single card holds an entire 70B-parameter model in FP16 without tensor parallelism, with ~100 GB of headroom remaining for KV cache and batching.

Infrastructure requirements are significant: 1,400W TDP per GPU mandates liquid cooling, and an 8-GPU DGX B300 draws approximately 14 kW peak. However, the density payoff is unmatched — 2.3 TB of total GPU memory in an 8-GPU configuration enables 400B+ parameter models with zero model parallelism, dramatically simplifying deployment architecture. NVLink 5 provides 1.8 TB/s of GPU-to-GPU bandwidth, and the integrated ConnectX-8 SmartNIC delivers 1.6 Tbps of networking, doubling inter-node bandwidth compared to B200.

QS Compute supplies NVIDIA B300 across SXM6 (DGX B300, HGX B300) and GB300 NVL72 rack-scale configurations, with liquid-cooled infrastructure support. Pricing is available on request.

Need NVIDIA B300 Blackwell Ultra?

Contact QS Compute for DGX B300, HGX B300, and GB300 NVL72 configurations.

Request Quote