Specifications
Switching Capacity
51.2 Tbps
Latency
250 ns switch latency at full 51.2 Tbps throughput
Small-Packet Performance
Line-rate switching at the minimum 64-byte packet size — up to 77 billion packets per second
Protocol
Scale-Up Ethernet (SUE)
Process Node
TSMC 5 nm
Header Efficiency
Reduces header overhead from 46 bytes down to as low as 10 bytes while remaining fully Ethernet compliant
Advanced Features
Link Layer Retry, In-Network Collectives and the AI Fabric Header
Flow Control
Credit-based flow control for lossless fabric operation
Board Compatibility
Package and pin compatible with Tomahawk 5, enabling same-board upgrades
Maximum Accelerator Scale
Up to 1,024 accelerators
Port Configurations
64x 800/400/200/100GbE in a 2U switch, supporting 2x 400GbE, 4x 200GbE or 8x 100GbE breakout
Versus NVIDIA NVLink Switch
250 ns vs 9-18 microseconds latency, 51.2 Tbps vs 28.8 Tbps, up to 1,024 accelerators vs 72 GPUs per chassis
Fabric Role
Scale-up networking inside an XPU node as well as between nodes
Availability
Announced and shipping from July 2025
Ecosystem
Adopted by switch partners including Micas Networks, Wistron and Accton for AI and HPC system builds
Overview
Broadcom Tomahawk Ultra attacks a different problem from the general-purpose Tomahawk line. Instead of maximising throughput for mixed datacentre traffic, it optimises for small packets and low latency so that Ethernet can replace InfiniBand and NVLink inside a scale-up GPU domain. It holds 51.2 Tbps at just 250 ns of switch latency and maintains line rate at the minimum 64-byte packet size — up to 77 billion packets per second.
The enabling technology is Scale-Up Ethernet, a Broadcom-defined protocol backed by an optimised header that cuts per-packet overhead from the standard 46 bytes down to as little as 10 bytes while staying fully Ethernet compliant. Link Layer Retry, In-Network Collectives and the AI Fabric Header provide the reliability and collective-operation offload that accelerator fabrics depend on, and credit-based flow control keeps the fabric lossless.
Against NVIDIA's NVLink Switch the comparison is striking on paper: 250 ns versus 9-18 microseconds of latency, 51.2 Tbps versus 28.8 Tbps of capacity, and scale to 1,024 accelerators rather than 72 GPUs in a single chassis. It is package and pin compatible with Tomahawk 5, so existing boards can be upgraded in place, and switch partners including Micas Networks, Wistron and Accton are building systems around it. QS Compute supplies AI cluster networking, switch platforms and scale-up interconnect hardware — request a quote.
Key Benefits
Ultra-low latency: 250 ns at full 51.2 Tbps with line rate at 64-byte packets. Open alternative to NVLink: Scale-Up Ethernet scales to 1,024 accelerators. Lower header overhead: 46 bytes down to as little as 10 bytes. Drop-in upgrade: pin compatible with Tomahawk 5 boards.
Applications
GPU and XPU scale-up fabrics, AI training clusters, HPC interconnect, replacing InfiniBand or proprietary scale-up fabrics, and in-node accelerator networking.
Request a Quote — BROADCOM TOMAHAWK ULTRA — 51.2 TBPS SCALE-UP ETHERNET SWITCH
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →