Specifications

Family

MemryX Cascade 100 Series

Silicon

MemryX MX3 dataflow AI accelerator

Cascade 100M

M.2 2280 M-Key module, 4x MX3 accelerators, 25 TFLOPS

Cascade 100M Power

1 W - 12 W

Cascade 100M Storage

42 MB on-chip model weight storage

Cascade 100P

PCI Express x8 full-height half-length card, 16x MX3 accelerators, 100 TFLOPS

Cascade 100P Power

5 W - 50 W

Cascade 100P Storage

168 MB on-chip model weight storage

Cascade 100R

Raspberry Pi 5 HAT+ module built around the 4-chip Cascade 100M, 25 TFLOPS

Cascade 100R Power

1 W - 12 W

Cascade 100R Host Link

M.2 slot on the HAT+ connects through the Raspberry Pi 5 PCIe interface

Cascade 100U

USB Type-C AI accelerator module, two or four MX3 chips, up to 25 TFLOPS

Cascade 100U Power

1 W - 12 W

Cascade 100U Form

Portable USB Type-C module, plug-and-play

Architecture

At-memory compute with pipelined dataflow

Numerics

BFloat16 activations - no retraining or hand-tuning required

Latency

Deterministic batch-1 latency, no batching required

Host CPU Offload

Dedicated AI offload; host compute preserved for other workloads

Software

MemryX SDK, compiler toolchain and model zoo

Unified Toolchain

Same software workflow from single-chip to 16-chip configurations

Distribution

Cascade modules ship through MemryX distribution partners (WPG-A and others)

Award

Cascade 100M - 2025 Edge AI and Vision Product of the Year

Overview

The MemryX Cascade 100 Series is a family of production edge AI accelerator modules built on the MemryX MX3 dataflow architecture. Rather than shipping a single form factor, MemryX covers the four integration paths customers actually design around: an M.2 2280 M-Key module (Cascade 100M), a full-height half-length PCIe card (Cascade 100P), a Raspberry Pi 5 HAT+ (Cascade 100R) and a USB Type-C module (Cascade 100U).

The Cascade 100M, 100R and 100U each carry four MX3 accelerators and deliver 25 TFLOPS of INT8-class inference within a 1 W to 12 W envelope, with 42 MB of on-chip model weight storage. The Cascade 100P scales the same silicon to sixteen MX3 devices for 100 TFLOPS and 168 MB of weight storage in a 5 W to 50 W PCIe card aimed at high-volume servers and multi-stream inference.

Because MX3 is a dataflow design rather than a batching accelerator, inference runs at deterministic batch-1 latency with the host processor free for the rest of the application. BFloat16 activations preserve model accuracy by direct recompilation, so existing CNN, transformer and hybrid models can be moved onto the accelerator without retraining.

Key Benefits

One toolchain across every form factor. The Cascade 100 Series runs the same MemryX compile-and-runtime workflow whether a design uses an M.2 module, a PCIe card, a Raspberry Pi 5 HAT+ or a USB module. Efficiency first. 25 TFLOPS within 1 W - 12 W on the module-class parts and 100 TFLOPS within 5 W - 50 W on the PCIe card keeps power budgets flat while adding dedicated AI capacity. Deterministic low latency. Pipelined dataflow keeps multiple inputs in flight, delivering consistent real-time latency without large batch sizes. No retraining. BFloat16 activations and direct recompilation preserve accuracy on existing models.

Applications

Embedded and edge vision, industrial inspection, robotics and autonomous platforms, smart retail analytics, portable and laptop AI acceleration, Raspberry Pi 5 edge AI deployment, multi-camera video analytics and production server inference offload.

Request a Quote — MEMRYX CASCADE 100 SERIES — M.2, PCIE, RASPBERRY PI 5 HAT+ AND USB-C MX3 ACCELERATORS

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

MemryX MX3 Edge AI Accelerator Hailo-8L M.2 Kit Hailo-10H M.2 Starter Kit Raspberry Pi AI HAT+ 2