Specifications
Family
MemryX Cascade 100 Series
Silicon
MemryX MX3 dataflow AI accelerator
Cascade 100M
M.2 2280 M-Key module, 4x MX3 accelerators, 25 TFLOPS
Cascade 100M Power
1 W - 12 W
Cascade 100M Storage
42 MB on-chip model weight storage
Cascade 100P
PCI Express x8 full-height half-length card, 16x MX3 accelerators, 100 TFLOPS
Cascade 100P Power
5 W - 50 W
Cascade 100P Storage
168 MB on-chip model weight storage
Cascade 100R
Raspberry Pi 5 HAT+ module built around the 4-chip Cascade 100M, 25 TFLOPS
Cascade 100R Power
1 W - 12 W
Cascade 100R Host Link
M.2 slot on the HAT+ connects through the Raspberry Pi 5 PCIe interface
Cascade 100U
USB Type-C AI accelerator module, two or four MX3 chips, up to 25 TFLOPS
Cascade 100U Power
1 W - 12 W
Cascade 100U Form
Portable USB Type-C module, plug-and-play
Architecture
At-memory compute with pipelined dataflow
Numerics
BFloat16 activations - no retraining or hand-tuning required
Latency
Deterministic batch-1 latency, no batching required
Host CPU Offload
Dedicated AI offload; host compute preserved for other workloads
Software
MemryX SDK, compiler toolchain and model zoo
Unified Toolchain
Same software workflow from single-chip to 16-chip configurations
Distribution
Cascade modules ship through MemryX distribution partners (WPG-A and others)
Award
Cascade 100M - 2025 Edge AI and Vision Product of the Year
Overview
The MemryX Cascade 100 Series is a family of production edge AI accelerator modules built on the MemryX MX3 dataflow architecture. Rather than shipping a single form factor, MemryX covers the four integration paths customers actually design around: an M.2 2280 M-Key module (Cascade 100M), a full-height half-length PCIe card (Cascade 100P), a Raspberry Pi 5 HAT+ (Cascade 100R) and a USB Type-C module (Cascade 100U).
The Cascade 100M, 100R and 100U each carry four MX3 accelerators and deliver 25 TFLOPS of INT8-class inference within a 1 W to 12 W envelope, with 42 MB of on-chip model weight storage. The Cascade 100P scales the same silicon to sixteen MX3 devices for 100 TFLOPS and 168 MB of weight storage in a 5 W to 50 W PCIe card aimed at high-volume servers and multi-stream inference.
Because MX3 is a dataflow design rather than a batching accelerator, inference runs at deterministic batch-1 latency with the host processor free for the rest of the application. BFloat16 activations preserve model accuracy by direct recompilation, so existing CNN, transformer and hybrid models can be moved onto the accelerator without retraining.
Key Benefits
One toolchain across every form factor. The Cascade 100 Series runs the same MemryX compile-and-runtime workflow whether a design uses an M.2 module, a PCIe card, a Raspberry Pi 5 HAT+ or a USB module. Efficiency first. 25 TFLOPS within 1 W - 12 W on the module-class parts and 100 TFLOPS within 5 W - 50 W on the PCIe card keeps power budgets flat while adding dedicated AI capacity. Deterministic low latency. Pipelined dataflow keeps multiple inputs in flight, delivering consistent real-time latency without large batch sizes. No retraining. BFloat16 activations and direct recompilation preserve accuracy on existing models.
Applications
Embedded and edge vision, industrial inspection, robotics and autonomous platforms, smart retail analytics, portable and laptop AI acceleration, Raspberry Pi 5 edge AI deployment, multi-camera video analytics and production server inference offload.
Request a Quote — MEMRYX CASCADE 100 SERIES — M.2, PCIE, RASPBERRY PI 5 HAT+ AND USB-C MX3 ACCELERATORS
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →