Models & Pricing

SAKURA-II M.2 2280 Module

Quote
  • 60 TOPS (INT8) / 30 TFLOPS (BF16)
  • 8W typical power
  • Up to 32GB LPDDR4x, 68 GB/s
  • 20MB on-chip SRAM
  • M.2 2280 form factor
  • -40 to 105°C option

SAKURA-II PCIe Card

Quote
  • Up to 120 TOPS (single and dual)
  • Low-profile single-slot
  • PCIe host interface
  • Sparse computing support
  • Batch=1 low-latency inference
  • Enterprise and industrial servers

SAKURA-II System

Quote
  • Up to 240 TOPS when deployed in systems
  • Multi-accelerator scaling
  • MERA heterogeneous inference
  • Ubuntu, Docker, Jupyter ready
  • Edge server deployments
  • Defense, robotics, smart city

Specifications

Accelerator

EdgeCortix SAKURA-II with Dynamic Neural Accelerator (DNA) architecture

AI Performance

60 TOPS (INT8) / 30 TFLOPS (BF16)

Power Consumption

8W typical

Compute Efficiency

Up to 90% accelerator utilisation

DRAM Support

Dual 64-bit LPDDR4x — 8GB / 16GB / 32GB

DRAM Bandwidth

68 GB/s

On-chip SRAM

20MB

Temperature Range

-40 to 105°C (industrial); 40 to 85°C commercial variant

Package

19 × 19 mm BGA

Module Form Factor

M.2 2280 (60 TOPS)

Card Form Factor

Low-profile single-slot PCIe, up to 120 TOPS single/dual

System Scale

Up to 240 TOPS across multiple accelerators

Generative AI

Llama 2, Stable Diffusion, DETR and ViT within 8W

Precision

Software-enabled mixed precision with near-FP32 accuracy

Software Stack

MERA compiler and framework — HuggingFace, TensorFlow Lite, ONNX, TVM, MLIR

Runtime Environment

Ubuntu, Docker, Jupyter

Certifications

FCC, CE, UKCA, ICES, VCCI, BSMI, REACH, RoHS

Recognition

2024 Edge AI Product of the Year winner

Ecosystem Partners

MegaChips, SoftBank, BittWare, Renesas RZ/V, Arm/Raspberry Pi 5

Overview

EdgeCortix positions SAKURA-II as the accelerator for the generative-AI era of edge computing, and the numbers support the claim: 60 TOPS of INT8 and 30 TFLOPS of BF16 inference inside an 8W typical power envelope, with support for multi-billion-parameter models. That is a power budget a fanless industrial enclosure can absorb, unlike the discrete GPUs normally required to run comparable workloads.

The architectural foundation is EdgeCortix’s Dynamic Neural Accelerator (DNA), a run-time reconfigurable and highly parallelised array built for Batch=1, low-latency inference rather than throughput-batched datacenter work. The company pairs it with up to 32GB of LPDDR4x at 68 GB/s and 20MB of on-chip SRAM, claiming up to four times the DRAM bandwidth of competing edge accelerators plus sparse-computing support that further reduces memory pressure. The result is up to 90% sustained utilisation, which is where the effective performance gap usually opens up.

Deployment spans three tiers from a single silicon design. The M.2 2280 module delivers 60 TOPS for space-constrained systems. Single and dual low-profile PCIe cards reach 120 TOPS for industrial and enterprise servers. Multi-card systems scale to 240 TOPS. Every tier runs the same MERA compiler and framework, which accepts models from Hugging Face, TensorFlow Lite, ONNX, TVM and MLIR and deploys them across heterogeneous hosts — including EdgeCortix’s own announced Raspberry Pi 5 support and partner integrations with Renesas RZ/V and BittWare.

The silicon is specified from –40 to 105°C and carries FCC, CE, UKCA, ICES, VCCI and BSMI certifications, with a full REACH and RoHS materials package — unusual rigour for an accelerator of this class. EdgeCortix has also validated a radiation-resilient configuration for orbital and lunar missions, which says something about the design’s robustness. For defense, robotics, drone and smart-manufacturing programmes that need generative and vision AI at the edge without dragging in a datacenter power and cooling budget, SAKURA-II is one of the few credible options at 8W.

Target Applications

Defense and security systems, robotics and drones, smart manufacturing inspection, smart city video analytics, automotive sensing, edge AI servers running multimodal models, and any deployment needing Llama 2 or Stable Diffusion inference inside an 8W thermal envelope.

Related Products

Request a Quote — EdgeCortix SAKURA-II

QS Compute — global B2B sourcing for EdgeCortix SAKURA-II M.2 modules, PCIe cards and multi-accelerator systems. Evaluation units and volume allocation on request.

Get Your Quote →