Specifications

Product Class

Fully programmable neural processing unit (NPU) IP core

ISA

RISC-V — CPU, vector and tensor unified in one core

TOPS per Core

Configurable 8 TOPS to 64 TOPS

Maximum Configuration

Scales to 256 TOPS

Clock Configurations

1 GHz and 2 GHz options for INT8 and INT4

Activations

INT8, INT16, INT32, INT64, FP16, FP32, FP64

Convolutions

INT4, INT8, INT16, FP16, BF16

Vector Unit

RVV1.0 compliant; DLEN 128–2048 bits, VLEN 128–4096 bits

Vector Cores

Up to 32 vector cores per vector unit

Tensor Unit

Built on the RW1.0 vector unit and existing vector registers

Memory Architecture

Gazzillion Misses memory controller for high-rate tensor data fetch

Clustering

C32 clustered IP for throughput scaling under one ISA

Target Workloads

LLMs, deep learning, recommendation systems, edge AI, AI datacenter

Programming

Vendor-lock-free — standard RISC-V toolchain

Overview

Semidynamics Cervell is an all-in-one RISC-V NPU IP core that unifies scalar CPU, vector and tensor processing in a single core built on the RISC-V instruction set. A configuration runs from 8 to 64 TOPS per core and the architecture scales to 256 TOPS, with 1 GHz and 2 GHz variants for INT8 and INT4 operation.

The design argument against conventional accelerators is data movement. A stand-alone NPU forces activations and weights to be shuffled between the CPU, the vector engine and the tensor engine, and often needs a DMA engine to keep the tensor unit fed. Cervell keeps all three in one core with one register file and one ISA, so there are no handovers between units, and the tensor unit reuses the vector registers to hold its matrices. That is what Semidynamics means by zero-latency AI compute.

The data-type support is unusually broad for a single core — INT4 through INT64 and FP16 through FP64, plus BF16 for convolutions. The vector unit is RVV1.0 compliant with data path lengths from 128 to 2048 bits and vector lengths from 128 to 4096 bits, allowing up to 32 vector cores to be chained for a wide power/performance/area trade-off. Because it is standard RISC-V, software is portable and there is no vendor lock-in. The C32 clustered IP scales throughput under a single ISA.

Key Benefits

CPU + vector + tensor in one RISC-V core — no data handovers between separate accelerators. 8 to 256 TOPS scalable from a single architecture across edge and datacenter. INT4 to FP64 and BF16 data-type coverage in one core. RVV1.0 with DLEN to 2048 bits and up to 32 vector cores. Standard RISC-V toolchain means no vendor lock-in.

Applications

Large-language-model inference acceleration, deep-learning and recommendation workloads, edge AI subsystems, AI datacenter accelerator designs, and silicon teams building custom AI SoCs who need a programmable RISC-V NPU rather than a fixed-function accelerator.

Request a Quote — SEMIDYNAMICS CERVELL — ALL-IN-ONE RISC-V NPU IP FROM 8 TO 256 TOPS

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

ESWIN EIC7700X — RISC-V Edge AI SoC EdgeCortix SAKURA-II — Edge AI Accelerator DEEPX DX-M2 — Edge AI NPU Accelerator Hailo-8 M.2 Accelerator — 26 TOPS Edge AI