Specifications
Product Class
Fully programmable neural processing unit (NPU) IP core
ISA
RISC-V — CPU, vector and tensor unified in one core
TOPS per Core
Configurable 8 TOPS to 64 TOPS
Maximum Configuration
Scales to 256 TOPS
Clock Configurations
1 GHz and 2 GHz options for INT8 and INT4
Activations
INT8, INT16, INT32, INT64, FP16, FP32, FP64
Convolutions
INT4, INT8, INT16, FP16, BF16
Vector Unit
RVV1.0 compliant; DLEN 128–2048 bits, VLEN 128–4096 bits
Vector Cores
Up to 32 vector cores per vector unit
Tensor Unit
Built on the RW1.0 vector unit and existing vector registers
Memory Architecture
Gazzillion Misses memory controller for high-rate tensor data fetch
Clustering
C32 clustered IP for throughput scaling under one ISA
Target Workloads
LLMs, deep learning, recommendation systems, edge AI, AI datacenter
Programming
Vendor-lock-free — standard RISC-V toolchain
Overview
Semidynamics Cervell is an all-in-one RISC-V NPU IP core that unifies scalar CPU, vector and tensor processing in a single core built on the RISC-V instruction set. A configuration runs from 8 to 64 TOPS per core and the architecture scales to 256 TOPS, with 1 GHz and 2 GHz variants for INT8 and INT4 operation.
The design argument against conventional accelerators is data movement. A stand-alone NPU forces activations and weights to be shuffled between the CPU, the vector engine and the tensor engine, and often needs a DMA engine to keep the tensor unit fed. Cervell keeps all three in one core with one register file and one ISA, so there are no handovers between units, and the tensor unit reuses the vector registers to hold its matrices. That is what Semidynamics means by zero-latency AI compute.
The data-type support is unusually broad for a single core — INT4 through INT64 and FP16 through FP64, plus BF16 for convolutions. The vector unit is RVV1.0 compliant with data path lengths from 128 to 2048 bits and vector lengths from 128 to 4096 bits, allowing up to 32 vector cores to be chained for a wide power/performance/area trade-off. Because it is standard RISC-V, software is portable and there is no vendor lock-in. The C32 clustered IP scales throughput under a single ISA.
Key Benefits
CPU + vector + tensor in one RISC-V core — no data handovers between separate accelerators. 8 to 256 TOPS scalable from a single architecture across edge and datacenter. INT4 to FP64 and BF16 data-type coverage in one core. RVV1.0 with DLEN to 2048 bits and up to 32 vector cores. Standard RISC-V toolchain means no vendor lock-in.
Applications
Large-language-model inference acceleration, deep-learning and recommendation workloads, edge AI subsystems, AI datacenter accelerator designs, and silicon teams building custom AI SoCs who need a programmable RISC-V NPU rather than a fixed-function accelerator.
Request a Quote — SEMIDYNAMICS CERVELL — ALL-IN-ONE RISC-V NPU IP FROM 8 TO 256 TOPS
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →