Specifications

Architecture

Kalray MPPA (Massively Parallel Processor Array), KVX VLIW cores

Coolidge v2 Processor

MPPA Coolidge v2 @ 1 GHz, 80 KVX cores (standalone design-in SoC)

Coolidge v2 Acceleration

Integrated AI accelerator, cryptography, FEC and data compression

Coolidge v2 Interfaces

Ethernet, PCIe Gen 4, DDR4

K200-LP Processor

1× MPPA Coolidge v1 @ 1 GHz, 80 cores

K200-LP Memory

8 GB DDR4 @ 3200 MT/s

K200-LP Network

2× 100G Ethernet

K200-LP Host

PCIe Gen 4, 16-lane

K200-LP Offload

NVMe-oF (RoCE / TCP), encryption, compression, erasure coding

K300 Processor

1× MPPA Coolidge v2 @ 1 GHz, 80 cores

K300 Memory

16 GB DDR4 @ 3200 MT/s

K300 Network

4× 10/25G Ethernet

K300 Performance

40 TOPS (INT8) · 20 TFLOPS (FP16)

K300 Offload

FEC (LDPC), cryptography, data compression

TC4 Processor

4× MPPA Coolidge v2 @ 1 GHz, 320 cores

TC4 Memory

32 GB DDR4, PCIe Gen 4 16-lane host

TC4 Performance

160 TOPS (INT8) · 80 TFLOPS (FP16)

Turbo Server TS1

2U, 2× AMD EPYC 9004 (168 cores), 8× TC4 cards

Turbo Server Performance

1280 TOPS (INT8) · 640 TFLOPS (FP16)

Turbo Server Memory

24 slots of 12-channel RDIMM DDR5

Turbo Server Expansion

8× PCIe Gen 5 ×16 FHFL slots

Software

Kalray SDK with 80+ AI libraries, Docker, Triton Inference Server API

Overview

Kalray's MPPA (Massively Parallel Processor Array) DPUs are production-deployed accelerators built around large arrays of KVX VLIW cores. Unlike GPUs, they execute very large numbers of parallel, asynchronous operations — an architecture that suits mixed classical and AI processing workloads such as smart vision, storage offload and network packet processing.

The family covers a range of deployment shapes: the Coolidge v2 processor as a standalone design-in SoC; the K200-LP for storage, security and networking offload with NVMe-oF over RoCE/TCP; the K300 for AI and smart vision at 40 TOPS INT8; and the TC4 with four Coolidge v2 processors delivering 160 TOPS INT8 for demanding industrial AI inference.

For data-intensive smart vision, the 2U Turbo Server TS1 combines two AMD EPYC 9004 CPUs with eight TC4 cards for 1280 TOPS INT8 across eight PCIe Gen 5 ×16 FHFL slots, and ships with the open Kalray SDK — standard toolchains, Docker and a Triton Inference Server API, with no proprietary framework lock-in.

Key Benefits

Open programmability: standard toolchains, 80+ AI libraries, Docker and a Triton Inference Server API — no vendor lock-in. Complementary to GPUs: massive parallel asynchronous execution suits vision, storage and packet workloads alongside GPUs. Production-ready: deployed today in OEM and infrastructure vendor systems with pre-validated performance.

Applications

AI inference at the edge and in industrial vision · NVMe-oF storage acceleration · Network and security offload · Industry 4.0 machine vision · Telecom infrastructure processing

Request a Quote — KALRAY MPPA DPU CARDS — COOLIDGE V2, K200-LP, K300 & TC4 ACCELERATORS

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

AMD Pensando Salina 400 DPU Marvell OCTEON 10 CN103 DPU NVIDIA BlueField-3 B3140H SuperNIC