Specifications

Product

Kinara Ara-2 discrete neural processing unit (DNPU), from the NXP Kinara acquisition

Architecture

Discrete NPU - dataflow execution with dense MAC arrays, tightly coupled on-chip memory and deterministic scheduling

Performance

Approximately 40 eTOPS on Ara-2; Ara-1 rated around 6 eTOPS for vision workloads

Power

Under 5 W typical on Ara-2

Model Coverage

Convolutional networks, transformer models, multimodal LLM and vision-language models

Generative AI

Runs 2B-parameter-class LLMs locally, enabling document summarisation, voice assistants and interactive applications without cloud latency

Generational Gain

Roughly 8x the performance of the first-generation Kinara NPU

Key Advantage

Significantly higher performance per watt than general CPUs or GPUs for edge neural inference

Scalability

Ara-1 and Ara-2 let designers scale AI performance independently of the host microprocessor

Host Independence

DNPU sits alongside an existing host - x86, Arm or NXP applications processor - rather than replacing it

Module Form

Available as an M.2 module (Ara240) for standard edge hosts with PCIe M.2 slots

Target Markets

Retail, warehouse robotics, industrial automation, elderly monitoring, home and building security, PC and laptop edge AI

Applications

Multimodal models extracting features and context from images and video, with text and voice inputs driving actions on visual data

NXP Integration

Long-standing member of the NXP Partner programme with existing deployments on NXP applications processors

Software

Model optimisation, quantisation and compilation workflow; support for standard frameworks

Deployment

Industrial safety analysis, factory hazard detection, retail analytics, smart building and security systems

Form Factor Ecosystem

From M.2 add-in modules up to embedded board designs

Reliability Target

Deterministic scheduling for predictable latency in real-time vision pipelines

Cross-Brand

Hailo-8 / Hailo-10H, MemryX MX3, DEEPX DX-M1, EdgeCortix SAKURA-II, Axelera Metis

Overview

Kinara's argument is that edge inference should not require a GPU. A discrete NPU (DNPU) is a standalone accelerator that sits next to whatever host processor the system already uses, and it is architected for inference only - dense MAC arrays, on-chip memory and deterministic scheduling rather than the cache hierarchy and general-purpose flexibility of a CPU or GPU.

The Ara-2 is the second-generation part. Rated at roughly 40 eTOPS and drawing under 5 W, it is around eight times the performance of the first-generation Kinara NPU and, more importantly for current designs, it targets transformer and multimodal workloads rather than only convolutional networks. That makes local LLM and vision-language inference practical without a cloud round trip.

Because the accelerator is independent of the host, it can be added to an existing x86 or Arm embedded system as an M.2 module, or built into a custom board. NXP acquired Kinara in 2025 and now positions the Ara family alongside its own applications processors, with Ara-1 covering vision-only workloads and Ara-2 covering generative and multimodal models.

Key Benefits

Under 5 W for approximately 40 eTOPS of edge inference, transformer and multimodal model support rather than CNN-only acceleration, and host-independent scaling so AI performance can be chosen separately from the main processor.

Applications

Industrial safety and hazard detection, warehouse and logistics robotics, retail analytics and loss prevention, elderly and patient monitoring, smart building and access control, factory quality inspection, and on-device assistants running local language models.

Request a Quote — NXP KINARA ARA-2 DISCRETE NPU — GENAI INFERENCE AT THE EDGE

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

MemryX MX3 — Edge AI Hailo-8 M.2 Accelerator — Edge AI DEEPX DX-M1 — Edge AI