Specifications
Product
Kinara Ara-2 discrete neural processing unit (DNPU), from the NXP Kinara acquisition
Architecture
Discrete NPU - dataflow execution with dense MAC arrays, tightly coupled on-chip memory and deterministic scheduling
Performance
Approximately 40 eTOPS on Ara-2; Ara-1 rated around 6 eTOPS for vision workloads
Power
Under 5 W typical on Ara-2
Model Coverage
Convolutional networks, transformer models, multimodal LLM and vision-language models
Generative AI
Runs 2B-parameter-class LLMs locally, enabling document summarisation, voice assistants and interactive applications without cloud latency
Generational Gain
Roughly 8x the performance of the first-generation Kinara NPU
Key Advantage
Significantly higher performance per watt than general CPUs or GPUs for edge neural inference
Scalability
Ara-1 and Ara-2 let designers scale AI performance independently of the host microprocessor
Host Independence
DNPU sits alongside an existing host - x86, Arm or NXP applications processor - rather than replacing it
Module Form
Available as an M.2 module (Ara240) for standard edge hosts with PCIe M.2 slots
Target Markets
Retail, warehouse robotics, industrial automation, elderly monitoring, home and building security, PC and laptop edge AI
Applications
Multimodal models extracting features and context from images and video, with text and voice inputs driving actions on visual data
NXP Integration
Long-standing member of the NXP Partner programme with existing deployments on NXP applications processors
Software
Model optimisation, quantisation and compilation workflow; support for standard frameworks
Deployment
Industrial safety analysis, factory hazard detection, retail analytics, smart building and security systems
Form Factor Ecosystem
From M.2 add-in modules up to embedded board designs
Reliability Target
Deterministic scheduling for predictable latency in real-time vision pipelines
Cross-Brand
Hailo-8 / Hailo-10H, MemryX MX3, DEEPX DX-M1, EdgeCortix SAKURA-II, Axelera Metis
Overview
Kinara's argument is that edge inference should not require a GPU. A discrete NPU (DNPU) is a standalone accelerator that sits next to whatever host processor the system already uses, and it is architected for inference only - dense MAC arrays, on-chip memory and deterministic scheduling rather than the cache hierarchy and general-purpose flexibility of a CPU or GPU.
The Ara-2 is the second-generation part. Rated at roughly 40 eTOPS and drawing under 5 W, it is around eight times the performance of the first-generation Kinara NPU and, more importantly for current designs, it targets transformer and multimodal workloads rather than only convolutional networks. That makes local LLM and vision-language inference practical without a cloud round trip.
Because the accelerator is independent of the host, it can be added to an existing x86 or Arm embedded system as an M.2 module, or built into a custom board. NXP acquired Kinara in 2025 and now positions the Ara family alongside its own applications processors, with Ara-1 covering vision-only workloads and Ara-2 covering generative and multimodal models.
Key Benefits
Under 5 W for approximately 40 eTOPS of edge inference, transformer and multimodal model support rather than CNN-only acceleration, and host-independent scaling so AI performance can be chosen separately from the main processor.
Applications
Industrial safety and hazard detection, warehouse and logistics robotics, retail analytics and loss prevention, elderly and patient monitoring, smart building and access control, factory quality inspection, and on-device assistants running local language models.
Request a Quote — NXP KINARA ARA-2 DISCRETE NPU — GENAI INFERENCE AT THE EDGE
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →