Specifications

NPU

Mobilint ARIES — 8 NPU cores at 1.25 GHz

AI Performance

80 TOPS (INT8)

Memory

16 GB LPDDR4X (32 GB optional), onboard

Memory Bandwidth

66.7 GB/s

Host Interface

PCI Express Gen4, 8-lane

Flash Memory

32 MB

UART

1 port

LED Indicators

Power and status

Thermal Design Power

25 W

Input Voltage

DC +12 V

Cooling

Passive heatsink / active fan options

Supported OS

Linux, Windows

Supported Frameworks

PyTorch, TensorFlow, TFLite, ONNX, Keras

SDK

Mobilint SDK qb (driver + runtime)

Benchmark — MobileNetV2

11,551 FPS

Benchmark — ResNet-50

3,082 FPS

Benchmark — YOLO-11s

784 FPS

Benchmark — YOLO26m

343 FPS

Benchmark — EXAONE-4.0-1.2B

31.62 TPS/user

Benchmark — Llama-3.2-3B

12.16 TPS/user

Model Coverage

490+ open and closed-source models validated on the NPU architecture

Price

Quote upon request

Availability

Available

Overview

The Mobilint MLA100 is a PCIe Gen4 x8 add-in card built around the ARIES neural processing unit. It delivers 80 TOPS of INT8 inference from 16 GB of onboard LPDDR4X (32 GB optional) at a 25 W thermal design power — roughly a quarter of the board power of a mainstream datacenter GPU, which makes it suitable for dense multi-stream video analytics and on-premises LLM serving in standard 1U and 2U servers.

Because ARIES runs DNN architectures without requiring model restructuring, existing PyTorch, TensorFlow, TFLite, ONNX and Keras graphs can be compiled through the Mobilint SDK qb driver and runtime. Mobilint publishes throughput figures for the card across vision and language workloads: MobileNetV2 at 11,551 FPS, ResNet-50 at 3,082 FPS, YOLO-11s at 784 FPS, YOLO26m at 343 FPS, and on-device generative inference at 31.62 TPS/user for EXAONE-4.0-1.2B and 12.16 TPS/user for Llama-3.2-3B.

The card draws +12 V directly from the slot, so it drops into existing chassis without auxiliary power cabling. Both Linux and Windows are supported, making the MLA100 usable both in Linux-based video-analytics appliances and in Windows industrial deployments.

Key Benefits

80 TOPS at 25 W — high inference density per watt for fanless and sealed industrial enclosures. No GPU dependency — a purpose-built inference NPU rather than a general-purpose graphics card. Framework-compatible — PyTorch, TensorFlow, TFLite, ONNX and Keras models compile without rewriting. Slot-powered — +12 V from the PCIe slot, no auxiliary power connector.

Applications

Multi-channel video analytics and NVR/VMS appliances; on-premises small-language-model serving; smart-city and traffic inference; retail loss prevention and shelf analytics; medical and industrial inspection stations; edge AI boxes for utilities and manufacturing.

Request a Quote — MOBILINT ARIES MLA100 PCIE CARD — 80 TOPS EDGE AI INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

DEEPX DX-H1 Quattro — 100 TOPS INT8 PCIe NPU Card Axelera Edge 130P — 214 TOPS PCIe AI Accelerator Tenstorrent Blackhole p100a — 120 Tensix PCIe Card