Specifications
NPU
Mobilint ARIES — 8 NPU cores at 1.25 GHz
AI Performance
80 TOPS (INT8)
Memory
16 GB LPDDR4X (32 GB optional), onboard
Memory Bandwidth
66.7 GB/s
Host Interface
PCI Express Gen4, 8-lane
Flash Memory
32 MB
UART
1 port
LED Indicators
Power and status
Thermal Design Power
25 W
Input Voltage
DC +12 V
Cooling
Passive heatsink / active fan options
Supported OS
Linux, Windows
Supported Frameworks
PyTorch, TensorFlow, TFLite, ONNX, Keras
SDK
Mobilint SDK qb (driver + runtime)
Benchmark — MobileNetV2
11,551 FPS
Benchmark — ResNet-50
3,082 FPS
Benchmark — YOLO-11s
784 FPS
Benchmark — YOLO26m
343 FPS
Benchmark — EXAONE-4.0-1.2B
31.62 TPS/user
Benchmark — Llama-3.2-3B
12.16 TPS/user
Model Coverage
490+ open and closed-source models validated on the NPU architecture
Price
Quote upon request
Availability
Available
Overview
The Mobilint MLA100 is a PCIe Gen4 x8 add-in card built around the ARIES neural processing unit. It delivers 80 TOPS of INT8 inference from 16 GB of onboard LPDDR4X (32 GB optional) at a 25 W thermal design power — roughly a quarter of the board power of a mainstream datacenter GPU, which makes it suitable for dense multi-stream video analytics and on-premises LLM serving in standard 1U and 2U servers.
Because ARIES runs DNN architectures without requiring model restructuring, existing PyTorch, TensorFlow, TFLite, ONNX and Keras graphs can be compiled through the Mobilint SDK qb driver and runtime. Mobilint publishes throughput figures for the card across vision and language workloads: MobileNetV2 at 11,551 FPS, ResNet-50 at 3,082 FPS, YOLO-11s at 784 FPS, YOLO26m at 343 FPS, and on-device generative inference at 31.62 TPS/user for EXAONE-4.0-1.2B and 12.16 TPS/user for Llama-3.2-3B.
The card draws +12 V directly from the slot, so it drops into existing chassis without auxiliary power cabling. Both Linux and Windows are supported, making the MLA100 usable both in Linux-based video-analytics appliances and in Windows industrial deployments.
Key Benefits
80 TOPS at 25 W — high inference density per watt for fanless and sealed industrial enclosures. No GPU dependency — a purpose-built inference NPU rather than a general-purpose graphics card. Framework-compatible — PyTorch, TensorFlow, TFLite, ONNX and Keras models compile without rewriting. Slot-powered — +12 V from the PCIe slot, no auxiliary power connector.
Applications
Multi-channel video analytics and NVR/VMS appliances; on-premises small-language-model serving; smart-city and traffic inference; retail loss prevention and shelf analytics; medical and industrial inspection stations; edge AI boxes for utilities and manufacturing.
Request a Quote — MOBILINT ARIES MLA100 PCIE CARD — 80 TOPS EDGE AI INFERENCE ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →