Specifications

NPU

Quad-ARIES — 4x Mobilint ARIES NPUs, 32 NPU cores

AI Performance

320 TOPS

NPU Frequency

1.25 GHz

Memory Type

LPDDR4X, onboard

Memory Capacity

64 GB

Memory Bandwidth

Up to 266.8 GB/s

Flash Memory

32 MB

Host Interface

PCIe Gen5 x16

Thermal Design Power

135 W

Input Voltage

DC +12 V (slot + AUX 8-pin)

Dimensions

268 x 112 x 19 mm

Weight

834 g

Cooling Options

Passive heatsink or heatsink + active fan (passive slim design available)

Target Deployment

Sustained throughput in ventilated rack environments

UART

1 port

LED Indicators

Power x3, Status x4

Supported OS

Linux, Windows

Supported Frameworks

PyTorch, TensorFlow, TFLite, ONNX, Keras

SDK

Mobilint SDK qb

Price

Quote upon request

Availability

Available

Overview

The MLA400 packs four ARIES NPUs onto a single PCIe Gen5 x16 card, producing 320 TOPS of AI inference within a 135 W power budget. With 64 GB of onboard LPDDR4X running at up to 266.8 GB/s, the card is dimensioned for multi-stream video analytics, multi-user LLM inference and multimodal workloads that would otherwise require several conventional GPU cards and far more rack power.

The card is 268 mm long, 112 mm high and 19 mm thick, and draws from both the PCIe slot and an auxiliary 8-pin connector. Mobilint offers both a passive slim cooling design for high-airflow servers and a heatsink-plus-active-fan variant for chassis where directed airflow is limited, with a stated deployment profile of sustained throughput in ventilated environments.

Software support mirrors the rest of the ARIES family: the SDK qb driver and runtime, Linux and Windows operating systems, and direct compilation of PyTorch, TensorFlow, TFLite, ONNX and Keras models.

Key Benefits

320 TOPS in a single slot-class card for density-sensitive inference racks. 266.8 GB/s memory bandwidth keeps large vision and language models fed. Passive or active cooling to match server airflow. One software stack across PCIe, MXM and box-level ARIES products.

Applications

Rack-scale video analytics and smart-city inference; on-premises multi-user LLM serving; multimodal document and video understanding; factory-wide visual quality inspection; telecom and broadcast content inference; sovereign and private-cloud AI deployment.

Request a Quote — MOBILINT ARIES MLA400 PCIE CARD — 320 TOPS QUAD-ARIES INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

Mobilint ARIES MLA100 PCIe Card — 80 TOPS Accelerator Axelera Server 150P — 856 TOPS Inference Card DEEPX DX-H1 V-NPU — 50 TOPS Video Analytics Card