Specifications

Vendor

Untether AI

Architecture

Second-generation at-memory compute

Compute cores

1,456 RISC-V processors per device

Peak throughput (speedAI240 device)

2 PFLOPS

Energy efficiency (speedAI240 device)

up to 30 TFLOPS/W

Thermal design power

75 W (Slim variant)

Form factor

Low-profile PCIe accelerator card

Model support

Any neural network type (CNN, transformer, VL model)

Software stack

imAIgine SDK push-button model compiler

Benchmark standing

Class-leading latency and throughput on MLPerf inference

Host compatibility

x86 and Arm-based CPU servers

Deployment tier

Cloud, regional datacenter and edge

Reference deployments

Ola-Krutrim (India and US), J-squared agtech

Target markets

Automotive vision, aerospace and defence, machine vision, agriculture

Overview

The Untether AI speedAI 240 Slim is a low-profile PCIe inference accelerator built on the company's at-memory compute architecture, in which processing elements sit directly alongside memory banks rather than shuttling data across a conventional bus. That structural choice removes the memory movement that dominates energy use in inference, which is why the family's headline figure is efficiency rather than raw peak throughput.

The 75 W Slim card is the deployment-oriented member of the speedAI 240 line. Where full-power accelerator cards need dedicated server airflow and rack power budgets, a 75 W low-profile part slots into standard commercial and industrial hosts, including Arm-based CPU servers - the configuration Ola-Krutrim runs in production. Untether AI positions the device for automotive vision, aerospace and defence object detection, manufacturing machine vision and agricultural autonomy, markets where determinism, latency and power draw matter more than peak FLOPS.

Software is delivered through the imAIgine SDK, a push-button flow that converts trained neural network models into optimised, inference-ready form without hand-written kernel work. The card supports any neural network model type rather than restricting buyers to a fixed operator set.

Key Benefits

75 W in a low-profile PCIe slot lets dense inference run in standard hosts without liquid cooling or specialised power delivery. At-memory compute yields up to 30 TFLOPS/W, cutting the energy cost per inference. 1,456 RISC-V cores give fine-grained parallelism without GPU-class memory traffic. Any-model support via imAIgine avoids retraining or model surgery at deployment.

Applications

Automotive ADAS and in-cabin vision, aerospace and defence object detection, factory machine-vision inspection, agricultural autonomous machinery, and regional cloud inference where power cost dominates.

Request a Quote — UNTETHER AI SPEEDAI 240 SLIM — 75W LOW-PROFILE PCIE INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

Axelera Europa AIPU Hailo-8 Century PCIe Card FuriosaAI RNGD Inference Accelerator Tenstorrent Blackhole p150