Specifications
Vendor
Untether AI
Architecture
Second-generation at-memory compute
Compute cores
1,456 RISC-V processors per device
Peak throughput (speedAI240 device)
2 PFLOPS
Energy efficiency (speedAI240 device)
up to 30 TFLOPS/W
Thermal design power
75 W (Slim variant)
Form factor
Low-profile PCIe accelerator card
Model support
Any neural network type (CNN, transformer, VL model)
Software stack
imAIgine SDK push-button model compiler
Benchmark standing
Class-leading latency and throughput on MLPerf inference
Host compatibility
x86 and Arm-based CPU servers
Deployment tier
Cloud, regional datacenter and edge
Reference deployments
Ola-Krutrim (India and US), J-squared agtech
Target markets
Automotive vision, aerospace and defence, machine vision, agriculture
Overview
The Untether AI speedAI 240 Slim is a low-profile PCIe inference accelerator built on the company's at-memory compute architecture, in which processing elements sit directly alongside memory banks rather than shuttling data across a conventional bus. That structural choice removes the memory movement that dominates energy use in inference, which is why the family's headline figure is efficiency rather than raw peak throughput.
The 75 W Slim card is the deployment-oriented member of the speedAI 240 line. Where full-power accelerator cards need dedicated server airflow and rack power budgets, a 75 W low-profile part slots into standard commercial and industrial hosts, including Arm-based CPU servers - the configuration Ola-Krutrim runs in production. Untether AI positions the device for automotive vision, aerospace and defence object detection, manufacturing machine vision and agricultural autonomy, markets where determinism, latency and power draw matter more than peak FLOPS.
Software is delivered through the imAIgine SDK, a push-button flow that converts trained neural network models into optimised, inference-ready form without hand-written kernel work. The card supports any neural network model type rather than restricting buyers to a fixed operator set.
Key Benefits
75 W in a low-profile PCIe slot lets dense inference run in standard hosts without liquid cooling or specialised power delivery. At-memory compute yields up to 30 TFLOPS/W, cutting the energy cost per inference. 1,456 RISC-V cores give fine-grained parallelism without GPU-class memory traffic. Any-model support via imAIgine avoids retraining or model surgery at deployment.
Applications
Automotive ADAS and in-cabin vision, aerospace and defence object detection, factory machine-vision inspection, agricultural autonomous machinery, and regional cloud inference where power cost dominates.
Request a Quote — UNTETHER AI SPEEDAI 240 SLIM — 75W LOW-PROFILE PCIE INFERENCE ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →