Published: August 8, 2026 | Category: Buying Guide | QSCompute
Embedded AI isn't one thing. It spans a $3 Arm Cortex-M85 running keyword spotting to a $15,000 dual-GPU edge server running multi-camera defect detection. The wrong tier choice means either underpowered hardware that can't run your model — or a $15K server idling at 8% utilization burning power on a factory floor. This guide maps five 嵌入式 compute tiers with real specs, power numbers, and deployment examples so you land in the right tier the first time.
| Tier | Platform Examples | AI Throughput | Power | Unit Cost | Best For |
|---|---|---|---|---|---|
| Tier 1: MCU + NPU | STM32N6 (Cortex-M85 + Neural-ART 600 GOPS), NXP i.MX RT700, ESP32-S3 | 0.1–0.6 TOPS | 0.05–0.5W | $3–$15 | Keyword spotting, vibration anomaly, person detection on battery |
| Tier 2: NPU SoC | Rockchip RK3588 (6 TOPS), Amlogic A311D2 (5 TOPS), MediaTek Genio 1200 (4.8 TOPS) | 4–6 TOPS | 3–8W | $25–$80 | Single-camera object detection, facial recognition, smart display |
| Tier 3: SBC / Dev Kit | Jetson Orin Nano 8 GB (40 TOPS), Raspberry Pi 5 + Hailo-8L (13 TOPS), Luckfox Lyra (1 TOPS) | 1–40 TOPS | 5–15W | $70–$499 | Multi-model pipelines, prototype edge AI, educational dev |
| Tier 4: Industrial SBC | Jetson Orin NX 16 GB (100 TOPS), Rockchip RK3588 industrial (-40~85 °C), Aaeon BOXER (AGX Orin) | 6–275 TOPS | 10–45W | $399–$2,799 | Factory-floor vision, autonomous mobile robots, outdoor gateways |
| Tier 5: Edge Server | Supermicro SYS-221H-TNR (dual Xeon + 2× L40S), Neousys Nuvo-10208GC (Xeon + RTX 5090) | 300–2,900 TOPS | 150–1,200W | $4,000–$15,000 | Multi-camera AOI, fleet inference, video analytics clusters |
| Question | → Tier 1 | → Tier 2 | → Tier 3 | → Tier 4 | → Tier 5 |
|---|---|---|---|---|---|
| Cameras? | 0 | 1 (≤1080p) | 1–2 (4K) | 2–8 (4K) | 8–64 (4K) |
| Model size? | <1M params | MobileNet, EfficientDet | YOLOv8-nano, ResNet-50 | YOLOv10-m, Llama 3.2 1B | Llama 3.1 8B–70B, ViT-L |
| Latency target? | <10 ms | 10–50 ms | 5–30 ms | 3–20 ms | <5 ms (batched) |
| Environment? | Indoor, battery | Indoor, PoE | Lab, prototype | Factory, outdoor, −40~85 °C | Server room, rack |
| Deployment units? | 100K–1M+ | 10K–100K | 10–100 | 10–500 | 1–10 |
STMicro's STM32N6 is the breakthrough here: a Cortex-M85 at 800 MHz with a proprietary Neural-ART accelerator delivering 600 GOPS at under 500 mW. That's enough for real-time keyword spotting (Hey Siri-class), vibration anomaly detection on motor bearings, and person presence via low-res IR sensor — all on a coin cell for months. NXP's i.MX RT700 (360 GOPS NPU, M33 + Cadence Tensilica DSP) competes at a similar price point. These chips are eating Tier 2's lunch for ultra-low-power sensor AI.
The Rockchip RK3588 dominates here with 6 TOPS NPU, quad A76 + quad A55 CPU, 8K video decode, and dual GbE at $25–$50 in volume. It runs YOLOv8n at 30 fps on a single 1080p stream at 6W. Amlogic A311D2 (5 TOPS) and MediaTek Genio 1200 (4.8 TOPS, integrated 5G modem option) compete directly. The Genio 1200's integrated ISP handles 48 MP sensors natively — no external ISP chip needed — which cuts total BOM by ~$8. This tier is ideal for smart cameras, access control terminals, and retail analytics boxes.
The Jetson Orin Nano 8 GB ($249 module, $499 dev kit) is the reference platform: 40 TOPS, CUDA/TensorRT ecosystem, runs any model you throw at it. But for cost-sensitive deployments, a Raspberry Pi 5 ($60) paired with a Hailo-8L M.2 accelerator ($70) delivers 13 TOPS at ~$130 total — practical for single-stream YOLOv8 at 30 fps. The Luckfox Lyra (RV1106, 1 TOPS, $15) is the ultra-budget option for simple classification. Tier 3 is for prototyping; production moves to Tier 2 or 4.
Same silicon as Tier 3, but with wide-temperature validation (-40~85 °C), industrial I/O (CAN bus, RS-485, isolated GPIO), and rugged enclosures (IP65/IP67). The Jetson Orin NX 16 GB at 100 TOPS is the workhorse — runs YOLOv10-m at 22 ms on 4K streams, handles Llama 3.2 1B for edge NLP, and fits a 10-year lifecycle. Aaeon, Aetina, and Seeed offer pre-validated systems that skip the 6-month carrier-board bring-up phase. Expect to pay a 40–60% premium over Tier 3 for industrial qualification.
A Supermicro SYS-221H-TNR with dual Intel Xeon 6780E (144 cores each) and 2× NVIDIA L40S puts 2,912 TOPS on an assembly line — enough for 32 concurrent 4K AOI streams running YOLOv10-x. The Neousys Nuvo-10208GC with a single RTX 5090 (1,637 TOPS at INT8) handles 16 streams at half the rack space. For fleet-level inference where you're processing all cameras from one node, Tier 5 is the only option. Liquid cooling becomes necessary above 800W of GPU — QSCompute pre-builds these systems with CDU and cold plates.
Building an 嵌入式 AI system? QSCompute stocks every tier — from STM32N6 to dual-GPU edge servers.
Pre-configured systems, burn-in tested, with 2-day dispatch from Shenzhen. Volume pricing from 100+ units.
Contact: +86 137-1464-6179 | sherry@qscompute.com