Published: August 15, 2026 | Category: Buying Guide | QSCompute
Retrieval-augmented generation (RAG) is moving off the public cloud. Enterprises want to query their own documents, schematics, and support tickets with a private LLM that never leaves the building — for data-sovereignty, cost, and latency reasons. The question is which 开发套件 actually runs a usable local LLM and RAG stack, not just a toy demo.
We benchmarked three dev kit tiers on the two workloads that matter for RAG — generation (Llama 3.1 8B, INT4) and embedding (BGE-M3 / bge-small-en-v1.5) — using the llama.cpp and Ollama runtimes most teams already deploy in production.
| Specification | Jetson Orin Nano Super | AMD Ryzen AI 9 HX 370 | Snapdragon X Elite (Dev Kit) |
|---|---|---|---|
| AI Compute | 67 TOPS (INT8) GPU | 50 TOPS NPU + Radeon 890M iGPU | 45 TOPS Hexagon NPU |
| CPU | 6× Cortex-A78AE (1.7 GHz) | 12× Zen 5 (up to 5.1 GHz) | 12× Oryon (up to 4.0 GHz) |
| RAM | 8 GB LPDDR5 (shared) | 32–64 GB LPDDR5x | 64 GB LPDDR5x |
| Storage | 64 GB eMMC + microSD | 512 GB–1 TB NVMe | 512 GB NVMe |
| LLM Runtime | Ollama, llama.cpp, MLC, TensorRT-LLM | Ollama, llama.cpp, LM Studio, ONNX | llama.cpp, ExecuTorch, ONNX |
| TDP | 7–25 W | 28–54 W | 15–30 W |
| Dev Kit Price (Q3 2026) | $249 | $799 (mini-PC) | $899 |
Test: Ollama/llama.cpp generating 256 tokens from a 4K-token prompt, measured in tokens/sec. This is the number your users feel during an interactive RAG session — anything below ~8 tok/s reads as "laggy".
| Metric | Jetson Orin Nano Super | AMD Ryzen AI 9 HX 370 | Snapdragon X Elite |
|---|---|---|---|
| Llama 3.1 8B INT4 (tok/s) | 11.2 | 27.8 | 22.4 |
| Llama 3.1 8B FP16 (tok/s) | N/A (OOM) | 8.9 | 6.1 |
| Max model size (INT4) | 8B | 13B | 13B |
| Context window (practical) | 4K | 16K | 16K |
| Power during generation | 18.4 W | 38.2 W | 24.6 W |
Winner for raw throughput: AMD Ryzen AI 9 HX 370. Its Radeon 890M iGPU (with 16 GB+ allocated from system RAM) sustains 27.8 tok/s on 8B INT4 — more than double the Jetson. The 8 GB Jetson is hard-limited to 8B-class models and a 4K context; the 64 GB Snapdragon and 32–64 GB Ryzen kits comfortably run 13B and long-context workloads. For a single-user or small-team RAG assistant, the Ryzen AI mini-PC is the sweet spot.
RAG ingestion — chunking a document corpus and writing vectors to a local store (Chroma, Qdrant, or LanceDB) — is the workload that actually dominates wall-clock time. We measured documents embedded per second (bge-small-en-v1.5, 384-dim, batch 32).
| Metric | Jetson Orin Nano Super | AMD Ryzen AI 9 HX 370 | Snapdragon X Elite |
|---|---|---|---|
| Embedding throughput (docs/s) | 86 | 312 | 204 |
| 10K-doc corpus ingest time | 1 min 56 s | 32 s | 49 s |
| Reranker support (bge-reranker-v2) | Marginal (8 GB) | YES | YES |
Winner: AMD Ryzen AI. CPU-bound embedding is dominated by core count and memory bandwidth — the 12-core Zen 5 part pulls ahead decisively. The Jetson's 6 Cortex-A78AE cores and 8 GB ceiling make it a poor choice for anything beyond a few hundred documents. If your RAG index grows past ~50K chunks, skip all three dev kits and move to a small edge server (see related articles below).
| Deployment | Corpus Size | Recommended Kit | Reason |
|---|---|---|---|
| Single-user code/doc assistant | <5K docs | Jetson Orin Nano Super | Cheapest, 11 tok/s is acceptable solo, 25 W ceiling |
| Small-team internal RAG | 5K–50K docs | AMD Ryzen AI 9 HX 370 | 27.8 tok/s, fast ingest, 13B capable |
| Battery/vehicle RAG | <10K docs | Snapdragon X Elite | Best tok/s per watt, 64 GB for long context |
| Enterprise RAG (100K+ docs) | >50K docs | Edge server (L40S / RTX 6000 Ada) | Dev kits hit memory and throughput walls |
$349
Jetson Orin Nano Super 8 GB · 128 GB NVMe (boot + vector store) · Pre-loaded Ollama + ChromaDB + Llama 3.1 8B INT4 · 30 W industrial PSU · Getting-started RAG notebook
$1,049
Ryzen AI 9 HX 370 · 64 GB LPDDR5x · 1 TB NVMe · Pre-loaded Ollama + Qdrant + bge-M3 · Llama 3.1 8B & 13B quantized · Fanless 1.2 L chassis
All three local LLM / RAG dev kits in stock — pre-loaded with models and a working RAG stack, ready to query your own documents.
Tell us your corpus size and latency target; we'll benchmark your exact model before you buy.
Contact: +86 137-1464-6179 | sherry@qscompute.com