Published: September 5, 2026 | Category: Technical | QSCompute
A vision-language model (VLM) fuses an image encoder with a large language model: one network that can read a gauge, describe a weld defect in a sentence, answer "why did this part fail?" from a photo, and even drive robots through vision-language-action policies. In 2026 the open-model ecosystem — Qwen2.5-VL, NVIDIA VILA, LLaVA, Gemma 3, MiniCPM-V — has matured enough that OEMs are replacing multi-model pipelines (detector + OCR + classifier + rules) with a single multimodal model. But VLMs change the hardware math: they are memory-hungry, token-hungry and quantization-sensitive. This guide sizes Jetson modules, workstation GPUs and data-center accelerators for VLM workloads, anchored on NVIDIA's own measured Orin-vs-Thor numbers.
Classic edge vision is a chain: a fast detector (YOLO-class) finds regions, a separate OCR or classifier labels them, and business logic decides. Each stage is small, cheap and runs at 30+ FPS on a 20–100 TOPS device. VLMs collapse that chain into one open-vocabulary model — prompt "describe any anomaly you see and read the serial number" without training a custom class. That suits low-volume, high-variety inspection (random defects, worn tooling), document and display reading, safety monitoring, and natural-language search over surveillance video (NVIDIA's Metropolis VSS pattern).
The trade-off is real: VLMs cost 10–100× more compute and memory per image than a detector, and they do not replace it in every role. A 60 FPS, 16-camera PCB line still belongs to a small YOLO-class model; the VLM earns its place as the reasoning layer that catches rare, unmodeled failures and writes the human-readable report. The architecture that wins in practice is hybrid: detector at the edge, VLM gatekeeper on anomalies, both on the same Jetson or GPU box.
VLM memory = model weights + KV cache for every token (text and image tokens) + runtime headroom. Weights dominate, and the rule of thumb is unchanged: ~2 GB per billion parameters at FP16, ~1 GB at INT8, ~0.6 GB at INT4. What surprises first-time buyers is the image side: every image is expanded into hundreds or thousands of tokens before it reaches the language layers, and each of those tokens occupies KV-cache memory for the whole generation.
| VLM class | Representative models | FP16 | INT8 | INT4 (W4A16) | Typical task |
|---|---|---|---|---|---|
| 1–3B | Qwen2.5-VL-3B, PaliGemma-3B, Moondream2, SmolVLM2 | 2–6 GB | 1–3 GB | 1–2 GB | Captioning, simple VQA, OCR-lite |
| 7–8B | Qwen2.5-VL-7B, VILA-1.5-8B, LLaVA-NeXT-8B, MiniCPM-V 2.6 | 14–16 GB | 7–8 GB | 5–6 GB | Defect description, document reading, visual QA |
| 11–13B | Llama 3.2 11B Vision, Gemma 3 12B, VILA-1.5-13B | 22–26 GB | 11–13 GB | 7–8 GB | Fine-grained reasoning over single images |
| 27–32B | Qwen2.5-VL-32B, Gemma 3 27B, InternVL2.5-26B | 54–64 GB | 27–32 GB | 17–20 GB | Agentic vision, complex scene reasoning |
| 72–90B | Qwen2.5-VL-72B, Llama 3.2 90B Vision | 144–180 GB | 72–90 GB | 43–55 GB | Highest accuracy, multi-image and video |
Vision-token counts vary by architecture: Gemma 3 encodes each image into 256 tokens, PaliGemma's SigLIP encoder produces 1,024 tokens at 448×448, and Qwen2.5-VL's dynamic resolution can pass 1,500 tokens for a high-resolution frame. Video scales linearly: a 30-second clip sampled at 2 FPS is 60 frames — at 500–1,000 tokens per frame that is 30,000–60,000 tokens of context. Budget 1–3 GB of headroom above weights at the 7–8B scale, and 4–16 GB for 32B-class models with long contexts or video. "The model fits in 24 GB" is not the same as "the deployment fits in 24 GB."
NVIDIA publishes output-tokens-per-second numbers for the same VLM on Jetson AGX Orin and AGX Thor — the cleanest apples-to-apples edge data available. On Qwen2.5-VL-3B (W4A16 on Orin, FP4 on Thor), Orin sustains ~216 output tokens/s and Thor ~357 tokens/s (1.65×). Qwen2.5-VL-7B on Thor with FP4 plus Eagle-style speculative decoding is up to 3.5× faster than Orin at W4A16, and Thor holds 16 concurrent VLM + LLM requests with time-to-first-token under 200 ms and time-per-output-token under 50 ms.
| Platform | Memory | Realistic VLM class | Precision sweet spot | Anchor & notes | 2026 price context |
|---|---|---|---|---|---|
| Jetson Orin Nano Super | 8 GB | 1–3B | INT4 / INT8 | Camera-side captioning and VQA; no room for long video | Dev kit $249 (67 TOPS) |
| Jetson Orin NX | 16 GB | 3–8B | INT4 / INT8 | 8B INT4 fits with ~10 GB free for vision tokens + KV | Module pricing at volume |
| Jetson AGX Orin | 64 GB | 7–13B; 32B at INT4 | W4A16 / INT8 | Qwen2.5-VL-3B ≈216 out tok/s; 70B-class FP16 does not fit | Industrial 64 GB module ≈$2,799 one-unit, ≈$2,239 at 1,000 |
| Jetson AGX Thor (T5000) | 128 GB @ 273 GB/s | 27–32B; 70B-class at FP4 | NVFP4 / FP4 (Blackwell) | Qwen2.5-VL-3B ≈357 out tok/s; 16 concurrent requests, TTFT <200 ms; 2,070 FP4 TFLOPS at 40–130 W | Dev kit $3,499 |
| RTX 4090 workstation | 24 GB | 7–13B; 32B INT4 tight | INT8 / INT4 | Good single-stream dev and demo box; no MIG-style sharing | Street, config-dependent |
| RTX 6000 Ada / L40S | 48 GB | 32B INT8; 72B INT4 | INT8 / INT4 / FP8 | Multi-stream serving, VLM fine-tuning, 4–8 camera gates per card | 48 GB class, quote-based |
| H100 / H200 class | 80 GB | 72B INT8/FP16; video VLMs | FP8 / INT8 / FP16 | Scale-out serving and fine-tuning; needed only above ~48 GB working set | Server platform, quote-based |
Three conclusions follow. First, Orin is a 7–13B platform: with W4A16 it runs the models most inspection and document tasks need, but Ampere has no FP4 tensor cores, so squeezing 32B-class accuracy onto Orin means INT4 or INT8 with a real quality check. Second, Thor (Blackwell) is the first edge module where the 27–32B tier is comfortable and 70B-class quantized models are plausible — NVFP4/FP4, 128 GB unified memory and 273 GB/s are exactly what video and agentic VLM workloads starve on without. Third, workstation and server GPUs still win on concurrency: vLLM-style batching on an L40S-class 48 GB card can serve several inspection lines at once, which no single Jetson can match.
Quantization is the biggest deployment lever, and it is silicon-dependent. W4A16 (INT4 weights, FP16 activations) is the default for Orin and RTX GPUs; modern 4-bit schemes keep most VLM QA accuracy within a point or two of FP16 — but verify on your own defect set, because OCR and fine-grained attribute questions degrade faster than captioning. On Blackwell parts (Thor, newer data-center GPUs), NVFP4 delivers near-FP8 accuracy at half the memory traffic of INT4, and FP8 spans both generations through NVIDIA's NGC containers. Run with the runtime you will ship: vLLM and TensorRT-LLM containers on Jetson, llama.cpp for light single-model cases. One field lesson from the Thor launch: VLM prompts are prefix-heavy (system prompt + image tokens repeat on every frame), and enabling vLLM's automatic prefix caching produced the largest single speedup seen — bigger than model or hardware changes alone.
VLMs have made edge AI genuinely multimodal in 2026, but they reward buyers who do the memory math before the hardware order. Match the VLM class to the silicon's precision capabilities, keep classic detectors for the high-FPS work, and verify throughput on the actual inference runtime — that combination separates deployments that demo well from deployments that ship.
Need a VLM-capable edge platform sized for your workload?
QSCompute configures Jetson AGX Thor and Orin systems, RTX/L40S GPU servers and hybrid detector+VLM inspection boxes — and we will run the memory and throughput math on your own images before you commit. Send your model class, camera count and frame-rate target for a validated BOM within 48 hours.
Contact: +86 137-1464-6179 | info@qscompute.com