Specifications
Manufacturer
Skymizer Taiwan Inc.
Model
HTX301
Platform
HyperThought (LPU IP)
Chips Per Card
6 × HTX301 LPU
On-Card Memory
384 GB standard LPDDR-class DRAM
Memory Range Across Configs
32 GB – 384 GB
Model Reach
4B – 700B parameters, no over-provisioning
Instruction Set
LISA
Total Board Power
≈ 240 W
Process Node
28 nm mature node
Host Interface
Standard PCIe add-in card
Requires NVLink / NVSwitch
No
Efficiency Point
30 tokens/s at 100 GB/s bandwidth and 0.5 TOPs
Llama2 7B Prefill (octa-core)
240 tokens/s
Multi-Chip Scaling
Up to 1,200 tokens/s (Llama 7B prefill)
Weight Compression
9% – 17.8% better than llama.cpp baseline
KV-Cache Compression
0.06% – 3.52% perplexity loss
Architecture Features
Prefill / decode disaggregation, decode-first
Deployment
On-premises, air-cooled
Target Workload
Enterprise and agentic AI inference
Overview
The Skymizer HTX301 takes an unusual route to large-model inference: instead of HBM and a cutting-edge process node, it uses six HTX301 LPU chips and 384GB of standard LPDDR4/LPDDR5 DRAM on a single PCIe card, built on mature 28nm silicon. The result is a card that runs 700B-parameter models locally at roughly 240W — less than half the power of a single RTX PRO 6000 Blackwell — with no GPU cluster, no NVLink or NVSwitch fabric, and none of the accompanying cooling infrastructure.
The architecture is built on the LISA instruction set with prefill and decode disaggregation and a decode-first bias, which is where LLM inference actually spends its time. Skymizer quotes an efficiency point of 30 tokens per second from just 0.5 TOPs at 100 GB/s of memory bandwidth, 240 tokens per second on Llama 2 7B prefill from the octa-core configuration, and up to 1,200 tokens per second when multiple chips scale together. Weight compression runs 9% to 17.8% better than a llama.cpp baseline, and KV-cache compression is claimed at 0.06% to 3.52% perplexity loss.
Because the card is a plain PCIe add-in rather than a rack-scale system, it deploys into an existing server and keeps inference, weights and prompts on-premises — the data-sovereignty argument for regulated industries that cannot send prompts to a public cloud. The same HyperThought platform scales down to 32GB and up from 4B to 700B parameters, so a deployment can be right-sized to the model actually being served rather than provisioned for the worst case.
Key Benefits
700B-parameter inference from one card, 384GB of memory, ~240W total board power, no NVLink/NVSwitch or multi-node orchestration, on-prem data sovereignty, and right-sizing from 32GB/4B up to 384GB/700B.
Applications
On-premises enterprise LLM serving and agentic AI workflows, regulated-industry inference where prompts and weights cannot leave the premises, private RAG pipelines, edge-to-mini-datacentre deployment, and cost-sensitive inference where power and cooling are the binding constraint.
Request a Quote — SKYMIZER HTX301 — 384GB SINGLE-CARD LLM INFERENCE ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →