Specifications

Manufacturer

Skymizer Taiwan Inc.

Model

HTX301

Platform

HyperThought (LPU IP)

Chips Per Card

6 × HTX301 LPU

On-Card Memory

384 GB standard LPDDR-class DRAM

Memory Range Across Configs

32 GB – 384 GB

Model Reach

4B – 700B parameters, no over-provisioning

Instruction Set

LISA

Total Board Power

≈ 240 W

Process Node

28 nm mature node

Host Interface

Standard PCIe add-in card

Requires NVLink / NVSwitch

No

Efficiency Point

30 tokens/s at 100 GB/s bandwidth and 0.5 TOPs

Llama2 7B Prefill (octa-core)

240 tokens/s

Multi-Chip Scaling

Up to 1,200 tokens/s (Llama 7B prefill)

Weight Compression

9% – 17.8% better than llama.cpp baseline

KV-Cache Compression

0.06% – 3.52% perplexity loss

Architecture Features

Prefill / decode disaggregation, decode-first

Deployment

On-premises, air-cooled

Target Workload

Enterprise and agentic AI inference

Overview

The Skymizer HTX301 takes an unusual route to large-model inference: instead of HBM and a cutting-edge process node, it uses six HTX301 LPU chips and 384GB of standard LPDDR4/LPDDR5 DRAM on a single PCIe card, built on mature 28nm silicon. The result is a card that runs 700B-parameter models locally at roughly 240W — less than half the power of a single RTX PRO 6000 Blackwell — with no GPU cluster, no NVLink or NVSwitch fabric, and none of the accompanying cooling infrastructure.

The architecture is built on the LISA instruction set with prefill and decode disaggregation and a decode-first bias, which is where LLM inference actually spends its time. Skymizer quotes an efficiency point of 30 tokens per second from just 0.5 TOPs at 100 GB/s of memory bandwidth, 240 tokens per second on Llama 2 7B prefill from the octa-core configuration, and up to 1,200 tokens per second when multiple chips scale together. Weight compression runs 9% to 17.8% better than a llama.cpp baseline, and KV-cache compression is claimed at 0.06% to 3.52% perplexity loss.

Because the card is a plain PCIe add-in rather than a rack-scale system, it deploys into an existing server and keeps inference, weights and prompts on-premises — the data-sovereignty argument for regulated industries that cannot send prompts to a public cloud. The same HyperThought platform scales down to 32GB and up from 4B to 700B parameters, so a deployment can be right-sized to the model actually being served rather than provisioned for the worst case.

Key Benefits

700B-parameter inference from one card, 384GB of memory, ~240W total board power, no NVLink/NVSwitch or multi-node orchestration, on-prem data sovereignty, and right-sizing from 32GB/4B up to 384GB/700B.

Applications

On-premises enterprise LLM serving and agentic AI workflows, regulated-industry inference where prompts and weights cannot leave the premises, private RAG pipelines, edge-to-mini-datacentre deployment, and cost-sensitive inference where power and cooling are the binding constraint.

Request a Quote — SKYMIZER HTX301 — 384GB SINGLE-CARD LLM INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

A100 Pcie A100 Sxm4 Adlink Egx Mxm Blackwell