MPN: Altus XE4318GT-KVC
Penguin MemoryAI KV Cache Server
11 TB Memory · 4U · CXL KV Cache

Dual AMD EPYC 9005 · 3TB DDR5 + 8× 1TB CXL AICs · 88 DIMMs @ 6400 MT/s · Industry's first production-ready CXL KV cache server

Overview

The Penguin Solutions MemoryAI KV Cache Server is the industry's first production-ready KV cache server built on CXL memory, announced by Penguin Solutions (NASDAQ: PENG) to attack the memory wall that now governs AI inference economics. The premise is straightforward: training and tuning are episodic and compute-bound, but always-on inference and agentic workloads are memory-bound — Penguin puts the split at roughly 30% compute-driven and 70% memory-driven. By storing and reusing computed key/value pairs in a dedicated high-capacity CXL memory appliance rather than inside GPU memory, the server expands the RAM available to GPUs across the cluster, enables memory disaggregation so each node draws what it needs, and removes the re-compute penalty that shows up as GPU idle time. The measured result is faster time-to-first-token, less time per output token and higher end-to-end token throughput, with more consistent latency against tight SLA targets. The production configuration, model Altus XE4318GT-KVC, is a 4U rackmount built on dual AMD EPYC 9005 processors with 88 DDR5 DIMM slots at 6400 MT/s. It carries 3 TB of DDR5 main memory and up to eight 1 TB CXL add-in cards for 8 TB of additional CXL-attached DDR5 memory, with total system memory reaching 11 TB; the CXL AIC in the reference configuration is the SMART CXA-8F2W. Expansion is eight PCIe Gen5 x16 full-height full-length slots plus two low-profile PCIe Gen5 x16 slots, which keeps the appliance usable as both a memory target and a PCIe host for accelerators or storage. Penguin states customers are already deploying the MemoryAI KV cache server in production clusters, and has shown it at NVIDIA GTC in San Jose (March 2026) and at AI Infra Summit 2026 in Santa Clara. It sits inside Penguin's Full-Stack AI Factory Platform alongside the company's compute systems, infrastructure software and design-and-build services for enterprise, sovereign and neocloud data centers.

Specifications

Manufacturer

Penguin Solutions, Inc.
NASDAQ: PENG, Fremont, California

Model

Altus XE4318GT-KVC
MemoryAI KV Cache Server

Form Factor

4U rackmount
Memory appliance

Processor

Dual AMD EPYC 9005 Series

Total Memory

Up to 11 TB DDR5
@ 6400 MT/s, 88 DIMM slots

Main Memory

3 TB DDR5
system memory

CXL Memory

Up to 8 TB DDR5
via eight 1 TB CXL AICs

CXL Add-in Card

SMART CXA-8F2W
reference configuration

PCIe Expansion

8× PCIe Gen5 x16 FHFL
2× PCIe Gen5 x16 low-profile

Core Function

KV cache offload
and memory disaggregation

Latency Benefit

Faster time-to-first-token
lower time per output token

Throughput Benefit

Higher end-to-end
token throughput

GPU Efficiency

Reduced GPU idle time
less redundant re-compute

Inference Profile

≈30% compute driven (GPU)
≈70% memory driven (RAM)

Memory Pooling

Disaggregated shared memory
accessible across nodes

Target Workloads

Agentic AI, LLM inference
RAG, vector databases

Deployment Status

Production-ready
customers already deploying

Platform

Penguin Full-Stack
AI Factory Platform

Shown At

NVIDIA GTC 2026, booth 1031
AI Infra Summit 2026, booth 520

Request a Quote — pricing available upon inquiry.

Contact Sales

Availability

Available. Contact our sales team for current lead times, volume pricing, and integration support.