Specifications
Architecture
Hybrid Von Neumann plus 2D SIMD matrix execution, unified pipeline
Family
Chimera QB series: QB1, QB4 and QB16 single cores, clusterable
Single-core ML
1 TOPS (QB1), 4 TOPS (QB4), 16 TOPS (QB16)
DSP capability
64 GOPS (QB1), 256 GOPS (QB4), 1 TOPS (QB16)
Scalable range
1 TOPS to 6,912 TOPS across multi-core and multi-cluster builds
MAC configurations
512 MAC to 8K MAC per core configuration
Pipeline
7-stage in-order, single issue per clock, 64-bit instruction word
Data types
INT8 ML MACs with optional FP16; 32-bit ALU DSP operations
On-chip memory
Configurable L2 data memory from 1 MB to 16 MB
Register storage
Distributed local register memories (LRM) with data broadcast networks
Instruction cache
Configurable 64 KB / 128 KB / 256 KB
Host interface
Configurable AXI interfaces for independent data and instruction access
Process node
1 GHz in mainstream 16 nm and 7 nm; portable to any foundry and process
Implementation
Conventional standard-cell flow with single-ported SRAM
Programmability
100% C++ programmable, no code partitioning across cores
Graph support
ONNX, TensorFlow and PyTorch graphs compiled into one binary
Toolchain
Chimera SDK plus Chimera Compute Library (CCL) and cycle-accurate simulation
Custom operators
Added as C++ kernels in software — no silicon respin required
Execution model
Deterministic and non-speculative with granular predication
Power management
Compiler-driven fine-grained clock gating and hierarchical memory
Target applications
Vision and imaging, radar and LiDAR, baseband, sensor fusion
Overview
The Quadric Chimera GPNPU is a licensable processor architecture that merges the throughput of a neural accelerator with the full programmability of a modern DSP. Instead of pairing a fixed-function NPU with a CPU or DSP fallback, Chimera executes matrix, vector and scalar instructions in one 7-stage in-order pipeline, so an entire inference graph plus its pre- and post-processing runs as a single binary on a single core.
That design removes the multi-core partitioning that dominates heterogeneous edge SoC work. New operators are written as C++ kernels against the Chimera Compute Library rather than taped out in silicon, which extends the useful life of a device across new network architectures, including transformers and large language models, through software updates alone.
Chimera scales from a single 1 TOPS QB1 core to clustered configurations rated at 6,912 TOPS on the same architecture and toolchain, and is portable across foundries, process nodes and cell libraries. It is designed for silicon teams building edge vision, radar and LiDAR, communications and sensor-fusion products.
Key Benefits
Unified core and toolchain — matrix, vector and scalar work share one pipeline, one compiler and one debug console, eliminating inter-core data shuffling and the extra memory subsystems that go with it. Future-proofed silicon — new operators and libraries arrive as software, avoiding costly respins. Efficient by construction — configurable on-chip memory of 1 MB to 16 MB with compiler-driven clock gating minimises power-hungry DDR traffic. Scales without redesign — one architecture covers 1 TOPS edge sensors through 6,912 TOPS clustered inference.
Applications
Edge AI vision and imaging subsystems, industrial camera and inspection pipelines, automotive and robotics perception with radar and LiDAR pre-processing, communications baseband where ML meets DSP, multi-sensor fusion modules, and any SoC program that needs custom operators, deterministic latency and a compact power envelope.
Request a Quote — QUADRIC CHIMERA GPNPU — 1 TO 6,912 TOPS NEURAL PROCESSOR IP
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →