Specifications

Architecture

Hybrid Von Neumann plus 2D SIMD matrix execution, unified pipeline

Family

Chimera QB series: QB1, QB4 and QB16 single cores, clusterable

Single-core ML

1 TOPS (QB1), 4 TOPS (QB4), 16 TOPS (QB16)

DSP capability

64 GOPS (QB1), 256 GOPS (QB4), 1 TOPS (QB16)

Scalable range

1 TOPS to 6,912 TOPS across multi-core and multi-cluster builds

MAC configurations

512 MAC to 8K MAC per core configuration

Pipeline

7-stage in-order, single issue per clock, 64-bit instruction word

Data types

INT8 ML MACs with optional FP16; 32-bit ALU DSP operations

On-chip memory

Configurable L2 data memory from 1 MB to 16 MB

Register storage

Distributed local register memories (LRM) with data broadcast networks

Instruction cache

Configurable 64 KB / 128 KB / 256 KB

Host interface

Configurable AXI interfaces for independent data and instruction access

Process node

1 GHz in mainstream 16 nm and 7 nm; portable to any foundry and process

Implementation

Conventional standard-cell flow with single-ported SRAM

Programmability

100% C++ programmable, no code partitioning across cores

Graph support

ONNX, TensorFlow and PyTorch graphs compiled into one binary

Toolchain

Chimera SDK plus Chimera Compute Library (CCL) and cycle-accurate simulation

Custom operators

Added as C++ kernels in software — no silicon respin required

Execution model

Deterministic and non-speculative with granular predication

Power management

Compiler-driven fine-grained clock gating and hierarchical memory

Target applications

Vision and imaging, radar and LiDAR, baseband, sensor fusion

Overview

The Quadric Chimera GPNPU is a licensable processor architecture that merges the throughput of a neural accelerator with the full programmability of a modern DSP. Instead of pairing a fixed-function NPU with a CPU or DSP fallback, Chimera executes matrix, vector and scalar instructions in one 7-stage in-order pipeline, so an entire inference graph plus its pre- and post-processing runs as a single binary on a single core.

That design removes the multi-core partitioning that dominates heterogeneous edge SoC work. New operators are written as C++ kernels against the Chimera Compute Library rather than taped out in silicon, which extends the useful life of a device across new network architectures, including transformers and large language models, through software updates alone.

Chimera scales from a single 1 TOPS QB1 core to clustered configurations rated at 6,912 TOPS on the same architecture and toolchain, and is portable across foundries, process nodes and cell libraries. It is designed for silicon teams building edge vision, radar and LiDAR, communications and sensor-fusion products.

Key Benefits

Unified core and toolchain — matrix, vector and scalar work share one pipeline, one compiler and one debug console, eliminating inter-core data shuffling and the extra memory subsystems that go with it. Future-proofed silicon — new operators and libraries arrive as software, avoiding costly respins. Efficient by construction — configurable on-chip memory of 1 MB to 16 MB with compiler-driven clock gating minimises power-hungry DDR traffic. Scales without redesign — one architecture covers 1 TOPS edge sensors through 6,912 TOPS clustered inference.

Applications

Edge AI vision and imaging subsystems, industrial camera and inspection pipelines, automotive and robotics perception with radar and LiDAR pre-processing, communications baseband where ML meets DSP, multi-sensor fusion modules, and any SoC program that needs custom operators, deterministic latency and a compact power envelope.

Request a Quote — QUADRIC CHIMERA GPNPU — 1 TO 6,912 TOPS NEURAL PROCESSOR IP

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

SiMa.ai Modalix MLSoC — Edge AI Silicon DEEPX DX-M1 — 25 TOPS NPU Accelerator Ambarella N1-655 — Edge GenAI SoC Synaptics Astra SL2610 — Edge AI Processor