The Problem
Separate NPU, DSP, CPU. Multiple vendors. Months porting models.
Traditional NPU IP Architecture
Separate NPU, DSP, and CPU components from different vendors that must be integrated, synchronized, debugged, and maintained independently.
Each block requires its own compiler and debugger, forcing AI workloads to fragment across cores.
Hardware designed for last year's models means new operators require costly silicon updates and respins.
The Quadric Approach
Single Core. 100% C++ Programmable. Single Binary.
Chimera GPNPU Core Architecture
Matrix, vector, and scalar operations execute in a single pipeline without partitioning.
One codestream, toolchain, and debug environment where ONNX and C++ merge seamlessly.
New operators can be added via C++ kernels instead of needing silicon updates.
Chimera simplifies SoC design, saves power and silicon, and accelerates porting of new AI models.
One core can handle an entire AI/ML workload, including the typical digital signal processing and conditioning workloads often intermixed with inference. Because it’s a unified design, hardware integration is simpler, and profiling and optimizing memory, power, and performance is easier.
Graph code from common training toolsets (TensorFlow, PyTorch, ONNX) is compiled by the Chimera SDK and can be merged with signal processing code written in C++, condensed into a single code stream running on a single processor core. The entire system can be debugged in a single debug console.
Run any AI/ML graph that can be captured in ONNX or written in C++. The chip’s useful life is dramatically extended because new neural network operators and libraries are implemented in code, not silicon.
Accelerator performance with full processor flexibility.
The Quadric Chimera GPNPU was designed from the ground up to address the constantly evolving AI inference deployment challenges facing system-on-chip (SoC) developers, featuring a powerful and elegant architecture with better matrix-computation performance than traditional NPUs.
Matrix, vector, and scalar code in one execution pipeline with no partitioning.
Continuously optimize performance throughout a device's lifecycle with software updates.
Run today's classic backbones, Transformers, and LLMs on architecture ready for networks not yet invented.
Chimera GPNPU Block Diagram
A hybrid Von Neumann + 2D SIMD architecture that unifies matrix, vector, and scalar operations in a single execution pipeline.
The Chimera GPNPU is driven by code instead of fixed-function silicon, empowering developers to continuously optimize performance of models and algorithms throughout the device's lifecycle. And because it's code-driven, it's ready to support novel operators and future networks and models without silicon respins.
Modern system-on-chip architectures deploy complex algorithms that mix traditional C++ based code with constantly emerging and rapidly evolving machine learning inference code. This combination is found in numerous chip subsystems: most prominently in vision and imaging subsystems, radar and LiDAR processing, communications based band subsystems, and many other data-rich processing pipelines.
Unlike alternatives that require splitting AI/ML graph execution across two or three heterogeneous cores, Chimera allows for simple expression of complex parallel workloads by operating as a single software-controlled core.
A hybrid Von Neumann + 2D SIMD architecture optimized for AI/ML inference.
Inference is bottlenecked by data and memory movement optimization, not compute efficiency.
ML/AI inference solutions are most often choked by memory system bandwidth utilization. With most state‑of‑the‑art models having millions or billions of parameters, fitting an entire model into on‑chip memory within an advanced system‑on‑chip is often impossible, necessitating smart management of available on‑chip data storage of weights and activations to achieve high efficiency.
Chimera Graph Compiler (CGC) manages data movement across the memory hierarchy.
Key Insight: Compiler optimizations that keep data resident in the Register File or LRM yield significant power savings.
Many second-generation NPUs are hardwired finite-state machines (FSMs) that handle performance intensive building-block AI operators. They can deliver high efficiency, but only if the network never wavers from the limited scope of operators built into the silicon, and they don't allow for fine-tuning of memory management strategies as network workloads evolve.
Understanding the key differences between traditional NPUs and the Chimera general-purpose NPU.
From single-core QC-Nano to large QC-Multi clusters and fully synthesizable for any process node.
*TOPS range varies by process node (12nm @ 1.0GHz to 3nm @ 1.7GHz) and configuration
The Quadric Chimera™ GPNPU family spans from 1 all the way up to 6,912 TOPS of INT8 ML inference performance. Chimera was designed as IP first and is fully synthesizable, able to be implemented on any process technology from mature 12nm to current 3nm nodes. Scale from ultra-efficient local devices to high-performance automotive and datacenter applications, all with the same core architecture.
Get the datasheet. Talk to our architects. See the benchmarks.