Chimera scales from a single core at 1 TOPS up to a 64-core multi-cluster solution capable of up to 6,912 TOPS.
Maximize performance by applying spatial, data, pipeline, or task parallelism across cores and workloads.
High resolution Vision in real time. Tile 4K images across cores for low-latency inference at full resolution.
Chiplet and multi‑chip ready. Easily bridge clusters across dies while using the same compiler‑managed toolchain.
Scale array size, add cores, bridge chiplets, all with the same software stack.
Up to 108 TOPS per core
Up to 8 cores, up to 864 TOPS
Up to 64 cores, up to 6,912 TOPS
Chimera GPNPU supports spatial, data, pipeline, and task parallelism: you choose the pattern that fits.
Large input is partitioned into tiles, with adjacent cores exchanging edge data via MLS. Processing is done in parallel.
Run the same model on each core, processing different batches, with perfect scaling.
Model layers are partitioned across cores as pipeline stages, maximizing weight bandwidth.
Each core runs a separate model or workload independently with no synchronization overhead.
Homogeneous clusters of 2, 4, or 8 cores with direct L2↔L2 sharing.
QC-Multi
Key: MLS enables direct L2↔L2 access between cores; AXI Coalescer optimizes external memory bandwidth.
For Processor Architects
For Software Architects
Scale Chimera QC-Multi GPNPU clusters through your Network-on-Chip.
Multi-Cluster Architecture
Component Overview
Per Cluster
TinyML / IoT
1× QC-Nano
Up to7TOPS
High-Volume Vision
2× QC-Perform
Up to28TOPS
Edge LLM
4× QC-Perform
Up to112TOPS
ADAS L2+
8× QC-Ultra (1 chiplet)
Up to864TOPS
Autonomous / L4+
64× QC-Ultra (8 chiplets)
6,912TOPS
8-Chiplet System Architecture
Chimera scales from 1 TOPS edge devices up to 6,912 TOPS autonomous systems, all on the same architecture and software.