Chimera delivers on-device LLM inference with eight-core QC-Multi implementations or multi-clusters and chiplets enabling up to 64 cores. Licensed by customers building LLM-capable chips today, with up-to-date benchmarks available in our DevStudio.
0B
QC-Multi
0+
Tokens/sec
2 - 8
Cores per QC-Multi
0B
Parameters via Multi-Cluster
Watch QWEN 0.5B inference on a working FPGA demo of Chimera at 50MHz, running in real time.
Get a personalized walkthrough of Chimera's LLM capabilities
The Chimera GPNPU has been licensed by customers for LLM workloads, validated on customer emulation platforms and proven production ready.
Licensing
License the Chimera GPNPU IP for an LLM-capable chip design.
Development
Easily port your LLM using the Chimera SDK with attention, pre-fill, and KV cache building blocks.
Validation
The full model is validated on your emulation platform with correct outputs.
The RTL has been validated on real customer emulation environments and de‑risked for your design.
The Chimera Graph Compiler handles attention layers, position encoding, and quantization. New LLM models compile from ONNX to optimized C++ in just weeks.
When the next breakthrough LLM is released, your chip stays competitive because the model is ported in software, with no silicon respin required.
Chimera scales seamlessly from single-core edge deployments to multi-core, multi-chiplet configurations that widen the external bus width, increase maximum external memory capacity and bandwidth, and enable LLMs with up to 30B parameters.

















































Performance validated on emulation with INT4 quantization. Results shown for various LLM sizes on multi‑core configurations.
0.5B parameters · for example, QWEN 2.5
0.6B parameters · for example, QWEN 3
1.7B parameters · for example, QWEN 3
4B parameters · for example, QWEN 3
8B parameters · for example, QWEN 3, LLaMA
Smaller models running on single-core Chimera. 8B model on 4-core QC-Multi cluster at 1 GHz.
Validated Performance
8B parameter LLM running on 4-core QC-Multi cluster with INT4 weights, validated on customer emulation platform.
~500 ms
Time to first token (512 ctx)
4 GB
INT4 weight footprint
W8A8
8-bit weights, 8-bit activations
W4A8
4-bit weights, 8-bit activations
INT4
Full INT4 computation
Memory‑Bound Optimization
LLM autoregressive inference is memory‑bandwidth bound. Chimera's architecture maximizes bandwidth utilization with optimized weight streaming and KV cache management.
Chimera's programmable architecture supports any transformer-based LLM. When a new model is released, easily port it in software without having to respin silicon.
Architecture Support
Chimera supports standard transformer architectures used by modern LLMs. Port any model from ONNX to optimized C++.
Alibaba
QWEN 2.5 and QWEN 3 models validated on customer emulation, with an 8B model running at 20+ tokens/second.
Meta
Meta's foundational LLM family with multi-core implementation available.
Custom
Bring your own model: Chimera Graph Compiler converts ONNX to optimized C++ automatically.
Core architectural features supported for modern LLMs
GQA support for efficient key/value sharing across query heads
Rotary position embeddings via optimized custom operations
Efficient autoregressive decoding with optimized cache management
Support for 150K+ token vocabularies with efficient gather operations
New models in weeks: Chimera Graph Compiler handles transformer architectures automatically. Port from ONNX without RTL changes.
Whether you're designing edge AI devices, automotive systems, or consumer electronics, Chimera GPNPU delivers the on-device LLM inference your customers demand.
Get detailed documentation, discuss your use case, and learn how Chimera can accelerate your LLM‑enabled product roadmap.
Request More InformationSign in to your existing account or create a new one to explore Chimera SDK and run models in our cloud development environment.
Launch DevStudio