Enter the Arena: How Our Toolchain Started Improving Itself
It started with me prompting an agent to shave cycles off a customer's model, one region at a time. It became an internal arena where engineers and their agents compete to make Chimera kernels faster, an oracle that checks every result bit for bit, and a pipeline that feeds the best ideas back into our compiler.
Written by Shayan Yassami

A lot of my work this year has been making models run faster on Chimera. A customer sends us a network, the Chimera Graph Compiler turns it into C++ that runs on the GPNPU, and then someone goes looking for the cycles that are still left on the table. That last step has always been slow, expert work. A few people would gather at a whiteboard, sketch out ideas, and then coordinate the change across the stack, from the graph compiler to the kernel library to the backend, before anyone could find out whether it helped. Then you profile, form a new theory, rewrite a kernel, check that nothing changed, and do it all again.
At some point this summer I noticed that agents were really good at that loop. The work is specific and unglamorous: read a profile, find the region where the time goes, figure out why a stall is happening, propose a fix, implement it, and prove the output didn't move by a single bit.
What came next was a lot bigger than I'd planned.
The single-player version
It started out very hands-on. There was no written method, so every session began from scratch and I typed the whole thing out again. Here are the regions. Here's where the cycles are going. Look at this stall and work out why it's there before you change anything. Tell me what you're going to do and why the output stays identical. Then do it, measure it, and keep it only if it's faster and still byte-exact.
It worked better than I expected. Benchmarks for a customer I was supporting kept getting faster, one region at a time. And the instructions I kept retyping slowly hardened into something more useful: a set of engineering skills that write down how we actually optimize for Chimera. How to read a profile. How the memory hierarchy behaves. How to write kernels against the Chimera Compute Library. How to prove that a change is correct. Those same skills now ship to customers through our Developer Studio MCP server, so their agents start from the playbook ours learned on.
Two agents in a Slack channel
The turning point came with a real-time speech-enhancement network a customer had asked us to speed up. I'd made good progress on it and sent the whole thing to our CTO: the emitted C++, the weights, everything needed to run it. He pointed his own agent at it.
Before long my agent and his (who goes by Dr. Dre) were trading results back and forth in Slack. His would find something, mine would build on it, and his would build on that. The model kept getting faster, well past the point where either of us would have stopped on our own.
Looking back, we'd built a game without meaning to. Two players were working one problem and keeping score. It just had no board, no referee and no way for anyone else to join.
Multiplayer autoresearch
Andrej Karpathy's autoresearch gives one agent a small LLM training setup and leaves it running: change the code, train for five minutes, keep the change only if the validation score improves, repeat. His README mentions, almost as an aside, that it's obvious "how you'd add more agents to the mix." That was more or less where we'd landed.
The Kernel Arena is multiplayer autoresearch for kernels: cycles in place of a validation score, a lot of people and their agents in place of one, and an oracle that throws out anything that changes a single output bit.
A task in the arena is a single downloadable package: a starting kernel, usually exactly what our toolchain emits, a frozen harness with the inputs and the expected output, and the same engineering skills I'd been building up. One command fetches it. From there you optimize however you like, with whatever agent you like, and upload the result.
Every upload goes to an oracle, a separate machine that rebuilds the submission from source on a pinned version of the toolchain and runs it again on the Chimera instruction-set simulator. The score is cycles. Correctness is a gate: the output has to match the reference byte for byte, with no tolerance and no partial credit. Because the simulator is deterministic, the same source produces the same number, so anyone can reproduce a result on the board.
The part I didn't expect to matter as much is that you start from the current leader's code, not from scratch. Matching the best entry earns nothing; only what you add on top moves the board. Every accepted submission also carries a note explaining why the change is correct, so the reasoning builds up along with the cycle counts.
It also turns out good ideas come from all over the company, not just from the people whose job title says compiler. Hardware engineers have climbed the leaderboard with hand-written assembly. One of our product managers used features from upcoming releases to land optimizations on the board before they'd even reached the SDK. Even our CEO's agent has been joining in on the fun, and it isn't there to make up the numbers. Anyone with an agent and an idea can play, and the oracle doesn't care what team you're on.
The first task I put on the board was that same speech-enhancement network, the version our CTO and I had already worked over and considered pretty much done. The arena has since found another 3.3× on top of it.
Many roads to the same answer
I think this works as well as it does because of what Chimera is. A GPNPU isn't a fixed-function accelerator with a list of supported operations. Every processing element is a full processor, with its own multiply-accumulate units, a 32-bit scalar ALU, local memory and links to its neighbors, and every byte of data movement is under software control. There's no cache making decisions for you.
That means there are many correct ways to implement the same computation. A fixed-function block usually has one way of doing a given operation. On Chimera, any program that produces the right bytes is a valid answer, and some of those programs are far faster than others. That's a large search space, and searching large spaces is what agents turn out to be good at.
Some of what they found, we hadn't seen before. One task is a small piece of BEVFormer, the bird's-eye-view perception transformer: it takes 900 reference points and broadcasts them out into the wider layout the decoder's cross-attention expects. It's exactly the kind of glue operation a compiler handles generically.
The agents stopped treating it as a computation. They programmed the DMA engines directly, writing the transfer descriptors by hand instead of going through the library's general setup path. They picked the strides themselves, including zero strides that make the engine repeat the same data as it moves it, and they staged the loads, gathers and stores so each one overlaps the next instead of waiting on it. The data is never really broadcast in the usual sense. The memory system lays it out in the right shape on the way in. None of our generated code did this, and none of us would have thought to write it. It was our Move 37, and it's now one of the largest improvements on the board.
You can see what the arena does to a task over time. This is the history of an LLM prefill task, causal attention at the Qwen3-8B shape. The task set a target of 70% MAC utilization, shown as the dotted line. Every point is a verified submission that beat the best result before it. The board went straight through the target. The best entry now runs at about 84%, and we're still going.

Each step down is somebody's idea, checked by the oracle and built on the step before it.
The arena turned on itself
Kernels were the obvious first thing to put in the arena, but they weren't the last. What makes something a good task is simple: a score a machine can compute the same way twice, and a strict definition of correct. A lot of our own infrastructure meets that bar.
The example I like most is the simulator. Every performance number we publish comes from the Chimera instruction-set simulator, and every arena score is measured on it. So we made the simulator a task. The goal is to make it run faster while reporting exactly what it reported before: every cycle count, every stall category and every output byte has to be identical on a set of real networks. Agents have made it more than five times faster.
Then we closed the loop. When a faster simulator build passes that gate, the arena's own oracle starts running on it, but only after checking that it reproduces existing results exactly. The arena made the tool that scores the arena faster, so it verifies more submissions in a day, so the board moves faster.
We've started applying the same idea to numerical accuracy too. One task scores quantization quality rather than speed: how closely a language model's output logits on the device track a floating-point reference. The number being optimized is different, but the setup is the same.
That points at something bigger. We get to choose the score. Cycles are where we started, and logit error is the second. It could just as easily be power, memory footprint, or latency at a particular batch size. Whatever matters most to a given customer, we can turn into a task, and the agents will go after it.
From the board back into the toolchain
A faster kernel on the board is great, but by itself it only helps that one model. What we really want is for the idea behind it to help every model a customer compiles.
That's the job of the harvester. The leaderboard measures how many cycles a submission saved. The harvester reads the same submission and asks what its author learned about generating code for Chimera. Submissions that rely on the same mechanism are grouped into a single finding, kept next to the diffs that prove it, and sent to whichever part of the toolchain owns the fix: the Chimera Graph Compiler, the Chimera Compute Library headers that kernels are built on, or the LLVM backend that turns all of it into machine code.
A finding is evidence, not a work order. The engineers who own that code review it, generalize it, measure it across many networks and decide what to ship. The first of these have already landed in the graph compiler and the compute library, so the next model anyone compiles starts from a faster baseline before the arena ever sees it.

A toolchain that improves itself
Put those pieces together and you get a loop. Agents find wins. The harvester turns the general ones into compiler and library changes. Those changes give every new task a better starting point. A faster simulator makes each turn of the loop quicker. Every part feeds the others.
That's what I'm most excited about. A programmable processor gives agents enough room to find new ideas, a deterministic simulator lets us trust what they find, and the harvester makes sure those ideas don't stay stuck on a leaderboard. Some of the improvements in our toolchain now come from the toolchain's own arena.
Anything we can score the same way twice can become a task, and we haven't come close to running out of those.