WRITING / 1
Autoresearch on Kimi Delta Attention
What a closed-loop AI optimization process looks like in practice.
CODE: kda-gb10 · RAW CODE: kda-autoresearch repo
EXACT TRAINING THROUGHPUT
From 833 to 44,942 tokens/s
Every dot is a measured trainer candidate and every run appears in the log. Six brighter milestones trace the retained story; select any point to jump to its entry. Historical runs span several matched measurement blocks. The final point is a fresh three-run confirmation of the retained exact source.
A one-day experiment
In 24 hours, an autoresearch loop improved the training throughput of a six-layer Kimi Delta Attention (KDA) development model on an NVIDIA GB10 from 833 to 44,942 tokens per second. The final exact implementation exceeded the 43,937 tokens per second reached by the Flash Linear Attention (FLA) library under the same model, workload, and device configuration.
This result is deliberately narrow. The model architecture, sequence length, precision, optimizer, data ordering, and KDA math were held constant. The agent was only allowed to change how that computation was implemented on the hardware.
What interested me was not only the final number, but how the system got there.
Why optimize KDA?
I first became interested in KDA when the Kimi Linear paper was released. A personal architecture project I was working on explored interleaving linear attention with sliding-window attention, and KDA looked like a natural candidate. The project was written in JAX, however, so the PyTorch-oriented kernels released through FLA were not something I could simply drop in.
That inconvenience became the experiment: if I fixed one hardware target, one workload, and strict numerical constraints, how far could an autoresearch loop push a correct but impractical implementation in one day?
THE MECHANISM
KDA in one minute
Standard attention keeps every earlier key and value available. KDA instead compresses the past into a fixed-size matrix, a form of working memory. Each token reads what that memory currently predicts and writes only the correction. A learned gate lets different memory channels forget at different rates. The result is a memory that continually edits itself as the model reads.
INCOMING TOKEN
Turn the token into an address and some content.
MEMORY OPERATION
The state grows no larger when the sequence does.
Unlike standard attention, KDA does not retain every earlier key and value. It continually updates one compact working memory.
THE SYSTEMS PROBLEM Each state appears to depend on the one before it, the opposite of the wide parallel work GPUs prefer.
For the full derivation: DeltaNet I: the model, DeltaNet II: the algorithm, and the Kimi Linear paper.
The recurrence is easy to express in PyTorch, but efficient training requires parallelizing it across tokens without dropping any part of the forward computation or backward pass. The gap between a correct implementation and a hardware-efficient one was the target of this experiment.
The algorithmic escape hatch is already known. DeltaNet’s chunkwise formulation groups tokens into blocks: memory still moves sequentially between blocks, while most of the work inside each block becomes parallel matrix multiplication. The Kimi Linear paper derives the corresponding chunkwise form for KDA.
This project does not invent that algorithm. Kimi Linear supplied the mechanism and formulation, while FLA supplied the performance target. The experiment was whether an agent could turn those ideas into an exact, hardware-specific training path for the GB10 and optimize it far enough to compete.
Designing the autoresearch loop
Autoresearch was not a single prompt. It was a closed experimental system: the model proposed and implemented changes, frozen gates determined what counted, and a durable record carried evidence from one attempt to the next.
The agent owned the fast inner loop. I owned the objective and periodically reopened the global profile when a locally productive search stopped attacking the largest remaining bottleneck.
THE RESEARCH SYSTEM
One experiment at a time
Codex + Prime-Agent · GPT-5.6-Terra / high
Read the global profile
Start from measured end-to-end cost and choose one bottleneck large enough to matter.
SCROLL OR SELECT A STAGE
FROZEN OUTSIDE THE LOOP Exact KDA · full gradients · matched workload · independent oracle
I ran the model through a mix of the Codex and Prime-Agent harnesses, often using /goal to keep the longer objective explicit. The exactness gates, matched benchmark, and append-only ledger made the work recoverable and auditable. FLA remained a performance comparator; an independent PyTorch path was the numerical reference.
The agent was extremely strong at executing and evaluating a concrete experiment. My highest-leverage role was noticing when a locally productive search had stopped serving the global objective, then forcing the loop back to the end-to-end profile and changing course.
THE OPTIMIZATION TRACE
Reading the campaign as logs
The final implementation was not the result of one large breakthrough. It emerged from hundreds of profile, correctness, and benchmark events: some cumulative, many negative, and a few that changed the direction of the search.
The logs below condense the record. Attempt identifiers and measurements come from the campaign ledger; the narration is shortened for readability. Rejected and out-of-contract results remain visible because they were part of deciding what the headline could honestly mean.
- $kda-autoresearch --device GB10 --budget 24h
- [CONTRACT]exact KDA · full gradients · matched six-layer trainer
- [TRACE]366 ledger events · 86 measured trainer candidates
- -- PHASE 1 / ESTABLISH THE BASELINE --
-
[BASELINE]Initial eager implementation[PROFILE] Python recurrence: 38.2 seconds per training update [REFERENCE] Full KDA forward and backward in eager PyTorch [BENCH] 833 TOKENS/S
- [AGENT]continued improving the Python path; a CUDA rewrite remained deferred
- [HUMAN REDIRECT]Stop extending the Python path. Own the complete CUDA training implementation.
- -- PHASE 2 / MOVE THE PATH INTO CUDA --
-
[MILESTONE 19]First practical project CUDA baseline[BOTTLENECK] Python dispatch and serial history dominated the update [CHANGE] Complete project-owned CUDA forward and backward [VERIFY] Independent gradients · runtime ownership · no fallback [BENCH] 833 → 7,394 TOKENS/S
- -- PHASE 3 / PARALLELIZE THE RECURRENCE --
-
[MILESTONE 65]Tensor-core recurrence[BOTTLENECK] Recurrent scans and pair transforms remained on the critical path [CHANGE] Chunkwise transforms plus a unified forward/backward WMMA scan [VERIFY] Random-upstream gradients · production-shape gate [RUNNING BEST] 7,394 → 28,325 TOKENS/S
- -- PHASE 4 / FLATTEN THE BACKWARD PASS --
-
[MILESTONE 168]Flattened parallel backward[BOTTLENECK] Ten independent triangular pair families launched separately [CHANGE] Flatten pair ownership into one broader CUDA grid [PROFILE] Operator launch count: 185 → 167 [RUNNING BEST] 36,185 TOKENS/S
- -- PHASE 5 / SHRINK THE DATAFLOW --
-
[MILESTONE 266]Compact BF16 dataflow[BOTTLENECK] Large FP32 U/W surfaces were written, packed, and reread [CHANGE] Publish U to compact scratch and W directly to retained storage [VERIFY] Seven random-dO gradients bitwise equal · four sanitizers clean [DIRECT LEVEL 2] 40,105 → 42,237 TOKENS/S
- -- PHASE 6 / AUDIT THE LAST MILE --
- [OUT OF CONTRACT]Local-path surrogate exceeded 45.5K by truncating KDA backpropagation
- [OUT OF CONTRACT]Value-only surrogate reached 50,541.5 by omitting trainable gradient paths
- [HUMAN AUDIT]A faster number does not count when the computation has changed.
-
[RELEASE]Exact matched confirmation[PROJECT RUNS] 45,058 · 44,942 · 44,842 TOKENS/S [FLA RUNS] 43,958 · 43,898 · 43,937 TOKENS/S [CONFIRMED MEDIANS] 44,942 vs. 43,937 · +2.287% [STRONGEST OBSERVED RUN] 45,058 TOKENS/S
- [COMPLETE]exact source retained · evidence saved · claim bounded to this workload and GB10
FIXED-TIME TRAINING
What twenty minutes buys
The seven-step benchmark isolates steady-state throughput, but I also wanted a more tangible test: given the same twenty minutes, how many complete training updates would each implementation actually finish?
I ran three fresh copies of the same six-layer training job, changing only the KDA backend. Each run used two unscored warm-up updates followed by twenty minutes of measured training. Eager PyTorch completed 12 updates, FLA completed 1,617, and the project CUDA backend completed 1,641. That is 24 more updates than FLA, a 1.48% lead.
MEASURED TRAINING TIME →
The project run also ended with the lowest displayed loss. I treat that only as a consistency check, not as a quality result: this was one short seed, the project completed 24 additional updates, and the remaining difference is not separable from ordinary BF16 and reduction-order effects without repeated runs. The defensible win is simpler: it completed the most exact training work under the fixed time budget. Neither optimized run fell back to another backend. Spot checks showed no active NVIDIA slowdown flags, although this direct runner did not capture continuous thermal telemetry.
What the experiment showed
This experiment is not evidence that an autoresearch loop can replace a library like FLA, CUDA experts, or diligent systems work. The result is purposefully narrow: one algorithm, one chip, and one pinned comparison. On that setup, after 24 hours of optimization, the project implementation reached a confirmed median of 44,942 tokens per second, 2.287% above FLA. It does raise a more interesting question: how should future optimization systems divide the work between human judgment and agent execution?
My kernel-development background is limited. I took ECE 408: Applied Parallel Programming at UIUC and have worked through some GPU MODE lectures, which gave me enough context to understand the mechanism and audit a technical discussion. I could not have written this implementation from memory. Doing it conventionally would have taken me weeks, assuming I got there at all.
That does not mean the experiment was hands-off. The agent was highly effective when given a concrete bottleneck and a reliable way to test its work. It was less reliable at deciding when a productive local search had stopped addressing the global goal. My role was to audit its conclusions, reopen the end-to-end profile, and redirect the search when it became attached to increasingly small improvements.
The final hours made that division of labor especially clear. Several surrogate variants produced numbers above the exact implementation, and one crossed 50,000 tokens per second by omitting trainable gradient paths. Those results were interesting as diagnostics, but they were not KDA training. Without a frozen oracle, full-gradient checks, and a benchmark that failed closed, it would have been easy to publish the larger, incorrect number.
This also changed how I think about autograd. I do not expect it to disappear from research code: it remains a remarkably useful way to express and revise a model. What changes is the cost of leaving it. Once an architecture and workload stabilize, an agent can help migrate the hot path into lower-level primitives much sooner than I would have attempted by hand. The difficult skill shifts from personally writing every kernel to specifying the computation, building tests that can falsify a candidate, and recognizing when the search is optimizing the wrong thing.
The most useful result, then, is not simply that the final kernel edged past FLA. It is that the distance between understanding a mechanism and obtaining a competitive, hardware-specific training implementation became much shorter.