VEER PAREEK

WRITING / 1

Autoresearch on Kimi Delta Attention

What a closed-loop AI optimization process looks like in practice.

CODE: kda-gb10 · RAW CODE: kda-autoresearch repo

EXACT TRAINING THROUGHPUT

From 833 to 44,942 tokens/s

KEPT DISCARDED

Every dot is a measured trainer candidate and every run appears in the log. Six brighter milestones trace the retained story; select any point to jump to its entry. Historical runs span several matched measurement blocks. The final point is a fresh three-run confirmation of the retained exact source.

A one-day experiment

In 24 hours, an autoresearch loop improved the training throughput of a six-layer Kimi Delta Attention (KDA) development model on an NVIDIA GB10 from 833 to 44,942 tokens per second. The final exact implementation exceeded the 43,937 tokens per second reached by the Flash Linear Attention (FLA) library under the same model, workload, and device configuration.

This result is deliberately narrow. The model architecture, sequence length, precision, optimizer, data ordering, and KDA math were held constant. The agent was only allowed to change how that computation was implemented on the hardware.

What interested me was not only the final number, but how the system got there.

Why optimize KDA?

I first became interested in KDA when the Kimi Linear paper was released. A personal architecture project I was working on explored interleaving linear attention with sliding-window attention, and KDA looked like a natural candidate. The project was written in JAX, however, so the PyTorch-oriented kernels released through FLA were not something I could simply drop in.

That inconvenience became the experiment: if I fixed one hardware target, one workload, and strict numerical constraints, how far could an autoresearch loop push a correct but impractical implementation in one day?

THE MECHANISM

KDA in one minute

Standard attention keeps every earlier key and value available. KDA instead compresses the past into a fixed-size matrix, a form of working memory. Each token reads what that memory currently predicts and writes only the correction. A learned gate lets different memory channels forget at different rates. The result is a memory that continually edits itself as the model reads.

INCOMING TOKEN

KEY ALPHA → VALUE 4

Turn the token into an address and some content.

FIXED-SIZE MEMORY WRITE
ADD ASSOCIATION

MEMORY OPERATION

address ALPHA now points toward 4

The state grows no larger when the sequence does.

Unlike standard attention, KDA does not retain every earlier key and value. It continually updates one compact working memory.

THE SYSTEMS PROBLEM Each state appears to depend on the one before it, the opposite of the wide parallel work GPUs prefer.

For the full derivation: DeltaNet I: the model, DeltaNet II: the algorithm, and the Kimi Linear paper.

The recurrence is easy to express in PyTorch, but efficient training requires parallelizing it across tokens without dropping any part of the forward computation or backward pass. The gap between a correct implementation and a hardware-efficient one was the target of this experiment.

The algorithmic escape hatch is already known. DeltaNet’s chunkwise formulation groups tokens into blocks: memory still moves sequentially between blocks, while most of the work inside each block becomes parallel matrix multiplication. The Kimi Linear paper derives the corresponding chunkwise form for KDA.

This project does not invent that algorithm. Kimi Linear supplied the mechanism and formulation, while FLA supplied the performance target. The experiment was whether an agent could turn those ideas into an exact, hardware-specific training path for the GB10 and optimize it far enough to compete.

Designing the autoresearch loop

Autoresearch was not a single prompt. It was a closed experimental system: the model proposed and implemented changes, frozen gates determined what counted, and a durable record carried evidence from one attempt to the next.

The agent owned the fast inner loop. I owned the objective and periodically reopened the global profile when a locally productive search stopped attacking the largest remaining bottleneck.

THE RESEARCH SYSTEM

One experiment at a time

Codex + Prime-Agent · GPT-5.6-Terra / high

AUTONOMOUS ATTEMPT / STEP 1

Read the global profile

Start from measured end-to-end cost and choose one bottleneck large enough to matter.

SCROLL OR SELECT A STAGE

FROZEN OUTSIDE THE LOOP Exact KDA · full gradients · matched workload · independent oracle

I ran the model through a mix of the Codex and Prime-Agent harnesses, often using /goal to keep the longer objective explicit. The exactness gates, matched benchmark, and append-only ledger made the work recoverable and auditable. FLA remained a performance comparator; an independent PyTorch path was the numerical reference.

The agent was extremely strong at executing and evaluating a concrete experiment. My highest-leverage role was noticing when a locally productive search had stopped serving the global objective, then forcing the loop back to the end-to-end profile and changing course.

THE OPTIMIZATION TRACE

Reading the campaign as logs

The final implementation was not the result of one large breakthrough. It emerged from hundreds of profile, correctness, and benchmark events: some cumulative, many negative, and a few that changed the direction of the search.

The logs below condense the record. Attempt identifiers and measurements come from the campaign ledger; the narration is shortened for readability. Rejected and out-of-contract results remain visible because they were part of deciding what the headline could honestly mean.

CONDENSED CAMPAIGN TRACE GB10 / 24 HOURS / EXACT TRAINING
READY
  1. $kda-autoresearch --device GB10 --budget 24h
  2. [CONTRACT]exact KDA · full gradients · matched six-layer trainer
  3. [TRACE]366 ledger events · 86 measured trainer candidates
  4. -- PHASE 1 / ESTABLISH THE BASELINE --
  5. [BASELINE]Initial eager implementation
    [PROFILE] Python recurrence: 38.2 seconds per training update [REFERENCE] Full KDA forward and backward in eager PyTorch [BENCH] 833 TOKENS/S
  6. [AGENT]continued improving the Python path; a CUDA rewrite remained deferred
  7. [HUMAN REDIRECT]Stop extending the Python path. Own the complete CUDA training implementation.
  8. -- PHASE 2 / MOVE THE PATH INTO CUDA --
  9. Backward state-history work was distributed instead of replayed serially.

  10. The corresponding forward history moved into the project-owned parallel path.

  11. The reverse recurrence was split across value tiles to expose more independent GPU work.

  12. [MILESTONE 19]First practical project CUDA baseline
    [BOTTLENECK] Python dispatch and serial history dominated the update [CHANGE] Complete project-owned CUDA forward and backward [VERIFY] Independent gradients · runtime ownership · no fallback [BENCH] 833 → 7,394 TOKENS/S
  13. -- PHASE 3 / PARALLELIZE THE RECURRENCE --
  14. Dependency structure was rewritten so chunk-local work could proceed in parallel.

  15. Independent rows of the pairwise vector-Jacobian product received separate owners.

  16. A partial tensor-core conversion was correct, but its measured gain was too small to advance.

  17. [MILESTONE 65]Tensor-core recurrence
    [BOTTLENECK] Recurrent scans and pair transforms remained on the critical path [CHANGE] Chunkwise transforms plus a unified forward/backward WMMA scan [VERIFY] Random-upstream gradients · production-shape gate [RUNNING BEST] 7,394 → 28,325 TOKENS/S
  18. -- PHASE 4 / FLATTEN THE BACKWARD PASS --
  19. Pair construction and its tensor-core consumers were brought into one retained path.

  20. Lower-precision history reduced storage but introduced enough extra work to regress the trainer.

  21. A guarded fast path accelerated the production shape while preserving the exact generic route.

  22. [MILESTONE 168]Flattened parallel backward
    [BOTTLENECK] Ten independent triangular pair families launched separately [CHANGE] Flatten pair ownership into one broader CUDA grid [PROFILE] Operator launch count: 185 → 167 [RUNNING BEST] 36,185 TOKENS/S
  23. -- PHASE 5 / SHRINK THE DATAFLOW --
  24. Only the group-level state needed by backward remained materialized.

  25. Retained factors were laid out for their consumers without changing the WY computation.

  26. Selected forward products crossed the producer-consumer boundary directly in BF16.

  27. [MILESTONE 266]Compact BF16 dataflow
    [BOTTLENECK] Large FP32 U/W surfaces were written, packed, and reread [CHANGE] Publish U to compact scratch and W directly to retained storage [VERIFY] Seven random-dO gradients bitwise equal · four sanitizers clean [DIRECT LEVEL 2] 40,105 → 42,237 TOKENS/S
  28. -- PHASE 6 / AUDIT THE LAST MILE --
  29. A locally positive graph result did not establish a durable retained improvement.

  30. Composing individually plausible captures produced a severe end-to-end regression.

  31. Producer layouts were aligned with the backward consumers that reused them.

  32. An exact vertical fusion removed a layer boundary and passed a three-by-three matched trainer comparison.

  33. [OUT OF CONTRACT]Local-path surrogate exceeded 45.5K by truncating KDA backpropagation
  34. [OUT OF CONTRACT]Value-only surrogate reached 50,541.5 by omitting trainable gradient paths
  35. [HUMAN AUDIT]A faster number does not count when the computation has changed.
  36. Bitwise exact, but the apparent microbenchmark gain disappeared in the matched trainer.

  37. Fewer launches lost to idle cluster residency and barriers; the exact layer became slower.

  38. [RELEASE]Exact matched confirmation
    [PROJECT RUNS] 45,058 · 44,942 · 44,842 TOKENS/S [FLA RUNS] 43,958 · 43,898 · 43,937 TOKENS/S [CONFIRMED MEDIANS] 44,942 vs. 43,937 · +2.287% [STRONGEST OBSERVED RUN] 45,058 TOKENS/S
  39. [COMPLETE]exact source retained · evidence saved · claim bounded to this workload and GB10

FIXED-TIME TRAINING

What twenty minutes buys

The seven-step benchmark isolates steady-state throughput, but I also wanted a more tangible test: given the same twenty minutes, how many complete training updates would each implementation actually finish?

I ran three fresh copies of the same six-layer training job, changing only the KDA backend. Each run used two unscored warm-up updates followed by twenty minutes of measured training. Eager PyTorch completed 12 updates, FLA completed 1,617, and the project CUDA backend completed 1,641. That is 24 more updates than FLA, a 1.48% lead.

20 MINUTES · SAME MODEL · SAME GB10 SMOOTHED LOSS × MEASURED TRAINING TIME
SHARED SCALES
1 / EAGER PYTORCH 12 updates
321 TOKENS/S
393,216 TOKENSFINAL LOSS 9.092
2 / FLA 1,617 updates
44,143 TOKENS/S
52,985,856 TOKENSFINAL LOSS 3.850
+24 UPDATES · +1.48%
3 / PROJECT CUDA 1,641 updates
44,783 TOKENS/S
53,772,288 TOKENSFINAL LOSS 3.811

MEASURED TRAINING TIME →

The project run also ended with the lowest displayed loss. I treat that only as a consistency check, not as a quality result: this was one short seed, the project completed 24 additional updates, and the remaining difference is not separable from ordinary BF16 and reduction-order effects without repeated runs. The defensible win is simpler: it completed the most exact training work under the fixed time budget. Neither optimized run fell back to another backend. Spot checks showed no active NVIDIA slowdown flags, although this direct runner did not capture continuous thermal telemetry.

What the experiment showed

This experiment is not evidence that an autoresearch loop can replace a library like FLA, CUDA experts, or diligent systems work. The result is purposefully narrow: one algorithm, one chip, and one pinned comparison. On that setup, after 24 hours of optimization, the project implementation reached a confirmed median of 44,942 tokens per second, 2.287% above FLA. It does raise a more interesting question: how should future optimization systems divide the work between human judgment and agent execution?

My kernel-development background is limited. I took ECE 408: Applied Parallel Programming at UIUC and have worked through some GPU MODE lectures, which gave me enough context to understand the mechanism and audit a technical discussion. I could not have written this implementation from memory. Doing it conventionally would have taken me weeks, assuming I got there at all.

That does not mean the experiment was hands-off. The agent was highly effective when given a concrete bottleneck and a reliable way to test its work. It was less reliable at deciding when a productive local search had stopped addressing the global goal. My role was to audit its conclusions, reopen the end-to-end profile, and redirect the search when it became attached to increasingly small improvements.

The final hours made that division of labor especially clear. Several surrogate variants produced numbers above the exact implementation, and one crossed 50,000 tokens per second by omitting trainable gradient paths. Those results were interesting as diagnostics, but they were not KDA training. Without a frozen oracle, full-gradient checks, and a benchmark that failed closed, it would have been easy to publish the larger, incorrect number.

This also changed how I think about autograd. I do not expect it to disappear from research code: it remains a remarkably useful way to express and revise a model. What changes is the cost of leaving it. Once an architecture and workload stabilize, an agent can help migrate the hot path into lower-level primitives much sooner than I would have attempted by hand. The difficult skill shifts from personally writing every kernel to specifying the computation, building tests that can falsify a candidate, and recognizing when the search is optimizing the wrong thing.

The most useful result, then, is not simply that the final kernel edged past FLA. It is that the distance between understanding a mechanism and obtaining a competitive, hardware-specific training implementation became much shorter.