Method and baselines
Each test program is one small function, written by hand in every language and timed on the same inputs. A result is checked against a checksum before it is timed. A time is the median of 7 interleaved runs, taken only while the 1-minute load average was at or below 10 (highest at a sample start: 9.78). Source: results/exec-benchmark-full.json.
Two baselines, because they answer different questions. Same calling convention: A0 and the best of C, Rust and Zig are all called out of line from a C driver. Own inlined driver: each language is timed with its own driver, which its compiler can inline; that is the harder baseline for A0.
Machines. Native speed: one Apple M3 with 8 cores (darwin-arm64), 2026-09-30 and 2026-10-01, Node v24.14.0, Apple clang 21.0.0, rustc 1.96.0, 7 samples per side, medians. Wasm: a different machine (Windows x64, 8 CPUs, 2026-10-07), Node v22.21.1, clang 22.1.8, 15 samples, results/wasm-benchmark-win32.json; its load is estimated from CPU utilisation because Windows has no load average. The two are never combined into one claim.
Reproduce. The exact commands, tools, run times, output file and field to read for each figure, and the experiments that need a model and cannot be re-run without one (their replies are committed), are in scripts/bench-repro/README.md. Every recorded loss is in results/loss-ledger.json.
How A0 was built. By AI agents under a gate, with the owner directing; the repository records who committed and which model co-authored a commit, not which lines a person wrote, and does not track what the work cost. Details: How A0 was built in the README.
Native speed on small test programs
The slowest test program is 4.31x slower on mat4 against baselines with the same calling convention, and 7.58x slower on mat4 against each language's own inlined driver. Same calling convention (A0 and the best of C, Rust and Zig, both called out of line): 2 wins, 9 ties, 8 losses of 19. Inlined drivers: 2 wins, 9 ties, 8 losses. Against hand-written C alone, A0 takes 1.16x as long, a geometric mean over the test programs; 1.00x would be equal.
Every speed result below uses the same test programs, 19 in all. A test program is one small function (benchmark authors call it a kernel), written by hand in every language and timed on the same inputs; each result is checked against a checksum before it is timed. A0 here is machine code from A0's own AArch64 code generator, with no runtime, no garbage collector and no C compiler in between. Measured with a load gate: the 1-minute load average stayed at or below 10 (highest at a sample start: 9.78).
A0 under both baselines, per test program
Same calling convention: 2 wins, 9 ties, 8 losses of 19, slowest 4.31x on mat4. Each language's own inlined driver: 2 wins, 9 ties, 8 losses, slowest 7.58x on mat4.
Two baselines, because they answer different questions. Same calling convention: A0 and the best of C, Rust and Zig are all called out of line from a C driver. Inlined: each language's fastest result with its own driver, which the compiler can inline. Speedup = the baseline's time divided by A0's: below 1.00x A0 is slower. Win = at least 1.2x faster; loss = more than 10% slower; the file gives verdicts only for the first baseline, the second uses the same win rule and a 10% tie band (results/exec-benchmark-full.json).
| Test program | Same calling convention: speedup, verdict | Own inlined driver: speedup, verdict |
|---|---|---|
| affine | 0.63x loss | 0.60x loss |
| rotl | 0.99x tie | 0.99x tie |
| clamp | 0.69x loss | 0.68x loss |
| mix | 0.96x tie | 0.96x tie |
| ident | 1.00x tie | 1.00x tie |
| noop | 1.00x tie | 1.00x tie |
| chain3 | 1.00x tie | 1.00x tie |
| branchy | 0.90x loss | 0.89x loss |
| arrfill | 0.98x tie | 0.99x tie |
| arrfill4k | 1.32x win | 1.33x win |
| loop64 | 0.99x tie | 0.99x tie |
| dot1k | 1.03x tie | 1.03x tie |
| prefix1k | 0.40x loss | 0.40x loss |
| hist256 | 0.67x loss | 0.67x loss |
| mat4 | 0.23x loss | 0.13x loss |
| fnv4k | 1.04x tie | 1.04x tie |
| xs4k | 0.74x loss | 0.73x loss |
| minmax1k | 0.75x loss | 0.75x loss |
| filter2 | 2.02x win | 2.00x win |
Every loss under either baseline is in bold. A0's code is called out of line from a C driver while the other languages' drivers are inlined, so the second baseline is the harder one for A0.
Where A0 places on each test program
First of the languages that ran each program on 4 of 19 test programs (a language is timed only on the programs it has source for, so fewer languages ran the later programs); on every one, A0 takes at most 7.58x as long as the fastest language.
Baseline: each language's own driver, inlined. Place 1 = fastest of the languages that ran that test program (every language ran the first ten; later ones only some). Time per call, median of 7 interleaved runs, every result checksum-verified (results/exec-benchmark-full.json).
measured: 49 of 49 languages
| Test program | A0's place (1 = fastest) | Fastest language other than A0 | A0 takes this many times as long |
|---|---|---|---|
| affine | 21 of 49 | Vala, 4.93 ns | 1.68x (A0 is slower) |
| rotl | 3 of 49 | Rust, 3.25 ns | 1.01x (A0 is slower) |
| clamp | 15 of 49 | Julia, 4.81 ns | 1.46x (A0 is slower) |
| mix | 15 of 49 | Zig, 3.21 ns | 1.04x (A0 is slower) |
| ident | 8 of 49 | D, 1.60 ns | 1.00x (same speed) |
| noop | 2 of 49 | Rust, 1.60 ns | 1.00x (same speed) |
| chain3 | 2 of 49 | C, 3.21 ns | 1.00x (same speed) |
| branchy | 14 of 49 | Odin, 3.21 ns | 1.12x (A0 is slower) |
| arrfill | 11 of 49 | Fortran, 3.21 ns | 1.01x (A0 is slower) |
| arrfill4k | 1 of 5 | C, 233.60 ns | 0.75x (A0 is faster) |
| loop64 | 6 of 49 | Zig, 78.19 ns | 1.01x (A0 is slower) |
| dot1k | 1 of 5 | C, 168.52 ns | 0.97x (A0 is faster) |
| prefix1k | 4 of 5 | C, 328.74 ns | 2.52x (A0 is slower) |
| hist256 | 3 of 5 | C, 1604.77 ns | 1.49x (A0 is slower) |
| mat4 | 4 of 5 | C, 18.30 ns | 7.58x (A0 is slower) |
| fnv4k | 1 of 5 | Rust, 20704.63 ns | 0.96x (A0 is faster) |
| xs4k | 5 of 5 | C, 6295.27 ns | 1.37x (A0 is slower) |
| minmax1k | 3 of 5 | C, 126.99 ns | 1.34x (A0 is slower) |
| filter2 | 1 of 5 | C, 653.13 ns | 0.50x (A0 is faster) |
Reading a row: the third column names the fastest of the other languages and its time per call; the last column is A0's time divided by that time. Below 1.00x, A0 is faster than every other language on that test program; above 1.00x, that language is faster and the number is the gap A0 has to close.
Speed against 48 other languages
8 of the other 48 languages are more than 5% faster than A0 (C, Rust, Crystal, Odin, D, C++, Nim, V); 3 are within 5% of A0; the rest are slower.
The same test programs, hand-written in each language and checksum-verified before they are timed; a language is timed on the programs it has source for. Each bar below is one language: its time per call divided by A0's on those programs, as a geometric mean.
How long each language takes, as a multiple of A0's time
1.00x = the same speed as A0; a longer bar = slower than A0. Bar length is logarithmic. Interleaved runs, medians; JIT rows warm; interpreters at their own iteration tier.
measured: 49 of 49 languages
Show all 49 languages
Shown first: A0 and the best-known languages. All 49 languages are behind the disclosure, in the same order. A0 here is machine code from A0's own AArch64 code generator, with no C compiler in between. Numbers, toolchains and iteration tiers are in results/exec-benchmark.json.
Startup
A0 takes 2.50 ms from launch to first result: place 9 of 48 languages (1 = fastest). Node takes 23 ms and Python 18 ms.
Startup is the time for one process launch to run one iteration of a test program and print its result; compile time is separate (1430 ms for the ten test programs with A0, 9536 ms with rustc).
Time from launch to first result
Milliseconds, shorter bar = faster start. Bar length is logarithmic. JVM and .NET rows include their runtime start.
measured: 48 of 49 languages
Show all 48 languages
Shown first: A0 and the best-known languages; all 48 are behind the disclosure, in the same order. The A0 binary for startup is the C-path build; startup of the direct AArch64 binary is not yet measured.
Tokens
Writing the 10 token test programs, A0 (canonical) takes 559 tokens: place 41 of 49 (1 = fewest), fewer than 8 of the other 48 languages, equal to 0 and more than 40. A0 (dense) takes 202: place 1, fewer than 48, equal to 0, more than 0; its lossless form, with the same ids and node order, takes 286: place 1. The dense totals are over the same 10 programs (results/dense-tokens.json); the speed tests below time more programs (19), but only 10 of them are counted for tokens.
A token is the unit a model reads and writes. Fewer tokens for the same program means less to read, write and pay for. The count below is the source of the same test programs in each language, with the o200k tokenizer. Both forms of A0 are plotted and highlighted. Dense is the same program in a shorter surface form; it converts losslessly to the canonical form. How many tokens a whole edit costs, measured with real models, is in the Cost section.
Source tokens of the 10 token test programs, per language
Sum over the test programs of the o200k tokens of each program as written in that language, shorter bar = fewer tokens. Bar length is proportional to the count.
measured: 49 of 49 languages
Show all 50 rows (49 languages, A0 in two forms)
Program length is one input; what a whole edit costs, measured with models, is in the Cost section. Shown first: A0 and the best-known languages; all 50 rows are behind the disclosure, in the same order. Source: results/lang-axes.json, and the dense totals in results/dense-tokens.json.
Tokens to make one small edit, A0 against other languages
To change one operator in a one-function file, A0 reads 32 tokens and writes a 9-token reply. TypeScript reads 35 and writes 20 as a line edit, 42 as a search-and-replace block or 79 as a unified diff; C reads 29 and writes 64 as a unified diff.
o200k tokens for the same edit, in each language's usual form. Shorter bar = fewer tokens. A0 reads more than C on this tiny file; the scoped view pays off as programs grow (Cost section).
Read: the code in front of the model
Write: the model's reply
Bars are linear within each group. The fixtures are the edit of affine in results/tokens.json; a model's real replies (Cost section) include the instructions it must be given first.
How A0 shrinks an edit: A0 reading a whole file vs A0's scoped view
For one edit in life.a0 (17 functions), A0's scoped view is 7.3x fewer tokens than reading the whole life.a0 file in A0: 239 against 1746. Both bars are A0; this is not a comparison with another language.
Tokens a model reads to edit one function, o200k tokenizer. Shorter bar = fewer tokens.
The view: one function, its callees' signatures, one handle.
Cost
At 4000 functions an A0 edit costs 706 tokens against 143854 for TypeScript, which is 203.8x fewer than TypeScript shown the whole numbered file (A0's scoped view against a whole-file workflow, not both with scoped views). At 1 function A0 costs 2.42x TypeScript's tokens and 2.29x Rust's. Sonnet, cache-adjusted.
Cost here is every token a model reads and writes to make one edit: the instructions it is first given (the primer), the code it reads, and its reply. A model edits one function through a scoped view: the function, the signatures it depends on, and its callers. The view stays the same size as the program grows; a numbered whole file does not, so against a whole-file workflow A0's advantage grows with program size and reverses on a one-function file, where the primer dominates. That is an advantage of the workflow, not of the language: when TypeScript, Rust, Python, Go, Java, C and Ruby are given an equal view (a parse-derived function and its callees, numbered line edits), pooled over the four program sizes (96 trials per cell) A0 canonical costs 318 tokens for one cold task and 176 in an unbounded session, against 386 and 138 for TypeScript and 375 and 115 for Ruby, so it is cheaper cold and dearer once the primer is cached; A0 dense costs 273 and 99. results/ai-edit-scoped.json.
Edit cost across 49 languages, with A0 in two forms
Over a session of 10 edits, A0 (canonical) costs 184 tokens per task: place 12 of 49 (1 = fewest), cheaper than 36 of the other 48 languages, equal to 1, dearer than 11. A0 (dense, lean view with callee bodies) costs 117: place 1, cheaper than 48, equal to 0, dearer than 0.
Dense is the same program in a shorter surface form; it converts losslessly to the canonical form. A0 (dense) is shown with the lean view: only the function to edit, with its direct callees as dense text. Both A0 subjects are measured on the same twelve tasks of set b (one function to edit) as the other languages, one fresh Haiku and one fresh Sonnet subject per task, one shot and one repair. Losses first: the dense primer is larger (145 tokens against 118 for canonical), so in a session of 1 task the two cost 274 and 311; each cell is 24 tasks, one task is 4.2 points of acceptance, and differences of about one task are within noise. This chart is set b only.
Tokens per task, session of 10 tasks (primer paid on the first call at 1.25x and on later calls at 0.05x; view and replies at 1x, as recorded in results/ai-edit-b48-dense.json); shorter bar = fewer tokens. Bar length is proportional.
measured: 49 of 49 languages
Show all 50 rows (49 languages, A0 in two forms)
| Measure | A0 (canonical): place | A0 (canonical): cheaper than, equal to, dearer than | A0 (dense): place | A0 (dense): cheaper than, equal to, dearer than |
|---|---|---|---|---|
| one-shot | 21 of 49 (13 tied) | 15, 13, 20 | 1 of 49 (5 tied) | 43, 5, 0 |
| after repair | 1 of 49 (33 tied) | 15, 33, 0 | 1 of 49 (33 tied) | 15, 33, 0 |
| cost 1 | 1 of 49 (0 tied) | 48, 0, 0 | 1 of 49 (0 tied) | 48, 0, 0 |
| cost 10 | 12 of 49 (1 tied) | 36, 1, 11 | 1 of 49 (0 tied) | 48, 0, 0 |
| cost inf | 24 of 49 (3 tied) | 22, 3, 23 | 1 of 49 (0 tied) | 48, 0, 0 |
| Subject | Accepted first try | Accepted after one repair | Primer | Tokens read | Tokens written | Session of 1 task | Session of 10 tasks | Long session |
|---|---|---|---|---|---|---|---|---|
| A0 (canonical) | 22 of 24 | 24 of 24 | 118 | 115 | 33 | 311 | 184 | 170 |
| A0 (dense) | 24 of 24 | 24 of 24 | 145 | 70 | 23 | 274 | 117 | 100 |
Measures: one-shot = accepted on the first try; after repair = accepted after one repair; cost 1, cost 10 and cost inf = tokens per task over a session of 1 task, of 10 tasks and of unbounded length (the primer's share vanishes). Place 1 = fewest tokens or most accepted; every rank is the subject against the same 48 other languages, and the three numbers beside it count the languages it beats, ties and loses to. Tokens in the second table are averages per task. The other dense variants measured (dense view with program view, lean view, lean view with a shorter primer) are in results/ai-edit-b48-dense.json. Numbers: results/ai-edit-b48-dense.json.
Tokens a model reads and writes per edit, by program size, Sonnet
Cache-adjusted; shorter bar = fewer tokens. The number on the right is TypeScript's tokens divided by A0's: above 1.00x A0 uses fewer tokens than TypeScript, below 1.00x (red) A0 uses more.
measured: 3 of 49 languages
These size charts are A0 (canonical); A0 (dense) was measured at one function only, in the chart above. Bars are linear within each row. The per-size charts below add every other language that has data at that size. Languages with no model-edit data: C, JavaScript, Objective-C, Kotlin, Swift, Zig, Ruby, PHP, Lua, Perl, Tcl, Fortran, F#, Visual Basic, Dart, Scala, Clojure, Groovy, Elixir, Erlang, Haskell, OCaml, Julia, R, Nim, Crystal, D, Pascal, Racket, Common Lisp, COBOL, Prolog, Scheme (Guile), Scheme (CHICKEN), Smalltalk, Forth, Haxe, V, Odin, Vala, Gleam. Each needs a hand-written translation of every edit task, a build-and-run acceptance check and paid model calls; only the languages shown have them.
Edits in a 1-function program: tokens read and written, 8 languages (a loss for A0)
At 1 function, A0 uses 688 tokens per edit: place 8 of 8 languages (1 = fewest). The fewest is TypeScript with 284.
Sonnet, cache-adjusted, shorter bar = fewer tokens; bar length is logarithmic. A language is listed only if it was measured at this size.
measured: 8 of 49 languages
| Language | Primer (instructions) | Code read | Reply written | Total tokens | Total as a multiple of A0's | Sonnet edits accepted | Haiku edits accepted |
|---|---|---|---|---|---|---|---|
| A0 | 550 | 104 | 34 | 688 | 1.00x | 100% | 100% |
| TypeScript | 159 | 98 | 28 | 284 | 0.41x | 100% | 92% |
| Rust | 178 | 97 | 25 | 300 | 0.44x | 100% | 100% |
| Python | 172 | 87 | 25 | 284 | 0.41x | 100% | 25% |
| Go | 179 | 103 | 24 | 306 | 0.44x | 100% | 100% |
| Java | 181 | 108 | 26 | 314 | 0.46x | 100% | 67% |
| C# | 169 | 107 | 23 | 299 | 0.43x | 100% | 75% |
| C++ | 175 | 120 | 25 | 320 | 0.47x | 100% | 92% |
Primer: language and workflow instructions, first read at the 1.25x cache-write rate. Code read: the view or numbered file. Reply written: the model's edit. Total as a multiple of A0's: the row total divided by A0's total; below 1.00x (red) that language needs fewer tokens than A0. Edits accepted: the share of the 12 tasks whose edit passed the tests on the first try, for each model.
Edits in a 40-function program: tokens read and written, 3 languages
At 40 functions, A0 uses 1174 tokens per edit: place 1 of 3 languages (1 = fewest). The fewest is A0 with 1174.
Sonnet, cache-adjusted, shorter bar = fewer tokens; bar length is logarithmic. A language is listed only if it was measured at this size.
measured: 3 of 49 languages
| Language | Primer (instructions) | Code read | Reply written | Total tokens | Total as a multiple of A0's | Sonnet edits accepted | Haiku edits accepted |
|---|---|---|---|---|---|---|---|
| A0 | 550 | 590 | 34 | 1174 | 1.00x | 100% | 92% |
| TypeScript | 159 | 1966 | 28 | 2152 | 1.83x | 100% | 75% |
| Rust | 178 | 2075 | 26 | 2278 | 1.94x | 100% | 92% |
Primer: language and workflow instructions, first read at the 1.25x cache-write rate. Code read: the view or numbered file. Reply written: the model's edit. Total as a multiple of A0's: the row total divided by A0's total; below 1.00x (red) that language needs fewer tokens than A0. Edits accepted: the share of the 12 tasks whose edit passed the tests on the first try, for each model.
Edits in a 400-function program: tokens read and written, 8 languages
At 400 functions, A0 uses 706 tokens per edit: place 1 of 8 languages (1 = fewest). The fewest is A0 with 706.
Sonnet, cache-adjusted, shorter bar = fewer tokens; bar length is logarithmic. A language is listed only if it was measured at this size.
measured: 8 of 49 languages
| Language | Primer (instructions) | Code read | Reply written | Total tokens | Total as a multiple of A0's | Sonnet edits accepted | Haiku edits accepted |
|---|---|---|---|---|---|---|---|
| A0 | 550 | 122 | 34 | 706 | 1.00x | 100% | 83% |
| TypeScript | 159 | 14102 | 28 | 14289 | 20.2x | 100% | 67% |
| Rust | 178 | 14169 | 26 | 14372 | 20.3x | 100% | 58% |
| Python | 172 | 12463 | 25 | 12660 | 17.9x | 100% | 8% |
| Go | 179 | 13164 | 24 | 13367 | 18.9x | 92% | 42% |
| Java | 181 | 13073 | 26 | 13280 | 18.8x | 100% | 67% |
| C# | 169 | 13724 | 24 | 13916 | 19.7x | 100% | 58% |
| C++ | 175 | 12972 | 25 | 13171 | 18.6x | 100% | 75% |
Primer: language and workflow instructions, first read at the 1.25x cache-write rate. Code read: the view or numbered file. Reply written: the model's edit. Total as a multiple of A0's: the row total divided by A0's total; below 1.00x (red) that language needs fewer tokens than A0. Edits accepted: the share of the 12 tasks whose edit passed the tests on the first try, for each model.
Edits in a 4000-function program: tokens read and written, 3 languages
At 4000 functions, A0 uses 706 tokens per edit: place 1 of 3 languages (1 = fewest). The fewest is A0 with 706.
Sonnet, cache-adjusted, shorter bar = fewer tokens; bar length is logarithmic. A language is listed only if it was measured at this size.
measured: 3 of 49 languages
| Language | Primer (instructions) | Code read | Reply written | Total tokens | Total as a multiple of A0's | Sonnet edits accepted | Haiku edits accepted |
|---|---|---|---|---|---|---|---|
| A0 | 550 | 123 | 34 | 706 | 1.00x | 100% | 92% |
| TypeScript | 159 | 143666 | 29 | 143854 | 203x | 100% | 92% |
| Rust | 178 | 143276 | 26 | 143479 | 203x | 100% | 83% |
Primer: language and workflow instructions, first read at the 1.25x cache-write rate. Code read: the view or numbered file. Reply written: the model's edit. Total as a multiple of A0's: the row total divided by A0's total; below 1.00x (red) that language needs fewer tokens than A0. Edits accepted: the share of the 12 tasks whose edit passed the tests on the first try, for each model.
Tasks are the same edits in every language; sets 400 and 4000 are one shared program grown to that size. Subjects are fresh Sonnet and Haiku contexts that see only the primer. In the 400 and 4000 rows the other languages read the whole numbered file and A0 its scoped view (a workflow comparison); the equal-context comparison is results/ai-edit-scoped.json. Numbers: results/ai-edit-experiment.{b,c,c400,c4000}.*.json.
Validation
After an edit, A0's native checker answers in 2.83 ms: place 1 of 34 languages that have a separate check step (1 = fastest). Checking and running the edit takes A0 2.49 ms: place 1 of 49 languages.
Validation is how long it takes to know an edit is right. The check is the build or type-check of the edited test program. Check and run adds compiling to an executable when the language needs it, and one run whose checksum must match. Every language uses its own toolchain on the same test programs, from a cold process.
Time to check one edited test program
A0 (native) checks in 2.83 ms: place 1 of 34 (1 = fastest). Fastest: A0, 2.83 ms. Median language: 281.42 ms. A0 (Node CLI), the previous path, took 125.23 ms.
Median milliseconds over the test programs, shorter bar = faster. Bar length is logarithmic.
measured: 34 of 49 languages
Show all 34 languages
Languages with no separate check step (they run the file directly) are only in the next chart. A0 here is the native self-hosted checker, run as a cold process with no Node (A0 (Node CLI) is the previous path). Shown first: A0 and the best-known languages; all 34 are behind the disclosure, in the same order. results/lang-axes.json.
Model edits: time from a reply to a type-checked program
On the model edits below, A0 applies the edit and type-checks the whole program in 0.51 ms at the median; TypeScript with a warm compiler takes 15.1 ms.
Milliseconds from reply received to program type-checked, median per edit, over the accepted set-C Sonnet edits on the 40-function program; shorter bar = faster; bar length is logarithmic. Go uses hand translations of the reference edits.
measured: 4 of 49 languages
Measured on a loaded machine (load average up to 97 on 8 cores); absolute times will move on a quiet run, which will replace these. Cold rows start the compiler per edit; warm rows reuse a running one. Only these languages have model-written edits to replay, because the model-edit experiment exists only for them (Cost section); every language is checked on a fixed edit in the two charts above. results/edit-loop.json.
Parallel folds
A0 --parallel is the fastest implementation on 2 of 7 test programs; on the others, at least one other implementation is faster (listed under the chart).
A fold whose step is an associative reduction can be split across cores without changing its result. a0 emit c --parallel does that from a cost model; every result is checked exact against serial A0 and the reference interpreter. Below, each test program is run by A0 --parallel and by hand-parallel versions in other languages.
Each implementation's time as a multiple of A0 --parallel's, per test program
1.00x = the same speed as A0 --parallel; above 1.00x A0 --parallel is faster; below 1.00x (red) the other implementation is faster. Bar length is logarithmic. 8 cores, median of 7 interleaved samples.
measured: 8 of 49 languages
sum24: 16777216 iterations, A0 2.59 ms
xor24: 16777216 iterations, A0 3.78 ms
count20: 1048576 iterations, A0 0.31 ms
mr22: 4194304 iterations, A0 3.06 ms
dot64k: 65536 iterations, A0 0.05 ms
max64k: 65536 iterations, A0 0.03 ms
hash64k: 65536 iterations, A0 6.20 ms
Measured on a loaded machine (load average up to 212 on 8 cores when timing started); ratios are interleaved, absolute times will move on a quiet run. Hand-parallel C is OpenMP parallel-for with a reduction. Where another implementation is faster, its time as a multiple of A0 --parallel's is in parentheses: xor24 (C + OpenMP 0.78x); mr22 (C + OpenMP 0.50x); mr22 (Rust 0.92x); mr22 (C 0.95x); dot64k (C 0.84x); dot64k (Rust 0.90x); max64k (Rust 0.78x); max64k (C 0.80x); hash64k (C + OpenMP 0.95x). On the 64K-element kernels the cost model keeps A0 serial. results/parallel.json.
Hardware
The 48 corpus functions compile to 49 synthesized modules (the shared divider counts as one): 55530 generic cells in all, 9325 in the largest, simulated on 5262 oracle cases.
A0 compiles the same programs to clocked SystemVerilog. A cell is one gate-level element after Yosys generic synthesis, so fewer cells means a smaller circuit; cell counts are relative size, not area on a real chip. Every number here is read from results/hardware.json.
synthesized modules
generic cells in all modules
cells in the largest module
oracle cases simulated
Of the 36 modules simulated, 13 are clocked: they take a median of 66.20 clock cycles per case, from 34.00 to 554.47. The rest are combinational and answer within the same cycle.
A cycle count is how many clock ticks a module needs from receiving its inputs (start) to raising its done signal, averaged over the simulated oracle cases. Division, remainder and loops make a module clocked, because they run over several cycles.
fewest mean cycles, any clocked module
median mean cycles across clocked modules
most mean cycles, any clocked module
Targets
One source, checked against one oracle: 10 paths (the interpreter, the optimizer, JavaScript, native C, C++, WebAssembly and the JVM) ran all 5262 generated cases; the native assembly targets ran the 4297 cases that need no io (5 targets); SystemVerilog is simulated separately on 5262 cases and is unverified in this run (1 target).
AArch64
A0's own code generator, no C in between
x86-64
A0's own code generator, verified under Rosetta
RISC-V, ARM32, AVR
A0's own code generators for 64-bit RISC-V, 32-bit ARM, and 8-bit AVR
wasm32
A0's own wasm backend, or the C path through Clang: this site
Native C
through clang or gcc, UBSan-clean, parity with hand-written C; --parallel for threads
JavaScript
typed arrays, in-place updates, boundary guards only
JVM
Java source, compiled and verified with javac
.NET
C# source, verified on .NET 10
GPU
Metal Shading Language kernels, verified on Apple silicon
FPGA / ASIC
clocked SystemVerilog, simulated and synthesized
Where A0 loses
Speed: the AArch64 backend is up to 4.31x behind the best of C, Rust and Zig (the mat4 kernel), and up to 7.58x behind each language's own inlined driver (results/exec-benchmark-full.json).
Wasm against clang on Windows x64: 8 losses on run time and 14 on load time, the load-time ratio moving by up to 0.27 between two runs (results/wasm-benchmark-win32.json, results/loss-ledger.json).
Emitted JavaScript: one recorded loss against hand-written JavaScript (the noop kernel); a quiet Windows rerun tied, and the loss stays in the ledger (results/exec-benchmark-noop-win32.json).
Tokens: A0 costs more than TypeScript on single-function tasks (377 of the 586 recorded losses are kernel token counts). Counts use OpenAI's o200k_base tokenizer, not a Claude tokenizer (results/lang-axes.json, results/loss-ledger.json).
Editing a real front end, Haiku had more edits accepted in TypeScript than in A0 (13 against 10; Sonnet ties at 14) (results/app-edit-keys.json).
Evidence: each edit cell uses fresh model sessions, one subject per cell and 24 trials per cell (STATUS.md, Known limits); the native timings come from one machine, and the compiler versions of the other languages are not recorded in the results file.
The language has no floating point, no heap, no recursion; no GPU or x86-64 speed claim is made (STATUS.md, Known limits).
Every loss, in full: results/loss-ledger.json. It can only shrink.