The ET-SoC-1 energy manual
A workload's energy on this card is the time it takes, times what the card draws doing nothing at the temperature the workload holds it at, plus a sum over everything it does of a per-event cost. This page is the table of per-event costs, with the two things it needs alongside them: the idle power at a given temperature (section 1) and the cost of an awake core (section 2).
At rest the card draws 35.9 W at 80 °C and 0.65 W more per degree (aifoundry2), 20–29 W of it leakage. An awake minion is 2 mW. On zeros an integer add is 6 pJ, a float add 24 and an eight-lane multiply-add 27; on random data 10, 26 and 56. A byte from the shire's own scratchpad is 2–9 pJ; from DRAM, 95–140; written back through the L1 to DRAM, 250–340.
Checked on three cards (26 September 2026): this page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 (both firmware 1.3.1; aifoundry3 held at 600 MHz) and aifoundry1 card 1 (firmware 1.2.0). Of 52 claims tested here, this page counts 25 held, 15 corrected, 10 differ by card and 2 not confirmed; the hub’s scoreboard, 19 “proven on the cards”, 16 “a test behind it failed”, 15 “differs by card” and 2 “fewer than three repeats”. Several that held have sharper numbers; the corrections are the cards' idle against the law, the card-to-card scale and what explains it, the DRAM write against the read and the unsensed blocks' temperature slope; where the cards differ, the text gives each card's value. The catalogue, the rings, the levels, the relay and the tensor unit were measured again on all three cards, and every table and chart here now pools those runs; the 18–23 September sessions they replace stay only where the text compares with them. Record: docs/reports/data/2026-09-25-claims-v3/. The same three cards then measured the gathers and scatters (E48, section 4.4), each tested item decided card by card.
What the range in brackets means
Brackets are the range over every pass on all three cards. A figure with no card named held on every card (three passes or more on each, or the same value on each); where one rests on one card, on fewer runs, or differs between the cards, the text says so. Section 9 says how each bar was made.
Terms used on this page
Minions are the chip's small in-order RISC-V cores, each with two hardware threads (harts), an 8-lane vector unit and a tensor unit whose TensorFMA instruction multiplies 16×16×16 fp32 tiles; eight minions form a neighbourhood, 32 a shire, and 32 shires (1,024 minions) run kernels. Each shire's 4 MB of SRAM is split into a 512 KB L2, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad that any shire can address over the mesh, and eight memory shires drive the 32 GB of DRAM. The PMIC is the board's power controller: it meters board power and three regulated rails (the minions, the SRAM and the mesh), which the service processor (SP), the on-die management core, reports. aifoundry2 and aifoundry3 (a2, a3) are lab machines with one card each; aifoundry1 holds two, and its card 1 (aifoundry1-c1, a1c1, on older firmware) is the third card here (its card 0 overheats and is left out); more in the hub's glossary.
1. The card at rest
How the fit was made, and idle power by rail
1.1 The SRAM arrays at rest
Fit details
2. A core that is awake
How to read "mean [lo–hi]" and the per-card column
Bars from here on: mean [lo–hi] is the mean over every pass on every card and the range those passes spanned; the per-card column is each card's own mean ± its pass-to-pass standard error.
Why a nop or a fence isn't the floor
3. Instructions
Energy per instruction retired, above idle, both harts of all 1,024 minions running it flat out, at 600 MHz and the minion rail's 0.5 V. Three operand sets: zeros, one constant everywhere, random values in [0.5, 2).
How the chart was measured
The 13 as a table: constant operand, per lane, issue rate
3.1 Every instruction the core executes
The full catalogue: every instruction the assembler accepts and the silicon executes in U-mode (user mode, where kernels run), measured three times in shuffled order on each of three cards, on zeros and on random operands. Each dot is one instruction on random data, one row per class; focus or hover a dot for its zeros figure, its bar, its issue rate and the other cards, or find one by name. Every figure is also in the table below the chart, by class. Thirteen instructions trapped in U-mode in a one-off check while the catalogue was written, and so have no energy (the card and the log of that check were not kept): .
How the chart was measured
The table, by class
Gathers, scatters and packed atomics from the L1, as a table (E48)
3.2 The tensor unit, per multiply-add
The nine rows as a table: marginal and loaded energy, watts and rates
What "marginal" and "loaded" mean, and the bars
The tensor unit is the one place on the chip resolved below the instruction, by simulating its RTL on the operands aifoundry2 ran (The Horace experiment): four kinds of event, four energies fitted on that card — — and 1.8 mW per minion of state machines whatever the data. A random-data tile priced this way comes to 27.2 W on 1,024 minions; aifoundry2 measured 27.6 in the 21 September runs the energies were fitted to. Fourteen structured matrices were priced before they ran on the same card, to 0.9 W rms (one session, two runs each).
4. Bytes through the memory hierarchy
Where should a workload's bytes live? The map puts every path of sections 4 and 5 at the energy per byte it costs and the aggregate bandwidth the card reached on it; the diagonals are constant power over idle.
What the map's numbers are
4.1 Reads and writes, measured together
What each row measures
Reads by level, L1 to DRAM: the version-3 passes on every card
4.2 Reads by level
4.3 Finer grain: wires, lines and rows
What a byte costs is not one number. These are the parts of it the card's instruments can separate.
How the wire chart was measured
Line fills, row patterns and neighbourhood reads, as tables
Gathers by pattern, mask and element size, and one minion against the chip, as a table (E48)
4.4 Irregular access: gathers and scatters by level
How the chart was measured
The table, level by level, with each card's values
4.5 Where the current flows, by meter
How the split was measured
5. Bytes between cores and shires
Register file to register file over the tensor network (TensorSend and TensorRecv), hart 0 of every minion sending and receiving around rings. Shire IDs do not follow the mesh, so each ring between shires is labelled by how many IDs apart its shires are and by how many mesh hops that is on average, on the shire map of On-chip communication (checked against latency on all three cards: the version-3 check found the same map from the round-trip times in every pass).
How the rings and the relay were measured
The same data as a chart, against mesh distance
6. Synchronisation
Notes on the sync table
The same E48 rows, pooled over the cards, are the scatter-add chart of Influence functions on the ET-SoC-1 (S3), in the same colours and marks. The contended atomic is measured in One hot line stops a shire; the barriers, the reduction trees and every other primitive in On-chip communication.
7. Composition: where a workload's joules go
7.1 Build a workload's energy
The equation above, applied: pick a die temperature, the data, and up to three things the workload does at a rate, and the tables of sections 1–5 price its board power and its energy. Start from a measured workload and move the levers.
What the workload does: up to three events at a rate (a preset fills them; change any).
How the calculator prices a row
7.2 The relay, priced from section 4
How the relay's price was computed
8. Three cards
What the chart's numbers are
9. Method, and what the manual cannot tell you
Caveats and method in full
- Board power is the PMIC's reading, 10 mW resolution, sampled at 10 Hz; under that sampling it takes a new value about every 156 ms on aifoundry2 and aifoundry1's card 1 and about every 263 ms on aifoundry3. Each measurement is a burst of back-to-back launches with 4.5 s of idle either side (the catalogue; the reruns of the rings and the levels leave 10 s); its idle is the mean of the two bracketing stretches, and the extra leakage of a burst that runs slightly warmer than its brackets is taken out with the section 1 slope — a correction of . Rates come from the device's cycle counter, not the wall clock. Some DRAM-read bursts slowed the sampler itself; the catalogue keeps them (Limits of observability, §4.1).
- "Per instruction" includes the awake core that issued it (4.4–5.1 pJ per slot with both harts, what a fence or a nop costs). Only the tensor unit is resolved below the instruction, and a "flip" there is an event in a simulation of the design, not a transistor.
- Everything is at 600 MHz; the minion rail reads 0.518 V on aifoundry2 and 0.523 V on aifoundry3. Thirteen instructions trapped in U-mode in a one-off check while the catalogue was written (listed in 3.1; the card and the log were not kept), among them every float and vector divide and square root.
- The unsensed remainder, about 15 W on aifoundry2 and 13 W on aifoundry3, the largest single component of idle on every card, cannot be split with any instrument here. What could be done about that, and every other limit of the card's instruments, is the subject of the limits of observability.
- Confidence bars.
Reproduce this
Every table is rebuilt from the committed data by the commands in docs/energy-manual/README.md; the measurement runs are docs/findings/03-experiments.md, E26–E30, and the version-3 check (its record is linked above).
Version history
Versions: three editions on 23 September (the first; every instruction on two cards and section 4.3; confidence bars and the reruns); corrected 24 September after a review; revised 25 September after a second review (DRAM rows, the stalled minion and the chip barrier, the unmetered attribution moved to the hub, and the maps and the calculator of sections 3.1, 4 and 7.1). 25 September (version 3): every claim checked for proof on both cards; the idle law is led by its measured slope, with its split into fixed and leakage given as a range; figures that rest on one card, on fewer than three runs, or that differ between the cards now say so; rankings and differences within the noise removed (record: docs/reports/data/2026-09-25-claims-v3/). 26 September: the version-3 check's measurements on three cards (aifoundry1's card 1 joins aifoundry2 and aifoundry3) replace the catalogue, the rings, the levels, the relay and the tensor rows, and add each card's idle. 27 September: the gathers, scatters and packed atomics measured on the same three cards after the check (E48): section 3.1's last table, 4.3's gather patterns, the new section 4.4 (irregular access by level; the rail split becomes 4.5) and six rows of section 6. Later on 27 September: two charts, the tensor unit's multiply-add by precision and data on every card (3.2, its table now folded) and every way to add into a table by rate and energy (6); section 4.3's pointer to Heat per millimetre takes that page's three-card figure (2.17 pJ/B per hop was its 24 September run's), and section 3's operating point and chart caption are computed from the data. 28 September: section 8 gives the hot passes' die temperature on aifoundry2 and the 90–103 °C pass at which nothing on the card acted; later, the review's cuts (captions of sections 1, 4.2 and 9, and Related reports).
10. Related reports
The manual is the table of costs; these are the reports that measured its parts, by the section they feed, and the hub that says how far the instruments can be trusted.
- Limits of observability — the hub: what each instrument can and cannot see, the unmetered remainder attributed, and the glossary.
- Section 1: The DVFS loop and its leakage — aifoundry2's governor, its states and thresholds and its three operating points, the rail split of idle, and why a warm card sits at 600 MHz; Power and temperature — the first power and thermal measurements: what the PMIC and the service processor report and how often, and the 34-shire voltage map.
- Sections 2 and 3.2: The Horace experiment — the same matmul on fourteen operand patterns launched from the same die temperature, the tensor unit's energy resolved to four kinds of RTL event, the thermal model, and structured matrices priced before they ran; Why is the ET-SoC-1 low power? — an ablation of the dense matmul, term by term of C·V²·f plus leakage, against an A100; Matmul efficiency and Sparse compute — the 18 September tensor-unit rates and the first zero-skip power measurements, which section 3.2 re-measures under temperature control.
- Section 4: Heat per millimetre — section 4.3's wires measured on their own: the bits on the links chosen, a hop's energy split into bits that differ between flits, ones carried and a fixed part, free against shared links, and the result per bit·mm set against Dally's rule of thumb; use its figures for the mesh; Anatomy of a memory access and Memory hierarchy — the 18–19 September reports that mapped the L3 homes and the DRAM banks and first measured the levels that section 4.2 now carries with bars; Ridge points — the reuse each memory level demands, with energy balance points built from sections 3.2, 4.2 and 5.
- Section 5: On-chip communication — the 18 September report that checked the shire map against latency and first measured the rings; Hand it to the next shire — the relay of sections 5 and 7.2: a multi-stage computation whose intermediate lives in another shire's scratchpad instead of DRAM, and where that stops paying.
- Section 6: One hot line stops a shire — the contended atomic: fair to its requesters, fatal to the host shire's own loads at a steady 600 MHz, with the errata that describe it.
- Spatial temperature: a brief — where the 35 temperature sensors sit and why the card reports one number.
docs/findings/in yaroslavvb/et-soc1-prototyping — every claim traced to its experiment and its file;docs/energy-manual/is this manual as markdown.