Where the work sits: placement and the thermal trip

27–28 September 2026 · development on aifoundry3, calibration and validation on aifoundry1 card 1, aifoundry2 for the clock · every prediction frozen before the validation () · part of the ET-SoC-1 measurement reports

Which theories survived.

Does it matter where on the chip a computation runs, if the goal is to stay under the temperature at which the clock is throttled? Two questions. First, what does the throttle look at? Second, can the same work run longer before that trip in some places than in others?

Terms used on this page

A shire is a tile of 32 small RISC-V cores (minions); the chip runs kernels on 32 shires laid out, with the master and spare shires and the I/O and PCIe shires, on a 6 × 6 grid. Each of the 34 minion shires has one temperature sensor. The mean is the whole-degree mean of those 34 sensors, each truncated to a whole degree first; the governor compares it with 65 °C and acts at 66. A placement is a set of shires and a number of minions per shire running the same random-data multiply-add heater; INT16@32 is 16 interior shires with all 32 minions each, PER16@32 the 16 perimeter shires, UNI32@16 all 32 shires with 16 minions each: 512 minions in every case. t66 is the time from a run's first kernel launch to the first 10 Hz sample in which the mean reads 66 °C. A block runs each placement once, in a balanced order; a card's value is the mean over its blocks with a 99% Student t interval. Development chose the parameters and the predictions on one card and tests nothing; the validation runs the frozen protocol on another card.

Perimeter against interior, same 512 minions: time to the trip
What the governor compares with 65 °C
Power, perimeter less interior
Validation on aifoundry1 card 1

1. The question, and how it was tested

The owner asked on 27 September (quoted in part):

“run experiments which confirm the spatial character of heat dissipation … for the same computation it matters where you place it if you care about keeping the temperature within a certain limit … Perhaps the corners are better cooled than the center … does voltage frequency scaling kick in when the average temperature exceeds the limit or when one of the dies exceeds the limit? It would make sense it should be the latter … run the same heat-intensive computation in different parts of the chip and demonstrate, if possible, that you can actually run it longer in some parts of the chip than in others before the frequency-voltage scaling kicks in … All the iterations should happen on one card and the validation should happen on another one or two cards. I don't want iterations after that … I want to see the summary of which of the theories survived this testing.”

That became two questions and one method. Q1: does the governor act on the mean of the die sensors or on one sensor? Q2: does the same work, placed on the edges of the die, run longer before the trip than placed in the interior? The method:

2. Q1: the governor compares the mean

From the firmware:

On a card:

What this experiment could add, and why it could not:

The probe of each card, and aifoundry2's two attempts

So Q1's answer rests on the source and on the DVFS page's development night; a governor that acts on the mean is also why Q2 is asked of the mean: the trip is the first moment the mean reads 66 °C.

3. The placements

Every placement runs the same kernel (sparsity_host --test fma --type fp32 --values randn: random fp32 multiply-adds on the tensor unit, no memory traffic after set-up) on a chosen set of shires and minions; unselected minions return at once. The three registered placements put 512 minions to work each, on the interior, on the perimeter, or half of every shire, so that total power and total work are the same and only their place differs.

The die's 6 × 6 grid: minions at work in each shire

Every placement, its mask and where it sits

4. The metric: time to 66 °C from the same start

  1. Preheat: : .
  2. Wait for the edge: with only the sampler open, until the mean first reads the start edge S on its way down (the S+1 → S step), so that every run starts from the same whole-degree reading, falling.
  3. Launch within one sample of the edge:
  4. t66 is the first 10 Hz sample with the mean at 66 °C, less the first launch's own start time;

One block, three placements: the mean the governor compares, from each run's first launch

5. Development on aifoundry3

Perimeter against interior, block by block: the log of the ratio of their times to 66 °C

Every development block: t66, power and the edge-cooling time of each run

What ran, block by block, on every card (times PDT)

6. The theories, and what was registered

7. Card 1's calibration

Calibration chains of 768 minions from each start edge: t66 against the acceptance band

8. Verdicts

Card 1's validation blocks: each placement's time to 66 °C

Which theories survived

9. What was dropped, and why

10. Departures from the design

All departures, as the tools' README records them

    11. Method, data, and how to reproduce

    Method notes and limits

    Data. Everything is in docs/reports/data/2026-09-28-heat-placement: every block's raw record (raw/: 10 Hz telemetry, heater output, launches, runs, marks, plan, code and binary hashes), the queue logs, the reductions (reductions/), the design and its critique (plan/), and this page's data, heat.json, which build_heat_data.py writes . The tools are tools/claims-v3/hp; the frozen predictions are its prereg/PREREG.md (SHA-256 ).

    V=docs/reports/data/2026-09-28-heat-placement
    python3 tools/claims-v3/hp/reduce.py --data $V/raw/aifoundry3 --card aifoundry3 --dev --out dev.json    # = $V/reductions/dev-r3.json
    python3 tools/claims-v3/hp/reduce.py --data $V/raw/aifoundry3 --card aifoundry3 --p3 \
        --cal-card1 $V/raw/aifoundry1-c1 --out reg.json                                                   # = $V/reductions/reg.json
    bash $V/collect_val.sh SRC LOGS                                                                      # card 1's validation blocks (README: the verdict step)
    python3 tools/claims-v3/hp/reduce.py --data $V/raw/aifoundry1-c1 --card aifoundry1-c1 --val \
        --prereg tools/claims-v3/hp/prereg/prereg.json --out $V/val.json                                  # the verdicts
    python3 $V/build_heat_data.py                                                                         # heat.json
    python3 scripts/build-report.py heat-placement $V/heat.json docs/reports/2026-09-28-et-soc1-heat-placement.html