| 1 |
|
- |
# Control Bench BOM (bench-v1)
|
| 2 |
|
- |
|
| 3 |
|
- |
Settled 2026-07-25. The machine actually being bought: an **all-new** host that is the
|
| 4 |
|
- |
always-on x86_64 build machine now and a single-to-dual-GPU inference box later.
|
| 5 |
|
- |
|
| 6 |
|
- |
This is not `mm-v1-bom.md`. That file is the aspirational spec, kept for what it is worth;
|
| 7 |
|
- |
this is the one with parts in it.
|
| 8 |
|
- |
|
| 9 |
|
- |
## What this machine is for, in priority order
|
| 10 |
|
- |
|
| 11 |
|
- |
1. **Always-on build host, fire-and-forget.** Alloy images and ISOs, the MNW build gate,
|
| 12 |
|
- |
scheduled rebuilds. The ecosystem has no always-on x86_64 machine: astra is always-on and
|
| 13 |
|
- |
aarch64, fw13 is x86_64 and sleeps, and production never builds binaries. Nobody waits on
|
| 14 |
|
- |
this machine interactively — the edit/compile loop stays in Helix on fw13 — so it is sized
|
| 15 |
|
- |
for **throughput, not latency**.
|
| 16 |
|
- |
|
| 17 |
|
- |
**Sando does not live here.** The original BOM assigned this box the Sando host role back
|
| 18 |
|
- |
when it was going to be one big always-on machine. That does not survive the box becoming
|
| 19 |
|
- |
a card-swapping bench: a production deploy controller should not share a chassis with
|
| 20 |
|
- |
unproven hardware, and a single non-redundant root NVMe makes it worse. Sando moves to
|
| 21 |
|
- |
**astra** — always-on, never opened, and arch-agnostic for a controller.
|
| 22 |
|
- |
2. **The known-good control for testing used parts one at a time.** Every part here is new.
|
| 23 |
|
- |
That is the point: when a used Collection card misbehaves on the bench, the host is not a
|
| 24 |
|
- |
suspect. One variable at a time, against a reference that does not move.
|
| 25 |
|
- |
3. **Single-GPU (and, via bifurcation, dual-GPU) inference work.** EveryCycle roadmap A1–A2
|
| 26 |
|
- |
and B3–B5 all run on one card. See the ceiling section for where it stops.
|
| 27 |
|
- |
|
| 28 |
|
- |
It is also **explicitly a stepping stone**: this machine gets sold and replaced before
|
| 29 |
|
- |
anything more powerful is sold to a customer. So overbuying for a future workload is not
|
| 30 |
|
- |
prudence here, it is waste. Buy the smallest credible control; put the money into cards.
|
| 31 |
|
- |
|
| 32 |
|
- |
## Component list
|
| 33 |
|
- |
|
| 34 |
|
- |
Prices below are a **summary observed 2026-07-25**, kept here for readability. They are not the
|
| 35 |
|
- |
source of truth: every observation lives in the `tm-mcp` store (`_private/tm/pricing.db`, tables
|
| 36 |
|
- |
`asks` and `sells`) with its route, condition and date. Ask it rather than trusting this column —
|
| 37 |
|
- |
`compare <sku>` — and record the real paid number there at purchase. Analysis of what the numbers
|
| 38 |
|
- |
mean is in `_private/docs/hardware/bench-v1/price-baseline.md`.
|
| 39 |
|
- |
|
| 40 |
|
- |
| Component | Choice | Approx | Notes |
|
| 41 |
|
- |
|---|---|---|---|
|
| 42 |
|
- |
| CPU | **AMD EPYC 4565P** (16c/32t "Zen 5", AM5) | **$589** list | 4.30 base / 5.70 boost, 64 MB L3, 170 W, 2× DDR5-5600, 28 PCIe Gen 5 lanes. Sized for throughput: clean builds, container builds, the ISO's zstd pass. See the CPU note below for why this and not a Ryzen 9950X. **Ceiling: EPYC 4005 stops at 16 cores**, so if throughput ever binds the answer is a new machine, not a new CPU — acceptable for a declared stepping stone |
|
| 43 |
|
- |
| Board | **ASRock Rack B650D4U** (base variant) | ~$420 | mATX server board. **AST2600 BMC** with IPMI 2.0, iKVM, vMedia, dedicated management LAN. Base variant chosen over -2L2T/-2L2T/BCM because **it is the only one with two M.2 slots**; see the networking note for why its 1 GbE is not a compromise |
|
| 44 |
|
- |
| RAM | **2× 32 GB Micron MTC20C2085S1EC56BD1** (64 GB) | ~$430 | DDR5-5600 ECC UDIMM, 2Rx8. **On the board QVL**, and the only 32 GB ECC part on it at 5600. 1 DPC runs 5600; 2 DPC drops to 3600, so two slots stay empty on purpose. The 48 GB tier is QVL-reachable too (SK Hynix HMCGY8MGBEB213N) but roughly doubles the line for capacity 16 cores do not need — see the memory section |
|
| 45 |
|
- |
| Storage 1 | 4 TB NVMe on **M2_1** (PCIe 5.0 ×4) | ~$300 | Raw XFS **scratch**: kv-cache overflow later, container build scratch from day one. The fast slot goes to the only thing whose point is bandwidth. Buy a Gen4 drive now while NAND is expensive; the Gen5 slot is headroom |
|
| 46 |
|
- |
| Storage 2 | 2 TB NVMe on **M2_2** (PCIe 4.0 ×4) | ~$200 | **xfs root**, and not ZFS — see the note in `mm-v1-bom.md`. Root gains nothing from Gen5 |
|
| 47 |
|
- |
| PSU | **1200 W Platinum, ATX 3.1 with native 12V-2x6** | ~$300 | Was 850 W, which contradicted this document's own claim that two cards are reachable: 2× 250 W passive + a 170 W CPU is ~750 W, or 88% load. 1200 W also covers one modern ~450 W card. **ATX 3.1 is deliberate**: it handles the transient spikes modern cards produce and provides 12V-2x6 natively, so no adapters when the current-gen vertical gets tested. The 500–600 W-class vertical triggers a PSU review |
|
| 48 |
|
- |
| Enclosure | **Open frame** — see the enclosure section | ~$150–250 | Optimized for access. Airflow is solved at the card, not by the case |
|
| 49 |
|
- |
| CPU cooler | AM5 tower, any height | ~$100 | Air. No lid to clear on an open frame |
|
| 50 |
|
- |
| Card airflow | Per-card **test sled**: server blower + shroud, per vertical | ~$150 setup | The actual thermal instrument. See the enclosure section |
|
| 51 |
|
- |
| Boot GPU | **none** | $0 | The AST2600 has its own VGA and iKVM. `mm-v1-bom.md` carried a GT 1030 only because it was unsure the BMC would suffice; the manual settles it |
|
| 52 |
|
- |
| **Total** | | **~$2,680** | vs ~$10,280 for the `mm-v1` host platform. Line-by-line observations, spreads and market notes: `_private/docs/hardware/bench-v1/price-baseline.md` |
|
| 53 |
|
- |
|
| 54 |
|
- |
Vendor manuals for the decided parts are in `_private/docs/hardware/bench-v1/`.
|
| 55 |
|
- |
|
| 56 |
|
- |
## CPU: why the EPYC and not a Ryzen 9950X
|
| 57 |
|
- |
|
| 58 |
|
- |
Decided on merit rather than price (Max: "$150 is marginal, whichever is the better part").
|
| 59 |
|
- |
From AMD's EPYC 4005 datasheet, held locally, against the 9950X:
|
| 60 |
|
- |
|
| 61 |
|
- |
| | Cores | Base | Boost | TDP | L3 | Channels | Gen 5 lanes |
|
| 62 |
|
- |
|---|---|---|---|---|---|---|---|
|
| 63 |
|
- |
| EPYC 4565P | 16/32 | 4.30 | 5.70 | 170 W | 64 MB | 2× DDR5-5600 | 28 |
|
| 64 |
|
- |
| Ryzen 9950X | 16/32 | 4.30 | 5.70 | 170 W | 64 MB | 2× DDR5-5600 | 28 |
|
| 65 |
|
- |
|
| 66 |
|
- |
Same silicon, same numbers. **There is no performance argument either way**, so this comes
|
| 67 |
|
- |
down to everything else, and the deciding factor is **officially validated ECC UDIMM**.
|
| 68 |
|
- |
|
| 69 |
|
- |
Ryzen 9000 ECC works on boards that wire it up, and ASRock Rack does — but board-vendor
|
| 70 |
|
- |
support is not CPU-vendor validation. This machine exists to be a trustworthy reference; its
|
| 71 |
|
- |
characterization gate below requires ECC counters that are both correct and readable; and a
|
| 72 |
|
- |
silent memory error corrupting a build artifact is exactly the failure ECC is bought to
|
| 73 |
|
- |
prevent. Validated beats de-facto for that job.
|
| 74 |
|
- |
|
| 75 |
|
- |
Secondary, both real: the embedded/server line carries a longer supply and firmware-support
|
| 76 |
|
- |
window, and ASRock Rack's BIOS and QVL work on this board targets EPYC 4004/4005 first. And
|
| 77 |
|
- |
since this file doubles as a TM product template, "server CPU with validated ECC on a server
|
| 78 |
|
- |
board with a BMC" is coherent to sell under a Care Policy in a way that a consumer part with
|
| 79 |
|
- |
unofficial ECC is not.
|
| 80 |
|
- |
|
| 81 |
|
- |
The 9950X's only genuine edge is its iGPU, which the AST2600's own VGA already covers.
|
| 82 |
|
- |
**Fallback:** the EPYC 4005 is a niche part, and if availability is bad the 9950X drops into
|
| 83 |
|
- |
this board with zero performance change. Preference with a documented alternative, not a hard
|
| 84 |
|
- |
requirement.
|
| 85 |
|
- |
|
| 86 |
|
- |
**Considered and rejected: the 4545P.** Also 16c/32t Zen 5, but 65 W instead of 170 W
|
| 87 |
|
- |
(3.0 base / 5.4 boost) and cheaper — tempting for an always-on box in a living space. Rejected
|
| 88 |
|
- |
because a mostly-idle CI host spends most of its hours at **idle**, where both parts draw
|
| 89 |
|
- |
about the same. The 170 W part only draws more while actually building, and it finishes
|
| 90 |
|
- |
sooner, so energy per build is roughly a wash while wall-clock is materially better. If heat
|
| 91 |
|
- |
or fan noise later becomes the actual complaint, the 4545P is the drop-in answer.
|
| 92 |
|
- |
|
| 93 |
|
- |
**Considered and rejected: the 4585PX** (16c, 128 MB via 3D V-Cache). Clocks are listed TBD in
|
| 94 |
|
- |
the datasheet, so it may not be shipping; compilation gains from doubled L3 are modest, and
|
| 95 |
|
- |
X3D parts usually trade all-core clock for the cache.
|
| 96 |
|
- |
|
| 97 |
|
- |
## Memory: the QVL settles it, and it is not the part that was priced
|
| 98 |
|
- |
|
| 99 |
|
- |
The manual's "Memory support is to be validated" hedge was the last hard blocker on ordering.
|
| 100 |
|
- |
It is answered: ASRock Rack publishes a **22-row memory QVL** for this board, pulled
|
| 101 |
|
- |
2026-07-25. Local copy: `_private/docs/hardware/bench-v1/b650d4u-memory-qvl.md`.
|
| 102 |
|
- |
|
| 103 |
|
- |
**What it says, for the two things that were actually in doubt:**
|
| 104 |
|
- |
|
| 105 |
|
- |
- **Dual-rank ECC UDIMM at 5600 is validated.** Both of the board's 5600-grade ECC entries
|
| 106 |
|
- |
are 2Rx8. So the worry that "ECC UDIMM + a server board + a not-quite-Ryzen CPU" would fall
|
| 107 |
|
- |
outside validated territory does not survive the list. Nothing about dual-rank is the risk
|
| 108 |
|
- |
it was assumed to be.
|
| 109 |
|
- |
- **48 GB at 5600 is on the list**, so the manual's 48 GB-per-DIMM ceiling is real capacity
|
| 110 |
|
- |
and not just a spec-sheet maximum.
|
| 111 |
|
- |
|
| 112 |
|
- |
**Only three ECC parts on the QVL run at 5600, and each capacity has exactly one:**
|
| 113 |
|
- |
|
| 114 |
|
- |
| Size | Vendor | Part | Rank |
|
| 115 |
|
- |
|---|---|---|---|
|
| 116 |
|
- |
| 48 GB | SK Hynix | HMCGY8MGBEB213N | 2Rx8 |
|
| 117 |
|
- |
| 32 GB | Micron | MTC20C2085S1EC56BD1 | 2Rx8 |
|
| 118 |
|
- |
| 16 GB | Micron | MTC10C1084S1EC56BD1 | 1Rx8 |
|
| 119 |
|
- |
|
| 120 |
|
- |
Everything else on the list is 4800 or 5200. Buying off-QVL is therefore not a small
|
| 121 |
|
- |
compromise here: it costs a third of the memory clock, which is the same penalty as filling
|
| 122 |
|
- |
all four slots.
|
| 123 |
|
- |
|
| 124 |
|
- |
**The Kingston KSM56E46BD8KM-48HM is not on the QVL.** That is the part the $270–$1,006
|
| 125 |
|
- |
spread was found on, and it was the presumed buy. Kingston appears on the list exactly once
|
| 126 |
|
- |
in ECC, as `KSM48E40BD8KM-32HM`, which is a **4800** part. Kingston's own compatibility tool
|
| 127 |
|
- |
does list the 48 GB module against this board — but that is vendor self-certification, not
|
| 128 |
|
- |
board-vendor validation, and this file already decided that distinction in the CPU section.
|
| 129 |
|
- |
Having chosen the EPYC over the 9950X specifically to get validated ECC, buying an off-QVL
|
| 130 |
|
- |
DIMM would spend that choice for nothing.
|
| 131 |
|
- |
|
| 132 |
|
- |
**Take 2× 32 GB Micron, not 2× 48 GB Hynix.** Both are QVL and both run 5600 at 1 DPC, so
|
| 133 |
|
- |
this is purely capacity against cost — roughly **$430 for 64 GB versus ~$990 for 96 GB**. The
|
| 134 |
|
- |
BOM already held that 64 GB is enough for 16 cores, the machine is a declared stepping stone,
|
| 135 |
|
- |
and memory is in an up-cycle where deferring spend is the cheaper bet. The two free DIMM
|
| 136 |
|
- |
slots remain the escape hatch, at the 3600 penalty.
|
| 137 |
|
- |
|
| 138 |
|
- |
One channel note: `HMCGY8MGBEB213N` is an OEM part number, so retail carries rebrands (Axiom
|
| 139 |
|
- |
`AX55600E46I/48G` at ~$868) rather than the SK Hynix module itself. If the 48 GB tier is ever
|
| 140 |
|
- |
wanted, that gap is part of its cost.
|
| 141 |
|
- |
|
| 142 |
|
- |
## Enclosure: open frame, and airflow solved at the card
|
| 143 |
|
- |
|
| 144 |
|
- |
The earlier draft of this file argued for a 4U rackmount because a passive card "only cools
|
| 145 |
|
- |
in a chassis with a front-to-back static-pressure path." **That argument does not survive
|
| 146 |
|
- |
scrutiny.** `mm-v1-bom.md` itself budgets a fan shroud for the P40 *inside* the 4U, and it is
|
| 147 |
|
- |
right to: passive datacenter cards are engineered for 1U/2U server fans producing static
|
| 148 |
|
- |
pressure that no 140 mm case fan approaches. The shroud is needed either way, so the 4U was
|
| 149 |
|
- |
buying repeatability, not cooling — a weaker claim than the one it was sold on.
|
| 150 |
|
- |
|
| 151 |
|
- |
Once noise is not a constraint, there is a better answer available, and it inverts the
|
| 152 |
|
- |
problem.
|
| 153 |
|
- |
|
| 154 |
|
- |
**Airflow is solved at the card, with real server blowers.** 40–80 mm server blowers at
|
| 155 |
|
- |
10k+ RPM produce the static pressure these cards were designed around. That is a thing a
|
| 156 |
|
- |
quiet build physically cannot do, and it is the single biggest quality difference in a bench
|
| 157 |
|
- |
test of a passive card. With sound deprioritized, use them.
|
| 158 |
|
- |
|
| 159 |
|
- |
**Once airflow lives at the card, the enclosure only has to be good to work in.** So:
|
| 160 |
|
- |
|
| 161 |
|
- |
- **Open aluminium frame** (mining-rig style, or an open bench-frame such as a Core P3 /
|
| 162 |
|
- |
BC1-class fixture). Nothing to unscrew, nothing to unrack, no lid, no cable gymnastics.
|
| 163 |
|
- |
Cards go in and out in seconds, which is the operation performed most.
|
| 164 |
|
- |
- **Unlimited card clearance.** Modern 3–3.5 slot, 350 mm cards fit trivially. A 4U would
|
| 165 |
|
- |
have constrained exactly the vertical most likely to need bench time later.
|
| 166 |
|
- |
- **Everything visible.** You can see fan spin, LEDs, and scorch marks on an unknown card
|
| 167 |
|
- |
before they become a smell.
|
| 168 |
|
- |
|
| 169 |
|
- |
### The test sled is the instrument
|
| 170 |
|
- |
|
| 171 |
|
- |
Repeatability comes from a fixture, not a case. Build one **test sled** per card class: a
|
| 172 |
|
- |
rigid bracket that holds the card, its shroud, and its blower at a **fixed** geometry, with
|
| 173 |
|
- |
the blower on a controller at a **fixed, recorded RPM**. Every card in that class is then
|
| 174 |
|
- |
tested in an identical thermal environment regardless of what is around it — better
|
| 175 |
|
- |
repeatability than a shared case gives, because a case's airflow changes with every other
|
| 176 |
|
- |
card and cable in it.
|
| 177 |
|
- |
|
| 178 |
|
- |
Consequences to hold:
|
| 179 |
|
- |
|
| 180 |
|
- |
- **One sled per vertical.** A passive P40, a 2-fan consumer card and a 3.5-slot current-gen
|
| 181 |
|
- |
card have nothing thermally in common. Sleds get built as verticals open, not up front.
|
| 182 |
|
- |
- **Log RPM and ambient with every test.** A thermal trace without them is not comparable.
|
| 183 |
|
- |
Until `bmc-agent` exists (and it may not — task `59335767`), blower RPM is set on a manual
|
| 184 |
|
- |
fan controller and written down, and BIOS-set chassis fan curves are recorded.
|
| 185 |
|
- |
- **Dust is the accepted cost.** An open frame running 24/7 in a living space collects dust
|
| 186 |
|
- |
in the CPU cooler and PSU intake. Mitigated by the blowers only running during tests, and
|
| 187 |
|
- |
by putting periodic cleaning in the standing ops list. This is a real trade and the reason
|
| 188 |
|
- |
a closed case would otherwise win.
|
| 189 |
|
- |
- **It is not shippable as-is.** When this machine is sold, it either gets rehoused in a case
|
| 190 |
|
- |
or sold as a parts bundle. Budget that, and do not let the frame make the machine
|
| 191 |
|
- |
unsellable by surprise.
|
| 192 |
|
- |
|
| 193 |
|
- |
## The GPU ceiling, stated exactly
|
| 194 |
|
- |
|
| 195 |
|
- |
From the board manual, section 2.6:
|
| 196 |
|
- |
|
| 197 |
|
- |
| Slot | Gen | Mechanical | Electrical |
|
| 198 |
|
- |
|---|---|---|---|
|
| 199 |
|
- |
| PCIE6 | 5.0 | x16 | x16 |
|
| 200 |
|
- |
| PCIE7 | 5.0 | **x4** | x4 |
|
| 201 |
|
- |
| PCIE4 | 4.0 | x1 | x1 |
|
| 202 |
|
- |
|
| 203 |
|
- |
**One GPU seats natively.** PCIE7 is x4 *mechanically*, so a card does not physically fit;
|
| 204 |
|
- |
the idea of parking a second Pascal card there is dead.
|
| 205 |
|
- |
|
| 206 |
|
- |
**Two GPUs are reachable via bifurcation.** BIOS exposes "Configure PCIE6 Link Width" with
|
| 207 |
|
- |
`[x16]`, `[x8x8]`, `[x8x4x4]`. Gen5 x8 is roughly twice the bandwidth a Gen3 x16 card such
|
| 208 |
|
- |
as a P40 can consume, so splitting costs those cards nothing. What it costs is a bifurcation
|
| 209 |
|
- |
riser and a physical mounting problem: two dual-slot cards on risers off an mATX board is a
|
| 210 |
|
- |
rig to solve, not a slot to populate. The open frame helps here — a frame is a mounting
|
| 211 |
|
- |
surface, where a case would have been a constraint. **Unproven until someone builds it.**
|
| 212 |
|
- |
|
| 213 |
|
- |
So against the EveryCycle roadmap:
|
| 214 |
|
- |
|
| 215 |
|
- |
- **A1** (1× P40) — fits natively.
|
| 216 |
|
- |
- **A2** — fits, single card.
|
| 217 |
|
- |
- **A3** (P40 + P100, mixed-arch in one executor) — needs the x8x8 riser rig. Reachable, not
|
| 218 |
|
- |
guaranteed.
|
| 219 |
|
- |
- **A4/A5** (3–4 cards) — **does not fit.** That is the replacement machine's job, and the
|
| 220 |
|
- |
2000 W PSU in `mm-v1-bom.md` exists for exactly that.
|
| 221 |
|
- |
- **B3, B5** (Qwen 8B, then 72B Q4 saturating a P40) — single card, fits.
|
| 222 |
|
- |
- **B6** (frontier-size with CPU offload) — **does not fit.** Two DDR5 channels is ~80–90 GB/s
|
| 223 |
|
- |
against ~358 GB/s for 8-channel DDR5-5600. CPU-offload work belongs to the replacement.
|
| 224 |
|
- |
|
| 225 |
|
- |
## Networking: 1 GbE now, 10 GbE deferred into PCIE7
|
| 226 |
|
- |
|
| 227 |
|
- |
The board comes in three variants and the choice is a three-way trade between M.2 slots,
|
| 228 |
|
- |
10 GbE and cost. Only the base `B650D4U` has **two** M.2 slots; both 10 GbE variants have one,
|
| 229 |
|
- |
because the 10 GbE controller eats the FCH lanes M2_2 would use.
|
| 230 |
|
- |
|
| 231 |
|
- |
Measured on the LAN 2026-07-25 rather than assumed:
|
| 232 |
|
- |
|
| 233 |
|
- |
- **fw13**: USB adapter, linked at **1000 Mb/s**. A Framework 13 has no easy 10 GbE path
|
| 234 |
|
- |
short of a Thunderbolt adapter.
|
| 235 |
|
- |
- **astra**: `enP3p3s0f0` and `f1` are on the **ixgbe** driver, so that is Intel 10 GbE
|
| 236 |
|
- |
silicon — but `f0` is linked at **1000 Mb/s** and `f1` is down.
|
| 237 |
|
- |
|
| 238 |
|
- |
So 10 GbE hardware exists on the LAN and nothing is actually running at 10 GbE. More to the
|
| 239 |
|
- |
point, the traffic that would justify it is moving multi-GB ISOs and images to wherever a USB
|
| 240 |
|
- |
stick gets written, which is fw13, which is 1 GbE-limited regardless.
|
| 241 |
|
- |
|
| 242 |
|
- |
**Decision: take the base board, keep both native M.2, and leave PCIE7 empty.** If a 10 GbE
|
| 243 |
|
- |
path later proves worth having, PCIE7 (Gen 5 ×4) takes a NIC with room to spare — 10 GbE needs
|
| 244 |
|
- |
a fraction of that. Deferring costs nothing and avoids paying a board premium plus an M.2
|
| 245 |
|
- |
adapter for a link that currently has no peer running at speed.
|
| 246 |
|
- |
|
| 247 |
|
- |
**Constraint on that later NIC: it must be ×4 *mechanical*.** PCIE7 is ×4 mechanically, so an
|
| 248 |
|
- |
Intel X550-T2 (×4) fits and an X710-T2L (×8) does not. An X550 also pairs with the same
|
| 249 |
|
- |
`ixgbe` driver astra already runs, which is a known-good story rather than a new one.
|
| 250 |
|
- |
|
| 251 |
|
- |
## The bench is not a control until it is characterized
|
| 252 |
|
- |
|
| 253 |
|
- |
"All-new means known-good" is the right instinct with the wrong mechanism. New parts have
|
| 254 |
|
- |
infant mortality and firmware bugs; a used part with hours on it can be *more* reliable than
|
| 255 |
|
- |
a fresh one. What new actually buys is an **RMA path** — when a host part does fail, it gets
|
| 256 |
|
- |
replaced rather than debugged, and it never becomes a suspect in a card's test result. That
|
| 257 |
|
- |
holds up. The "known-good" part does not, until it is measured.
|
| 258 |
|
- |
|
| 259 |
|
- |
**Before the first used card goes in**, record a baseline and keep it in the runbook:
|
| 260 |
|
- |
|
| 261 |
|
- |
- memtest86+ clean pass over all installed DIMMs.
|
| 262 |
|
- |
- Sustained all-core `stress-ng` run: clocks held, package temp, no throttle.
|
| 263 |
|
- |
- `fio` on both NVMes: throughput and latency, so a future slow build has a reference.
|
| 264 |
|
- |
- Idle and full-load wall power, and the ambient temperature both were taken at.
|
| 265 |
|
- |
- ECC corrected-error count at zero, **and proof you can read it** — see below.
|
| 266 |
|
- |
|
| 267 |
|
- |
**Designate a reference card per vertical.** To tell "this card is bad" from "the slot, riser,
|
| 268 |
|
- |
PSU, sled or BIOS is bad" you need a known-good card to swap in. Buying a new one in each
|
| 269 |
|
- |
class is not worth it. Instead: **the first used card that passes cleanly in a vertical
|
| 270 |
|
- |
becomes that vertical's reference** — its numbers are recorded, and it is not sold. Cheap,
|
| 271 |
|
- |
and it makes every later test in that class a comparison rather than an absolute.
|
| 272 |
|
- |
|
| 273 |
|
- |
**ECC you cannot read is not ECC.** Rising correctable-error counts are the early warning the
|
| 274 |
|
- |
memory is paid for, and Alloy ships no `rasdaemon`, `edac-utils`, `smartmontools` or
|
| 275 |
|
- |
`ipmitool`. Those belong in the server profile (GoingsOn alloy task `d1fed0d7`) and the
|
| 276 |
|
- |
baseline above cannot be taken without them.
|
| 277 |
|
- |
|
| 278 |
|
- |
## Card verticals, opened one at a time
|
| 279 |
|
- |
|
| 280 |
|
- |
The Collection wants at least one cut in several verticals, and cannot fund testing all of
|
| 281 |
|
- |
them at once. So verticals open in sequence, and each one opens a small equipment bill:
|
| 282 |
|
- |
|
| 283 |
|
- |
| Vertical | Bench implications |
|
| 284 |
|
- |
|---|---|
|
| 285 |
|
- |
| **Passive datacenter pulls** (P40, P100, MI50) | Shroud + server blower sled. Some use CPU-style EPS 8-pin, not PCIe. Large BARs need Above 4G + ReBAR. First, because the roadmap names it and the cards are cheapest |
|
| 286 |
|
- |
| **Used consumer / gaming** | Self-cooling, standard PCIe power, trivial sled. Needs a monitor on the card under test, since display output is part of what is being sold |
|
| 287 |
|
- |
| **B-stock / open-box current-gen** | Standard cooling, may retain manufacturer warranty. Bench time still required to prove it |
|
| 288 |
|
- |
| **Modern high-TDP** (500–600 W class) | 12V-2x6 native (the ATX 3.1 PSU covers this), real transient spikes, 3.5 slots, 350 mm. **Triggers a PSU review past 1200 W** and its own sled |
|
| 289 |
|
- |
|
| 290 |
|
- |
Sequence the verticals; do not buy sleds or PSU headroom for a vertical that is not open yet.
|
| 291 |
|
- |
|
| 292 |
|
- |
## What was deliberately not bought, and why
|
| 293 |
|
- |
|
| 294 |
|
- |
- **8-channel memory bandwidth.** The single real loss. It costs a $3,900 CPU tier plus
|
| 295 |
|
- |
$2,500 of DDR5 on WRX90, and pays back only on B6, the last item on the model thread. Buy
|
| 296 |
|
- |
it when B6 is close, on whatever platform is cheap by then.
|
| 297 |
|
- |
- **PCIe 5.0 x16 slot count.** The roadmap's cards are Gen3 Pascal. Seven Gen5 x16 slots is
|
| 298 |
|
- |
paying for signalling nothing in the plan can use.
|
| 299 |
|
- |
- **512 GB RAM.** Justified in `mm-v1-bom.md` as "405B Q4 fits in RAM if ever wanted", which
|
| 300 |
|
- |
is an aspiration priced at $2,500. This board tops out at 192 GB anyway, and past 96 GB the
|
| 301 |
|
- |
memory clock drops from 5600 to 3600.
|
| 302 |
|
- |
- **A 2000 W PSU and 4 GPU bays.** Sized for a machine this is explicitly not.
|
| 303 |
|
- |
- **Used parts.** Deliberate, and the reason the bench is credible. The used parts are the
|
| 304 |
|
- |
cards under test.
|
| 305 |
|
- |
|
| 306 |
|
- |
## Sourcing channels
|
| 307 |
|
- |
|
| 308 |
|
- |
All-new parts, so these are retail and distribution channels, not the used-GPU channels in the
|
| 309 |
|
- |
TailoredMachines sourcing playbook.
|
| 310 |
|
- |
|
| 311 |
|
- |
| Part | Channels | Watch for |
|
| 312 |
|
- |
|---|---|---|
|
| 313 |
|
- |
| EPYC 4565P | ShopBLT, Provantage, NeweggBusiness, Wiredzone, CDW | Niche server SKU, so **availability is the risk, not price**. Check **tray vs boxed**: tray is OEM, ships without a cooler, and the warranty runs through the reseller |
|
| 314 |
|
- |
| ASRock Rack B650D4U | NeweggBusiness, Newegg, Wiredzone, ServerSupply, Provantage | Confirm it is the **base** variant, not -2L2T or -2L2T/BCM — only the base has two M.2 slots |
|
| 315 |
|
- |
| DDR5 ECC UDIMM | Walmart, Amazon, Memory.net, ServerSupply, OEMPCWorld, eBay | **Buy the QVL part number and nothing else**: `MTC20C2085S1EC56BD1`. Same-part asks ran $215 to $2,270 on one day, so the channel matters more than the part does. Watch for the `...BD1R` suffix, which is the same module in a different retail pack |
|
| 316 |
|
- |
| 2× NVMe | Microcenter in person, Newegg, B&H | Microcenter in-store usually wins on price and removes shipping risk |
|
| 317 |
|
- |
| 1200 W ATX 3.1 PSU | Microcenter, Newegg | Seasonic PRIME/Vertex, Corsair, Super Flower. Confirm **native 12V-2x6**, not an adapter in the box |
|
| 318 |
|
- |
| Open frame | Mining-frame vendors (Veddha, Kingwin), or **8020 aluminium extrusion** for a custom build | Extrusion lets the sled geometry be designed in rather than worked around |
|
| 319 |
|
- |
| Server blowers, shrouds, fan controller | eBay/Amazon for Delta and San Ace 40–80 mm; shrouds printed or bought | Bench equipment, bought per vertical rather than up front |
|
| 320 |
|
- |
|
| 321 |
|
- |
**Microcenter Denver is already a standing trip** — the sourcing playbook has Saturday mornings
|
| 322 |
|
- |
there for open-box GPUs. Commodity parts on the same trip means no shipping risk and no wait.
|
| 323 |
|
- |
|
| 324 |
|
- |
## Open before ordering
|
| 325 |
|
- |
|
| 326 |
|
- |
- ~~Does the board QVL cover DDR5 ECC UDIMM?~~ **settled 2026-07-25 by pulling the QVL** —
|
| 327 |
|
- |
see the memory section. Dual-rank ECC UDIMM at 5600 is validated; the part is Micron
|
| 328 |
|
- |
`MTC20C2085S1EC56BD1`, not the Kingston module that was priced. **Caveat: the QVL carries no
|
| 329 |
|
- |
CPU column**, and its weighting toward 4800 parts suggests it was built in the EPYC
|
| 330 |
|
- |
4004/Ryzen 7000 era. So it validates the *board*, and "with an EPYC 4005 specifically"
|
| 331 |
|
- |
remains unproven until the characterization gate runs. That is a bring-up risk, not an
|
| 332 |
|
- |
ordering blocker.
|
| 333 |
|
- |
- ~~1 GbE vs 10 GbE~~ **settled 2026-07-25 by measuring the LAN** — see the networking
|
| 334 |
|
- |
section. Base board, PCIE7 left empty for a future ×4-mechanical NIC.
|
| 335 |
|
- |
- **OpenBMC.** `crates/bmc-agent` says it runs on "an OpenBMC fork on Aspeed AST2600-class
|
| 336 |
|
- |
hardware". This board has the AST2600 but ships ASRock's vendor IPMI firmware, and vendor
|
| 337 |
|
- |
BMC images are typically signed. If it cannot be replaced, `bmc-agent` retargets
|
| 338 |
|
- |
IPMI/Redfish or wants a dedicated dev board. **This question is not specific to this
|
| 339 |
|
- |
board — it applies to the WRX90 spec identically**, and answering it on a $400 board is
|
| 340 |
|
- |
cheaper.
|
| 341 |
|
- |
- ~~48 GB ECC UDIMM availability and price~~ **settled 2026-07-25** — 48 GB is real and
|
| 342 |
|
- |
QVL-listed, but it is an OEM part sold through rebrands at roughly double the 32 GB tier.
|
| 343 |
|
- |
Decision: 2× 32 GB. See the memory section.
|
| 344 |
|
- |
- **Where the frame physically sits.** There is no rack, and an open frame in a living space
|
| 345 |
|
- |
has different constraints than a case: dust, curiosity, and anything that can be knocked
|
| 346 |
|
- |
into it. Pick the spot before buying the frame.
|
| 347 |
|
- |
- **The x8x8 riser rig**, which an open frame makes easier than any case would: two dual-slot
|
| 348 |
|
- |
cards on risers need mounting, and a frame is a mounting surface.
|
| 349 |
|
- |
|
| 350 |
|
- |
## Sequence the bring-up
|
| 351 |
|
- |
|
| 352 |
|
- |
Too many unknowns arrive at once otherwise: a new hardware platform, Alloy in a role it has
|
| 353 |
|
- |
never held, no settled sshd story, bench duty, and CI. Add one at a time — the same
|
| 354 |
|
- |
discipline the bench itself exists to enforce.
|
| 355 |
|
- |
|
| 356 |
|
- |
1. Bare bring-up on the frame. BIOS, BMC, iKVM, management LAN.
|
| 357 |
|
- |
2. OS install, and find out what `alloy install` does on non-laptop hardware.
|
| 358 |
|
- |
3. Characterization baseline (above). This is the gate that makes it a control.
|
| 359 |
|
- |
4. Build load: images, ISOs, the MNW gate.
|
| 360 |
|
- |
5. First used card, in a sled, with the reference procedure.
|
| 361 |
|
- |
|
| 362 |
|
- |
`/var/lib/containers` mounts on the 4 TB scratch, not on root. On a bootc host `/var` is
|
| 363 |
|
- |
stateroot, so this is a deliberate step; get it wrong and every container build lands on the
|
| 364 |
|
- |
2 TB root.
|
| 365 |
|
- |
|
| 366 |
|
- |
## BIOS at first boot
|
| 367 |
|
- |
|
| 368 |
|
- |
- **PCIE6 Link Width**: leave `[x16]` while there is one card. Revisit at A3.
|
| 369 |
|
- |
- **Above 4G Decoding** and **Resizable BAR**: on. Harmless now, required by large-BAR
|
| 370 |
|
- |
datacenter cards.
|
| 371 |
|
- |
- **IOMMU**: on. Needed for podman CDI GPU passthrough.
|
| 372 |
|
- |
- **Memory**: JEDEC DDR5-5600 at 1 DPC. Not pushed.
|
| 373 |
|
- |
- Configure the **dedicated management LAN** and confirm iKVM works before the machine goes
|
| 374 |
|
- |
anywhere it is inconvenient to reach.
|