|
1 |
+ |
# Control Bench BOM (bench-v1)
|
|
2 |
+ |
|
|
3 |
+ |
Settled 2026-07-25. The machine actually being bought: an **all-new** host that is the
|
|
4 |
+ |
always-on x86_64 build machine now and a single-to-dual-GPU inference box later.
|
|
5 |
+ |
|
|
6 |
+ |
This is not `mm-v1-bom.md`. That file is the aspirational spec, kept for what it is worth;
|
|
7 |
+ |
this is the one with parts in it.
|
|
8 |
+ |
|
|
9 |
+ |
## What this machine is for, in priority order
|
|
10 |
+ |
|
|
11 |
+ |
1. **Always-on build host.** Alloy images and ISOs, the MNW build gate, Sando. The ecosystem
|
|
12 |
+ |
has no always-on x86_64 machine: astra is always-on and aarch64, fw13 is x86_64 and sleeps,
|
|
13 |
+ |
and production never builds binaries.
|
|
14 |
+ |
2. **The known-good control for testing used parts one at a time.** Every part here is new.
|
|
15 |
+ |
That is the point: when a used Collection card misbehaves on the bench, the host is not a
|
|
16 |
+ |
suspect. One variable at a time, against a reference that does not move.
|
|
17 |
+ |
3. **Single-GPU (and, via bifurcation, dual-GPU) inference work.** EveryCycle roadmap A1–A2
|
|
18 |
+ |
and B3–B5 all run on one card. See the ceiling section for where it stops.
|
|
19 |
+ |
|
|
20 |
+ |
It is also **explicitly a stepping stone**: this machine gets sold and replaced before
|
|
21 |
+ |
anything more powerful is sold to a customer. So overbuying for a future workload is not
|
|
22 |
+ |
prudence here, it is waste. Buy the smallest credible control; put the money into cards.
|
|
23 |
+ |
|
|
24 |
+ |
## Component list
|
|
25 |
+ |
|
|
26 |
+ |
Prices are 2026-07 estimates, not quotes. Snapshot the real number and date at purchase.
|
|
27 |
+ |
|
|
28 |
+ |
| Component | Choice | Approx | Notes |
|
|
29 |
+ |
|---|---|---|---|
|
|
30 |
+ |
| CPU | **AMD EPYC 4565P** (16c/32t "Zen 5", AM5) | ~$700–800 | Up to 5.7 GHz. Chosen for single-thread: incremental `cargo` builds are latency-bound, and this beats a 32-core Milan at the thing done fifty times a day. 2× DDR5-5600, 28 PCIe Gen 5 lanes, 170 W class |
|
|
31 |
+ |
| Board | **ASRock Rack B650D4U** (base variant) | ~$400–450 | mATX server board. **AST2600 BMC** with IPMI 2.0, iKVM, vMedia, dedicated management LAN. Base variant chosen over -2L2T/-2L2T/BCM specifically because **it is the only one with two M.2 slots** |
|
|
32 |
+ |
| RAM | 2× 48 GB DDR5 ECC UDIMM (96 GB) | ~$450 | 1 DPC runs 5600; 2 DPC drops to 3600. Two slots stay empty on purpose |
|
|
33 |
+ |
| Storage 1 | 2 TB NVMe on M2_1 (PCIe 5.0 x4) | ~$180 | **xfs** root. Not ZFS — see the note in `mm-v1-bom.md` |
|
|
34 |
+ |
| Storage 2 | 4 TB NVMe on M2_2 (PCIe 4.0 x4) | ~$280 | Raw XFS scratch: kv-cache overflow later, container build scratch from day one |
|
|
35 |
+ |
| PSU | 850 W Platinum ATX | ~$150 | One 250–300 W card plus a 170 W CPU. Sized for the machine that exists, not a 4-GPU machine that does not |
|
|
36 |
+ |
| Chassis | 4U rackmount, ≥2 dual-slot bays (Sliger CX4712 or Rosewill RSV-L4500U) | $120–400 | Undecided. 4U for airflow, not for form-factor prototyping — see below |
|
|
37 |
+ |
| CPU cooler | AM5 tower, ≤160 mm | ~$100 | Air. Height limit is the 4U lid |
|
|
38 |
+ |
| Case fans | 4–6× 120/140 mm, high static pressure | ~$120 | Load-bearing for the bench role, see below |
|
|
39 |
+ |
| Boot GPU | **none** | $0 | The AST2600 has its own VGA and iKVM. `mm-v1-bom.md` carried a GT 1030 only because it was unsure the BMC would suffice; the manual settles it |
|
|
40 |
+ |
| **Total** | | **~$2,500–2,900** | vs ~$10,280 for the `mm-v1` host platform |
|
|
41 |
+ |
|
|
42 |
+ |
Vendor manuals for the decided parts are in `_private/docs/hardware/bench-v1/`.
|
|
43 |
+ |
|
|
44 |
+ |
## Why 4U, when a tower would be nicer to work in
|
|
45 |
+ |
|
|
46 |
+ |
A tower with a side panel is plainly better for swapping cards, and swapping cards is what
|
|
47 |
+ |
this machine does. It still loses.
|
|
48 |
+ |
|
|
49 |
+ |
Collection cards are datacenter pulls, and a Tesla P40 is **passively cooled** — no fan at
|
|
50 |
+ |
all. It cools only in a chassis with a front-to-back static-pressure path. A tower needs a
|
|
51 |
+ |
strapped-on blower per card, which is exactly the sort of ad-hoc rig that makes two bench
|
|
52 |
+ |
results incomparable. A fixed 4U airflow path means every card is tested in the same thermal
|
|
53 |
+ |
environment, which is the only thing that makes a Care Tag's thermal trace worth printing.
|
|
54 |
+ |
|
|
55 |
+ |
The case fans are therefore a test instrument, not cooling. Do not substitute quiet fans.
|
|
56 |
+ |
|
|
57 |
+ |
## The GPU ceiling, stated exactly
|
|
58 |
+ |
|
|
59 |
+ |
From the board manual, section 2.6:
|
|
60 |
+ |
|
|
61 |
+ |
| Slot | Gen | Mechanical | Electrical |
|
|
62 |
+ |
|---|---|---|---|
|
|
63 |
+ |
| PCIE6 | 5.0 | x16 | x16 |
|
|
64 |
+ |
| PCIE7 | 5.0 | **x4** | x4 |
|
|
65 |
+ |
| PCIE4 | 4.0 | x1 | x1 |
|
|
66 |
+ |
|
|
67 |
+ |
**One GPU seats natively.** PCIE7 is x4 *mechanically*, so a card does not physically fit;
|
|
68 |
+ |
the idea of parking a second Pascal card there is dead.
|
|
69 |
+ |
|
|
70 |
+ |
**Two GPUs are reachable via bifurcation.** BIOS exposes "Configure PCIE6 Link Width" with
|
|
71 |
+ |
`[x16]`, `[x8x8]`, `[x8x4x4]`. Gen5 x8 is roughly twice the bandwidth a Gen3 x16 card such
|
|
72 |
+ |
as a P40 can consume, so splitting costs those cards nothing. What it costs is a bifurcation
|
|
73 |
+ |
riser and a physical mounting problem: two dual-slot cards hanging off risers above an mATX
|
|
74 |
+ |
board in a 4U is a rig to solve, not a slot to populate. **Unproven until someone builds it.**
|
|
75 |
+ |
|
|
76 |
+ |
So against the EveryCycle roadmap:
|
|
77 |
+ |
|
|
78 |
+ |
- **A1** (1× P40) — fits natively.
|
|
79 |
+ |
- **A2** — fits, single card.
|
|
80 |
+ |
- **A3** (P40 + P100, mixed-arch in one executor) — needs the x8x8 riser rig. Reachable, not
|
|
81 |
+ |
guaranteed.
|
|
82 |
+ |
- **A4/A5** (3–4 cards) — **does not fit.** That is the replacement machine's job, and the
|
|
83 |
+ |
2000 W PSU in `mm-v1-bom.md` exists for exactly that.
|
|
84 |
+ |
- **B3, B5** (Qwen 8B, then 72B Q4 saturating a P40) — single card, fits.
|
|
85 |
+ |
- **B6** (frontier-size with CPU offload) — **does not fit.** Two DDR5 channels is ~80–90 GB/s
|
|
86 |
+ |
against ~358 GB/s for 8-channel DDR5-5600. CPU-offload work belongs to the replacement.
|
|
87 |
+ |
|
|
88 |
+ |
## What was deliberately not bought, and why
|
|
89 |
+ |
|
|
90 |
+ |
- **8-channel memory bandwidth.** The single real loss. It costs a $3,900 CPU tier plus
|
|
91 |
+ |
$2,500 of DDR5 on WRX90, and pays back only on B6, the last item on the model thread. Buy
|
|
92 |
+ |
it when B6 is close, on whatever platform is cheap by then.
|
|
93 |
+ |
- **PCIe 5.0 x16 slot count.** The roadmap's cards are Gen3 Pascal. Seven Gen5 x16 slots is
|
|
94 |
+ |
paying for signalling nothing in the plan can use.
|
|
95 |
+ |
- **512 GB RAM.** Justified in `mm-v1-bom.md` as "405B Q4 fits in RAM if ever wanted", which
|
|
96 |
+ |
is an aspiration priced at $2,500. This board tops out at 192 GB anyway, and past 96 GB the
|
|
97 |
+ |
memory clock drops from 5600 to 3600.
|
|
98 |
+ |
- **A 2000 W PSU and 4 GPU bays.** Sized for a machine this is explicitly not.
|
|
99 |
+ |
- **Used parts.** Deliberate, and the reason the bench is credible. The used parts are the
|
|
100 |
+ |
cards under test.
|
|
101 |
+ |
|
|
102 |
+ |
## Open before ordering
|
|
103 |
+ |
|
|
104 |
+ |
- **Does the board QVL cover DDR5 ECC UDIMM with EPYC 4005?** The manual hedges: "Memory
|
|
105 |
+ |
support is to be validated." ECC UDIMM plus a server board plus a not-quite-Ryzen CPU is
|
|
106 |
+ |
exactly where QVL matters. Resolve before buying DIMMs.
|
|
107 |
+ |
- **1 GbE vs 10 GbE.** The base variant is 2× 1 GbE. The existing LAN (192.168.1.x, consumer
|
|
108 |
+ |
gateway) is almost certainly 1 GbE, so 10 GbE would have nothing to talk to, and PCIE7's
|
|
109 |
+ |
Gen5 x4 can take a 10 GbE NIC later if that changes. Confirm the LAN before paying for
|
|
110 |
+ |
-2L2T.
|
|
111 |
+ |
- **OpenBMC.** `crates/bmc-agent` says it runs on "an OpenBMC fork on Aspeed AST2600-class
|
|
112 |
+ |
hardware". This board has the AST2600 but ships ASRock's vendor IPMI firmware, and vendor
|
|
113 |
+ |
BMC images are typically signed. If it cannot be replaced, `bmc-agent` retargets
|
|
114 |
+ |
IPMI/Redfish or wants a dedicated dev board. **This question is not specific to this
|
|
115 |
+ |
board — it applies to the WRX90 spec identically**, and answering it on a $400 board is
|
|
116 |
+ |
cheaper.
|
|
117 |
+ |
- **Chassis choice**, and with it whether an mATX board plus a bifurcation riser can actually
|
|
118 |
+ |
mount two dual-slot cards.
|
|
119 |
+ |
|
|
120 |
+ |
## BIOS at first boot
|
|
121 |
+ |
|
|
122 |
+ |
- **PCIE6 Link Width**: leave `[x16]` while there is one card. Revisit at A3.
|
|
123 |
+ |
- **Above 4G Decoding** and **Resizable BAR**: on. Harmless now, required by large-BAR
|
|
124 |
+ |
datacenter cards.
|
|
125 |
+ |
- **IOMMU**: on. Needed for podman CDI GPU passthrough.
|
|
126 |
+ |
- **Memory**: JEDEC DDR5-5600 at 1 DPC. Not pushed.
|
|
127 |
+ |
- Configure the **dedicated management LAN** and confirm iKVM works before the machine goes
|
|
128 |
+ |
anywhere it is inconvenient to reach.
|