Skip to main content

max / everycycle

Move the hardware BOMs out of the repo and into the wiki Both BOMs carried TailoredMachines commercial content: per-part pricing, sourcing channels and routes, and the House Cut Platform SKU framing sold to external customers. This repo is public on makenot.work, and git is the public layer while the wiki is not. That is the line the files were on the wrong side of, and it is TM business material rather than everycycle documentation either way. Now [[tm-bench-v1-bom]] and [[tm-mm-v1-bom]], copied verbatim with frontmatter and cross-references converted to wikilinks. tm-naming-decision's instruction to keep them in place is marked superseded so this does not get undone; it predates the repo being public. roadmap.md's two references now point at the wiki without restating BOM contents.
Co-Authored-By
Claude Opus 5 (1M context) <noreply@anthropic.com>
Author: Max Johnson <me@maxj.phd> · 2026-07-26 20:33 UTC
Signed with PGP, not checked
Commit: b22597e9c707ef287c09ec26f94a38e1bfe6df21
Parent: d98134f
3 files changed, +2 insertions, -505 deletions
M docs/roadmap.md +2 -2
@@ -97,7 +97,7 @@
97 97
98 98 ### Thread H — Appliance line (the business that funds the software)
99 99
100 - The hardware line is how the operation pays for EveryCycle's development. The first machine (BOM at `docs/hardware/mm-v1-bom.md` — filename retained; not renamed per the TailoredMachines naming decision) is in service as Sando host and EveryCycle dev box; MNW buys machines from this line for its own workloads at fair internal-transfer pricing.
100 + The hardware line is how the operation pays for EveryCycle's development. Machine BOMs, sourcing routes and pricing live in the TailoredMachines wiki rather than in this repo.
101 101
102 102 **Note (2026-07-10):** The hardware line was rebranded to **TailoredMachines** and narrowed to a small-batch atelier cutting workstations for **1–6 programmers doing bursty pair-programming and audit/fuzz workloads**. The product architecture is a single Platform SKU (v1: "The House Cut") plus three explicit GPU paths (Alterations / Off-the-Rack / The Collection). See the TailoredMachines plan and brand docs for the current framing. The tier list below (H1–H4, previously "MakeMachines") is historical — H3 (hobbyist edge) and H4 (rack appliance) fall outside TailoredMachines' scope entirely, and H1/H2 are collapsed into the House Cut framing.
103 103
@@ -162,7 +162,7 @@
162 162 - **Cross-vendor layer placement: belongs to A4.** The executor trait is forward-compatible from day one (activation-handoff signatures present even if no impl uses them); the inter-executor protocol gets built when A4 is reached.
163 163 - **Project structure: parallel threads, not numbered phases.** Each thread has its own pace; project state is the snapshot across all threads.
164 164 - **First GPU: Tesla P40.** Walks the defensive-curation talk on day one. Pascal facing imminent CUDA-legacy transition makes this card the canonical example of what EveryCycle exists to support.
165 - - **Host platform: dual-use with the TailoredMachines House Cut.** See `docs/hardware/mm-v1-bom.md` (filename retained per naming-decision). Threadripper Pro 7975WX + WRX90D8 + 512 GB ECC + Gen5 NVMe is the substrate; GPUs are fungible.
165 + - **Host platform: dual-use with the TailoredMachines line.** The substrate is a workstation-class host with ECC memory, a BMC, and Gen5 NVMe; GPUs are fungible above it. Exact BOMs live in the TailoredMachines wiki.
166 166 - **First model family: Qwen, starting small.** B1 is the smallest Qwen variant that loads; subsequent milestones step up. Choice driven by open weights, broad quant availability, and clean size laddering.
167 167 - **Client API stability is a sacred commitment.** Once Thread I reaches v1 freeze, the client-facing API follows strict semver: minor versions are additive, patch versions are bug-fix-only, major versions require a migration document and a deprecation window of at least one full release cycle on every supported distro. Before v1 freeze, the API is explicitly marked unstable and may break on any release. The freeze is the line that turns EveryCycle from a tool into a platform; everything downstream (compatibility shims, third-party clients, distro defaults) depends on this commitment being honored.
168 168 - **Multi-client is the default, not a feature.** The daemon assumes concurrent clients from A1, even when only one is connected. The internal data model, scheduler, and audit surface are designed around N clients sharing M devices under policy; the N=1 case is just the limit. This is the structural difference between a system service and a runtime library.
@@ -1,374 +1,0 @@
1 - # Control Bench BOM (bench-v1)
2 -
3 - Settled 2026-07-25. The machine actually being bought: an **all-new** host that is the
4 - always-on x86_64 build machine now and a single-to-dual-GPU inference box later.
5 -
6 - This is not `mm-v1-bom.md`. That file is the aspirational spec, kept for what it is worth;
7 - this is the one with parts in it.
8 -
9 - ## What this machine is for, in priority order
10 -
11 - 1. **Always-on build host, fire-and-forget.** Alloy images and ISOs, the MNW build gate,
12 - scheduled rebuilds. The ecosystem has no always-on x86_64 machine: astra is always-on and
13 - aarch64, fw13 is x86_64 and sleeps, and production never builds binaries. Nobody waits on
14 - this machine interactively — the edit/compile loop stays in Helix on fw13 — so it is sized
15 - for **throughput, not latency**.
16 -
17 - **Sando does not live here.** The original BOM assigned this box the Sando host role back
18 - when it was going to be one big always-on machine. That does not survive the box becoming
19 - a card-swapping bench: a production deploy controller should not share a chassis with
20 - unproven hardware, and a single non-redundant root NVMe makes it worse. Sando moves to
21 - **astra** — always-on, never opened, and arch-agnostic for a controller.
22 - 2. **The known-good control for testing used parts one at a time.** Every part here is new.
23 - That is the point: when a used Collection card misbehaves on the bench, the host is not a
24 - suspect. One variable at a time, against a reference that does not move.
25 - 3. **Single-GPU (and, via bifurcation, dual-GPU) inference work.** EveryCycle roadmap A1–A2
26 - and B3–B5 all run on one card. See the ceiling section for where it stops.
27 -
28 - It is also **explicitly a stepping stone**: this machine gets sold and replaced before
29 - anything more powerful is sold to a customer. So overbuying for a future workload is not
30 - prudence here, it is waste. Buy the smallest credible control; put the money into cards.
31 -
32 - ## Component list
33 -
34 - Prices below are a **summary observed 2026-07-25**, kept here for readability. They are not the
35 - source of truth: every observation lives in the `tm-mcp` store (`_private/tm/pricing.db`, tables
36 - `asks` and `sells`) with its route, condition and date. Ask it rather than trusting this column —
37 - `compare <sku>` — and record the real paid number there at purchase. Analysis of what the numbers
38 - mean is in `_private/docs/hardware/bench-v1/price-baseline.md`.
39 -
40 - | Component | Choice | Approx | Notes |
41 - |---|---|---|---|
42 - | CPU | **AMD EPYC 4565P** (16c/32t "Zen 5", AM5) | **$589** list | 4.30 base / 5.70 boost, 64 MB L3, 170 W, 2× DDR5-5600, 28 PCIe Gen 5 lanes. Sized for throughput: clean builds, container builds, the ISO's zstd pass. See the CPU note below for why this and not a Ryzen 9950X. **Ceiling: EPYC 4005 stops at 16 cores**, so if throughput ever binds the answer is a new machine, not a new CPU — acceptable for a declared stepping stone |
43 - | Board | **ASRock Rack B650D4U** (base variant) | ~$420 | mATX server board. **AST2600 BMC** with IPMI 2.0, iKVM, vMedia, dedicated management LAN. Base variant chosen over -2L2T/-2L2T/BCM because **it is the only one with two M.2 slots**; see the networking note for why its 1 GbE is not a compromise |
44 - | RAM | **2× 32 GB Micron MTC20C2085S1EC56BD1** (64 GB) | ~$430 | DDR5-5600 ECC UDIMM, 2Rx8. **On the board QVL**, and the only 32 GB ECC part on it at 5600. 1 DPC runs 5600; 2 DPC drops to 3600, so two slots stay empty on purpose. The 48 GB tier is QVL-reachable too (SK Hynix HMCGY8MGBEB213N) but roughly doubles the line for capacity 16 cores do not need — see the memory section |
45 - | Storage 1 | 4 TB NVMe on **M2_1** (PCIe 5.0 ×4) | ~$300 | Raw XFS **scratch**: kv-cache overflow later, container build scratch from day one. The fast slot goes to the only thing whose point is bandwidth. Buy a Gen4 drive now while NAND is expensive; the Gen5 slot is headroom |
46 - | Storage 2 | 2 TB NVMe on **M2_2** (PCIe 4.0 ×4) | ~$200 | **xfs root**, and not ZFS — see the note in `mm-v1-bom.md`. Root gains nothing from Gen5 |
47 - | PSU | **1200 W Platinum, ATX 3.1 with native 12V-2x6** | ~$300 | Was 850 W, which contradicted this document's own claim that two cards are reachable: 2× 250 W passive + a 170 W CPU is ~750 W, or 88% load. 1200 W also covers one modern ~450 W card. **ATX 3.1 is deliberate**: it handles the transient spikes modern cards produce and provides 12V-2x6 natively, so no adapters when the current-gen vertical gets tested. The 500–600 W-class vertical triggers a PSU review |
48 - | Enclosure | **Open frame** — see the enclosure section | ~$150–250 | Optimized for access. Airflow is solved at the card, not by the case |
49 - | CPU cooler | AM5 tower, any height | ~$100 | Air. No lid to clear on an open frame |
50 - | Card airflow | Per-card **test sled**: server blower + shroud, per vertical | ~$150 setup | The actual thermal instrument. See the enclosure section |
51 - | Boot GPU | **none** | $0 | The AST2600 has its own VGA and iKVM. `mm-v1-bom.md` carried a GT 1030 only because it was unsure the BMC would suffice; the manual settles it |
52 - | **Total** | | **~$2,680** | vs ~$10,280 for the `mm-v1` host platform. Line-by-line observations, spreads and market notes: `_private/docs/hardware/bench-v1/price-baseline.md` |
53 -
54 - Vendor manuals for the decided parts are in `_private/docs/hardware/bench-v1/`.
55 -
56 - ## CPU: why the EPYC and not a Ryzen 9950X
57 -
58 - Decided on merit rather than price (Max: "$150 is marginal, whichever is the better part").
59 - From AMD's EPYC 4005 datasheet, held locally, against the 9950X:
60 -
61 - | | Cores | Base | Boost | TDP | L3 | Channels | Gen 5 lanes |
62 - |---|---|---|---|---|---|---|---|
63 - | EPYC 4565P | 16/32 | 4.30 | 5.70 | 170 W | 64 MB | 2× DDR5-5600 | 28 |
64 - | Ryzen 9950X | 16/32 | 4.30 | 5.70 | 170 W | 64 MB | 2× DDR5-5600 | 28 |
65 -
66 - Same silicon, same numbers. **There is no performance argument either way**, so this comes
67 - down to everything else, and the deciding factor is **officially validated ECC UDIMM**.
68 -
69 - Ryzen 9000 ECC works on boards that wire it up, and ASRock Rack does — but board-vendor
70 - support is not CPU-vendor validation. This machine exists to be a trustworthy reference; its
71 - characterization gate below requires ECC counters that are both correct and readable; and a
72 - silent memory error corrupting a build artifact is exactly the failure ECC is bought to
73 - prevent. Validated beats de-facto for that job.
74 -
75 - Secondary, both real: the embedded/server line carries a longer supply and firmware-support
76 - window, and ASRock Rack's BIOS and QVL work on this board targets EPYC 4004/4005 first. And
77 - since this file doubles as a TM product template, "server CPU with validated ECC on a server
78 - board with a BMC" is coherent to sell under a Care Policy in a way that a consumer part with
79 - unofficial ECC is not.
80 -
81 - The 9950X's only genuine edge is its iGPU, which the AST2600's own VGA already covers.
82 - **Fallback:** the EPYC 4005 is a niche part, and if availability is bad the 9950X drops into
83 - this board with zero performance change. Preference with a documented alternative, not a hard
84 - requirement.
85 -
86 - **Considered and rejected: the 4545P.** Also 16c/32t Zen 5, but 65 W instead of 170 W
87 - (3.0 base / 5.4 boost) and cheaper — tempting for an always-on box in a living space. Rejected
88 - because a mostly-idle CI host spends most of its hours at **idle**, where both parts draw
89 - about the same. The 170 W part only draws more while actually building, and it finishes
90 - sooner, so energy per build is roughly a wash while wall-clock is materially better. If heat
91 - or fan noise later becomes the actual complaint, the 4545P is the drop-in answer.
92 -
93 - **Considered and rejected: the 4585PX** (16c, 128 MB via 3D V-Cache). Clocks are listed TBD in
94 - the datasheet, so it may not be shipping; compilation gains from doubled L3 are modest, and
95 - X3D parts usually trade all-core clock for the cache.
96 -
97 - ## Memory: the QVL settles it, and it is not the part that was priced
98 -
99 - The manual's "Memory support is to be validated" hedge was the last hard blocker on ordering.
100 - It is answered: ASRock Rack publishes a **22-row memory QVL** for this board, pulled
101 - 2026-07-25. Local copy: `_private/docs/hardware/bench-v1/b650d4u-memory-qvl.md`.
102 -
103 - **What it says, for the two things that were actually in doubt:**
104 -
105 - - **Dual-rank ECC UDIMM at 5600 is validated.** Both of the board's 5600-grade ECC entries
106 - are 2Rx8. So the worry that "ECC UDIMM + a server board + a not-quite-Ryzen CPU" would fall
107 - outside validated territory does not survive the list. Nothing about dual-rank is the risk
108 - it was assumed to be.
109 - - **48 GB at 5600 is on the list**, so the manual's 48 GB-per-DIMM ceiling is real capacity
110 - and not just a spec-sheet maximum.
111 -
112 - **Only three ECC parts on the QVL run at 5600, and each capacity has exactly one:**
113 -
114 - | Size | Vendor | Part | Rank |
115 - |---|---|---|---|
116 - | 48 GB | SK Hynix | HMCGY8MGBEB213N | 2Rx8 |
117 - | 32 GB | Micron | MTC20C2085S1EC56BD1 | 2Rx8 |
118 - | 16 GB | Micron | MTC10C1084S1EC56BD1 | 1Rx8 |
119 -
120 - Everything else on the list is 4800 or 5200. Buying off-QVL is therefore not a small
121 - compromise here: it costs a third of the memory clock, which is the same penalty as filling
122 - all four slots.
123 -
124 - **The Kingston KSM56E46BD8KM-48HM is not on the QVL.** That is the part the $270–$1,006
125 - spread was found on, and it was the presumed buy. Kingston appears on the list exactly once
126 - in ECC, as `KSM48E40BD8KM-32HM`, which is a **4800** part. Kingston's own compatibility tool
127 - does list the 48 GB module against this board — but that is vendor self-certification, not
128 - board-vendor validation, and this file already decided that distinction in the CPU section.
129 - Having chosen the EPYC over the 9950X specifically to get validated ECC, buying an off-QVL
130 - DIMM would spend that choice for nothing.
131 -
132 - **Take 2× 32 GB Micron, not 2× 48 GB Hynix.** Both are QVL and both run 5600 at 1 DPC, so
133 - this is purely capacity against cost — roughly **$430 for 64 GB versus ~$990 for 96 GB**. The
134 - BOM already held that 64 GB is enough for 16 cores, the machine is a declared stepping stone,
135 - and memory is in an up-cycle where deferring spend is the cheaper bet. The two free DIMM
136 - slots remain the escape hatch, at the 3600 penalty.
137 -
138 - One channel note: `HMCGY8MGBEB213N` is an OEM part number, so retail carries rebrands (Axiom
139 - `AX55600E46I/48G` at ~$868) rather than the SK Hynix module itself. If the 48 GB tier is ever
140 - wanted, that gap is part of its cost.
141 -
142 - ## Enclosure: open frame, and airflow solved at the card
143 -
144 - The earlier draft of this file argued for a 4U rackmount because a passive card "only cools
145 - in a chassis with a front-to-back static-pressure path." **That argument does not survive
146 - scrutiny.** `mm-v1-bom.md` itself budgets a fan shroud for the P40 *inside* the 4U, and it is
147 - right to: passive datacenter cards are engineered for 1U/2U server fans producing static
148 - pressure that no 140 mm case fan approaches. The shroud is needed either way, so the 4U was
149 - buying repeatability, not cooling — a weaker claim than the one it was sold on.
150 -
151 - Once noise is not a constraint, there is a better answer available, and it inverts the
152 - problem.
153 -
154 - **Airflow is solved at the card, with real server blowers.** 40–80 mm server blowers at
155 - 10k+ RPM produce the static pressure these cards were designed around. That is a thing a
156 - quiet build physically cannot do, and it is the single biggest quality difference in a bench
157 - test of a passive card. With sound deprioritized, use them.
158 -
159 - **Once airflow lives at the card, the enclosure only has to be good to work in.** So:
160 -
161 - - **Open aluminium frame** (mining-rig style, or an open bench-frame such as a Core P3 /
162 - BC1-class fixture). Nothing to unscrew, nothing to unrack, no lid, no cable gymnastics.
163 - Cards go in and out in seconds, which is the operation performed most.
164 - - **Unlimited card clearance.** Modern 3–3.5 slot, 350 mm cards fit trivially. A 4U would
165 - have constrained exactly the vertical most likely to need bench time later.
166 - - **Everything visible.** You can see fan spin, LEDs, and scorch marks on an unknown card
167 - before they become a smell.
168 -
169 - ### The test sled is the instrument
170 -
171 - Repeatability comes from a fixture, not a case. Build one **test sled** per card class: a
172 - rigid bracket that holds the card, its shroud, and its blower at a **fixed** geometry, with
173 - the blower on a controller at a **fixed, recorded RPM**. Every card in that class is then
174 - tested in an identical thermal environment regardless of what is around it — better
175 - repeatability than a shared case gives, because a case's airflow changes with every other
176 - card and cable in it.
177 -
178 - Consequences to hold:
179 -
180 - - **One sled per vertical.** A passive P40, a 2-fan consumer card and a 3.5-slot current-gen
181 - card have nothing thermally in common. Sleds get built as verticals open, not up front.
182 - - **Log RPM and ambient with every test.** A thermal trace without them is not comparable.
183 - Until `bmc-agent` exists (and it may not — task `59335767`), blower RPM is set on a manual
184 - fan controller and written down, and BIOS-set chassis fan curves are recorded.
185 - - **Dust is the accepted cost.** An open frame running 24/7 in a living space collects dust
186 - in the CPU cooler and PSU intake. Mitigated by the blowers only running during tests, and
187 - by putting periodic cleaning in the standing ops list. This is a real trade and the reason
188 - a closed case would otherwise win.
189 - - **It is not shippable as-is.** When this machine is sold, it either gets rehoused in a case
190 - or sold as a parts bundle. Budget that, and do not let the frame make the machine
191 - unsellable by surprise.
192 -
193 - ## The GPU ceiling, stated exactly
194 -
195 - From the board manual, section 2.6:
196 -
197 - | Slot | Gen | Mechanical | Electrical |
198 - |---|---|---|---|
199 - | PCIE6 | 5.0 | x16 | x16 |
200 - | PCIE7 | 5.0 | **x4** | x4 |
201 - | PCIE4 | 4.0 | x1 | x1 |
202 -
203 - **One GPU seats natively.** PCIE7 is x4 *mechanically*, so a card does not physically fit;
204 - the idea of parking a second Pascal card there is dead.
205 -
206 - **Two GPUs are reachable via bifurcation.** BIOS exposes "Configure PCIE6 Link Width" with
207 - `[x16]`, `[x8x8]`, `[x8x4x4]`. Gen5 x8 is roughly twice the bandwidth a Gen3 x16 card such
208 - as a P40 can consume, so splitting costs those cards nothing. What it costs is a bifurcation
209 - riser and a physical mounting problem: two dual-slot cards on risers off an mATX board is a
210 - rig to solve, not a slot to populate. The open frame helps here — a frame is a mounting
211 - surface, where a case would have been a constraint. **Unproven until someone builds it.**
212 -
213 - So against the EveryCycle roadmap:
214 -
215 - - **A1** (1× P40) — fits natively.
216 - - **A2** — fits, single card.
217 - - **A3** (P40 + P100, mixed-arch in one executor) — needs the x8x8 riser rig. Reachable, not
218 - guaranteed.
219 - - **A4/A5** (3–4 cards) — **does not fit.** That is the replacement machine's job, and the
220 - 2000 W PSU in `mm-v1-bom.md` exists for exactly that.
221 - - **B3, B5** (Qwen 8B, then 72B Q4 saturating a P40) — single card, fits.
222 - - **B6** (frontier-size with CPU offload) — **does not fit.** Two DDR5 channels is ~80–90 GB/s
223 - against ~358 GB/s for 8-channel DDR5-5600. CPU-offload work belongs to the replacement.
224 -
225 - ## Networking: 1 GbE now, 10 GbE deferred into PCIE7
226 -
227 - The board comes in three variants and the choice is a three-way trade between M.2 slots,
228 - 10 GbE and cost. Only the base `B650D4U` has **two** M.2 slots; both 10 GbE variants have one,
229 - because the 10 GbE controller eats the FCH lanes M2_2 would use.
230 -
231 - Measured on the LAN 2026-07-25 rather than assumed:
232 -
233 - - **fw13**: USB adapter, linked at **1000 Mb/s**. A Framework 13 has no easy 10 GbE path
234 - short of a Thunderbolt adapter.
235 - - **astra**: `enP3p3s0f0` and `f1` are on the **ixgbe** driver, so that is Intel 10 GbE
236 - silicon — but `f0` is linked at **1000 Mb/s** and `f1` is down.
237 -
238 - So 10 GbE hardware exists on the LAN and nothing is actually running at 10 GbE. More to the
239 - point, the traffic that would justify it is moving multi-GB ISOs and images to wherever a USB
240 - stick gets written, which is fw13, which is 1 GbE-limited regardless.
241 -
242 - **Decision: take the base board, keep both native M.2, and leave PCIE7 empty.** If a 10 GbE
243 - path later proves worth having, PCIE7 (Gen 5 ×4) takes a NIC with room to spare — 10 GbE needs
244 - a fraction of that. Deferring costs nothing and avoids paying a board premium plus an M.2
245 - adapter for a link that currently has no peer running at speed.
246 -
247 - **Constraint on that later NIC: it must be ×4 *mechanical*.** PCIE7 is ×4 mechanically, so an
248 - Intel X550-T2 (×4) fits and an X710-T2L (×8) does not. An X550 also pairs with the same
249 - `ixgbe` driver astra already runs, which is a known-good story rather than a new one.
250 -
251 - ## The bench is not a control until it is characterized
252 -
253 - "All-new means known-good" is the right instinct with the wrong mechanism. New parts have
254 - infant mortality and firmware bugs; a used part with hours on it can be *more* reliable than
255 - a fresh one. What new actually buys is an **RMA path** — when a host part does fail, it gets
256 - replaced rather than debugged, and it never becomes a suspect in a card's test result. That
257 - holds up. The "known-good" part does not, until it is measured.
258 -
259 - **Before the first used card goes in**, record a baseline and keep it in the runbook:
260 -
261 - - memtest86+ clean pass over all installed DIMMs.
262 - - Sustained all-core `stress-ng` run: clocks held, package temp, no throttle.
263 - - `fio` on both NVMes: throughput and latency, so a future slow build has a reference.
264 - - Idle and full-load wall power, and the ambient temperature both were taken at.
265 - - ECC corrected-error count at zero, **and proof you can read it** — see below.
266 -
267 - **Designate a reference card per vertical.** To tell "this card is bad" from "the slot, riser,
268 - PSU, sled or BIOS is bad" you need a known-good card to swap in. Buying a new one in each
269 - class is not worth it. Instead: **the first used card that passes cleanly in a vertical
270 - becomes that vertical's reference** — its numbers are recorded, and it is not sold. Cheap,
271 - and it makes every later test in that class a comparison rather than an absolute.
272 -
273 - **ECC you cannot read is not ECC.** Rising correctable-error counts are the early warning the
274 - memory is paid for, and Alloy ships no `rasdaemon`, `edac-utils`, `smartmontools` or
275 - `ipmitool`. Those belong in the server profile (GoingsOn alloy task `d1fed0d7`) and the
276 - baseline above cannot be taken without them.
277 -
278 - ## Card verticals, opened one at a time
279 -
280 - The Collection wants at least one cut in several verticals, and cannot fund testing all of
281 - them at once. So verticals open in sequence, and each one opens a small equipment bill:
282 -
283 - | Vertical | Bench implications |
284 - |---|---|
285 - | **Passive datacenter pulls** (P40, P100, MI50) | Shroud + server blower sled. Some use CPU-style EPS 8-pin, not PCIe. Large BARs need Above 4G + ReBAR. First, because the roadmap names it and the cards are cheapest |
286 - | **Used consumer / gaming** | Self-cooling, standard PCIe power, trivial sled. Needs a monitor on the card under test, since display output is part of what is being sold |
287 - | **B-stock / open-box current-gen** | Standard cooling, may retain manufacturer warranty. Bench time still required to prove it |
288 - | **Modern high-TDP** (500–600 W class) | 12V-2x6 native (the ATX 3.1 PSU covers this), real transient spikes, 3.5 slots, 350 mm. **Triggers a PSU review past 1200 W** and its own sled |
289 -
290 - Sequence the verticals; do not buy sleds or PSU headroom for a vertical that is not open yet.
291 -
292 - ## What was deliberately not bought, and why
293 -
294 - - **8-channel memory bandwidth.** The single real loss. It costs a $3,900 CPU tier plus
295 - $2,500 of DDR5 on WRX90, and pays back only on B6, the last item on the model thread. Buy
296 - it when B6 is close, on whatever platform is cheap by then.
297 - - **PCIe 5.0 x16 slot count.** The roadmap's cards are Gen3 Pascal. Seven Gen5 x16 slots is
298 - paying for signalling nothing in the plan can use.
299 - - **512 GB RAM.** Justified in `mm-v1-bom.md` as "405B Q4 fits in RAM if ever wanted", which
300 - is an aspiration priced at $2,500. This board tops out at 192 GB anyway, and past 96 GB the
301 - memory clock drops from 5600 to 3600.
302 - - **A 2000 W PSU and 4 GPU bays.** Sized for a machine this is explicitly not.
303 - - **Used parts.** Deliberate, and the reason the bench is credible. The used parts are the
304 - cards under test.
305 -
306 - ## Sourcing channels
307 -
308 - All-new parts, so these are retail and distribution channels, not the used-GPU channels in the
309 - TailoredMachines sourcing playbook.
310 -
311 - | Part | Channels | Watch for |
312 - |---|---|---|
313 - | EPYC 4565P | ShopBLT, Provantage, NeweggBusiness, Wiredzone, CDW | Niche server SKU, so **availability is the risk, not price**. Check **tray vs boxed**: tray is OEM, ships without a cooler, and the warranty runs through the reseller |
314 - | ASRock Rack B650D4U | NeweggBusiness, Newegg, Wiredzone, ServerSupply, Provantage | Confirm it is the **base** variant, not -2L2T or -2L2T/BCM — only the base has two M.2 slots |
315 - | DDR5 ECC UDIMM | Walmart, Amazon, Memory.net, ServerSupply, OEMPCWorld, eBay | **Buy the QVL part number and nothing else**: `MTC20C2085S1EC56BD1`. Same-part asks ran $215 to $2,270 on one day, so the channel matters more than the part does. Watch for the `...BD1R` suffix, which is the same module in a different retail pack |
316 - | 2× NVMe | Microcenter in person, Newegg, B&H | Microcenter in-store usually wins on price and removes shipping risk |
317 - | 1200 W ATX 3.1 PSU | Microcenter, Newegg | Seasonic PRIME/Vertex, Corsair, Super Flower. Confirm **native 12V-2x6**, not an adapter in the box |
318 - | Open frame | Mining-frame vendors (Veddha, Kingwin), or **8020 aluminium extrusion** for a custom build | Extrusion lets the sled geometry be designed in rather than worked around |
319 - | Server blowers, shrouds, fan controller | eBay/Amazon for Delta and San Ace 40–80 mm; shrouds printed or bought | Bench equipment, bought per vertical rather than up front |
320 -
321 - **Microcenter Denver is already a standing trip** — the sourcing playbook has Saturday mornings
322 - there for open-box GPUs. Commodity parts on the same trip means no shipping risk and no wait.
323 -
324 - ## Open before ordering
325 -
326 - - ~~Does the board QVL cover DDR5 ECC UDIMM?~~ **settled 2026-07-25 by pulling the QVL** —
327 - see the memory section. Dual-rank ECC UDIMM at 5600 is validated; the part is Micron
328 - `MTC20C2085S1EC56BD1`, not the Kingston module that was priced. **Caveat: the QVL carries no
329 - CPU column**, and its weighting toward 4800 parts suggests it was built in the EPYC
330 - 4004/Ryzen 7000 era. So it validates the *board*, and "with an EPYC 4005 specifically"
331 - remains unproven until the characterization gate runs. That is a bring-up risk, not an
332 - ordering blocker.
333 - - ~~1 GbE vs 10 GbE~~ **settled 2026-07-25 by measuring the LAN** — see the networking
334 - section. Base board, PCIE7 left empty for a future ×4-mechanical NIC.
335 - - **OpenBMC.** `crates/bmc-agent` says it runs on "an OpenBMC fork on Aspeed AST2600-class
336 - hardware". This board has the AST2600 but ships ASRock's vendor IPMI firmware, and vendor
337 - BMC images are typically signed. If it cannot be replaced, `bmc-agent` retargets
338 - IPMI/Redfish or wants a dedicated dev board. **This question is not specific to this
339 - board — it applies to the WRX90 spec identically**, and answering it on a $400 board is
340 - cheaper.
341 - - ~~48 GB ECC UDIMM availability and price~~ **settled 2026-07-25** — 48 GB is real and
342 - QVL-listed, but it is an OEM part sold through rebrands at roughly double the 32 GB tier.
343 - Decision: 2× 32 GB. See the memory section.
344 - - **Where the frame physically sits.** There is no rack, and an open frame in a living space
345 - has different constraints than a case: dust, curiosity, and anything that can be knocked
346 - into it. Pick the spot before buying the frame.
347 - - **The x8x8 riser rig**, which an open frame makes easier than any case would: two dual-slot
348 - cards on risers need mounting, and a frame is a mounting surface.
349 -
350 - ## Sequence the bring-up
351 -
352 - Too many unknowns arrive at once otherwise: a new hardware platform, Alloy in a role it has
353 - never held, no settled sshd story, bench duty, and CI. Add one at a time — the same
354 - discipline the bench itself exists to enforce.
355 -
356 - 1. Bare bring-up on the frame. BIOS, BMC, iKVM, management LAN.
357 - 2. OS install, and find out what `alloy install` does on non-laptop hardware.
358 - 3. Characterization baseline (above). This is the gate that makes it a control.
359 - 4. Build load: images, ISOs, the MNW gate.
360 - 5. First used card, in a sled, with the reference procedure.
361 -
362 - `/var/lib/containers` mounts on the 4 TB scratch, not on root. On a bootc host `/var` is
363 - stateroot, so this is a deliberate step; get it wrong and every container build lands on the
364 - 2 TB root.
365 -
366 - ## BIOS at first boot
367 -
368 - - **PCIE6 Link Width**: leave `[x16]` while there is one card. Revisit at A3.
369 - - **Above 4G Decoding** and **Resizable BAR**: on. Harmless now, required by large-BAR
370 - datacenter cards.
371 - - **IOMMU**: on. Needed for podman CDI GPU passthrough.
372 - - **Memory**: JEDEC DDR5-5600 at 1 DPC. Not pushed.
373 - - Configure the **dedicated management LAN** and confirm iKVM works before the machine goes
374 - anywhere it is inconvenient to reach.
@@ -1,129 +1,0 @@
1 - # Reference Bench BOM (aspirational spec)
2 -
3 - > **Status, 2026-07-25: this is not the machine being bought.** It is the aspirational
4 - > spec — the right host for roadmap **A4/A5** (3–4 GPUs) and **B6** (frontier-size with CPU
5 - > offload), and nothing else in the plan needs it. The machine actually being built is
6 - > `bench-v1-bom.md`: all-new parts, ~$2,500–2,900 against this one's ~$10,280 host platform,
7 - > serving as the always-on build host and the known-good control bench for testing used
8 - > Collection cards one at a time.
9 - >
10 - > What that re-examination found, so it is not re-litigated from scratch:
11 - >
12 - > - The **$3,900 CPU is not a core-count purchase.** AMD's Threadripper Pro 7000 design
13 - > document tiers memory speed: 7945WX/7955WX/7965WX are rated DDR5-4800, and 7975WX is the
14 - > cheapest SKU on the DDR5-5600 side. Dropping a core tier silently derates the 8-channel
15 - > bandwidth this whole platform exists for. There is no cheap way down this ladder.
16 - > - The **7× PCIe 5.0 ×16 slots do not serve the roadmap's cards.** A1–A3 are Tesla P40 and
17 - > P100, which are **PCIe 3.0**. A Gen3 ×16 card consumes the bandwidth of four Gen5 lanes.
18 - > - **8-channel bandwidth pays back only at B6**, the last item on the model thread.
19 - > - Note also AMD's own product page lists the 7975WX's memory at up to **5200** MT/s while
20 - > the design document says 5600. This file says "DDR5-5600 (JEDEC)"; that disagreement wants
21 - > settling before anyone buys DIMMs for this spec.
22 - >
23 - > Also see the ZFS-root note below: it applies to any unit running a bootc OS.
24 -
25 - Originally settled 2026-05-23. Top-of-line host platform; GPUs are fungible and live on the EveryCycle GPU thread (see `docs/roadmap.md`).
26 -
27 - The substrate is built once and kept stable; GPU experimentation happens above it without revisiting motherboard, CPU, or RAM.
28 -
29 - **This BOM serves two roles.** It is the personal EveryCycle dev box and Sando host, and it is the basis for the **TailoredMachines House Cut** Platform SKU sold to external customers (see the TailoredMachines plan and product-shape docs). The hardware spec is the same in both roles. Filename retained as `mm-v1-bom.md` per the TailoredMachines naming decision — the pre-rebrand name (MakeMachine) is preserved in the path to avoid churning cross-references.
30 -
31 - ## Component list
32 -
33 - | Component | Choice | Approx cost | Notes |
34 - |---|---|---|---|
35 - | CPU | AMD Threadripper Pro 7975WX | $3,900 | 32 cores, 5.3 GHz boost, 8-channel DDR5. Same memory controller and PCIe lanes as bigger SKUs; cores past ~32 starve on memory bandwidth for inference. |
36 - | Motherboard | ASRock Rack WRX90D8-2L/2T | $1,200 | 7× PCIe 5.0 ×16, 8 DIMM slots, **ASPEED AST2600 BMC (OpenBMC-friendly)**, dual 10 GbE, 4× M.2 + 4× SlimSAS. |
37 - | RAM | 8× 64 GB DDR5-5600 ECC RDIMM (512 GB total) | $2,500 | One DIMM per channel — 2 DIMMs per channel forces DDR5 down to ~4400 MT/s on WRX90, costing ~20% memory bandwidth (real impact on CPU-offload layers). Brand: Micron or Hynix off the board QVL. |
38 - | Storage 1 | 4 TB Gen5 NVMe (Samsung 9100 Pro or Crucial T705) | $600 | ZFS root pool: OS, models, Sando state, logs. **Read the ZFS-root note below before partitioning.** |
39 - | Storage 2 | 4 TB Gen5 NVMe (same) | $600 | **Raw XFS kv-scratch.** Not ZFS — ZFS caps Gen5 throughput; scratch is by definition disposable. EveryCycle uses this for kv-cache overflow. |
40 - | PSU | 2000 W Titanium-class (Super Flower Leadex Titanium or equivalent) | $600 | Headroom for any GPU combination the loose-parts experimentation hits. |
41 - | Chassis | 4U rackmount with ≥4 dual-slot GPU bays (Sliger CX4712 or similar) | $400 | Matches the eventual EveryCycle reference inference box; rackable from day one. |
42 - | CPU cooler | Silverstone XE360-TR5 (air) or Noctua NH-U14S TR5-SP6 | $200 | Air, not AIO — pump failure on a 24/7 box is worse than fan failure. |
43 - | Case fans | 6× 140 mm Noctua industrial | $200 | Front-to-back airflow; GPUs in a 4U breathe through these. |
44 - | Boot GPU | Nvidia GT 1030 (low-profile) | $80 | WRX90 has no iGPU; need a tiny card for console/boot. Also used as console GPU when datacenter cards (no display output) are installed. GT 710 originally specced but effectively EOL retail in 2026; GT 1030 is the current floor. Alternative: skip entirely if WRX90 BMC serial-over-LAN proves reliable. |
45 - | **Host subtotal** | | **~$10,280** | Before any compute GPU. |
46 -
47 - ## First compute GPU (Thread A1)
48 -
49 - | Component | Choice | Approx cost | Notes |
50 - |---|---|---|---|
51 - | GPU | 1× Nvidia Tesla P40 | $200 | 24 GB GDDR5, Pascal (cc 6.1). On-thesis: Pascal is the next architecture facing CUDA-legacy transition. |
52 - | Cooling adapter | 3D-printed or commercial fan shroud for P40 | $25 | Server card; passive cooling needs chassis airflow OR a strapped-on fan. |
53 - | Power adapter | EPS 8-pin to PCIe (or proper EPS routing) | $10 | P40 uses CPU-style 8-pin, not standard PCIe 8-pin. |
54 - | **GPU subtotal at A1** | | **~$235** | |
55 -
56 - ## Total
57 -
58 - | | |
59 - |---|---|
60 - | Host platform | ~$10,280 |
61 - | First GPU (A1) | ~$235 |
62 - | **Reference Bench v0 / House Cut v0 (assembled, runnable)** | **~$10,515** |
63 -
64 - Budget originally specced at $14–16K for the previous spec. Net savings: ~$4–5K. The savings are the GPU experimentation budget for advancing the GPU thread through A4–A5 over time.
65 -
66 - ## BIOS settings worth confirming at first boot
67 -
68 - - **Above 4G Decoding: enabled.** Required for datacenter cards with large BARs (Tesla P40, MI50, etc.).
69 - - **Resizable BAR: enabled.** Same reason.
70 - - **IOMMU: enabled.** Required for podman + CDI GPU passthrough.
71 - - **SR-IOV: as needed.** Not critical at A1; revisit at A5.
72 - - **Memory speed: DDR5-5600 (JEDEC).** Not pushed past spec.
73 - - **PCIe lane bifurcation: leave default initially; revisit for multi-GPU configurations.**
74 -
75 - ## Why this spec, summarized
76 -
77 - - **Threadripper Pro WRX90 platform** chosen over consumer TRX50 for the 7× PCIe 5.0 ×16 slots and the AST2600 BMC. EveryCycle's `bmc-agent` crate needs a real BMC to talk to.
78 - - **512 GB RAM** lets a 405B Q4 model fit fully in RAM if ever wanted; the immediate need is comfortable headroom for kv-cache experiments.
79 - - **Two Gen5 NVMes with different filesystems** because ZFS root is the right answer for the OS and model store, but ZFS would cap Gen5 bandwidth on the kv-scratch where bandwidth is the whole point. **The first half of that no longer holds if the unit runs a bootc OS — see below.**
80 -
81 - ### ZFS root does not survive a bootc OS
82 -
83 - Added 2026-07-25, when unit 1 was assigned the always-on build-host role and Alloy was picked as its OS (wiki `tm-first-unit-build-host`, GoingsOn alloy task `d1fed0d7`).
84 -
85 - A bootc image-based OS puts an **ostree** root on the boot disk, and ostree-on-ZFS is not a supported combination. There is also no ZFS anywhere in the Alloy image: on Fedora it is an out-of-tree kmod, and an immutable image is the worst place to carry one.
86 -
87 - So on any unit running Alloy (or Silverblue, or CoreOS, or anything else bootc):
88 -
89 - - **Storage 1 root becomes xfs**, the bootc default. Do not spend bring-up time building a ZFS root pool; it will not boot.
90 - - **Storage 2 is unaffected.** Raw XFS kv-scratch was already the answer and the reasoning still stands.
91 - - ZFS survives only as an optional **data** pool on additional disks, if something later actually wants snapshots and checksums for the model store. Nothing does today.
92 -
93 - This is a live conflict in a settled BOM rather than a resolved one: the storage table still reads "ZFS root pool" because a unit running a conventional distro could still do that, and a customer machine's OS is not decided by this file. The rule is the OS decides. Check which OS the unit is getting before partitioning.
94 - - **2000 W PSU** is sized for the worst-case 4-GPU configuration on the GPU thread, not for A1 alone.
95 - - **4U rackmount chassis** because it doubles as a prototype of the eventual EveryCycle reference appliance form factor.
96 -
97 - ## Sourcing notes
98 -
99 - Most of the spine is specialist channels; Microcenter covers the commodity parts.
100 -
101 - | Part | Channel |
102 - |---|---|
103 - | TR Pro 7975WX | Newegg, ShopBLT, Provantage |
104 - | WRX90D8-2L/2T | ASRock Rack direct, ShopBLT (longest lead — order first) |
105 - | DDR5-5600 ECC RDIMM (Micron/Hynix off board QVL) | Nemix, ServerSupply, Newegg |
106 - | Sliger CX4712 | Sliger direct (~6 wk historical lead). Fallback: Rosewill RSV-L4500U from Newegg if waiting is unacceptable. |
107 - | Tesla P40 + EPS-to-PCIe adapter + fan shroud | eBay |
108 - | 2× 4 TB Gen5 NVMe (9100 Pro / T705) | Microcenter |
109 - | PSU | Microcenter if Super Flower Leadex Titanium 2000 W in stock; otherwise Seasonic PRIME TX-1600 is the MC-reliable alternative (sufficient for A1–A3, marginal at A4 4-GPU). |
110 - | Noctua NF-A14 industrialPPC-3000 fans | Microcenter |
111 - | NH-U14S TR5-SP6 cooler | Noctua direct or Newegg; Microcenter rarely stocks |
112 - | GT 1030 boot GPU, paste, cables, M.2 heatsinks | Microcenter |
113 -
114 - Suggested order of operations: board first (longest lead) → CPU + RAM kit together (QVL match) → chassis → Microcenter run for commodity parts → P40 last (cheapest, most fungible).
115 -
116 - ## Open considerations (not blocking purchase)
117 -
118 - - **PSU redundancy.** Single 2000 W Titanium is fine for v0. Dual-redundant CRPS (e.g. FSP Twins) is the server-class move for v1.
119 - - **Cooler headroom.** NH-U14S TR5-SP6 is adequate for the 7975WX's 350 W TDP, not generous. XE360-TR5 (listed alt) gives more headroom for sustained all-core + GPU host load. Both are air.
120 - - **Boot GPU skip.** If WRX90 BMC serial-over-LAN proves reliable, GT 1030 line item is removable and a slot is reclaimed.
121 - - **kv-scratch PLP.** Consumer 9100 Pro / T705 have no power-loss protection. Fine for disposable scratch; revisit if EveryCycle ever checkpoints kv-cache through it.
122 -
123 - ## What this BOM deliberately does not include
124 -
125 - - **Multiple compute GPUs at purchase time.** Each GPU thread milestone is a discrete add; advancing the thread is its own decision moment.
126 - - **Liquid cooling.** More maintenance overhead than a single-operator shop should carry.
127 - - **Optane / SLC drives.** Optane is discontinued; Gen5 TLC is the right answer now.
128 - - **128-core CPU.** Memory-bandwidth-starved for inference; ~$6K extra for no inference gain.
129 - - **A second NIC card.** The 2× 10 GbE onboard is enough for v0.