stayz3ro.dev

Seven days on a power strip: the pve02 saga

An earlier post covers the first half of this story: a hardware transplant into a Dell OptiPlex 3070 Micro, a validation pass, and then a reset loop. This is the second half: the hypothesis that finally fit the evidence, the one-variable test, and the week of elapsed time it took to believe the result.

Where the story stood

On September 14 I moved the RAM and NVMe from an aging HP ProDesk into the Dell and booted the existing Proxmox installation on it. The disk carried the node identity, so the cluster came back without a rejoin dance, and every check passed, including an incremental backup against the new hardware. Six days later the node was resetting every few minutes. The closeout was reopened, because one good validation window does not prove replacement hardware is sound.

The evidence on September 19 was ugly. Monitoring showed 75 observed new boots on that one day, within 79 boot-timestamp changes over the prior seven days. Observed intervals ranged from about 2.5 to 83.7 minutes, median around 11. The journal never captured a reset happening; pstore was empty. A hard power fault leaves no note.

What the diagnostics cleared, and what they did not

memtest86+ first rebooted the machine before it could finish. After I swapped the two SODIMMs between their slots, it completed a full pass in about 37 minutes with 0 errors. I have mixed feelings about that pass. It came only after a hardware change, so it implicates the change as much as it clears the memory. The reseating may have fixed a seating problem, or the memory may never have been the problem.

One observation was more interesting. The machine sat in the BIOS setup screen for hours without a reset, but reset on any other screen. That points at hardware or firmware behavior rather than the operating system, since the OS was not doing anything in either case. The Dell pre-boot diagnostics and the memtest pass both came back clean, and neither made the fault reproduce on demand. Clean diagnostics on an intermittent fault rule out very little.

The power-path hypothesis

The clue that reframed everything was not about this node at all. The UPS that pve02 plugged into also powered the network equipment, and that equipment had been shutting down randomly too. Meanwhile pve01 and pve03 sat on a plain power strip and had run more than 12 days without a single reset. The UPS had a history of dropping its outputs even with a full battery.

Two unrelated kinds of device misbehaving on one UPS, while two nodes on a cheap strip stayed up, made the UPS the common factor. I put the confidence at about 85 percent, and I want to be clear that this was a hypothesis, not a diagnosis. Nothing had proven the UPS did anything.

One variable, one move

The test was deliberately boring: move pve02 from the UPS to a dedicated power strip, change nothing else, and watch. Microcode, the C-state cap, the kernel, and the swapped RAM all stayed exactly as they were. The move happened on September 20 at 16:22. The node came back at 18:58:35 that evening.

Watching did not mean staring at it. A small reset-watcher script on a separate machine polled the node’s boot timestamp on a schedule, and a monitoring alert posted to Discord whenever any node’s boot time changed. Either would catch a reset without anyone looking.

The clock as the verdict

Thirty-three continuous hours later, the watcher had recorded zero resets. That number matters because of the baseline: the longest quiet stretch in the earlier data was roughly 84 to 90 minutes. The node had already gone more than twenty times its previous best by the time the first checkpoint came.

That strengthened the hypothesis but was still only elapsed-time evidence, so I kept waiting. The acceptance line was seven clean days. On September 27 the count stood: up since September 20 18:58:35, zero resets through seven full days on the power strip. Stability accepted.

What I am willing to claim

pve02 is stable on the power strip. The move is the only change that correlates with the fault stopping, across a week of continuous operation. That is the honest summary.

Root cause is still unproven. I never tested the UPS in isolation, never put pve02 back on it to see the fault return, and never completed the formal Dell pre-boot diagnostic, which is now optional rather than a gate. A skeptic can still say the RAM swap fixed it, or that the fault was intermittent and paused on its own. The difference between “the fault stopped after the move” and “the move fixed the fault” is exactly one experiment I have not run, and I am not interested in running it on production. The node stays on the strip.

The deferral list tells the rest: a BIOS update waits until after the network cutover, which has since happened, so it can be scheduled on its own now. The retired HP chassis is still on the shelf as the documented fallback, and deciding its fate is an open question. The UPS itself is still uninspected: make, model, load, battery age, fault log, self-test. Nothing else gets powered from it until that happens.

What the week taught

  • A closed phase can regress. The stability claim in a closeout is provisional until an observation window passes with no recurrence. The HP chassis earned its shelf spot by being kept, not assumed unnecessary.
  • Check the power path as a system. I tested the node with diagnostics before I tested what fed the node. The cheapest test, plugging it into a different outlet path, was the one that mattered, and it came after memtest and pre-boot diagnostics.
  • Elapsed time is evidence, but only against a baseline. Thirty-three hours means something because the previous best was 90 minutes. Seven days means something because the fault used to fire every eleven.
  • One variable at a time is what makes the conclusion usable. The move carried microcode, kernel, and RAM constants with it precisely so the result could not be explained by them.