Skip to content

feat(install): Holoport install script, runbook and a SeaBIOS VM check - #61

Merged
Soushi888 merged 7 commits into
mainfrom
feat/holoport-install
Sep 28, 2026
Merged

Soushi888 merged 7 commits into
mainfrom
feat/holoport-install

Conversation

@Soushi888

Copy link
Copy Markdown
Contributor

SoushAI analysis. Drafted by Soushi's AI assistant, reviewed and posted by @Soushi888.

Refs #6, #8

Why

The next session at the Sensorica lab installs the event node on one real Holoport, and a Holoport boots legacy BIOS only (ADR-017). hosts/common.nix already configures GRUB for both halves and says the BIOS half comes from a runbook, and the fleet README said that runbook "is not yet written up (tracked in #6)". This PR writes it as one script, documents it, and proves it in a VM that boots the way the Holoport does.

What

  • scripts/holoport-install.sh, exposed as packages.x86_64-linux.holoport-install with every tool it calls pinned. holoport-install DISK SOURCE does GPT with a 1 MiB bios_grub partition, a vfat ESP labelled boot, an ext4 root labelled nixos and 8 GiB of swap labelled swap; it mounts root at /mnt and the ESP at /mnt/efi-boot, runs nixos-install, then grub-install --target=i386-pc --boot-directory=/mnt/boot DISK. The sequence follows holochain/wind-tunnel-runner (installer.nix, base-install.nix).
    • It never picks a disk (a HoloPort+ has two). With no disk or no source it exits 2. It shows the disk it will erase and the disks it leaves alone, and erases nothing unless you type the disk name back. It refuses a partition, a disk with anything mounted, and a second disk that already carries one of its labels, because the installed system mounts by label.
    • SOURCE is either a flake reference (...#edgenode-01), installed with nixos-install --flake and the Holochain cache passed as --option, since the target has no nix.conf yet; or a /nix/store/... system built elsewhere. If that system is not on the box yet, the script prints nix copy --to 'ssh://root@IP?remote-store=/mnt' PATH and waits for the copy. This is the laptop path: a HoloPort's disk is slow and its CPU is weak, and the installer's store lives in RAM.
  • docs/deployment.md § "Installing on a Holoport (legacy BIOS)": both machine variants (disks, BIOS keys), boot the workshop or stock ISO, check the network, install from the box (2a) or from a laptop (2b), write the Grafana password under /mnt before the first boot, reboot, verify (conductor, hApp installer journal, Grafana). Every command is on one line.
  • checks.x86_64-linux.vmTestHoloportInstall: an installer VM with a tmpfs root and one empty 40 GB SATA disk behind AHCI (so it is /dev/sda, reached by ahci from the fleet's own hardware-configuration.nix, not by virtio) runs the package. The disk then boots in a second VM under QEMU's default firmware, SeaBIOS, with no kernel handed in and no OVMF. The target is the real event node: hosts/edgenode-01/configuration.nix with common.nix, the 0.6 line and all three fleet hApps (hREA, Kando, Requests & Offers), Plasma included. The test asserts the layout and labels, that the booted system is the installed store path with root on /dev/sda3 and swap on /dev/sda4, that both GRUB halves are present, that the conductor is active, that each hApp is listed once and all three are enabled, and that Grafana serves the provisioned dashboard with the password written before boot.
  • CI: a separate holoport-install job. The event node's closure is about 10 GiB (most of it the desktop), and the job holds it twice (runner store and qcow2), so it frees the preinstalled toolchains first rather than sharing the disk with the other VM tests.
  • The fleet README, the fleet template README and the common.nix comment now point at the script.

How to test

nix build -L .#checks.x86_64-linux.vmTestHoloportInstall
nix flake check --no-build --all-systems --accept-flake-config

What I ran, on the Builder's machine (Nix 2.25.4, 16 cores, KVM). nix build of a VM check cannot open /dev/kvm in the sandbox there, so each run built the .driver attribute and ran nixos-test-driver directly with KVM, under a lock shared with other agents' VM tests. CI runs the literal nix build.

Run Script Exit Where it stopped
1 as committed 1 test bug: parted was only inside the package, not on the installer's PATH (install itself done in 404 s); fixed in the test
2 as committed 0 all subtests; install 487 s, SeaBIOS to multi-user 84 s, 828 s total
falsifier BIOS grub-install line deleted 1 SeaBIOS boots the disk through the BIOS GRUB failed: action timed out after 600.16 seconds waiting for the kernel's first line
restored as committed (same driver hash as run 2) 0 all subtests; install 360 s, boot 55 s, 800 s total

On the first boot all three hApps went from install to enabled 76 s after the kernel started (hc-sandbox: Enabled app: "requests-and-offers" at 75.8 s).

Also: alejandra --check (4.0.0, the devShell's) on flake.nix and examples/sensorica-fleet/hosts/common.nix: exit 0. nix flake check --no-build --all-systems --accept-flake-config: exit 0. The example fleet's nix flake check --no-build --override-input nixos-holochain <checkout>: exit 0. shellcheck scripts/holoport-install.sh: exit 0.

Deviations

  • The test gives the script a prebuilt store path instead of a flake reference, so it runs nixos-install --system where the doc's path 2a runs nixos-install --flake: the sandbox has no network, as in nixpkgs' nixos/tests/installer.nix. Everything before and after that one call is the documented command.
  • The target system is built from this flake's nixpkgs, not the fleet's own lock, and carries the test driver's backdoor module (as nixpkgs' installer tests do). It adds nothing else.
  • The package and the check exist on x86_64-linux only (i386-pc GRUB, SeaBIOS). They use an attribute named null elsewhere rather than //, so the existing attribute sets are not re-indented.

Not covered

  • No real Holoport yet (Hardware: one Holoport boots vanilla NixOS from the workshop ISO (ADR-003 gate) #8). Unknown until the lab: the BIOS setup key and boot-menu key, whether the Holoport's BIOS boots a GPT disk through the protective MBR the way SeaBIOS does, the real disk's name and speed, and the time the base HoloPort takes.
  • The laptop path (2b) is not exercised by the check. Nix 2.25 accepts remote-store on an ssh:// store URL (an unknown setting warns; this one does not), but the copy into /mnt over SSH and the script's wait loop have not run end to end.
  • Path 2a's nixos-install --flake with the Holochain cache options has not run against the network either; the check covers --system only.
  • nix run ...#holoport-install on the stock ISO, and whether the workshop ISO's sshd starts at boot (docs: rescue a failed install over SSH, learned on the homelab #25 notes the same open question).
  • The CI job's disk budget and runtime on a GitHub runner are estimates until this PR's run lands.
  • docs: rescue a failed install over SSH, learned on the homelab #25's rescue section overlaps step 2b's SSH setup. This section repeats the two commands it needs, so it stands alone whichever merges first.

scripts/holoport-install.sh partitions the named disk the ADR-017 way (GPT, 1 MiB bios_grub, vfat ESP 'boot', ext4 root 'nixos', swap 'swap'), mounts root at /mnt and the ESP at /mnt/efi-boot, runs nixos-install and then grub-install --target=i386-pc. It refuses to run without an explicit disk and asks for the disk name back. The flake exposes it as packages.x86_64-linux.holoport-install. checks.x86_64-linux.vmTestHoloportInstall runs that package on an installer VM against an empty AHCI disk, then boots the disk under SeaBIOS and asserts the conductor, the fleet's three hApps and Grafana on edgenode-01. CI runs it in its own job because the closure is about 10 GiB.
Boot an installer, get network, run holoport-install from the box or from a laptop that copies the closure straight into /mnt, write the Grafana password before the first boot, verify. The fleet README, the fleet template and the common.nix comment now point at the script instead of saying the sequence is unwritten.
Soushi888 added a commit that referenced this pull request Sep 27, 2026
…a-event-node

Once #59 and #61 are both in, hosts/common.nix no longer reads fleetLine and fleetHapps: the fleet imports nixosModules.sensorica-event-node for them. holoportTarget still passed the old arguments, so the installed system had no hApps, no installer unit, and vmTestHoloportInstall failed ("holochain-happ-installer.service is inactive and there are no pending jobs"). It now imports the module the fleet imports, and the check passes: hrea, kando and requests-and-offers enabled 20 to 43 s after boot.
@Soushi888

Copy link
Copy Markdown
Contributor Author

SoushAI analysis. Drafted by Soushi's AI assistant, reviewed and posted by @Soushi888.

Merge note: this PR and #59 break vmTestHoloportInstall once both are in, whichever lands second.

#59 moves the fleet's line, hApps and seed into nixosModules.sensorica-event-node, and hosts/common.nix stops reading fleetLine and fleetHapps. This PR's holoportTarget still passes those two arguments, so on a tree with both, the installed system has no hApps and no installer unit. The check then fails with unit "holochain-happ-installer.service" is inactive and there are no pending jobs.

The fix is to replace the _module.args block in holoportTarget with self.nixosModules.sensorica-event-node. It is commit 92cf87f on lab/holoport-session, a branch that combines #61, #60, #59, #62, #34, #25 and the dashboard stack for the lab session. There vmTestHoloportInstall passes, with hREA, Kando and Requests & Offers enabled 20 to 43 s after boot. The same commit belongs on whichever of #59 and #61 merges second.

One other observation from these runs: one attempt failed at mount /dev/sda3 with "wrong fs type" right after mkfs.ext4, and the next two did not. It looks intermittent. If it shows up again, the script may need a udevadm settle between formatting and mounting.

# Conflicts:
#	docs/deployment.md
#	flake.nix
…id disk

A VM run once stopped at the root mount with 'wrong fs type' right after mkfs. The script now waits for udev after formatting and mounts the root as ext4 explicitly. A DISK given as a /dev/disk/by-id link is resolved to its kernel node first, so partition names and the label check compare /dev/sdX names.
Esc opens the base HoloPort's boot menu, the stick reads reliably only from a USB 2 port, and the graphical desktop freezes on the HD 610 (session log 2026-09-27). The root password prompt at the end of nixos-install appears only when its input is a terminal.
@Soushi888

Copy link
Copy Markdown
Contributor Author

SoushAI analysis. Drafted by Soushi's AI assistant, reviewed and posted by @Soushi888.

Brought up to date with main, plus two small fixes

The branch now merges cleanly on top of main after #25, #34 and #62 landed. Nothing was rebased or force-pushed. The review found nothing blocking, so the two extra commits only pick up non-blocking items from its list that sit in files this PR already owns.

Commits pushed

  • 29e1bca merges origin/main. Both conflicts were purely additive and keep both sides. In docs/deployment.md, the Holoport install section comes first and docs: rescue a failed install over SSH, learned on the homelab #25's "Rescuing an install from another machine" follows it. In flake.nix, the onX86/holoportInstall/holoportTarget/holoportDisk bindings sit next to feat: holochain-bootstrap module and relayAllowPlainText #62's bootstrapTest, and packages.holoport-install sits next to legacyPackages.falsifiers. Against main, the merge only adds lines and removes none.
  • fcaf173 fix(install): after formatting, the script runs udevadm settle and then mounts the root with -t ext4. This targets the intermittent "wrong fs type" at ==> mounting reported in the earlier comment. The DISK argument also goes through readlink -f, so a /dev/disk/by-id/... link gets correct partition names and a correct label check.
  • 0c9b423 docs(deployment): adds what the first real Holoport install on 2026-09-27 showed. Esc opens the base HoloPort's boot menu, and the BIOS setup key and the HoloPort+'s keys are still unknown. The stick reads reliably only from a USB 2 port. The graphical desktop freezes on the HD 610, and the page gives the way to a console. It also says the root password prompt appears only when nixos-install's input is a terminal (its source tests [ -t 0 ]) and how to set the password otherwise.

Checks

  • Locally, all passed: nix flake check --no-build --all-systems, the example fleet's nix flake check --no-build against this checkout, nix build .#holoport-install (which runs shellcheck), the options-doc diff (in sync) and alejandra --check flake.nix.
  • Locally, vmTestHoloportInstall passed the refusal subtests, formatting, the root mount, the install, the ADR-017 layout check and the BIOS GRUB boot. It then timed out at wait_for_unit("multi-user.target") (900 s). The host had a load average around 12 at the time, and the guest was logging soft lockups. The CI runner ran the same check on this head and it passed.
  • CI on 0c9b423: flake eval, nix flake check and Holoport install (SeaBIOS) are all green. The PR shows MERGEABLE / CLEAN.

Left open, deliberately

…a-event-node

Once #59 and #61 are both in, hosts/common.nix no longer reads fleetLine and fleetHapps: the fleet imports nixosModules.sensorica-event-node for them. holoportTarget still passed the old arguments, so the installed system had no hApps, no installer unit, and vmTestHoloportInstall failed ("holochain-happ-installer.service is inactive and there are no pending jobs"). It now imports the module the fleet imports, and the check passes: hrea, kando and requests-and-offers enabled 20 to 43 s after boot.
@Soushi888
Soushi888 merged commit ced9de1 into main Sep 28, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant