Repository navigation
feat: the lab session into main (Moss node service, Sensorica Holoport fleet, Grafana home page, mdBook) - #67
Conversation
…/holoport-session
…holoport-session # Conflicts: # .github/workflows/ci.yml
…oloport-session # Conflicts: # .github/workflows/ci.yml
…oport-session # Conflicts: # flake.nix
… lab/holoport-session # Conflicts: # flake.nix
…holoport-session # Conflicts: # docs/deployment.md
…a-event-node Once #59 and #61 are both in, hosts/common.nix no longer reads fleetLine and fleetHapps: the fleet imports nixosModules.sensorica-event-node for them. holoportTarget still passed the old arguments, so the installed system had no hApps, no installer unit, and vmTestHoloportInstall failed ("holochain-happ-installer.service is inactive and there are no pending jobs"). It now imports the module the fleet imports, and the check passes: hrea, kando and requests-and-offers enabled 20 to 43 s after boot.
A NixOS Holoport can run more than a Holochain edgenode (a Moss node, a bootstrap server, Grafana), so the machine is named for what it is and the edgenode stays a role. Renames the host directories, the flake outputs, the CI host list, the scrape-target examples and the docs.
One word rebuilds whichever Holoport it runs on, from the checkout the install leaves in /root/nixos-holochain, the way athanor's machines do.
hosts/desk.nix gives the sensorica session five launchers (fleet dashboard, this node, Holochain logs, Moss node, rebuild), a Plasma layout declared with plasma-manager (bottom panel, Breeze Dark, CPU and RAM monitor, no screen lock or suspend), tmux, btop and hc, and Avahi so sensorica-holoport-01.local resolves in the lab. sensorica.eventMode (off by default) logs sensorica in and opens the room dashboard full screen, with read-only anonymous Grafana on the monitor node. home-manager (release-26.05) and plasma-manager are inputs of the example only; the modules stay desktop-free.
desktop-file-validate rejects a '?' outside quotes in Exec, which failed the system build on sensorica-holoport-01.
…#28) nixosModules.holochain-moss-node runs wdocker's daemon under its own user, the conductor password passed as a systemd credential from a root-only file and read on stdin, so the node needs no terminal and no tmux. Joining the group stays one command at a terminal, since the invite link carries the network seed: moss-node join "INVITE_LINK" runs wdocker join-group as the node's user and restarts the daemon. The readings and the Sensorica Moss page are ported from athanor's mossNodeMetrics (athanor 539e648): the shared conductor exporter under conductor="Moss", names from Nix, the page in its own provider with a configurable title. vmTestMossNode starts the service headless, checks the readings and a restart, and fails when the password file is missing; moss-dashboard and moss-names come over with their broken-copy checks. The Sensorica fleet enables it on sensorica-holoport-01.
Plasma's power management suspended sensorica-holoport-01 from the login screen, taking its conductors and dashboards offline. Disabling the sleep, suspend, hibernate and hybrid-sleep targets makes suspend impossible whatever the desktop or logind asks.
The Holoports install from this branch, so the example reads the modules it carries (event node, Moss node) rather than main, which has neither yet. A fresh clone then needs only the operator key before the install. Point it back at main once the lab branch merges.
docs/moss-node.md opens with nixosModules.holochain-moss-node (options, the credential, moss-node join/status/logs, readings and page, checks) and keeps the by-hand path. The fleet README covers the host names, the pin to lab/holoport-session, the rebuild alias, never-sleep, the operator desk with event mode, and the Moss node on sensorica-holoport-01.
…ing node On sensorica-holoport-01 an empty password file made wdaemon create the conductor with an empty password and exit 0, which Restart=on-failure left down. ExecStartPre now refuses an empty credential with a message naming the file, Restart=always brings the node back however it ends, and moss-node join stops before its prompts when the node is not running (join-group otherwise asks everything, then fails on 'No port file found'). vmTestMossNode gains an empty-password machine that must log the refusal and create no conductor.
sensorica-holoport-01 drives two screens and the panel only reached the first.
Opens the monitor's Grafana home from the sensorica desktop.
In /root the sensorica operator could not open it. /etc/nixos-holochain is root:wheel with setgid, git marks it safe for every user and keeps its objects group-writable, so the checkout can be edited from the desk while rebuild still runs as root. The rebuild alias, the Rebuild launcher, the fleet README and the install steps follow.
Every unit a module lists in services.holochain-services.units can now declare the version of the package it runs (version) and, for a unit that runs a conductor, the Holochain that conductor is (holochainVersion). Both are published as labels of holochain_service_info, so they ride through holochain:service_watched into holochain:service_state. The modules fill them from their packages: the conductor and the app installer from the edgenode's package and hcPackage, the HTTP gateway and the bootstrap server from theirs, Prometheus, Grafana, node_exporter, sshd, Tailscale and the Nix daemon from theirs, and the Moss node from wdocker, whose package now passes through the Holochain version it brings. A unit without one publishes no version label. node_exporter also exports each unit's start time (--collector.systemd.enable-start-time-metrics) on the edgenode and the monitor.
A fifth dashboard, holochain-home, provisioned with the others and set as Grafana's home page, so / opens on it after login. At the top, the machine in words: its state, host name, NixOS release, kernel, uptime, and how busy its processor, memory and fullest disk are. Below, every service it runs, worst first, with its state in the node page's words and colours, the version Nix runs it from (a dash where a unit declares none) and when it last started; then its Holochain conductors, each with the service that runs it, its Holochain version and whether it answers, and every app on them in the six state words, each linking to the node page. A node picker shows any machine of the fleet. The module's rewrite defaults it to this machine: the name of the scrape target on a loopback address, else networking.hostName. Links lead to the fleet, node, network and room pages, and to any dashboard tagged moss, so the Moss page shows where the Moss module provisions it. A dashboards directory without the new page falls back to the room screen as home, as before. The label check now reads host name, system and kernel release, and the declared versions as names a person reads.
dashboardLabels and dashboardQueries take in the home page: its columns and descriptions, its uid, and, on the homelab fixtures with versions on two units, a services table that shows each version and a dash where there is none, and a conductors table with one row per conductor. Copies that lose the versions or repeat a claimed conductor must fail. grafanaProvisioning compares every version the modules declare with the package each unit runs, a Moss node's included, and requires the published list to carry each as a label; copies with one label removed must fail by the unit's name. It requires the home page to be the provisioned holochain-home, opening on the loopback target or the host name, and refuses a copy without it and a copy whose home page lost its uid. vmTestGrafana expects five dashboards, holochain-home at home on machine, the home page's state words, and the conductor's versions in Prometheus.
sensorica.desktop picks KDE Plasma, GNOME or no graphical session. The desk keeps a single Grafana launcher, on the home page, beside the Holochain and Moss logs and Rebuild; the stale tmux and per-page launchers go. Each host file opens with the switches an operator flips by hand before running rebuild.
Instant queries come back from Prometheus in no fixed order, so the machine tiles, the per-node bars and the app-by-node grids shuffled between refreshes. Each now sorts on the node label first.
book.toml builds docs/ into a gitignored book/, with a SUMMARY organised
for three readers: installing a Holoport, running the modules, and
contributing. The fleet example, workshop, template and happs READMEs and
CONTRIBUTING.md appear through {{#include}} wrapper pages without moving.
The Deploy Documentation workflow builds the book on pull requests that
touch its sources and deploys it to Pages from main, with mdBook pinned to
the devShell's 0.5.2. mdbook joins the devShell, the README links the
site, and deployment.md's link to scripts/holoport-install.sh, which
leaves docs/, becomes an absolute GitHub URL.
mdBook 0.5 expands {path} relative to the repo root, so the edit links
pointed at docs/docs/<page>.md. The template now ends in {path}.
The Pages artifact and deploy job now also require refs/heads/main, since
workflow_dispatch can run on any branch.
architecture.md's "Holochain conductor (<conductor>)" lost its placeholder
on the page, parsed as an unclosed HTML tag; it is now in backticks.
…s named A monitor that listed itself by its own host name or FQDN under a key of its own (lab-1.address = "monitor:9100") defaulted the home page's node to its host name, which no target carried, so Grafana opened on the first node in the list. A target at this machine's host name or FQDN is now this machine too, after a loopback target. grafanaProvisioning checks both forms, and fails on the previous module.
The shared config installs one app, so holochain-happ-installer.service is listed; its version must equal hcPackage's, and a copy of the published list without its version label must fail by its name.
"Each app across the fleet" carried the home page's node, so the network page came up for one machine and "how many run it" read 1. It no longer carries variables. dashboardLabels now refuses any link by address that carries a page's variables to a page whose node picker opens on All, and a copy that sets it back fails by the link's name.
…port The icon opened the monitor's bare /, whose home page defaults to the monitor itself, so on holoport-02 to 05 it showed holoport-01. It now opens /d/holochain-home with this machine's node, as the node launcher already does.
sensorica.remoteAccess.enable runs the Tailscale client against https://hs.sensorica.co with a pre-auth key read from a root-only file, and trusts the tailnet interface. Off on every Holoport until the Headscale server exists; each host's switches block carries it.
…session Merges feat/sensorica-holoport-names: "What is this machine running?" as Grafana's home page, the version each listed unit runs, their checks and docs, and the review fixes. One conflict, in examples/sensorica-fleet/hosts/desk.nix: the lab side rewrote the desk around the sensorica.desktop switch with one Grafana launcher and one desktop icon, while the home page branch pointed its icon at the home page on this Holoport. The lab structure is kept, and both the holoport-grafana launcher and Desktop/grafana.desktop now open the same page, /d/holochain-home with this machine's node, since the monitor's bare / opens the home page on the monitor itself.
Conflicts, resolved as main's reviewed code unless lab carries a later fix of the same lines: - flake.nix: the Moss node module export is lab's and is kept. The Holoport install block (onX86, holoportInstall, holoportTarget, holoportDisk) and edgenodeBinaryCache sit where main moved them; lab's copies at the old places are dropped. holoportTarget takes the renamed host, hosts/sensorica-holoport-01. - .github/workflows/ci.yml: main's build list, which already carries vmTestWdocker and adds the metrics checks. - docs/architecture.md: lab's vmTestGrafana paragraph (five dashboards, holochain-home) with main's fix that node state 2 is A service is down. - docs/deployment.md: main's install text (Esc boot menu, USB 2, the root password outside a terminal) with lab's host names, lab's absolute script link for the book, and lab's /etc/nixos-holochain checkout step, its code fence closed.
The Moss node module (#28) came with its VM test and two checks without a VM, and none of them was in the build list.
…sk switches The renamed host file sets sensorica.desktop and sensorica.eventMode, which hosts/desk.nix declares, and enables the Moss node. The install target imported neither, so vmTestHoloportInstall, and with it nix flake check, failed to evaluate: "The option `sensorica.desktop' does not exist". The target now imports the Moss node module, as the fleet does, and a stand-in for the desk that declares its two switches and keeps the Plasma session with SDDM for "plasma", the desktop main's target had from hosts/common.nix. desk.nix itself needs home-manager and plasma-manager, which this flake does not carry.
…er is not packaged The Moss facts (ee8dba5) evaluate the Moss node's package, and wdocker exists for x86_64-linux only, so nix flake check --all-systems failed on aarch64-linux with "attribute 'wdocker-0_15' missing". The facts and the three checks that read them now run on x86_64-linux only; every other system keeps the rest of the check.
… textfile Stopping the timer leaves a run it already started going, and that run can finish after the test writes the bad file and replace it with good readings, so "A metrics file could not be read" never appears and the test times out. Seen locally at 248 s: holochain-conductor-metrics.service finished right after the write. The service is now stopped with the timer.
…rst boot sensorica-holoport-01 now enables the Moss node, which reads its conductor password through LoadCredential. The runbook wrote only Grafana's password file, so a fresh install booted with moss-node.service failing on status=243/CREDENTIALS and the machine reading "A service is down". Step 3 now writes /mnt/var/lib/secrets/moss-node-password, step 4 checks the unit is active and points at the join step in docs/moss-node.md.
Review fixes before mergeTwo reviews ran on 3480cd9. One finding was blocking, and 1c591b4 fixes it. Fixed: the runbook never wrote the Moss node password. sensorica-holoport-01 now enables the Moss node, which reads its conductor password through
Checks at 1c591b4
Not blocking, left for after the merge. These are in addition to the follow-ups in the PR body:
One correction to the PR body: GitHub Pages is already enabled ( |
Refs #28, #43. Neither closes on merge: #28's metal acceptance (a Moss desktop showing the node online after a reboot) and #43's first Pages deploy both happen after this lands.
What and why
This brings the lab session (
lab/holoport-session, the branch the Holoports were installed from on 2026-09-27) intomain. The branch had 28 commits thatmaindoes not hold, and merges of the 17 feature branches thatmainnow holds in their reviewed form.origin/main(168cff1) is merged into it here, so the reviewed code is the base and the lab work sits on top.What the lab adds:
Moss node as a NixOS service (
modules/holochain-moss-node.nix,nixosModules.holochain-moss-node, Moss always-online node service #28): wdocker under systemd with its password from a credential, groups joined from invite links, a restart after each join, a refusal of an empty password, its readings asconductor="Moss"and its own Grafana page. Checks:vmTestMossNode,moss-dashboard,moss-names. Docs:docs/moss-node.md.Sensorica fleet example: hosts renamed
sensorica-holoport-01..05; arebuildalias; the operator desk (hosts/desk.nix) with onesensorica.desktopswitch (plasma,gnome,none), one Grafana entry, the panel on every screen and a desktop icon; event mode; Holoports never sleep; the checkout in/etc/nixos-holochain; per-host switches for the desktop, event mode, the conductor and remote access (hosts/remote-access.nix, for Sensorica's Headscale); sensorica-holoport-01 runs the Sensorica group's Moss node.Grafana home page
holochain-home, "What is this machine running?", opening on this machine however its scrape target is named; every listed unit carries the version Nix runs it from (version, andholochain_versionfor a conductor); dashboards list machines in node order; the fleet-wide network link opens on every node.The mdBook:
book.toml,docs/SUMMARY.md,docs/introduction.mdand.github/workflows/docs.yml, which builds the book on PRs that touch its sources and deploys it to Pages frommainonly.Commits added on this branch after the merge:
docs(book):docs/SUMMARY.mdlistsdocs/releasing.md(feat(release): CHANGELOG and a tag-driven release workflow #49) anddocs/adr/(docs(adr): move the design record from issues into docs/adr/ #51); the introduction said the project is AGPL-3.0, and chore: license the repository under MIT #36 licensed it MIT.ci:vmTestMossNode,moss-dashboardandmoss-namesjoin the build list, per the PR template's rule for a new module.style(flake): alejandra on the Moss checks'letblock, the one file the lab left unformatted.fix(tests):vmTestHoloportInstalldid not evaluate, on the lab branch as well, sonix flake checkfailed: the renamed host file setssensorica.desktopandsensorica.eventMode, whichhosts/desk.nixdeclares, and enables the Moss node. The install target now imports the Moss module and a stand-in for the desk that declares those two switches and keeps Plasma with SDDM for"plasma", the desktop main's target already had fromhosts/common.nix.desk.nixitself needs home-manager and plasma-manager, which the root flake does not carry, so the desk's launchers and panel layout are not part of the install rehearsal.fix(tests):grafanaProvisioningread the Moss node's package on every system, and wdocker is packaged for x86_64-linux only, sonix flake check --all-systemsfailed on aarch64-linux (attribute 'wdocker-0_15' missing), on the lab branch as well. The Moss facts and their three checks now run on x86_64-linux only.fix(tests): the knownvmTestGrafanarace. The test stopped the metrics timer before corrupting the textfile, but a run the timer had already started finished after the write and put good readings back, so "A metrics file could not be read" never appeared. It failed that way locally (the service finished at 248 s, right after the write); the service is now stopped with the timer, and the rerun passed.Conflicts and how each was resolved
The rule: main's version of code from a reviewed PR wins, unless lab carries a later fix of the same code; lab-only work is kept whole.
flake.nixnixosModulesexport: lab'sholochain-moss-nodeagainst nothing on mainflake.nixonX86,holoportInstall,holoportTarget,holoportDisk), which main moved earlier in theletholoportTargettakes lab's rename,hosts/sensorica-holoport-01.flake.nixedgenodeBinaryCache, which main moved to the top ofchecks.github/workflows/ci.ymlvmTestWdockerand addsmetricsHelpAgreement,metricsNameShape,edgenodeNamesWiring. The host loop auto-merged to lab'ssensorica-holoport-0Nnames.docs/architecture.mdvmTestGrafanaparagraphholochain-home, three service tables, versions), with main's later fix 846f1be:holochain:node_statereads 2, not 3, for A service is down.docs/deployment.mddocs/(7c594a2).docs/deployment.mddocs/deployment.md/etc/nixos-holochaincheckout step. That step's code fence was fused with the next sentence on the lab branch; it is closed here.docs/deployment.mdnix build,ssh,nix copy, the Grafana password)sensorica-holoport-01names.Host names
grep -rn 'edgenode-0' .github docs README.md examples testsfinds onlytests/fixtures/edgenode-0_6_3/(a fixture directory for the 0.6.3 edgenode shape, not a host) and two quotes in ADR-005 and ADR-011 of the CI failure of that time (hosts/edgenode-0*/configuration.nix). The ADRs copy issue text as it was written, so those quotes stay. Everything else, CI's host loop included, sayssensorica-holoport-0N.The fleet still pins the lab branch
examples/sensorica-fleet/flake.nixkeepsgithub:Sensorica/nixos-holochain/lab/holoport-session, locked at 4b1b722, for now: the Holoports rebuild from that input, and relocking tomainbefore this merge would point them at a commitmaindoes not have yet. CI evaluates the example with--override-input nixos-holochain "$GITHUB_WORKSPACE", so the pin does not change what CI checks. The relock tomainis the first follow-up.Commands run and what they printed
At 3480cd9, on the laptop (NixOS VM tests on KVM, max-jobs 4):
VM tests built
Run locally in parallel; each output below is the one for HEAD's derivation.
vmTestMossNode: passedvmTestServices: passedvmTestGrafana: failed on the known race ("A metrics file could not be read" timed out after 120 s), fixed in 3480cd9, then passedvmTestHoloportInstall: passed, with the stand-in desk (install, ADR-017 layout, SeaBIOS boot assensorica-holoport-01, conductor, three hApps installed once and enabled, Grafana with the pre-written password)The other VM tests are unchanged by the lab work and left to CI.
Option reference
docs/module-options.mdwas regenerated withcp "$(nix build .#options-doc --print-out-paths)" docs/module-options.mdin the same commit. (The Moss node module's options, the per-unit versions and the home page, each regenerated in its own lab commit; the diff above shows the merged tree in sync.)Documentation
docs/(andREADME.md, if it says anything about this) updated in this PR, or nothing there describes the changed behaviour.Follow-ups
Relock
examples/sensorica-fleettomainonce this merges: pointnixos-holochain.urlatgithub:Sensorica/nixos-holochainand update its lock, then rebuild the Holoports from it.Enable GitHub Pages (source: GitHub Actions) so
docs.ymlcan deploy the book frommain; then docs: documentation site (mdBook) #43 can close.CHANGELOG.mdunder[Unreleased]records the September stack only (refactor(flake): fleet as example, toolchain on 0.7, CI green (slice 1) #13 to chore(example): lock nixos-holochain to main after the stack #23). The feature PRs since (feat(edgenode): declare the Holochain Foundation binary cache #34, feat(grafana): an overview of every node and service on the fleet dashboard #35, feat(edgenode): per-DHT metrics from dump-network-metrics #47 to docs(adr): move the design record from issues into docs/adr/ #51, feat(modules): export the Sensorica event profile once (#33) #59 to feat(services): every service the modules install, on the dashboards by name #66) and this one have no entries yet.The install rehearsal runs without the operator desk's launchers and panel. Carrying
desk.nixinto the rehearsal would need home-manager and plasma-manager as root flake inputs; whether that is worth it is open.Moss always-online node service #28 stays open for its metal acceptance.