Skip to content

feat: the lab session into main (Moss node service, Sensorica Holoport fleet, Grafana home page, mdBook) - #67

Merged
Soushi888 merged 45 commits into
mainfrom
feat/lab-into-main
Sep 28, 2026
Merged

Soushi888 merged 45 commits into
mainfrom
feat/lab-into-main

Conversation

@Soushi888

Copy link
Copy Markdown
Contributor

SoushAI analysis. Drafted by Soushi's AI assistant, reviewed and posted by @Soushi888.

Refs #28, #43. Neither closes on merge: #28's metal acceptance (a Moss desktop showing the node online after a reboot) and #43's first Pages deploy both happen after this lands.

What and why

This brings the lab session (lab/holoport-session, the branch the Holoports were installed from on 2026-09-27) into main. The branch had 28 commits that main does not hold, and merges of the 17 feature branches that main now holds in their reviewed form. origin/main (168cff1) is merged into it here, so the reviewed code is the base and the lab work sits on top.

What the lab adds:

  • Moss node as a NixOS service (modules/holochain-moss-node.nix, nixosModules.holochain-moss-node, Moss always-online node service #28): wdocker under systemd with its password from a credential, groups joined from invite links, a restart after each join, a refusal of an empty password, its readings as conductor="Moss" and its own Grafana page. Checks: vmTestMossNode, moss-dashboard, moss-names. Docs: docs/moss-node.md.

  • Sensorica fleet example: hosts renamed sensorica-holoport-01..05; a rebuild alias; the operator desk (hosts/desk.nix) with one sensorica.desktop switch (plasma, gnome, none), one Grafana entry, the panel on every screen and a desktop icon; event mode; Holoports never sleep; the checkout in /etc/nixos-holochain; per-host switches for the desktop, event mode, the conductor and remote access (hosts/remote-access.nix, for Sensorica's Headscale); sensorica-holoport-01 runs the Sensorica group's Moss node.

  • Grafana home page holochain-home, "What is this machine running?", opening on this machine however its scrape target is named; every listed unit carries the version Nix runs it from (version, and holochain_version for a conductor); dashboards list machines in node order; the fleet-wide network link opens on every node.

  • The mdBook: book.toml, docs/SUMMARY.md, docs/introduction.md and .github/workflows/docs.yml, which builds the book on PRs that touch its sources and deploys it to Pages from main only.

Commits added on this branch after the merge:

  • docs(book): docs/SUMMARY.md lists docs/releasing.md (feat(release): CHANGELOG and a tag-driven release workflow #49) and docs/adr/ (docs(adr): move the design record from issues into docs/adr/ #51); the introduction said the project is AGPL-3.0, and chore: license the repository under MIT #36 licensed it MIT.

  • ci: vmTestMossNode, moss-dashboard and moss-names join the build list, per the PR template's rule for a new module.

  • style(flake): alejandra on the Moss checks' let block, the one file the lab left unformatted.

  • fix(tests): vmTestHoloportInstall did not evaluate, on the lab branch as well, so nix flake check failed: the renamed host file sets sensorica.desktop and sensorica.eventMode, which hosts/desk.nix declares, and enables the Moss node. The install target now imports the Moss module and a stand-in for the desk that declares those two switches and keeps Plasma with SDDM for "plasma", the desktop main's target already had from hosts/common.nix. desk.nix itself needs home-manager and plasma-manager, which the root flake does not carry, so the desk's launchers and panel layout are not part of the install rehearsal.

  • fix(tests): grafanaProvisioning read the Moss node's package on every system, and wdocker is packaged for x86_64-linux only, so nix flake check --all-systems failed on aarch64-linux (attribute 'wdocker-0_15' missing), on the lab branch as well. The Moss facts and their three checks now run on x86_64-linux only.

  • fix(tests): the known vmTestGrafana race. The test stopped the metrics timer before corrupting the textfile, but a run the timer had already started finished after the write and put good readings back, so "A metrics file could not be read" never appeared. It failed that way locally (the service finished at 248 s, right after the write); the service is now stopped with the timer, and the rerun passed.

Conflicts and how each was resolved

The rule: main's version of code from a reviewed PR wins, unless lab carries a later fix of the same code; lab-only work is kept whole.

File Conflict Resolution
flake.nix nixosModules export: lab's holochain-moss-node against nothing on main Lab's export kept.
flake.nix The Holoport install block (onX86, holoportInstall, holoportTarget, holoportDisk), which main moved earlier in the let Main's placement; lab's copy at the old place dropped. The two differed only in the host path, so holoportTarget takes lab's rename, hosts/sensorica-holoport-01.
flake.nix edgenodeBinaryCache, which main moved to the top of checks Main's placement; lab's copy dropped (identical).
.github/workflows/ci.yml VM test build list Main's list, which already carries vmTestWdocker and adds metricsHelpAgreement, metricsNameShape, edgenodeNamesWiring. The host loop auto-merged to lab's sensorica-holoport-0N names.
docs/architecture.md The vmTestGrafana paragraph Lab's text (five dashboards, home page holochain-home, three service tables, versions), with main's later fix 846f1be: holochain:node_state reads 2, not 3, for A service is down.
docs/deployment.md Script link Lab's absolute GitHub URL: the book cannot link outside docs/ (7c594a2).
docs/deployment.md Holoport hardware and booting the installer Main's text: Esc for the base HoloPort's boot menu, USB 2 against the squashfs error, the frozen graphical ISO.
docs/deployment.md Install from the Holoport (path 2a) Main's text on the root password prompt outside a terminal, with lab's host name and lab's /etc/nixos-holochain checkout step. That step's code fence was fused with the next sentence on the lab branch; it is closed here.
docs/deployment.md Paths 2b and 3 (nix build, ssh, nix copy, the Grafana password) Lab's sensorica-holoport-01 names.

Host names

grep -rn 'edgenode-0' .github docs README.md examples tests finds only tests/fixtures/edgenode-0_6_3/ (a fixture directory for the 0.6.3 edgenode shape, not a host) and two quotes in ADR-005 and ADR-011 of the CI failure of that time (hosts/edgenode-0*/configuration.nix). The ADRs copy issue text as it was written, so those quotes stay. Everything else, CI's host loop included, says sensorica-holoport-0N.

The fleet still pins the lab branch

examples/sensorica-fleet/flake.nix keeps github:Sensorica/nixos-holochain/lab/holoport-session, locked at 4b1b722, for now: the Holoports rebuild from that input, and relocking to main before this merge would point them at a commit main does not have yet. CI evaluates the example with --override-input nixos-holochain "$GITHUB_WORKSPACE", so the pin does not change what CI checks. The relock to main is the first follow-up.

Commands run and what they printed

At 3480cd9, on the laptop (NixOS VM tests on KVM, max-jobs 4):

$ nix run nixpkgs#alejandra -- --check .
(no file requires formatting)

$ nix flake check --no-build --all-systems
exit 0

$ diff <(cat "$(nix build .#options-doc --print-out-paths)") docs/module-options.md && echo "options-doc in sync"
options-doc in sync

$ (cd examples/sensorica-fleet && nix flake check --no-build --override-input nixos-holochain <worktree>)
exit 0

$ nix eval --json --override-input nixos-holochain <worktree> .#colmena --apply 'h: builtins.attrNames (removeAttrs h ["meta"])'
["sensorica-holoport-01","sensorica-holoport-02","sensorica-holoport-03","sensorica-holoport-04","sensorica-holoport-05"]

$ for t in minimal fleet; do nix flake init -t path:<worktree>#$t && nix flake check --no-build --override-input nixos-holochain path:<worktree>; done
template minimal exit 0
template fleet exit 0

$ nix shell --inputs-from . nixpkgs#mdbook -c mdbook build
 INFO Book building has started
 INFO Running the html backend
 INFO HTML book written to `.../book`
(mdbook v0.5.2, no warnings)

$ nix build --no-link -L .#checks.x86_64-linux.{conductorMetricsJq,dashboardLabels,dashboardQueries,dashboardWords,dhtMetricsJq,edgenodeBinaryCache,edgenodeConfigRender,edgenodeNamesWiring,grafanaProvisioning,holochainRules,metricsHelpAgreement,metricsNameShape,moss-dashboard,moss-names}
exit 0 (every non-VM check; grafanaProvisioning still runs the Moss part on x86_64: "fails as it should: moss-node.service has lost its holochain_version label")

VM tests built

Run locally in parallel; each output below is the one for HEAD's derivation.

  • vmTestMossNode: passed
  • vmTestServices: passed
  • vmTestGrafana: failed on the known race ("A metrics file could not be read" timed out after 120 s), fixed in 3480cd9, then passed
  • vmTestHoloportInstall: passed, with the stand-in desk (install, ADR-017 layout, SeaBIOS boot as sensorica-holoport-01, conductor, three hApps installed once and enabled, Grafana with the pre-written password)

The other VM tests are unchanged by the lab work and left to CI.

Option reference

  • No option was added, changed or removed.
  • Options changed, and docs/module-options.md was regenerated with cp "$(nix build .#options-doc --print-out-paths)" docs/module-options.md in the same commit. (The Moss node module's options, the per-unit versions and the home page, each regenerated in its own lab commit; the diff above shows the merged tree in sync.)

Documentation

  • docs/ (and README.md, if it says anything about this) updated in this PR, or nothing there describes the changed behaviour.

Follow-ups

…holoport-session

# Conflicts:
#	.github/workflows/ci.yml
…oloport-session

# Conflicts:
#	.github/workflows/ci.yml
… lab/holoport-session

# Conflicts:
#	flake.nix
…holoport-session

# Conflicts:
#	docs/deployment.md
…a-event-node

Once #59 and #61 are both in, hosts/common.nix no longer reads fleetLine and fleetHapps: the fleet imports nixosModules.sensorica-event-node for them. holoportTarget still passed the old arguments, so the installed system had no hApps, no installer unit, and vmTestHoloportInstall failed ("holochain-happ-installer.service is inactive and there are no pending jobs"). It now imports the module the fleet imports, and the check passes: hrea, kando and requests-and-offers enabled 20 to 43 s after boot.
A NixOS Holoport can run more than a Holochain edgenode (a Moss node, a
bootstrap server, Grafana), so the machine is named for what it is and the
edgenode stays a role. Renames the host directories, the flake outputs, the
CI host list, the scrape-target examples and the docs.
One word rebuilds whichever Holoport it runs on, from the checkout the
install leaves in /root/nixos-holochain, the way athanor's machines do.
hosts/desk.nix gives the sensorica session five launchers (fleet
dashboard, this node, Holochain logs, Moss node, rebuild), a Plasma
layout declared with plasma-manager (bottom panel, Breeze Dark, CPU and
RAM monitor, no screen lock or suspend), tmux, btop and hc, and Avahi so
sensorica-holoport-01.local resolves in the lab. sensorica.eventMode
(off by default) logs sensorica in and opens the room dashboard full
screen, with read-only anonymous Grafana on the monitor node.

home-manager (release-26.05) and plasma-manager are inputs of the
example only; the modules stay desktop-free.
desktop-file-validate rejects a '?' outside quotes in Exec, which failed
the system build on sensorica-holoport-01.
…#28)

nixosModules.holochain-moss-node runs wdocker's daemon under its own
user, the conductor password passed as a systemd credential from a
root-only file and read on stdin, so the node needs no terminal and no
tmux. Joining the group stays one command at a terminal, since the
invite link carries the network seed: moss-node join "INVITE_LINK" runs
wdocker join-group as the node's user and restarts the daemon.

The readings and the Sensorica Moss page are ported from athanor's
mossNodeMetrics (athanor 539e648): the shared conductor exporter under
conductor="Moss", names from Nix, the page in its own provider with a
configurable title. vmTestMossNode starts the service headless, checks
the readings and a restart, and fails when the password file is missing;
moss-dashboard and moss-names come over with their broken-copy checks.

The Sensorica fleet enables it on sensorica-holoport-01.
Plasma's power management suspended sensorica-holoport-01 from the login
screen, taking its conductors and dashboards offline. Disabling the
sleep, suspend, hibernate and hybrid-sleep targets makes suspend
impossible whatever the desktop or logind asks.
The Holoports install from this branch, so the example reads the modules
it carries (event node, Moss node) rather than main, which has neither
yet. A fresh clone then needs only the operator key before the install.
Point it back at main once the lab branch merges.
docs/moss-node.md opens with nixosModules.holochain-moss-node (options,
the credential, moss-node join/status/logs, readings and page, checks)
and keeps the by-hand path. The fleet README covers the host names, the
pin to lab/holoport-session, the rebuild alias, never-sleep, the
operator desk with event mode, and the Moss node on sensorica-holoport-01.
…ing node

On sensorica-holoport-01 an empty password file made wdaemon create the
conductor with an empty password and exit 0, which Restart=on-failure
left down. ExecStartPre now refuses an empty credential with a message
naming the file, Restart=always brings the node back however it ends,
and moss-node join stops before its prompts when the node is not running
(join-group otherwise asks everything, then fails on 'No port file
found'). vmTestMossNode gains an empty-password machine that must log
the refusal and create no conductor.
sensorica-holoport-01 drives two screens and the panel only reached the
first.
Opens the monitor's Grafana home from the sensorica desktop.
In /root the sensorica operator could not open it. /etc/nixos-holochain
is root:wheel with setgid, git marks it safe for every user and keeps
its objects group-writable, so the checkout can be edited from the desk
while rebuild still runs as root. The rebuild alias, the Rebuild
launcher, the fleet README and the install steps follow.
Every unit a module lists in services.holochain-services.units can now
declare the version of the package it runs (version) and, for a unit that
runs a conductor, the Holochain that conductor is (holochainVersion). Both
are published as labels of holochain_service_info, so they ride through
holochain:service_watched into holochain:service_state.

The modules fill them from their packages: the conductor and the app
installer from the edgenode's package and hcPackage, the HTTP gateway and
the bootstrap server from theirs, Prometheus, Grafana, node_exporter,
sshd, Tailscale and the Nix daemon from theirs, and the Moss node from
wdocker, whose package now passes through the Holochain version it brings.
A unit without one publishes no version label. node_exporter also exports
each unit's start time (--collector.systemd.enable-start-time-metrics) on
the edgenode and the monitor.
A fifth dashboard, holochain-home, provisioned with the others and set as
Grafana's home page, so / opens on it after login. At the top, the machine
in words: its state, host name, NixOS release, kernel, uptime, and how busy
its processor, memory and fullest disk are. Below, every service it runs,
worst first, with its state in the node page's words and colours, the
version Nix runs it from (a dash where a unit declares none) and when it
last started; then its Holochain conductors, each with the service that
runs it, its Holochain version and whether it answers, and every app on
them in the six state words, each linking to the node page.

A node picker shows any machine of the fleet. The module's rewrite defaults
it to this machine: the name of the scrape target on a loopback address,
else networking.hostName. Links lead to the fleet, node, network and room
pages, and to any dashboard tagged moss, so the Moss page shows where the
Moss module provisions it. A dashboards directory without the new page
falls back to the room screen as home, as before.

The label check now reads host name, system and kernel release, and the
declared versions as names a person reads.
dashboardLabels and dashboardQueries take in the home page: its columns and
descriptions, its uid, and, on the homelab fixtures with versions on two
units, a services table that shows each version and a dash where there is
none, and a conductors table with one row per conductor. Copies that lose
the versions or repeat a claimed conductor must fail.

grafanaProvisioning compares every version the modules declare with the
package each unit runs, a Moss node's included, and requires the published
list to carry each as a label; copies with one label removed must fail by
the unit's name. It requires the home page to be the provisioned
holochain-home, opening on the loopback target or the host name, and
refuses a copy without it and a copy whose home page lost its uid.

vmTestGrafana expects five dashboards, holochain-home at home on machine,
the home page's state words, and the conductor's versions in Prometheus.
sensorica.desktop picks KDE Plasma, GNOME or no graphical session. The
desk keeps a single Grafana launcher, on the home page, beside the
Holochain and Moss logs and Rebuild; the stale tmux and per-page
launchers go. Each host file opens with the switches an operator flips
by hand before running rebuild.
Instant queries come back from Prometheus in no fixed order, so the
machine tiles, the per-node bars and the app-by-node grids shuffled
between refreshes. Each now sorts on the node label first.
book.toml builds docs/ into a gitignored book/, with a SUMMARY organised
for three readers: installing a Holoport, running the modules, and
contributing. The fleet example, workshop, template and happs READMEs and
CONTRIBUTING.md appear through {{#include}} wrapper pages without moving.

The Deploy Documentation workflow builds the book on pull requests that
touch its sources and deploys it to Pages from main, with mdBook pinned to
the devShell's 0.5.2. mdbook joins the devShell, the README links the
site, and deployment.md's link to scripts/holoport-install.sh, which
leaves docs/, becomes an absolute GitHub URL.
mdBook 0.5 expands {path} relative to the repo root, so the edit links
pointed at docs/docs/<page>.md. The template now ends in {path}.

The Pages artifact and deploy job now also require refs/heads/main, since
workflow_dispatch can run on any branch.

architecture.md's "Holochain conductor (<conductor>)" lost its placeholder
on the page, parsed as an unclosed HTML tag; it is now in backticks.
…s named

A monitor that listed itself by its own host name or FQDN under a key of
its own (lab-1.address = "monitor:9100") defaulted the home page's node
to its host name, which no target carried, so Grafana opened on the
first node in the list. A target at this machine's host name or FQDN is
now this machine too, after a loopback target. grafanaProvisioning
checks both forms, and fails on the previous module.
The shared config installs one app, so holochain-happ-installer.service
is listed; its version must equal hcPackage's, and a copy of the
published list without its version label must fail by its name.
"Each app across the fleet" carried the home page's node, so the network
page came up for one machine and "how many run it" read 1. It no longer
carries variables. dashboardLabels now refuses any link by address that
carries a page's variables to a page whose node picker opens on All, and
a copy that sets it back fails by the link's name.
…port

The icon opened the monitor's bare /, whose home page defaults to the
monitor itself, so on holoport-02 to 05 it showed holoport-01. It now
opens /d/holochain-home with this machine's node, as the node launcher
already does.
sensorica.remoteAccess.enable runs the Tailscale client against
https://hs.sensorica.co with a pre-auth key read from a root-only file,
and trusts the tailnet interface. Off on every Holoport until the
Headscale server exists; each host's switches block carries it.
…session

Merges feat/sensorica-holoport-names: "What is this machine running?" as
Grafana's home page, the version each listed unit runs, their checks and
docs, and the review fixes.

One conflict, in examples/sensorica-fleet/hosts/desk.nix: the lab side
rewrote the desk around the sensorica.desktop switch with one Grafana
launcher and one desktop icon, while the home page branch pointed its
icon at the home page on this Holoport. The lab structure is kept, and
both the holoport-grafana launcher and Desktop/grafana.desktop now open
the same page, /d/holochain-home with this machine's node, since the
monitor's bare / opens the home page on the monitor itself.
Conflicts, resolved as main's reviewed code unless lab carries a later
fix of the same lines:

- flake.nix: the Moss node module export is lab's and is kept. The
  Holoport install block (onX86, holoportInstall, holoportTarget,
  holoportDisk) and edgenodeBinaryCache sit where main moved them;
  lab's copies at the old places are dropped. holoportTarget takes the
  renamed host, hosts/sensorica-holoport-01.
- .github/workflows/ci.yml: main's build list, which already carries
  vmTestWdocker and adds the metrics checks.
- docs/architecture.md: lab's vmTestGrafana paragraph (five dashboards,
  holochain-home) with main's fix that node state 2 is A service is down.
- docs/deployment.md: main's install text (Esc boot menu, USB 2, the
  root password outside a terminal) with lab's host names, lab's
  absolute script link for the book, and lab's /etc/nixos-holochain
  checkout step, its code fence closed.
…icense in its introduction

main added docs/adr/ (#51) and docs/releasing.md (#49), which the book's SUMMARY did not list, and licensed the repository MIT (#36), while the introduction still said AGPL-3.0.
The Moss node module (#28) came with its VM test and two checks without a VM, and none of them was in the build list.
…sk switches

The renamed host file sets sensorica.desktop and sensorica.eventMode,
which hosts/desk.nix declares, and enables the Moss node. The install
target imported neither, so vmTestHoloportInstall, and with it
nix flake check, failed to evaluate: "The option `sensorica.desktop'
does not exist".

The target now imports the Moss node module, as the fleet does, and a
stand-in for the desk that declares its two switches and keeps the
Plasma session with SDDM for "plasma", the desktop main's target had
from hosts/common.nix. desk.nix itself needs home-manager and
plasma-manager, which this flake does not carry.
…er is not packaged

The Moss facts (ee8dba5) evaluate the Moss node's package, and wdocker
exists for x86_64-linux only, so nix flake check --all-systems failed on
aarch64-linux with "attribute 'wdocker-0_15' missing". The facts and the
three checks that read them now run on x86_64-linux only; every other
system keeps the rest of the check.
… textfile

Stopping the timer leaves a run it already started going, and that run
can finish after the test writes the bad file and replace it with good
readings, so "A metrics file could not be read" never appears and the
test times out. Seen locally at 248 s: holochain-conductor-metrics.service
finished right after the write. The service is now stopped with the timer.
…rst boot

sensorica-holoport-01 now enables the Moss node, which reads its conductor password through LoadCredential. The runbook wrote only Grafana's password file, so a fresh install booted with moss-node.service failing on status=243/CREDENTIALS and the machine reading "A service is down". Step 3 now writes /mnt/var/lib/secrets/moss-node-password, step 4 checks the unit is active and points at the join step in docs/moss-node.md.
@Soushi888

Copy link
Copy Markdown
Contributor Author

SoushAI analysis. Drafted by Soushi's AI assistant, reviewed and posted by @Soushi888.

Review fixes before merge

Two reviews ran on 3480cd9. One finding was blocking, and 1c591b4 fixes it.

Fixed: the runbook never wrote the Moss node password. sensorica-holoport-01 now enables the Moss node, which reads its conductor password through LoadCredential. Step 3 of docs/deployment.md wrote only Grafana's password file. A fresh install would boot with moss-node.service failing on status=243/CREDENTIALS, retrying every 30 s, and the machine would read "A service is down".

  • Step 3 now writes /mnt/var/lib/secrets/moss-node-password (mode 0400, no trailing newline), using the same systemd-ask-password | install line as docs/moss-node.md.
  • Step 4 now checks that moss-node is active, and points at the moss-node join step (with the leading-space habit that keeps the invite link out of shell history).

Checks at 1c591b4

  • mdbook build: no warnings. The new links resolve in the built book: moss-node.html#as-a-nixos-service matches the heading id.
  • nix flake check --no-build --all-systems: exit 0.
  • The password line, run locally against a scratch directory with a stand-in password, wrote a 0400 file of exactly the password's bytes.

Not blocking, left for after the merge. These are in addition to the follow-ups in the PR body:

  • Relock the example fleet to main, and fix the README lines that say main does not carry the modules yet. This comes first.
  • Stale prose:
    • The README test table.
    • hosts/edgenode-XX at deployment.md:36.
    • "four provisioned dashboards".
    • "Holochain Fleet".
  • There is no path 2b checkout in /etc/nixos-holochain, and the operator key is left as an uncommitted edit in 2a.
  • services.holochain-moss-node.* is missing from the generated options reference.
  • passwordFile is types.str, not pathWith { inStore = false; }.
  • The docs workflow's concurrency group is shared by PR builds and the deploy.
  • The Moss page's PromQL is never run by a check.
  • The CHANGELOG has no entries.
  • The book's Contributing page has dead links.

One correction to the PR body: GitHub Pages is already enabled (build_type: workflow), so the book deploys when this merges.

@Soushi888
Soushi888 merged commit 666c5ae into main Sep 28, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant