Build stamp catch-up for 1 merged PR(s) since v1.0.0.0.21:
40fa250 refactor(tests): move the sanitization denylist to a data file both sides can read (#66)
The rules lived inside Sanitization.Tests.ps1, so the publisher in the
private toolkit kept its own second list -- and the two guarded different
things. The publisher's was about SECRETS: keys, private-key blocks, ssh
targets. These are about IDENTITY: internal project names, product domains,
private-only script names, operator paths.
So the publisher reported "clean" on files this suite rejects, and would
have published a tree that fails the public repo's own tests
(evo.scripts#106). Proven at the time by copying the private ZHelpers.ps1
in: two failures naming EvoCivilCode, EvoPlatform and three private-only
script names, against a scan that called the same file clean.
tests/sanitization-patterns.psd1 is now the one source. The suite reads it
and refuses to run if it is missing or empty, rather than passing vacuously
against no rules -- an empty denylist that reports success is the failure
this whole fix is about.
No rule changed. Only where they live.
.psd1 is not in the scanned extension list, which is deliberate and matches
why Sanitization.Tests.ps1 excludes itself: a file that necessarily contains
every pattern it looks for cannot also be scanned for them.
Pester: 240 passed, 0 failed. CHECKSUMS regenerated; changelog and its
plain-text twin updated.
The mirror's deploy pair had drifted ~280 lines behind: it lacked the
transactional .env preserve/restore (an interrupted deploy could destroy
server-side env files), the stderr-flattening step wrapper (a successful
deploy reported failure and skipped its own verification), and the
verification rework (channel re-picked every retry, edge only with a Host
to route by, verify.timeoutSeconds, honest split of "stale build" vs "no
channel answered").
The sync is byte-faithful to the private tree except where the mirror's
own Sanitization suite demands otherwise - and it caught the first copy:
three failures for private project names, a private domain, and a
private-only script name that rode along in comments. Each war story keeps
its lesson and loses its cast, per the convention already in the file
("EvoCivilCode: deploy/" was already published as "(deploy/, infra/,
...)"). That suite going red on an unsanitized copy is exactly what it
exists for.
Also in this change:
- tests/VerifyPlan.Tests.ps1 - the channel-selection rules are pure
functions and Pester pins them (no domain => no edge attempt; the PS 5.1
one-element-unroll trap). First verification tests in the mirror.
- README: the verify block now documents viaProxy/upstream (they shipped
in the docker-network read but were never in the README),
timeoutSeconds, and the channel order with why it re-resolves per retry.
- README.txt: generated plain-text twin, via scripts/readme_txt.py
(vendored from the fleet's reference implementation; the file is
generated, never edited by hand).
- CHANGELOG.md/.txt: entries merged into the existing Unreleased sections.
CHECKSUMS.txt refreshed (42 entries). Pester: 240 passed, 0 failed.
Both synced files parse clean.
Observed, untouched: the Unreleased section carries duplicate "### Fixed"
headings from earlier appends; folding them risks reordering entries whose
prose references their neighbours, so it is left for the next release cut
(zbump #110 rolls Unreleased into the version being cut).
zversion's usage block and its bump help line both said 'one per PR, one per
defect fix'. The build counter advances once per release: the number names
something that shipped, so a release carrying five PRs moves it by one, and
PRs that never shipped on their own were never separate builds.
This is help text rather than behaviour, but it is the wording that gets
followed - it is what stamped a single evo.www release as two builds. The
matching comments in ZHelpers.ps1 and zdeploy.ps1 are corrected with it.
CHECKSUMS.txt regenerated for the three edited scripts, since the manifest
tests fail the moment it drifts. CHANGELOG entry added under Unreleased, with
its .txt twin. The older CHANGELOG entry recording what the rule was when
zversion shipped is deliberately left as written - a changelog describes what
happened, not what is currently true.
Pester: 231 passed, 0 failed.
Build stamp catch-up for 1 merged PR(s) since v1.0.0.0.18:
e738b62 feat(version): read a live build from inside the docker network, not the public proxy (#59)
zdeploy, zec2 and zec2online now prefer
docker exec <viaProxy> curl http://<upstream>/api/build-version
when a project sets verify.viaProxy and verify.upstream.
Two problems it closes. A build stamp is something many sites deliberately
do not serve publicly, and a checker that reads it over the public URL stops
working the moment that endpoint is blocked -- reporting "unknown", which is
indistinguishable from "could not reach it". And the proxy answers from
whichever vhost matches the Host header, so a container with no public route
was getting another site's version back and failing deploys that had worked.
Reading it from a container on the shared network also exercises the real
HTTP path, so it proves the app is serving rather than that its database
knows a version. Purely additive: a project without those two keys behaves
exactly as before.
zec2 gains $PemKey and $SshTarget, which it had no need of until now -- the
read is inside a try/catch, so without them it would throw, be swallowed,
and fall through silently.
CHECKSUMS.txt regenerated (zchecksums -Update), zconfig.example.json
documents both shapes of the verify block, and CHANGELOG.md carries its
plain-text twin.
Pester: 231 passed, 0 failed -- including the sanitization suite.
Build stamp catch-up for 4 merged PR(s) since v1.0.0.0.14:
4230566 feat: two blank lines after every z-script run (#57)
432f210 fix(zdeploy): the pre-zip line states the label verification will require (#56)
fcbad44 feat(zdeploy): add the static deploy kind, and a test that keeps this repo sanitized (#55)
7b32bfa fix(zdeploy): docker kind now ships config subdirectories (#54)
zdeploy ended with two blank lines and nothing else did, so its output
was the only one that did not butt up against the next prompt. Applied
everywhere.
Implemented in Stop-ZTracking rather than per script, because every
tracked script ends by calling it - including the usage and guard paths
that do `Stop-ZTracking; exit 1`. One place therefore covers every exit
of sixteen scripts.
* Write-ZTrailer emits the two lines. Stop-ZTracking calls it on both
paths, including the early return when tracking never started - a
script that printed output still deserves the separation.
* -FinalNote prints one last line AFTER the tracking footer and BEFORE
the blanks. zdeploy's "Last deployed at ..." now goes through it, so
it stays the last thing on screen instead of being followed by the
token count.
* The standalone tools (zchecksums, zversion, zrelease) get a local
three-line copy at each exit rather than dot-sourcing ZHelpers -
they are deliberately dependency-free so they run from an extracted
release zip with nothing beside them.
* Redundant trailing blanks were removed where a script already
printed one before exiting, so the count is exactly two, not three.
Also fixes a related gap found while testing: the error exits inside
Get-ZConfig and Get-ZProject bypassed Stop-ZTracking entirely, so a run
that died on a missing zconfig.json printed no trailer AND no token
footer. Those three exits now route through Stop-ZTracking.
zkill/zrestart are one-line wrappers around scripts that already trail,
so they are untouched - adding it would double the blanks.
Verified by running each script and counting the trailing blank lines in
its captured output: 2 on every success path, every usage path, and the
config-error path. Suite 231/231. CHECKSUMS.txt refreshed.
Mirrors evo.scripts#79. Invoke-ViteDeploy announced "(server-side build will
bump +1)" and then Wait-VerifyStaticBuild asked for the committed stamp
unchanged, so the operator was told to expect one label and watched a check
pass on another.
The behavior was already correct; the message predates the removal of the
prebuild hook that used to self-bump inside the image. Wait-VerifyStaticBuild
documents it: "Expect the COMMITTED stamp, not +1: builds no longer
self-bump ... so 'is the build I just packed live?' means an exact match."
CHECKSUMS.txt regenerated, since zdeploy.ps1 changed.
Two things, both prompted by the same incident.
STATIC KIND
Plain static sites - no build, no container of their own; a shared web
container serves them off disk, so shipping the files IS the deploy.
Directories are staged to a sibling and swapped in with mv rather than
copied in place, because a large media file uploaded in place is served
half-written to anyone who requests it mid-copy. The swap is a rename,
so the switch is atomic.
Skips $JunkDirNames + .github + deploy.skipDirs, matching the docker
kind rather than inventing a third convention.
SANITIZATION TEST
This repo is public and must stay standalone, but the toolkit is
developed in a private checkout and copied here. Twice in three days a
wholesale copy landed carrying real project names, internal hostnames,
host disk figures, and references to scripts that exist only in the
private copy. Both times a human reading the diff caught it - the
control that fails exactly when a diff is 280 lines of good work with
three bad words buried in it.
So it is a test now. Seven rules: private project names, private product
domains, private-only script names, local drive paths, the operator home
path, the real key filename, and any IPv4 outside RFC 5737 documentation
space and the private ranges.
Two anti-vacuity guards, because a denylist that silently matches
nothing is worse than no denylist: one asserts the file scan is
non-empty, and one plants a known violation and requires the pattern to
find it.
The IP rule bounds on [\d.] rather than \d deliberately - this toolkit's
own 5-segment version (v1.0.0.0.14) contains "0.0.0.14", which a plain
digit boundary reads as an address.
Verified against the real unsanitized copy that slipped through: 3 of
the 7 rules fire, naming file and line.
Tests: 231/231 (222 + 9 new). CHECKSUMS.txt refreshed.
Invoke-DockerDeploy uploaded top-level files ONLY. For any stack that
keeps configuration in a directory, that is silently wrong in the worst
way: the compose file arrives, the containers restart, and the config
they read is whatever was already on the box. The deploy reports
success while shipping nothing that matters.
The sharp case is a provisioning directory read at container start -
alert rules, datasources, mounted config - where the restart makes it
look like the change was applied.
Same fix the edge kind already got: subdirectories ship recursively.
Skipped are $JunkDirNames, plus .github (CI config belongs in the repo,
never on a deploy target - it is not in $JunkDirNames because backups
DO want it), plus anything the project lists in the new
deploy.skipDirs.
deploy.skipDirs exists for SERVER-SIDE STATE that shares the tree: a
data/ holding mailboxes or a time-series database must never be
overwritten by whatever the local checkout has - usually nothing, which
is the dangerous case, since scp -r of an absent directory is not the
no-op you want to rely on.
CHECKSUMS.txt refreshed alongside. Tests: 222/222.
Build stamp catch-up for 3 merged PR(s) since v1.0.0.0.11:
6c30f40 feat(zdeploy): sync deploy hardening from the private toolkit (#53)
8959d9f fix(zdeploy): stop passing ssh -n to scp, which rejects it (#52)
d99f8a1 fix(zdeploy): stop mirroring the build stamp into the local checkout (#51)
Nine fixes that had accumulated only in the private copy. Each one is a
production failure that already happened:
* ssh -n on the deploy path. Without it ssh reads its stdin, inherits
the console handle under PowerShell, and can block forever. The
existing timeouts cannot catch it - they bound a connection that is
dying, and this one was never established. Get-Ec2ScpOpts derives
from the same list minus -n, because scp rejects it with a usage
error that reads like an unrelated failure.
* Invoke-DeployGitPull now switches TO the default branch instead of
pulling whatever is checked out. Pulling the current branch breaks
as soon as the remote deletes branches on merge, and worse, could
ship a feature branch to production. Refuses over local changes,
excluding the files the deploy itself stamps.
* nextjs handler resolves composeDir. It ran compose against
remote.path, so a project whose compose lives in a subdirectory
recycled whatever docker-compose.yml sat at the project root -
usually the local dev one shipped in the same archive. Two deploys
in a row exited 0 having shipped nothing, and verification passed
because the untouched old container still answered.
* BUILD BEFORE DOWN. The old order took the site offline for the whole
build and left it offline if the build failed - one deploy served
502 for 19 hours with no container running. The old image now keeps
serving until the new one is ready.
* compose build takes the named service, so a compose file that does
not define it fails loudly instead of succeeding with nothing to do.
* deploy.verifyHost overrides domain for verification only, for when a
domain is retired ahead of its replacement.
* deploy.stampCmd writes the build version from git before zipping,
then reverts the tree so the stamp cannot block the next deploy.
* Optional ztokens pseudo-project and passive usage records. Both
degrade to a printed skip when no sibling ztokens checkout exists.
Sanitized on the way across, per the public-repo rule: the war stories
that make these comments worth reading are kept, the private product
names, domains, internal script names, outage dates and host figures are
not. Also dropped a "See issue #28" pointer that resolves to an
unrelated PR in this repo, and two references to scripts that exist only
in the private toolkit.
CHECKSUMS.txt refreshed alongside, per the manifest rule.
Tests: 222/222 (the 4 failures before the refresh were the checksum
suite correctly flagging both edited files).
Mirrors evo.scripts #66, which added -n to Get-Ec2SshOpts so a deploy step
cannot block on inherited stdin, plus the follow-up that keeps that flag away
from scp. OpenSSH's scp has no -n: it exits 1 with 'unknown option -- n' and
prints its usage block, so every upload failed once the shared option array
reached it.
Get-Ec2ScpOpts is derived from Get-Ec2SshOpts with -n filtered out rather than
duplicated, so the connect and keepalive timeouts cannot drift apart between
the two transports. All four scp call sites here use it - three plain uploads
and the recursive directory upload, which the private tree does not have.
The upload failure message asserted 'Likely server disk space' without checking
anything; it now points at scp's own output, where the real diagnosis already
was.
Verified: both files parse clean, Get-Ec2ScpOpts returns the five -o pairs with
no -n, and the added lines carry no real hosts, keys, or paths.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Mirrors evomedia-net/evo.scripts#62. Every python-kind deploy wrote
build-version.json locally, leaving the tree dirty; committing it hit
branch protection, so a deploy either tripped the next deploy's
clean-tree guard or bypassed the rule.
The build number is bumped in the container, written to .build_version
on the server, and proven by /api/build-version. The repo file records
the stage baseline only.
Build stamp catch-up for 1 merged PR(s) since v1.0.0.0.10:
a1dd642 chore(changelog): drop the product name from the versioning entry; add the plain-text twin (#50)
'the SmartPlant 5-segment scheme' was the only product-name reference
left anywhere in the public tree (found by a full sanitization sweep:
paths, keys, domains, IPs, emails, ssh details, sibling-repo coupling
- everything else already clean). The scheme description stands on its
own without naming where it came from.
CHANGELOG.txt is the plain-text mirror the docs rule requires for any
touched .md - generated mechanically (headings underlined, markup
stripped, code blocks indented).
CHECKSUMS.txt is untouched: it covers .ps1/.cmd only.
With deploy.gitPull set, Invoke-DeployGitPull ran `git pull --ff-only`.
Pull merges whatever .git/FETCH_HEAD marks "for merge", and a
concurrent fetch in the same repo - an editor's background auto-fetch
racing the deploy - can leave duplicate for-merge lines. Pull then dies
with "Cannot fast-forward to multiple branches" even when both lines
name the same commit. Reproduced live on 2026-08-13: the deploy failed
at step 0, and by inspection time the same pull succeeded, with the
duplicated FETCH_HEAD entries as the only evidence.
Now the function fetches explicitly and fast-forwards against the
remote-tracking ref:
git fetch origin
git merge --ff-only origin/$branch
Identical fast-forward-or-refuse semantics, but nothing shared and
mutable in the path. A failed fetch and a failed merge now also throw
distinct messages, so the operator knows which half broke.
CHECKSUMS.txt refreshed alongside, per the manifest rule.
Tested: PS 5.1 parse clean; full Pester suite 222/222 after the
manifest refresh (the 3 pre-refresh failures were the checksum suite
correctly flagging the edited file). The same fix pattern was
exercised against live and scratch repos for the private twin
(evomedia-net/evo.scripts#58): behind -> fast-forwards, diverged ->
refuses and aborts the deploy.
Closes#48
tests/Invoke-Coverage.ps1 runs the suite with Pester's profiler-based
coverage collector and writes both JaCoCo XML and a coverage-summary.json
in the same shape vitest and jest emit, so one reader handles every tool
in the fleet.
CodeCoverage.Path is every *.ps1 in the repo root, including the ones no
test touches. They report 0% and drag the total down, which is correct -
measuring only the already-tested files answers "how well covered is the
covered code". The same mistake in vitest form had evo.www reporting 93%.
Baseline: 222 tests pass, 7.0% of commands (219/3,125 across 24 scripts).
Pester counts commands, not statements; the collector labels it as such
rather than filing it under a unit it is not.
Build stamp catch-up for 8 merged PR(s) since v1.0.0.0.0:
6c03138 fix(zdeploy): deploy edge first for any project list, not just 'all' (#46)
976eb6d chore: point repo URLs at evomedia-net/evo.* (#44)
0e4d9c3 fix(zdeploy): ship edge asset subdirectories, stop deploys hanging on ssh prompts, keep vendored archives (#43)
c8fcfee fix(zdeploy): verify Next.js deploys on the server instead of probing the public IP (stops the bogus security-group warning) (#41)
b84d24b fix(zdeploy): report where a domain-less project deployed to (#40)
4668404 fix(zdeploy): re-land the two blank lines after the timestamp (#38)
a868973 fix(zdeploy): verify against the committed build stamp, and read 5-segment versions (#37)
0ddffc1 feat(zdeploy): print 'Last deployed at' timestamp as the last line of a run (#36)
Closes#45.
'zdeploy evo edge' deployed evo first, then edge. The edge-first sort was
nested inside the 'all' branch, so an explicit list walked in whatever order
was typed.
That ordering is a correctness property, not a convenience of 'all': the proxy
has to route before the apps behind it ship, or there is a window where a new
app is live behind stale routing. 'zdeploy evo edge' reads as 'these two, edge
included' and quietly did the risky order.
Hoisted the sort out of the branch so it applies however the list was produced.
Order within each group is preserved, so an intentional sequence still holds -
and when the order does change, it now says so rather than reordering silently.
Verified across six shapes: evo+edge reorders and reports it; edge+evo is left
alone with no note; edge last in a longer list moves to front keeping the rest
in order; single project and no-edge lists are untouched.
All 16 repos moved to the evomedia-net org and were renamed into the evo.*
namespace, so every github.com/kellymichels/<old-name> reference in source
headers, CI badges, security links and docs pointed at a redirect.
Mechanical URL-only rewrite, applied longest-name-first so smartplantehs-docs
could not be clobbered by the smartplantehs rule. Nothing else changes: no
code, no product names, no behaviour. smartplantehs -> evo.ehs here is the
REPO url only; the product rename is separate and still pending.
* fix(zdeploy): edge deploy ships asset subdirectories, not just top-level files
A project self-hosting assets in folders (fonts/, vendor/) lost them on
every deploy: docker created empty root-owned mount points and nginx
served 404s from them, so fonts fell back silently and vendored JS
never loaded. Every subdirectory except nginx-logs/ and .git/ now ships
recursively, and the ensure-dir chown is recursive so scp into
docker-created root-owned dirs cannot fail.
Closes#42
* fix(zdeploy): stop deploys hanging on ssh prompts, keep vendored archives, make unzip install idempotent
- Get-Ec2SshOpts: every deploy-path ssh/scp now carries BatchMode=yes
plus connect/keepalive timeouts. Without BatchMode ssh prompts and
waits forever, and because the deploy pipes stderr through the
pipeline the prompt never reaches the screen - the run just stops
under the last step label with no explanation.
- Archive filter exempts files under vendor/: a vendored *.tgz is a
build input, and dropping it fails a Dockerfile COPY at image build.
- unzip install checks command -v first instead of an apt round-trip
on every deploy.
Ported from the private script; three tests cover the vendor/ exemption.
* fix(zdeploy): verify Next.js deploys on the server, not the public IP
The nextjs handler verified by requesting http://<ec2-ip>:<prod-port>/.
A compose stack behind the edge proxy normally publishes to 127.0.0.1
only, so that request can never be answered and every deploy ended with
"is the port open in the security group?" — pointing at a firewall rule
for an app that was already serving fine. No security-group change could
have made that probe succeed, which is what made the warning actively
harmful: it named the one fix guaranteed not to work, and opening the
port would have exposed the app directly, bypassing the proxy and TLS.
The nextjs handler now uses the precedence the python handler already
used: verify.port (curl localhost on the server) -> domain (Host-header
check through the edge proxy) -> an honest "NOT verified" instead of a
misleading warning.
Also follow redirects in the health check. An app whose "/" answers 307
(Next.js -> /login) returns a body of a few bytes that read as "not
ready"; -L fetches the page that actually renders. No effect on projects
that point verify.path at a plain 200 endpoint such as /health.
The example config now documents the verify block on the nextjs project,
and CHECKSUMS.txt is regenerated for the changed script.
* fix(zdeploy): truncate the health-check body in the PASS line
Pointing the nextjs handler at Test-DeployHealth means the body is now
often an HTML page, not a line of JSON, so the PASS line dumped ~10 KB
of markup into the deploy output and buried everything after it.
Match on the full body as before, but print at most 200 characters plus
the byte count. A /health endpoint still prints in full.
Closes#39.
A project without a public domain finished with no indication of where it went
- the 'Site:' line was gated on $Proj.domain in all four handlers, so an
internal service deployed successfully and said nothing about the result.
The information was never missing, only unused: a domain-less project that
deploys almost certainly has a verify block (the documented pattern for stacks
not published through the edge proxy), carrying the exact port and health path
the deploy had just checked. zdeploy verified against that endpoint and then
discarded it.
New Write-DeployLocation falls back through what the project actually declares:
domain -> public URL; else verify.port (+ path) -> server-local URL flagged as
not publicly routed; else ports.prod -> same without a path; else prints
nothing, as before. Replaces the four call sites, keeping the nextjs handler's
column alignment via -Pad.
Exercised against every shape: domain, verify-only, ports.prod-only, nothing at
all, and padded. CHECKSUMS regenerated; suite 219/219.
These were written and pushed, but landed on the PR #36 branch AFTER that PR
had already merged, so they never reached master. Private main has them; public
did not. The two copies of zdeploy.ps1 had quietly diverged on the last line of
output.
CHECKSUMS.txt regenerated. Tail now byte-identical to the private copy.
Deploy verification expected buildNumber + 1, which was only correct while
a prebuild hook self-incremented the stamp during the image build. That hook
is gone (it produced versions matching no commit and left every build with a
dirty tree), so the check waited out its timeout and warned on every deploy
even when the deploy had succeeded.
Both verifier call sites now expect the committed label, and
Get-LabelFromBuildJsonObj understands the 5-segment scheme
v{major}.{rc}.{beta}.{alpha}.{build} while still reading the legacy
{productVersion, buildNumber} stamp for projects that haven't migrated.
Verified: both files parse; label fn returns v0.0.0.1.8 (5-segment) and
v1.0.0.70 (legacy).
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(zdeploy): print 'Last deployed at' as the final line of a run
Scroll to the bottom of a deploy and you can see how long ago it happened,
without digging through the log or checking the server.
Placed after Stop-ZTracking so it is genuinely the last line, below the
token-tracking footer. Only reached on success - a failed deploy throws out of
the dispatch loop, so this never claims a deploy that did not happen. Prints
once per run rather than per project, so 'zdeploy all' ends with a single
timestamp.
24-hour clock (MM/dd/yyyy HH:mm:ss): the requested format carried no AM/PM
marker, and 24-hour is unambiguous for judging elapsed time at a glance.
CHECKSUMS.txt regenerated, since changing a covered script invalidates it.
* feat(zdeploy): use 12-hour clock with AM/PM in the timestamp
Kelly's preference: 'Last deployed at 08/03/2026 01:24:09 PM' rather than
24-hour. The AM/PM marker keeps it unambiguous.
CHECKSUMS.txt regenerated.
* feat(zchecksums): SHA-256 manifest so a download can be verified before it's run
CHECKSUMS.txt lists a SHA-256 for every top-level .ps1 and .cmd - the files a
user actually executes. zchecksums verifies them; zchecksums -Update
regenerates after an intentional edit.
The manifest is sha256sum format, so 'sha256sum -c CHECKSUMS.txt' works on
Linux/macOS/WSL as well as the PowerShell path on Windows. Hashes are identical
on every platform because .gitattributes pins .ps1/.cmd to CRLF everywhere -
that pin is now load-bearing, so it is commented as such.
Beyond changed and missing files it also reports a script that is on disk but
NOT in the manifest, so something added outside a commit still gets noticed.
Exits non-zero on any of the three.
Honest about its limits, in the header and the README: the manifest lives in
the same repo as the code, so it is an integrity check rather than a signature.
It catches a truncated clone, a forgotten local edit, or an unlisted file - not
a compromised repo.
CHECKSUMS.txt is pinned to LF: sha256sum treats a trailing CR as part of the
filename and would report every entry as missing on Linux.
tests/Checksums.Tests.ps1 keeps it from rotting - a stale manifest is worse
than none, since it either cries wolf until people ignore it or quietly stops
covering a new script. The tests assert the format, LF endings, sort order,
full coverage of on-disk scripts, current hashes, and that zchecksums itself
exits 1 on a tampered file (proved by appending a byte and restoring it).
* feat(zversion, zrelease): toolkit versioning + downloadable release zips
Implements the versioning rule (SmartPlant's 5-segment scheme, now the global
standard; currently only sp and zscripts are on it at v1.x):
v{major}.{rc}.{beta}.{alpha}.{build}
zversion: get / bump / bump-stage / set. A stage bump zeroes every lower
segment including build. 'bump' is one per PR and one per defect fix, not per
file. Any write rewrites three things together, because they are only useful
when they agree: build-version.json (source of truth), a '# Version:' line in
all 42 script headers (a lone copied script still says which release it came
from), and CHECKSUMS.txt (stamping changes every file).
zrelease: packages the current version as releases/zscripts-<version>.zip with
a sibling .sha256, for people who want the toolkit without cloning. One hash
verifies the download; the bundled CHECKSUMS.txt verifies the extracted
contents. Refuses to overwrite an existing version's zip (released = immutable;
bump instead), and refuses to package when zchecksums fails. tests/ excluded
from the zip; releases/ never packages itself.
First release included: releases/zscripts-v1.0.0.0.0.zip (42 scripts + 7
support files) and its .sha256.
.gitattributes: releases/*.sha256 pinned LF (sha256sum treats a trailing CR as
part of the filename), releases/*.zip marked binary.
Verified end-to-end as a downloader would experience it, in WSL: sha256sum -c
on the zip passes, unzip, sha256sum -c CHECKSUMS.txt inside gives 42 OK / 0
FAILED, and the extracted zdeploy.ps1 header and build-version.json both read
v1.0.0.0.0. Double-release guard and -Verify mode exercised. Full Pester suite
219/219 (the checksum tests absorb the new files automatically).
The bash port and its bats suite are removed from the tree while they get more
testing. Everything they were referenced from is cleaned up so nothing dangles:
- README drops the 'Linux / macOS / WSL' pointer to bash/README.md (a dead
link once the folder is gone) and the aside that ZCONFIG is honoured by
both ports.
- CHANGELOG drops the two bash mentions.
- ZHelpers.ps1's ZCONFIG comment no longer cites zhelpers.sh.
- .gitattributes drops the now-dead bash/**, tests/bash/** and *.bats rules,
keeping the PowerShell CRLF rules.
No behaviour change to the PowerShell scripts; the Pester suite is untouched
and still passes 169/169.
Deliberately NOT a history rewrite: the port stays in this repo's history and
in full in the private mirror, so it can be restored with a revert when the
testing is done. Nothing here is secret - it is unfinished, not sensitive.
* test: add bats suite for the bash port + fix underscore-key guard (phase 5)
Part of #26. Closes#30.
The bash port reimplements the exclude lists, config accessors and argument
parsing, so it can drift from PowerShell independently. 50 bats tests mirror
the Pester suites assertion-for-assertion where the two are meant to agree:
z_archive_excludes (per-kind lists and the deploy-vs-backup gating), config
accessors, z_path Windows->WSL translation, json_build_label, and argument
handling (bare invocation, unknown key, 'all' expansion, --port override).
Fixes#30 along the way, because the alternative was a test enshrining the bug:
zproj_require accepted underscore comment keys. It only checked the key was
non-null, and a comment is a non-null JSON string, so 'zkill _note' sailed
through and exited 0 having done nothing - the silent-success failure mode.
Now rejects any _-prefixed key and requires the value to be a JSON object.
Both checks earn their place: the type check catches string comments, the
prefix rule catches an object-valued _template key that PowerShell refuses and
the type check alone would allow.
Documents #31 rather than fixing it: a leading dash on a project key works in
every PowerShell script but only in bash/zdeploy - the others reject -myapp as
an unknown option. Stripping it everywhere would make a mistyped flag resolve
as a project key, so the tests pin current behaviour and bash/README.md now
states the difference instead of the README's blanket claim.
Verified by mutation testing: all 9 mutations turn the suite red - removing the
python and vite backup gates, reintroducing #23 in bash, unfiltering underscore
keys in zproj_keys and zproj_require, breaking zremote_compose_dir fallback and
z_path translation, and removing zkill's all-expansion and no-args guard. Both
mutated files confirmed restored byte-for-byte.
An early run also caught a bug in the tests themselves: the membership helper
used 'grep -qx' (regex), so the needle '.env' matched 'venv' and several
'excludes .env' assertions were false passes. Now uses -qxF.
* fix(gitattributes): keep .bats files LF so 'bats tests/bash' works on a Windows checkout
The LF rule was scoped to 'bash/**', which does not match tests/bash/. With
core.autocrlf a Windows working copy got CRLF .bats files, and bats fails on
them - so the command the README documents would not run on the machine the
suite was written on without stripping \r first.
Adds tests/bash/** and *.bats to the same eol=lf rule and renormalises.
Verified by running 'bats tests/bash' with no sed preprocessing: 50/50.
zstart only warns when a python project has no venv; zsetup is the
command that provisions one. For a python project it creates <root>/.venv
and installs deps; for vite/nextjs it runs npm install. Idempotent.
The pip install command comes from the project's optional "install"
config field (e.g. "-e backend" for deps in a subfolder, "-r reqs.txt"),
or is auto-detected from a root pyproject.toml/setup.py ("-e .") or
requirements.txt ("-r requirements.txt"). Bash + PowerShell + .cmd
wrapper, documented in the README and example configs.
Part of #26. Phase 2 asserted that Get-ArchiveExcludes returns the right list;
this asserts the archive that actually ships. A name can be on the exclude list
and still land in the zip - only opening the zip proves otherwise. Each test
builds a real temp source tree, archives it, and reads back the entries.
66 tests across: TopLevelExclude (files, directories, multiple names, absent
names), recursive junk-directory pruning at top level and any depth, junk
extensions, nested archives, OS junk filenames, script-file inclusion via
-IncludeScriptFiles, ExtraFiles (how zbackup bundles a pg_dump), zip mechanics
(overwrite, destination creation, forward-slash entry paths), and two
end-to-end shapes - a python deploy zip built from the real exclude list, and a
backup of the same tree that must retain .env and uploads.
Two behaviours are pinned deliberately:
- TopLevelExclude matches TOP-LEVEL names only, so a nested backend/.env is
NOT excluded by listing '.env'. That is current behaviour and the reason
deploys rely on server-side preservation; the test exists so changing it is
a decision rather than an accident.
- An all-junk tree throws instead of producing an empty zip - a silent empty
deploy would unpack to nothing on the server.
Verified by mutation testing, which paid for itself immediately: the first run
showed the junk-directory tests were vacuous. Inside a Where-Object, $_ rebinds
to the pipeline item and shadowed Pester's -ForEach value, so the pattern
matched nothing and the tests passed while asserting nothing. Fixed by
capturing the value first; re-run now catches all 8 mutations (disabling
TopLevelExclude, top-level and nested junk pruning, each file filter, and the
empty-archive throw). ZHelpers.ps1 restored byte-for-byte afterwards.
Part of #26. This layer has regressed more than any other - bare vs dashed
keys, 'all' expansion, and whether a bad key fails loudly or quietly selects
nothing.
42 tests over three guarantees:
- Running bare shows usage and exits non-zero, for all ten target-taking
scripts. Some of these used to mean 'do it to everything' when run with no
args, which is how an unintended full backup or deploy happens.
- An unknown key fails loudly. Exiting 0 having selected nothing is the
dangerous outcome: a typo'd key in a scheduled task looks like a
successful run that backed up nothing.
- A leading dash is stripped before the lookup, so -myapp == myapp. Asserted
via a dashed *unknown* key, so the error must name 'x' rather than '-x'.
Plus zkill's own resolution: bare key, dashed key, several keys, 'all' and
'-all' expanding to projects that have a ports.dev, edge/docker stacks skipped,
and -Port overriding the configured port.
Safety: the tests run the real scripts as child processes, so they are confined
to paths that exit before doing any work. Only zkill runs with a valid target,
because the fixture's dev ports (59990/59991) are deliberately unused and its
localRoots do not exist - it finds no listeners and kills nothing. zdeploy,
zbackup_ec2, zec2, zec2online, zrepair, zstop, zstart and zbackup are never
invoked with a real target; that is integration territory needing a disposable
server.
Each test runs against an isolated temp installation - the .ps1 files copied
beside a fixture zconfig.json - so nothing touches this repo, no real config is
read, and the suite passes on a machine that has never been configured. That
also makes it independent of the ZCONFIG seam in #27, so the PRs can merge in
any order.
Verified by mutation testing: removing zkill's all-expansion, dash tolerance
and no-args guard, making an unknown key exit 0, and letting underscore comment
keys leak in as projects each turn the suite red (2/1/2/16/4 tests). Both
mutated files confirmed restored byte-for-byte.
Part of #26. The toolkit had no automated tests at all - including for the
functions that decide what goes into a deploy zip, which is where a dev .env
reached production (#23).
Phase 1 - the seam. Get-ZConfig read a hardcoded $PSScriptRoot\zconfig.json,
so nothing config-dependent could be tested without touching the real config.
Adds Get-ZConfigPath honoring $env:ZCONFIG (the bash port has always had this,
so it also closes a parity gap) and Reset-ZConfigCache to drop the memoised
config between fixtures. Deliberately did NOT convert the exit 1 paths to
throw: that changes observed CLI output, and the pure functions don't need it.
Phase 2 - 61 tests over the functions with no side effects: Get-ArchiveExcludes
(common/python/vite/nextjs lists, deploy.exclude merging, dedupe, array shape),
Get-ZConfig / Get-ZConfigPath / Get-ZProjectKeys / Get-ZProject (dash tolerance,
underscore-key filtering, memoisation), Get-ZEdgeProject, Get-RemoteComposeDir,
Get-Ec2Target / Get-Ec2Home, Get-LabelFromBuildJsonObj, Read-JsonBuildVersion.
The suite is verified by mutation testing rather than assumed useful - six
deliberate regressions were each introduced and confirmed to turn it red,
including reintroducing the exact #23 bug and its inverse (backups silently
dropping .env/uploads, which would produce restore points that cannot restore).
Runs off a fixture config injected via ZCONFIG, so it never reads a real
zconfig.json and passes on a machine that has never been configured.
* feat(zec2_rotatekeys): rotate/reset server-side secrets without exposing values
New tool for the leaked/overwritten prod .env case: -Rotate KEY regenerates a
key ON THE SERVER (openssl rand -hex 32) so the value never leaves the box;
-Set KEY takes an operator-known value from a masked prompt and streams it over
SSH stdin (never a command arg, never echoed). Backs the server .env up to a
timestamped .bak first, updates keys atomically (match-or-append), auto-detects
backend/.env from deploy.preserve, restarts only with -Restart, and -WhatIf
previews the plan. Docs added to README + CHANGELOG.
* fix(zec2_rotatekeys): recreate container on -Restart so the new .env loads
A plain 'docker compose restart' reuses the container's existing environment
and would NOT pick up env_file changes, leaving the app on the old secrets
after a rotation. -Restart now runs 'up -d --force-recreate <svc>', the
reliable way to apply the new .env. Docs updated to match.
* feat(zstart): support uvicorn/ASGI apps via a startApp config field
Python projects could only be started as `python -m <startModule>`, so
FastAPI/ASGI apps that run under uvicorn (like evo-ai:
`uvicorn app.main:app`) couldn't be started by zstart in either port.
Add an optional `startApp` field. When set, zstart runs
`uvicorn <startApp> --host <bind-host> --port <ports.dev> --reload` via
the venv python's -m (no PATH juggling), integrating zstart's existing
bind-host and dev-port handling. startApp takes precedence over
startModule; a python project still needs one or the other. Applied to
bash and PowerShell, documented in the README + example configs.
* feat(zstart): warn when falling back to system python (no project venv)
A python project with no .venv (or only a Windows .venv when on WSL)
silently ran under the system interpreter, which usually lacks the
project's deps - producing a cryptic ModuleNotFoundError far from the
cause. Now zstart prints a clear warning naming the missing venv and the
one-liner to create it, before starting. Bash + PowerShell.
* fix(bash): zstart --detached no longer hangs on the tracking FIFO
Detached mode forked the long-lived server while it still inherited the
ztokens tracking fds (the capture FIFO on 1/2, saved stdout/stderr on
3/4). The parent's EXIT-trap footer runs `tee` on that FIFO and waits for
EOF, which never came while the server held it open - so `zstart
--detached` (and zstartd / zrestart --detached) hung instead of
returning. detach() now redirects stdin<-/dev/null, stdout/stderr->log
and closes fd 3/4 before exec'ing the server. Verified on WSL: detached
returns in 0s and the server still boots.
* fix(zstart): git-pull pre-step can't hang on a credential prompt
start.gitPull ran `git pull --ff-only` before starting the server; in an
environment with no cached git credentials (e.g. WSL against an HTTPS
GitHub remote) git prompted "Username for 'https://github.com':" and the
whole start blocked on stdin. Run the pull with GIT_TERMINAL_PROMPT=0 so
it fails fast, log a clear "auto-pull skipped" note, and start with the
current checkout. Bash + PowerShell.
* fix(ps): zbackup explicit target + robust DATABASE_URL parsing; ssh-stderr deploy fix
Restores parked, previously-uncommitted PowerShell improvements:
- zbackup / zbackup_and_sync require an explicit target: bare invocation
now prints usage instead of quietly backing up everything; 'all' does
what bare used to (matching zdeploy). setup_backup_schedule.ps1 passes
'all' to the scheduled task; both tolerate switch-style args.
- zbackup parses more DATABASE_URL styles: strips surrounding quotes
(Prisma convention), accepts postgres:// and postgresql+driver://
schemes, and treats the port as optional (defaults to 5432).
- Invoke-Ec2Step survives ssh stderr warnings: under ErrorActionPreference
'Stop', PS 5.1 turns any native stderr line (e.g. Docker's COMPOSE_BAKE
deprecation notice) into a terminating NativeCommandError, aborting a
deploy that actually succeeded. Drop to Continue locally and flatten
stderr so only the real exit code decides success.
* fix(ps): zbackup reports the real pg_dump failure, not "No DATABASE_URL"
Mirror of the bash fix. Invoke-LocalPgDump now owns all its messaging
(caller just captures success) and distinguishes the cases:
- no .env / no DATABASE_URL -> calm "No local DATABASE_URL".
- database not reachable (connection refused / could not connect / DNS /
timeout) -> calm "Local database not running at host:port" - a stopped
dev DB is a normal state.
- any other failure (version mismatch, auth, missing db) -> the loud,
full pg_dump error plus the host:port/db it tried, instead of a bare
"pg_dump failed (exit N)" followed by a misleading "No DATABASE_URL".
Captures pg_dump stderr (was 2>$null); drops ErrorActionPreference to
Continue locally so PS 5.1 doesn't turn that stderr into a terminating
NativeCommandError under the script's 'Stop' setting.
Only nextjs-kind excluded .env* from the deploy archive; python and vite
did not - so a project's local .env at its root got zipped and shipped to
the server on every deploy, planting local secrets over the server's own
(the operator-file restore only wins for files it preserved). Add
.env/.env.local/.env.production to the python and vite deploy excludes,
matching nextjs. Kept DEPLOY-only (like `uploads`): backups still capture
.env so a source backup stays complete. Bash + PowerShell.
Note: this covers a ROOT .env. A nested secret (e.g. DocketMail's
backend/.env) is handled separately via deploy.preserve in the project's
zconfig.
These three defaulted to "every project" when run with no arguments -
inconsistent with zbackup/zdeploy/zstart/etc. (which show usage), and a
surprising amount of work to kick off by accident: zbackup_ec2 pulls a
server backup of every project, and zec2online deep-checks everything AND
auto-starts any downed stacks. Now a bare invocation prints usage and
lists the projects; 'all' does what bare used to. Applied to both the
bash and PowerShell versions.
* fix(bash): zbackup reports the real pg_dump failure, not "No DATABASE_URL"
local_pg_dump returned 1 for four different cases (no .env, no
DATABASE_URL, unparseable URL, and an actual dump failure), and the
caller printed "No local DATABASE_URL - source-only backup" for every
one. So a project that HAS a DATABASE_URL whose dump failed was
mislabeled as having none — the two messages contradicted each other.
- Distinguish the cases with exit codes: 0 dumped, 1 attempted-but-failed,
2 no local database. The caller now prints the right line for each.
- Stop swallowing pg_dump's stderr (2>/dev/null); on failure, surface the
actual error (version mismatch, unreachable host, auth) plus the
host:port/db it tried, so failures are diagnosable.
* fix(bash): zbackup treats a not-running local DB as a calm skip
Fold all pg_dump messaging into local_pg_dump (caller just captures
success) and special-case an unreachable database. "connection refused"
/ "could not connect" / DNS / timeout now print a quiet
"Local database not running at host:port - source-only backup." instead
of a red multi-line error - a stopped dev DB is a normal state. Real
failures (version mismatch, auth, missing db) still print the full
pg_dump error so they're diagnosable.
* fix(bash): zbackup/zbackup_and_sync require an explicit target, matching PowerShell
Bare `zbackup` quietly backed up every project - inconsistent with the
PowerShell version (and with zdeploy), and an easy way to kick off a
huge unintended backup. Now a bare invocation prints usage and lists the
projects; `all` does what bare used to (every project + the scripts
folder). Same for `zbackup_and_sync`, and setup_backup_schedule's cron
line now passes `all` so the scheduled job still backs everything up.
zkill now accepts 'all', expanding to every project that has a ports.dev
(edge/docker stacks with no local dev server are skipped) - matching
zdeploy all / zbackup all. Ported to both the PowerShell (ZKillOnly.ps1)
and bash (bash/zkill) versions; README + CHANGELOG updated.
The README advertises macOS support, but the port used constructs that
fail on the userland macOS actually ships:
- `mapfile` (bash 4+) in 10 spots — macOS ships bash 3.2 as /usr/bin/bash,
so a mac user following the README hit "mapfile: command not found" and
silently got empty project lists. Replace each with a portable
`while IFS= read -r` loop (identical arrays; set -u safe on empty input).
- `_z_commafy` used the GNU-only `sed :a;...;ta` label/branch idiom, which
errors on BSD/macOS sed (the token footer's thousands separators).
Reimplement with awk (already a dependency).
- README: correct the macOS line — bash 4+ does NOT ship with macOS; the
stock 3.2 now works, and only jq needs brew.
Also two correctness nits found in the same review:
- zkill/zrestart `--kill-all` was parsed but silently ignored; now it
prints a "not implemented in the bash port" notice instead of no-op.
- zrepair's smoke-test line said "https://$domain" but probes
http://$HOST with a Host header; label now matches what it does.
Verified on WSL (bash 5): read-loops produce the same 11 project / 7
domain keys as mapfile; awk commafy matches across 0..1,234,567; empty
producer yields a 0-length array; footer renders. 3.2-compat is by static
analysis — no bash-4-only constructs remain.
* fix(bash): translate Windows config paths to WSL/Unix form
A shared zconfig.json (one file used from a Windows checkout and from
WSL via the symlink) holds Windows paths like "F:\evomedia.net\app".
The bash port passed those to local file ops verbatim, so every
path-using command failed on WSL (e.g. zbackup: "Root not found:
F:\evomedia.net\evomedia-docs").
Add a z_path helper that converts drive-letter paths to the host's
native form (wslpath, with a /mnt fallback) and is a no-op for Unix
paths and empty strings — so a bash-native config is unaffected. Route
every LOCAL path through it: localRoot (new zproj_root accessor),
paths.* (temp, backupsLocal, backupsEc2, scriptsRoot, oneDriveBackups),
ec2.pemKey (zec2_pem), and ztokens.dataDir. Server-side paths
(remote.path, composeDir, stackRoot, certsSource) are left verbatim.
Verified on WSL: all 11 project roots, all paths.*, and the C:/G: pem
and OneDrive paths resolve to existing dirs; zbackup evodocs (the
failing case) now runs clean end-to-end.
* fix(bash): keep z_path fallback bash-3.2 clean (tr, not ${x,,})
The wslpath-absent fallback lowercased the drive letter with ${drive,,},
a bash-4.0 construct that breaks on macOS's stock bash 3.2. Use tr
instead so the whole helper stays 3.2-compatible.