zdeploy ended with two blank lines and nothing else did, so its output
was the only one that did not butt up against the next prompt. Applied
everywhere.
Implemented in Stop-ZTracking rather than per script, because every
tracked script ends by calling it - including the usage and guard paths
that do `Stop-ZTracking; exit 1`. One place therefore covers every exit
of sixteen scripts.
* Write-ZTrailer emits the two lines. Stop-ZTracking calls it on both
paths, including the early return when tracking never started - a
script that printed output still deserves the separation.
* -FinalNote prints one last line AFTER the tracking footer and BEFORE
the blanks. zdeploy's "Last deployed at ..." now goes through it, so
it stays the last thing on screen instead of being followed by the
token count.
* The standalone tools (zchecksums, zversion, zrelease) get a local
three-line copy at each exit rather than dot-sourcing ZHelpers -
they are deliberately dependency-free so they run from an extracted
release zip with nothing beside them.
* Redundant trailing blanks were removed where a script already
printed one before exiting, so the count is exactly two, not three.
Also fixes a related gap found while testing: the error exits inside
Get-ZConfig and Get-ZProject bypassed Stop-ZTracking entirely, so a run
that died on a missing zconfig.json printed no trailer AND no token
footer. Those three exits now route through Stop-ZTracking.
zkill/zrestart are one-line wrappers around scripts that already trail,
so they are untouched - adding it would double the blanks.
Verified by running each script and counting the trailing blank lines in
its captured output: 2 on every success path, every usage path, and the
config-error path. Suite 231/231. CHECKSUMS.txt refreshed.
Mirrors evo.scripts#79. Invoke-ViteDeploy announced "(server-side build will
bump +1)" and then Wait-VerifyStaticBuild asked for the committed stamp
unchanged, so the operator was told to expect one label and watched a check
pass on another.
The behavior was already correct; the message predates the removal of the
prebuild hook that used to self-bump inside the image. Wait-VerifyStaticBuild
documents it: "Expect the COMMITTED stamp, not +1: builds no longer
self-bump ... so 'is the build I just packed live?' means an exact match."
CHECKSUMS.txt regenerated, since zdeploy.ps1 changed.
Two things, both prompted by the same incident.
STATIC KIND
Plain static sites - no build, no container of their own; a shared web
container serves them off disk, so shipping the files IS the deploy.
Directories are staged to a sibling and swapped in with mv rather than
copied in place, because a large media file uploaded in place is served
half-written to anyone who requests it mid-copy. The swap is a rename,
so the switch is atomic.
Skips $JunkDirNames + .github + deploy.skipDirs, matching the docker
kind rather than inventing a third convention.
SANITIZATION TEST
This repo is public and must stay standalone, but the toolkit is
developed in a private checkout and copied here. Twice in three days a
wholesale copy landed carrying real project names, internal hostnames,
host disk figures, and references to scripts that exist only in the
private copy. Both times a human reading the diff caught it - the
control that fails exactly when a diff is 280 lines of good work with
three bad words buried in it.
So it is a test now. Seven rules: private project names, private product
domains, private-only script names, local drive paths, the operator home
path, the real key filename, and any IPv4 outside RFC 5737 documentation
space and the private ranges.
Two anti-vacuity guards, because a denylist that silently matches
nothing is worse than no denylist: one asserts the file scan is
non-empty, and one plants a known violation and requires the pattern to
find it.
The IP rule bounds on [\d.] rather than \d deliberately - this toolkit's
own 5-segment version (v1.0.0.0.14) contains "0.0.0.14", which a plain
digit boundary reads as an address.
Verified against the real unsanitized copy that slipped through: 3 of
the 7 rules fire, naming file and line.
Tests: 231/231 (222 + 9 new). CHECKSUMS.txt refreshed.
Invoke-DockerDeploy uploaded top-level files ONLY. For any stack that
keeps configuration in a directory, that is silently wrong in the worst
way: the compose file arrives, the containers restart, and the config
they read is whatever was already on the box. The deploy reports
success while shipping nothing that matters.
The sharp case is a provisioning directory read at container start -
alert rules, datasources, mounted config - where the restart makes it
look like the change was applied.
Same fix the edge kind already got: subdirectories ship recursively.
Skipped are $JunkDirNames, plus .github (CI config belongs in the repo,
never on a deploy target - it is not in $JunkDirNames because backups
DO want it), plus anything the project lists in the new
deploy.skipDirs.
deploy.skipDirs exists for SERVER-SIDE STATE that shares the tree: a
data/ holding mailboxes or a time-series database must never be
overwritten by whatever the local checkout has - usually nothing, which
is the dangerous case, since scp -r of an absent directory is not the
no-op you want to rely on.
CHECKSUMS.txt refreshed alongside. Tests: 222/222.
Build stamp catch-up for 3 merged PR(s) since v1.0.0.0.11:
6c30f40 feat(zdeploy): sync deploy hardening from the private toolkit (#53)
8959d9f fix(zdeploy): stop passing ssh -n to scp, which rejects it (#52)
d99f8a1 fix(zdeploy): stop mirroring the build stamp into the local checkout (#51)
Nine fixes that had accumulated only in the private copy. Each one is a
production failure that already happened:
* ssh -n on the deploy path. Without it ssh reads its stdin, inherits
the console handle under PowerShell, and can block forever. The
existing timeouts cannot catch it - they bound a connection that is
dying, and this one was never established. Get-Ec2ScpOpts derives
from the same list minus -n, because scp rejects it with a usage
error that reads like an unrelated failure.
* Invoke-DeployGitPull now switches TO the default branch instead of
pulling whatever is checked out. Pulling the current branch breaks
as soon as the remote deletes branches on merge, and worse, could
ship a feature branch to production. Refuses over local changes,
excluding the files the deploy itself stamps.
* nextjs handler resolves composeDir. It ran compose against
remote.path, so a project whose compose lives in a subdirectory
recycled whatever docker-compose.yml sat at the project root -
usually the local dev one shipped in the same archive. Two deploys
in a row exited 0 having shipped nothing, and verification passed
because the untouched old container still answered.
* BUILD BEFORE DOWN. The old order took the site offline for the whole
build and left it offline if the build failed - one deploy served
502 for 19 hours with no container running. The old image now keeps
serving until the new one is ready.
* compose build takes the named service, so a compose file that does
not define it fails loudly instead of succeeding with nothing to do.
* deploy.verifyHost overrides domain for verification only, for when a
domain is retired ahead of its replacement.
* deploy.stampCmd writes the build version from git before zipping,
then reverts the tree so the stamp cannot block the next deploy.
* Optional ztokens pseudo-project and passive usage records. Both
degrade to a printed skip when no sibling ztokens checkout exists.
Sanitized on the way across, per the public-repo rule: the war stories
that make these comments worth reading are kept, the private product
names, domains, internal script names, outage dates and host figures are
not. Also dropped a "See issue #28" pointer that resolves to an
unrelated PR in this repo, and two references to scripts that exist only
in the private toolkit.
CHECKSUMS.txt refreshed alongside, per the manifest rule.
Tests: 222/222 (the 4 failures before the refresh were the checksum
suite correctly flagging both edited files).
Mirrors evo.scripts #66, which added -n to Get-Ec2SshOpts so a deploy step
cannot block on inherited stdin, plus the follow-up that keeps that flag away
from scp. OpenSSH's scp has no -n: it exits 1 with 'unknown option -- n' and
prints its usage block, so every upload failed once the shared option array
reached it.
Get-Ec2ScpOpts is derived from Get-Ec2SshOpts with -n filtered out rather than
duplicated, so the connect and keepalive timeouts cannot drift apart between
the two transports. All four scp call sites here use it - three plain uploads
and the recursive directory upload, which the private tree does not have.
The upload failure message asserted 'Likely server disk space' without checking
anything; it now points at scp's own output, where the real diagnosis already
was.
Verified: both files parse clean, Get-Ec2ScpOpts returns the five -o pairs with
no -n, and the added lines carry no real hosts, keys, or paths.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Mirrors evomedia-net/evo.scripts#62. Every python-kind deploy wrote
build-version.json locally, leaving the tree dirty; committing it hit
branch protection, so a deploy either tripped the next deploy's
clean-tree guard or bypassed the rule.
The build number is bumped in the container, written to .build_version
on the server, and proven by /api/build-version. The repo file records
the stage baseline only.
Build stamp catch-up for 1 merged PR(s) since v1.0.0.0.10:
a1dd642 chore(changelog): drop the product name from the versioning entry; add the plain-text twin (#50)
Build stamp catch-up for 8 merged PR(s) since v1.0.0.0.0:
6c03138 fix(zdeploy): deploy edge first for any project list, not just 'all' (#46)
976eb6d chore: point repo URLs at evomedia-net/evo.* (#44)
0e4d9c3 fix(zdeploy): ship edge asset subdirectories, stop deploys hanging on ssh prompts, keep vendored archives (#43)
c8fcfee fix(zdeploy): verify Next.js deploys on the server instead of probing the public IP (stops the bogus security-group warning) (#41)
b84d24b fix(zdeploy): report where a domain-less project deployed to (#40)
4668404 fix(zdeploy): re-land the two blank lines after the timestamp (#38)
a868973 fix(zdeploy): verify against the committed build stamp, and read 5-segment versions (#37)
0ddffc1 feat(zdeploy): print 'Last deployed at' timestamp as the last line of a run (#36)
Closes#45.
'zdeploy evo edge' deployed evo first, then edge. The edge-first sort was
nested inside the 'all' branch, so an explicit list walked in whatever order
was typed.
That ordering is a correctness property, not a convenience of 'all': the proxy
has to route before the apps behind it ship, or there is a window where a new
app is live behind stale routing. 'zdeploy evo edge' reads as 'these two, edge
included' and quietly did the risky order.
Hoisted the sort out of the branch so it applies however the list was produced.
Order within each group is preserved, so an intentional sequence still holds -
and when the order does change, it now says so rather than reordering silently.
Verified across six shapes: evo+edge reorders and reports it; edge+evo is left
alone with no note; edge last in a longer list moves to front keeping the rest
in order; single project and no-edge lists are untouched.
All 16 repos moved to the evomedia-net org and were renamed into the evo.*
namespace, so every github.com/kellymichels/<old-name> reference in source
headers, CI badges, security links and docs pointed at a redirect.
Mechanical URL-only rewrite, applied longest-name-first so smartplantehs-docs
could not be clobbered by the smartplantehs rule. Nothing else changes: no
code, no product names, no behaviour. smartplantehs -> evo.ehs here is the
REPO url only; the product rename is separate and still pending.
* fix(zdeploy): edge deploy ships asset subdirectories, not just top-level files
A project self-hosting assets in folders (fonts/, vendor/) lost them on
every deploy: docker created empty root-owned mount points and nginx
served 404s from them, so fonts fell back silently and vendored JS
never loaded. Every subdirectory except nginx-logs/ and .git/ now ships
recursively, and the ensure-dir chown is recursive so scp into
docker-created root-owned dirs cannot fail.
Closes#42
* fix(zdeploy): stop deploys hanging on ssh prompts, keep vendored archives, make unzip install idempotent
- Get-Ec2SshOpts: every deploy-path ssh/scp now carries BatchMode=yes
plus connect/keepalive timeouts. Without BatchMode ssh prompts and
waits forever, and because the deploy pipes stderr through the
pipeline the prompt never reaches the screen - the run just stops
under the last step label with no explanation.
- Archive filter exempts files under vendor/: a vendored *.tgz is a
build input, and dropping it fails a Dockerfile COPY at image build.
- unzip install checks command -v first instead of an apt round-trip
on every deploy.
Ported from the private script; three tests cover the vendor/ exemption.
* fix(zdeploy): verify Next.js deploys on the server, not the public IP
The nextjs handler verified by requesting http://<ec2-ip>:<prod-port>/.
A compose stack behind the edge proxy normally publishes to 127.0.0.1
only, so that request can never be answered and every deploy ended with
"is the port open in the security group?" — pointing at a firewall rule
for an app that was already serving fine. No security-group change could
have made that probe succeed, which is what made the warning actively
harmful: it named the one fix guaranteed not to work, and opening the
port would have exposed the app directly, bypassing the proxy and TLS.
The nextjs handler now uses the precedence the python handler already
used: verify.port (curl localhost on the server) -> domain (Host-header
check through the edge proxy) -> an honest "NOT verified" instead of a
misleading warning.
Also follow redirects in the health check. An app whose "/" answers 307
(Next.js -> /login) returns a body of a few bytes that read as "not
ready"; -L fetches the page that actually renders. No effect on projects
that point verify.path at a plain 200 endpoint such as /health.
The example config now documents the verify block on the nextjs project,
and CHECKSUMS.txt is regenerated for the changed script.
* fix(zdeploy): truncate the health-check body in the PASS line
Pointing the nextjs handler at Test-DeployHealth means the body is now
often an HTML page, not a line of JSON, so the PASS line dumped ~10 KB
of markup into the deploy output and buried everything after it.
Match on the full body as before, but print at most 200 characters plus
the byte count. A /health endpoint still prints in full.
Closes#39.
A project without a public domain finished with no indication of where it went
- the 'Site:' line was gated on $Proj.domain in all four handlers, so an
internal service deployed successfully and said nothing about the result.
The information was never missing, only unused: a domain-less project that
deploys almost certainly has a verify block (the documented pattern for stacks
not published through the edge proxy), carrying the exact port and health path
the deploy had just checked. zdeploy verified against that endpoint and then
discarded it.
New Write-DeployLocation falls back through what the project actually declares:
domain -> public URL; else verify.port (+ path) -> server-local URL flagged as
not publicly routed; else ports.prod -> same without a path; else prints
nothing, as before. Replaces the four call sites, keeping the nextjs handler's
column alignment via -Pad.
Exercised against every shape: domain, verify-only, ports.prod-only, nothing at
all, and padded. CHECKSUMS regenerated; suite 219/219.
These were written and pushed, but landed on the PR #36 branch AFTER that PR
had already merged, so they never reached master. Private main has them; public
did not. The two copies of zdeploy.ps1 had quietly diverged on the last line of
output.
CHECKSUMS.txt regenerated. Tail now byte-identical to the private copy.
Deploy verification expected buildNumber + 1, which was only correct while
a prebuild hook self-incremented the stamp during the image build. That hook
is gone (it produced versions matching no commit and left every build with a
dirty tree), so the check waited out its timeout and warned on every deploy
even when the deploy had succeeded.
Both verifier call sites now expect the committed label, and
Get-LabelFromBuildJsonObj understands the 5-segment scheme
v{major}.{rc}.{beta}.{alpha}.{build} while still reading the legacy
{productVersion, buildNumber} stamp for projects that haven't migrated.
Verified: both files parse; label fn returns v0.0.0.1.8 (5-segment) and
v1.0.0.70 (legacy).
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(zdeploy): print 'Last deployed at' as the final line of a run
Scroll to the bottom of a deploy and you can see how long ago it happened,
without digging through the log or checking the server.
Placed after Stop-ZTracking so it is genuinely the last line, below the
token-tracking footer. Only reached on success - a failed deploy throws out of
the dispatch loop, so this never claims a deploy that did not happen. Prints
once per run rather than per project, so 'zdeploy all' ends with a single
timestamp.
24-hour clock (MM/dd/yyyy HH:mm:ss): the requested format carried no AM/PM
marker, and 24-hour is unambiguous for judging elapsed time at a glance.
CHECKSUMS.txt regenerated, since changing a covered script invalidates it.
* feat(zdeploy): use 12-hour clock with AM/PM in the timestamp
Kelly's preference: 'Last deployed at 08/03/2026 01:24:09 PM' rather than
24-hour. The AM/PM marker keeps it unambiguous.
CHECKSUMS.txt regenerated.
* feat(zchecksums): SHA-256 manifest so a download can be verified before it's run
CHECKSUMS.txt lists a SHA-256 for every top-level .ps1 and .cmd - the files a
user actually executes. zchecksums verifies them; zchecksums -Update
regenerates after an intentional edit.
The manifest is sha256sum format, so 'sha256sum -c CHECKSUMS.txt' works on
Linux/macOS/WSL as well as the PowerShell path on Windows. Hashes are identical
on every platform because .gitattributes pins .ps1/.cmd to CRLF everywhere -
that pin is now load-bearing, so it is commented as such.
Beyond changed and missing files it also reports a script that is on disk but
NOT in the manifest, so something added outside a commit still gets noticed.
Exits non-zero on any of the three.
Honest about its limits, in the header and the README: the manifest lives in
the same repo as the code, so it is an integrity check rather than a signature.
It catches a truncated clone, a forgotten local edit, or an unlisted file - not
a compromised repo.
CHECKSUMS.txt is pinned to LF: sha256sum treats a trailing CR as part of the
filename and would report every entry as missing on Linux.
tests/Checksums.Tests.ps1 keeps it from rotting - a stale manifest is worse
than none, since it either cries wolf until people ignore it or quietly stops
covering a new script. The tests assert the format, LF endings, sort order,
full coverage of on-disk scripts, current hashes, and that zchecksums itself
exits 1 on a tampered file (proved by appending a byte and restoring it).
* feat(zversion, zrelease): toolkit versioning + downloadable release zips
Implements the versioning rule (SmartPlant's 5-segment scheme, now the global
standard; currently only sp and zscripts are on it at v1.x):
v{major}.{rc}.{beta}.{alpha}.{build}
zversion: get / bump / bump-stage / set. A stage bump zeroes every lower
segment including build. 'bump' is one per PR and one per defect fix, not per
file. Any write rewrites three things together, because they are only useful
when they agree: build-version.json (source of truth), a '# Version:' line in
all 42 script headers (a lone copied script still says which release it came
from), and CHECKSUMS.txt (stamping changes every file).
zrelease: packages the current version as releases/zscripts-<version>.zip with
a sibling .sha256, for people who want the toolkit without cloning. One hash
verifies the download; the bundled CHECKSUMS.txt verifies the extracted
contents. Refuses to overwrite an existing version's zip (released = immutable;
bump instead), and refuses to package when zchecksums fails. tests/ excluded
from the zip; releases/ never packages itself.
First release included: releases/zscripts-v1.0.0.0.0.zip (42 scripts + 7
support files) and its .sha256.
.gitattributes: releases/*.sha256 pinned LF (sha256sum treats a trailing CR as
part of the filename), releases/*.zip marked binary.
Verified end-to-end as a downloader would experience it, in WSL: sha256sum -c
on the zip passes, unzip, sha256sum -c CHECKSUMS.txt inside gives 42 OK / 0
FAILED, and the extracted zdeploy.ps1 header and build-version.json both read
v1.0.0.0.0. Double-release guard and -Verify mode exercised. Full Pester suite
219/219 (the checksum tests absorb the new files automatically).
The project-directory replacement preserved only ./.env, silently
destroying every other server-side file (.env.db, staged signing keys,
certs) on every deploy — and the vite kind preserved nothing at all.
- preserve all .env* files at the project root by default
- new deploy.preserve array for additional files/directories
- implemented via tar to the home dir before the wipe, extract after
the unzip; server-side copies win over zip contents (same semantics
./.env always had)
- helpers deliberately avoid embedded quotes and $( ): PowerShell 5.1
strips embedded double quotes when passing args to ssh.exe, which
silently corrupts remote commands (discovered when v1 of this fix
failed exactly that way)
Fixes#2
Projects not published through the edge proxy had a false-PASS problem:
the fallback reachability check hit http://<server-ip>/, which the
proxy's default vhost happily answers for apps that never started.
- new Test-DeployHealth: checks the app FROM the server over SSH
(curl localhost:<port><path>), optional expected substring
- opt in per project: "verify": { "port", "path", "expect" }
- projects with neither domain nor verify are reported NOT verified
instead of green-lighting the proxy's default page
- example config + README + changelog updated
Config-driven PowerShell scripts to run infrastructure tasks (deploy, restart, backup, diagnostics) yourself instead of having an AI agent orchestrate them, to save agent tokens. Environment specifics live in zconfig.json (gitignored).