mirror of
https://github.com/kellymichels/zscripts-token-savers
synced 2026-10-06 07:08:17 +00:00
docs(security): security notes, and a twin for every root .md (#78)
Defers to the org policy for how to report and covers what is particular to a repository that is a sanitised mirror: the most valuable report here is not a crash, it is something REAL that should not be here - a credential, an internal hostname, an operator path, an identifier naming a private project. Mail those rather than filing an issue, because a public issue about a leaked secret publishes it a second time. It also says what the automated check is and is not. The sanitisation suite is a DENYLIST: it proves the absence of known patterns, not the absence of secrets. Green tests are why a human report is still worth sending. And the ordinary warning for what these actually are - automation that archives a tree, uploads it, rebuilds containers and restarts services. Read before running, nothing here is a sandbox, the config is yours to replace. TWINS ARE NOW DISCOVERED, NOT LISTED. PAIRS was hand-kept and two files had outgrown it: ELEVATOR_PITCH.md and TOKEN_SAVINGS.md had no twin at all. Adding a document and remembering to add it to a list are two acts, and the second is the one that gets skipped. The Pester test had the same shape in reverse - it scraped PAIRS out of the generator's source, so it could only prove the list was self-consistent and a document nobody listed was invisible to it. It now asks the REPOSITORY what markdown it has. Proven by deleting SECURITY.txt and watching two tests fail. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
3b0194af93
commit
573e01021b
38
ELEVATOR_PITCH.txt
Normal file
38
ELEVATOR_PITCH.txt
Normal file
@ -0,0 +1,38 @@
|
||||
Evomedia.net Token Savers — https://github.com/evomedia-net/evo.zscripts
|
||||
Created by Kelly Michels · dev@evomedia.net
|
||||
Licensed under the MIT License. See LICENSE.
|
||||
|
||||
Elevator Pitch
|
||||
==============
|
||||
|
||||
The one-liner
|
||||
-------------
|
||||
|
||||
AI coding agents waste thousands of tokens a day on infrastructure orchestration. Token Savers gives you one-word commands to run those parts yourself — so your agent spends tokens on code, not on SSH.
|
||||
|
||||
The 30-second version
|
||||
---------------------
|
||||
|
||||
Every time your AI coding agent runs your infrastructure for you — a deploy, a restart, a health check — the full output lands in its context window: Docker layers, SSH banners, health-check chatter. We measured it per run — at a typical active-day cadence that's ~26,500 tokens of pure script output through the agent, and a single full Docker rebuild adds ~35,000 more. The dollars are small; the context is not — every line of infrastructure noise crowds out the code your agent is supposed to be reasoning about.
|
||||
|
||||
Token Savers collapses the infrastructure side into short, one-word commands you run yourself: zdeploy myapp, zrepair myapp, zstart myapp. Describe each project once in zconfig.json — where it lives, what kind it is, where it deploys — and every command just knows. You run the deploy; your agent edits the code. You run the health check; your agent reads the result and fixes whatever's wrong.
|
||||
|
||||
Measured per-run; ~26,500 tokens of script output per active development day at a typical cadence — kept out of your agent's context entirely when you run the commands yourself. See TOKEN_SAVINGS.md (TOKEN_SAVINGS.md) for the per-script measurements and method.
|
||||
|
||||
Why it's different
|
||||
------------------
|
||||
|
||||
- Built around the AI-agent workflow. The commands are short on purpose — fewer keystrokes for you, fewer tokens when an agent invokes them. But the real saving is the operations you don't hand to the agent at all.
|
||||
- The project name IS the command. zstart blog, zdeploy api, zbackup store — no flags to memorize, no switches to wire up.
|
||||
- One config file, zero secrets in git. Server IP, SSH key, paths, and project definitions live in one gitignored JSON. Clone it anywhere, drop in your config, go.
|
||||
- It verifies the deploy actually landed. Not "did the server return 200" (a stale cache does that too) — it checks that the build number went live, so you know the code you just shipped is the code that's running.
|
||||
|
||||
Who it's for
|
||||
------------
|
||||
|
||||
Solo developers and small teams running several containerized web apps (Python, Vite, Next.js, plus edge proxies and stock Docker images) on a single VPS or EC2 box, from a Windows dev machine, over SSH — and using AI coding agents to write the code.
|
||||
|
||||
The tagline
|
||||
-----------
|
||||
|
||||
Fewer keystrokes. Fewer tokens. One config to rule your fleet.
|
||||
68
SECURITY.md
Normal file
68
SECURITY.md
Normal file
@ -0,0 +1,68 @@
|
||||
# Security
|
||||
|
||||
How to report a vulnerability, what to expect, and what is in scope:
|
||||
**[the evomedia-net security policy](https://github.com/evomedia-net/.github/blob/main/SECURITY.md)**.
|
||||
Short version — email [dev@evomedia.net](mailto:dev@evomedia.net), not a public
|
||||
issue.
|
||||
|
||||
What follows is particular to this repository, which is unusual in one way
|
||||
worth stating plainly.
|
||||
|
||||
## This is a mirror, and the interesting bug is a leak
|
||||
|
||||
These scripts are published from a private tree. The copy here is sanitised:
|
||||
placeholder hosts, example configuration, dummy data. So the most valuable
|
||||
thing anyone can report about this repository is not a crash — it is
|
||||
**something real that should not be here**:
|
||||
|
||||
- a credential, key, token, or private-key block
|
||||
- an internal hostname, a product domain, or an operator's path
|
||||
- an identifier that names a private project or a private-only script
|
||||
|
||||
If you find one, treat it as a live secret and mail
|
||||
[dev@evomedia.net](mailto:dev@evomedia.net) rather than opening an issue. A
|
||||
public issue about a leaked secret publishes it a second time and pins it to
|
||||
the top of the page.
|
||||
|
||||
**The automated check is a denylist.** `tests/Sanitization.Tests.ps1` and
|
||||
`tests/sanitization-patterns.psd1` hold the patterns this repository must never
|
||||
contain, and CI enforces them. A denylist proves the absence of *known*
|
||||
patterns, not the absence of secrets — it is a regression net for a specific
|
||||
recurring mistake, not a substitute for reading what is published. That is why
|
||||
a report here is worth sending even though the tests are green.
|
||||
|
||||
## They are automation scripts, so read them before running them
|
||||
|
||||
Everything here drives real infrastructure: archives a working tree, uploads
|
||||
it, rebuilds containers, restarts services. That is the purpose, and it means
|
||||
the ordinary rules for running someone else's shell scripts apply with more
|
||||
force than usual.
|
||||
|
||||
- **Read a script before the first run**, and run it against something you can
|
||||
afford to break.
|
||||
- **Nothing here is a sandbox.** There is no dry-run guarantee unless a script
|
||||
documents one; the flag that exists on one command may not exist on the next.
|
||||
- **The configuration is yours.** The example config carries placeholders, and
|
||||
every host, key path and target in it has to be replaced with your own before
|
||||
anything is pointed at real infrastructure.
|
||||
- Addresses in examples use the ranges reserved for documentation, and
|
||||
loopback. They are placeholders, not somewhere to send anything.
|
||||
|
||||
Scripts that destroy or overwrite state are the ones to read twice. A report
|
||||
that one of them does something destructive **without saying so** is a good
|
||||
report; a report that a script named after a destructive act performs it is
|
||||
not.
|
||||
|
||||
## Release integrity
|
||||
|
||||
Releases carry checksums. They are an **integrity check, not a signature** —
|
||||
they catch a truncated download, a corrupted mirror and an accidental edit,
|
||||
and they do not catch a forger, because whoever can change an archive can
|
||||
change the manifest that travels with it.
|
||||
|
||||
## Not a finding here
|
||||
|
||||
- **Placeholder credentials and example configuration.** Fake values are the
|
||||
sanitisation working, not a leak.
|
||||
- **The private tree.** Only what is published here is in scope; the internal
|
||||
original is not public and cannot be reviewed.
|
||||
73
SECURITY.txt
Normal file
73
SECURITY.txt
Normal file
@ -0,0 +1,73 @@
|
||||
Security
|
||||
========
|
||||
|
||||
How to report a vulnerability, what to expect, and what is in scope:
|
||||
the evomedia-net security policy (https://github.com/evomedia-net/.github/blob/main/SECURITY.md).
|
||||
Short version — email dev@evomedia.net (mailto:dev@evomedia.net), not a public
|
||||
issue.
|
||||
|
||||
What follows is particular to this repository, which is unusual in one way
|
||||
worth stating plainly.
|
||||
|
||||
This is a mirror, and the interesting bug is a leak
|
||||
---------------------------------------------------
|
||||
|
||||
These scripts are published from a private tree. The copy here is sanitised:
|
||||
placeholder hosts, example configuration, dummy data. So the most valuable
|
||||
thing anyone can report about this repository is not a crash — it is
|
||||
something real that should not be here:
|
||||
|
||||
- a credential, key, token, or private-key block
|
||||
- an internal hostname, a product domain, or an operator's path
|
||||
- an identifier that names a private project or a private-only script
|
||||
|
||||
If you find one, treat it as a live secret and mail
|
||||
dev@evomedia.net (mailto:dev@evomedia.net) rather than opening an issue. A
|
||||
public issue about a leaked secret publishes it a second time and pins it to
|
||||
the top of the page.
|
||||
|
||||
The automated check is a denylist. tests/Sanitization.Tests.ps1 and
|
||||
tests/sanitization-patterns.psd1 hold the patterns this repository must never
|
||||
contain, and CI enforces them. A denylist proves the absence of known
|
||||
patterns, not the absence of secrets — it is a regression net for a specific
|
||||
recurring mistake, not a substitute for reading what is published. That is why
|
||||
a report here is worth sending even though the tests are green.
|
||||
|
||||
They are automation scripts, so read them before running them
|
||||
-------------------------------------------------------------
|
||||
|
||||
Everything here drives real infrastructure: archives a working tree, uploads
|
||||
it, rebuilds containers, restarts services. That is the purpose, and it means
|
||||
the ordinary rules for running someone else's shell scripts apply with more
|
||||
force than usual.
|
||||
|
||||
- Read a script before the first run, and run it against something you can
|
||||
afford to break.
|
||||
- Nothing here is a sandbox. There is no dry-run guarantee unless a script
|
||||
documents one; the flag that exists on one command may not exist on the next.
|
||||
- The configuration is yours. The example config carries placeholders, and
|
||||
every host, key path and target in it has to be replaced with your own before
|
||||
anything is pointed at real infrastructure.
|
||||
- Addresses in examples use the ranges reserved for documentation, and
|
||||
loopback. They are placeholders, not somewhere to send anything.
|
||||
|
||||
Scripts that destroy or overwrite state are the ones to read twice. A report
|
||||
that one of them does something destructive without saying so is a good
|
||||
report; a report that a script named after a destructive act performs it is
|
||||
not.
|
||||
|
||||
Release integrity
|
||||
-----------------
|
||||
|
||||
Releases carry checksums. They are an integrity check, not a signature —
|
||||
they catch a truncated download, a corrupted mirror and an accidental edit,
|
||||
and they do not catch a forger, because whoever can change an archive can
|
||||
change the manifest that travels with it.
|
||||
|
||||
Not a finding here
|
||||
------------------
|
||||
|
||||
- Placeholder credentials and example configuration. Fake values are the
|
||||
sanitisation working, not a leak.
|
||||
- The private tree. Only what is published here is in scope; the internal
|
||||
original is not public and cannot be reviewed.
|
||||
296
TOKEN_SAVINGS.txt
Normal file
296
TOKEN_SAVINGS.txt
Normal file
@ -0,0 +1,296 @@
|
||||
Evomedia.net Token Savers — https://github.com/evomedia-net/evo.zscripts
|
||||
Created by Kelly Michels · dev@evomedia.net
|
||||
Licensed under the MIT License. See LICENSE.
|
||||
|
||||
Token Savings: Why You Should Run These Scripts Yourself
|
||||
========================================================
|
||||
|
||||
Running these scripts manually keeps their output out of your AI coding agent's
|
||||
context window. Every line the agent doesn't have to read is a token you don't
|
||||
pay for — and a token the agent can spend on the actual problem instead of on
|
||||
deployment sequencing, SSH output, and Docker health checks.
|
||||
|
||||
This document reports two baselines side by side:
|
||||
|
||||
- You run it → Claude runs the script. What Claude ingests if it invokes the
|
||||
z-script as a single command. These figures are measured (see method below).
|
||||
- You run it → Claude orchestrates raw. What Claude would ingest if the scripts
|
||||
didn't exist and it drove scp / ssh / docker compose step by step itself.
|
||||
These figures are estimates — the same command output plus the agent's
|
||||
reasoning and retry logic across every discrete step.
|
||||
|
||||
The savings from running a script yourself is the first column: if you run it,
|
||||
Claude ingests zero. The extra value of having the scripts at all is the gap
|
||||
between the two columns.
|
||||
|
||||
Dollar equivalents use a blended input/output rate: Sonnet 5 ≈ $9/1M | Opus 4.8 ≈ $15/1M | Fable 5 ≈ $30/1M
|
||||
|
||||
> Measurement note: "Measured" figures come from token-count.ps1, which runs
|
||||
> each script under Start-Transcript and counts output characters ÷ 3.5
|
||||
> chars/token. Captured in Claude Code (Sonnet 4.6) against the sp project.
|
||||
> Strictly speaking that's measured output volume with estimated tokenization:
|
||||
> ÷3.5 is a prose heuristic, and code-heavy output (paths, JSON, container IDs)
|
||||
> fragments into more tokens per character under a real BPE tokenizer — so the
|
||||
> token figures here are likely conservative. "Estimated (raw)" figures are not
|
||||
> measured — they approximate manual orchestration and are marked est.
|
||||
> throughout. Other models/interfaces tokenize differently.
|
||||
|
||||
---
|
||||
|
||||
Measured per-run output (script-run baseline)
|
||||
---------------------------------------------
|
||||
|
||||
The bold figures below are real captures from token-count.ps1 sp; rows flagged est. are not:
|
||||
|
||||
| Script | Measured tokens/run | Notes |
|
||||
|--------|--------------------:|-------|
|
||||
| zec2online | 267 | reachability + version check |
|
||||
| zec2 | 331 | EC2 TCP/HTTP + build match |
|
||||
| zbackup_ec2 | 334 | pull server backup |
|
||||
| zrepair | 364 | clean audit; more if it restarts containers |
|
||||
| zkill | 377 | free the dev port |
|
||||
| zbackup | 436 | local project snapshot |
|
||||
| zrestart | 724 | kill + restart (detached) |
|
||||
| zstart | 762 | start dev server (detached) |
|
||||
| zsync | 769 | mirror backups offsite |
|
||||
| zdeploy (cached) | ~810 | 53s deploy, layers cached |
|
||||
| zdeploy (full rebuild) | ~34,600 est. | packages changed; streams full docker build |
|
||||
| zstart_docker | not measured | est. ~500–1,500 |
|
||||
|
||||
Cache state is what drives zdeploy. A cached deploy is ~810 tokens; the
|
||||
large number only appears on a full rebuild (dependencies changed), which streams
|
||||
the entire docker build. During rapid deploy → test → fix iteration almost every run
|
||||
is cached, so ~810 is the realistic per-run cost — with occasional spikes when you
|
||||
change packages.
|
||||
|
||||
---
|
||||
|
||||
Local Development Control
|
||||
-------------------------
|
||||
|
||||
zstart — Start dev servers
|
||||
--------------------------
|
||||
Measured: ~762 tokens/run | est. raw orchestration: ~1,500–3,000 | typical 2–3 runs/day
|
||||
|
||||
Run it yourself and Claude sees none of the version-bump, MOTD, and startup output.
|
||||
If Claude started the server raw, it would also wait on health checks and confirm
|
||||
the port is listening — reasoning the script does deterministically.
|
||||
|
||||
zstart viteapp # start Vite dev server on its configured port
|
||||
zstart pyapp -Port 3000 # override the port
|
||||
zstart nextapp -Detached # start in background, prompt returns
|
||||
|
||||
---
|
||||
|
||||
zkill — Stop dev servers
|
||||
------------------------
|
||||
Measured: ~377 tokens/run | est. raw orchestration: ~1,000–2,000 | typical 2–3 runs/day
|
||||
|
||||
Raw, Claude would enumerate processes, kill them, and re-check the port is free.
|
||||
The script collapses that to one command.
|
||||
|
||||
zkill viteapp
|
||||
zkill pyapp nextapp
|
||||
|
||||
---
|
||||
|
||||
zrestart — Restart in one command
|
||||
---------------------------------
|
||||
Measured: ~724 tokens/run | est. raw orchestration: ~2,500–4,500 | typical 10–15 runs/day
|
||||
|
||||
The most-used command during rapid iteration. Raw, it's stop → wait → start with
|
||||
error handling at each hop — several tool calls and their reasoning. As one script
|
||||
it's a single call, and the -Detached switch now propagates correctly through the
|
||||
kill→restart chain so the server backgrounds cleanly.
|
||||
|
||||
zrestart viteapp
|
||||
zrestart pyapp -Detached
|
||||
|
||||
---
|
||||
|
||||
Build & Deployment
|
||||
------------------
|
||||
|
||||
zdeploy — Deploy to EC2
|
||||
-----------------------
|
||||
Measured: ~810 tokens/run cached (spikes to ~34,600 on a full rebuild) | est. raw orchestration: ~5,000–12,000 cached, ~35,000+ full rebuild | typical 10–15 runs/day
|
||||
|
||||
The biggest lever — and the one where cache state matters most. The script streams
|
||||
the docker/SSH output whether Claude runs it or not, so a cached deploy really is only
|
||||
~810 tokens even through Claude. The raw-orchestration cost is higher not because of
|
||||
extra output but because Claude would reason between ~15 discrete steps (zip, preflight
|
||||
cleanup, scp, unzip, build, up, version bump, restart, verify) and handle retries
|
||||
itself. Running it yourself zeroes out all of that.
|
||||
|
||||
Measured cached: three runs at 808 / 858 / 808 tokens (53–54s each). The full-rebuild
|
||||
figure (~34,600) is an estimate for package-change deploys — treat it as the upper
|
||||
bound.
|
||||
|
||||
zdeploy pyapp -Note "Fix nav alignment"
|
||||
zdeploy edge # reload edge nginx config
|
||||
zdeploy all -Note "weekly release"
|
||||
|
||||
---
|
||||
|
||||
zstart_docker — Start local Docker stack
|
||||
----------------------------------------
|
||||
Not measured (est. ~500–1,500 tokens/run) | typical 1 run/day
|
||||
|
||||
One-time setup per session; doesn't need agent involvement.
|
||||
|
||||
---
|
||||
|
||||
Backup & Sync
|
||||
-------------
|
||||
|
||||
zbackup — Backup projects locally
|
||||
---------------------------------
|
||||
Measured: ~436 tokens/run | est. raw orchestration: ~1,200–2,500 | typical 1–2 runs/day
|
||||
|
||||
Raw, Claude enumerates files, decides exclusions, compresses, and stamps timestamps.
|
||||
You decide when to snapshot.
|
||||
|
||||
zbackup # everything + scripts folder
|
||||
zbackup pyapp -Tag "pre-refactor"
|
||||
|
||||
---
|
||||
|
||||
zsync — Sync backups offsite
|
||||
----------------------------
|
||||
Measured: ~769 tokens/run | est. raw orchestration: ~1,500–3,000 | typical 1 run/day
|
||||
|
||||
Raw, Claude tracks file diffs, runs robocopy, and verifies the copy. You manage
|
||||
cadence independently.
|
||||
|
||||
zsync
|
||||
zsync viteapp # build + mirror dist to $env:ZSYNC_DEST
|
||||
|
||||
---
|
||||
|
||||
zbackup_ec2 — Pull backups from the server
|
||||
------------------------------------------
|
||||
Measured: ~334 tokens/run | est. raw orchestration: ~1,000–2,000 | typical 1 run/day
|
||||
|
||||
Separates database/app backup from code changes. Claude focuses on code; you manage
|
||||
infrastructure snapshots.
|
||||
|
||||
zbackup_ec2
|
||||
|
||||
---
|
||||
|
||||
Diagnostics & Troubleshooting
|
||||
-----------------------------
|
||||
|
||||
zec2 — Check EC2 reachability
|
||||
-----------------------------
|
||||
Measured: ~331 tokens/run (zec2online: ~267) | est. raw orchestration: ~1,000–2,000 | typical 5–8 runs/day
|
||||
|
||||
When a deploy fails you run this first to confirm EC2 is reachable and the right
|
||||
build is live — before asking Claude to debug. Raw, that's blind network diagnostics
|
||||
over SSH. Runs frequently alongside zdeploy.
|
||||
|
||||
zec2 viteapp
|
||||
zec2 # check all projects
|
||||
zec2online sp # lightweight HTTP-only variant
|
||||
|
||||
---
|
||||
|
||||
zrepair — Audit & repair container routing
|
||||
------------------------------------------
|
||||
Measured: ~364 tokens/run (clean audit) | est. raw orchestration: ~2,000–4,000 | typical 1–2 runs/day
|
||||
|
||||
When a page 502s, this isolates routing vs. DNS vs. app logic across several
|
||||
containers — rather than handing Claude an SSH session to figure out blind. The
|
||||
364-token figure is a healthy run with nothing to repair; a run that actually
|
||||
restarts containers emits more. Raw, Claude would SSH per container and reason
|
||||
across each check.
|
||||
|
||||
zrepair viteapp
|
||||
|
||||
---
|
||||
|
||||
Daily Token Savings Summary
|
||||
---------------------------
|
||||
|
||||
Per-run × runs/day. The per-run figures are measured; the daily totals multiply
|
||||
them by assumed typical run counts (midpoints) — zdeploy and zrestart at
|
||||
10–15/day dominate the sum, so scale the total to your own cadence. The est. raw
|
||||
column approximates what Claude would burn orchestrating the same work with no
|
||||
scripts.
|
||||
|
||||
| Script | Measured/run | Runs/day | Measured/day | Est. raw/day |
|
||||
|--------|-------------:|:--------:|-------------:|-------------:|
|
||||
| zstart | 762 | 2–3 | ~1,900 | ~3,800–9,000 |
|
||||
| zkill | 377 | 2–3 | ~940 | ~2,500–6,000 |
|
||||
| zrestart | 724 | 10–15 | ~9,050 | ~31,000–68,000 |
|
||||
| zdeploy (cached) | ~810 | 10–15 | ~10,100 | ~62,000–180,000 |
|
||||
| zec2 (+online) | ~330 | 5–8 | ~2,200 | ~6,500–16,000 |
|
||||
| zbackup | 436 | 1–2 | ~650 | ~1,800–5,000 |
|
||||
| zsync | 769 | 1 | ~770 | ~1,500–3,000 |
|
||||
| zbackup_ec2 | 334 | 1 | ~330 | ~1,000–2,000 |
|
||||
| zrepair | 364 | 1–2 | ~550 | ~3,000–6,000 |
|
||||
| Total (active dev day) | | | ~26,500 | ~115,000–295,000 est. |
|
||||
|
||||
The ~26,500/day figure is measured per-run at an assumed typical cadence —
|
||||
reproducible on the per-run side, workflow-specific on the multiplier. It reflects an
|
||||
active tool-development day of mostly cached deploys. The
|
||||
~115k–295k est. upper figure is what it would cost to have Claude drive the raw
|
||||
ssh/docker sequences instead — dominated by per-step reasoning on zdeploy and
|
||||
zrestart, not by output volume. Treat that column as an **upper bound, not a
|
||||
prediction**: a capable agent asked to deploy might well write its own wrapper
|
||||
script and ingest very little — the counterfactual depends entirely on how the
|
||||
agent chooses to work. A day with several full-rebuild deploys pushes the measured
|
||||
figure higher too, since each rebuild streams ~34,600 tokens.
|
||||
|
||||
Daily dollar savings during active tool development:
|
||||
|
||||
Script output the agent ingests is billed at input rates, so the measured column
|
||||
uses input pricing. The est.-raw column keeps the blended rate, because raw
|
||||
orchestration also generates agent output (reasoning and tool calls between steps).
|
||||
|
||||
| Model | Measured/day @ input rate | Est. raw/day @ blended rate |
|
||||
|-------|--------------------------:|----------------------------:|
|
||||
| Sonnet 5 | ~$0.08 ($3/1M) | ~$1.04–$2.66 ($9/1M) |
|
||||
| Opus 4.8 | ~$0.13 ($5/1M) | ~$1.73–$4.43 ($15/1M) |
|
||||
| Fable 5 | ~$0.27 ($10/1M) | ~$3.45–$8.85 ($30/1M) |
|
||||
|
||||
One-time ingest slightly understates the true cost: tokens that enter the context are
|
||||
re-sent on every later turn of the session (at cheaper cache-read rates when prompt
|
||||
caching applies), so the cumulative figure is somewhat higher than a single ingest.
|
||||
|
||||
Over a ~22-day working month, the measured savings run ~$2–$6/mo (Sonnet →
|
||||
Fable); the raw-orchestration estimate runs ~$23–$195/mo. The honest dollar
|
||||
figure is small — the real currency is context: every infrastructure token kept
|
||||
out of the window is context your agent keeps for the actual problem, and that's
|
||||
worth more than the dollars suggest.
|
||||
|
||||
---
|
||||
|
||||
Claude Model Token Costs (July 2026)
|
||||
------------------------------------
|
||||
|
||||
| Model | Input | Output | Typical use |
|
||||
|-------|-------|--------|-------------|
|
||||
| Haiku 4.5 | $1/1M | $5/1M | Quick edits, small changes |
|
||||
| Sonnet 5 | $3/1M | $15/1M | Daily coding, medium complexity |
|
||||
| Opus 4.8 | $5/1M | $25/1M | Complex reasoning, multi-file refactors |
|
||||
| Fable 5 | $10/1M | $50/1M | Advanced reasoning, agentic workflows |
|
||||
|
||||
---
|
||||
|
||||
When to Run Scripts Yourself vs. Ask the Agent
|
||||
----------------------------------------------
|
||||
|
||||
Run yourself when:
|
||||
- ✅ You know exactly what action is needed
|
||||
- ✅ The script is deterministic (same input = same output)
|
||||
- ✅ You want to parallelize — run zstart while asking Claude for code
|
||||
- ✅ You're troubleshooting and need fast feedback loops
|
||||
|
||||
Ask the agent when:
|
||||
- ❌ You need conditional logic ("if this test fails, try X")
|
||||
- ❌ You're chaining operations that depend on each other's output
|
||||
- ❌ You want the agent to interpret script output and decide next steps
|
||||
|
||||
Bottom line: These scripts are optimized for you to run directly. Use them. Save
|
||||
tokens. Let Claude focus on coding.
|
||||
@ -28,11 +28,20 @@ from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
# Every markdown file that owes the repo a plain-text twin.
|
||||
PAIRS = (
|
||||
("README.md", "README.txt"),
|
||||
("CHANGELOG.md", "CHANGELOG.txt"),
|
||||
)
|
||||
def pairs() -> tuple[tuple[str, str], ...]:
|
||||
"""Every root-level markdown file, with the twin it owes.
|
||||
|
||||
Discovered rather than listed. The list was hand-kept, and two files had
|
||||
quietly outgrown it - ELEVATOR_PITCH.md and TOKEN_SAVINGS.md had no twin
|
||||
at all, because adding a document and remembering to add it here are two
|
||||
separate acts and the second one is the one that gets skipped. Discovery
|
||||
makes them one act.
|
||||
"""
|
||||
return tuple((f.name, f.with_suffix(".txt").name) for f in sorted(ROOT.glob("*.md")))
|
||||
|
||||
|
||||
#: Kept as a name because the Pester suite reads it to know what to check.
|
||||
PAIRS = pairs()
|
||||
|
||||
|
||||
def _inline(text: str) -> str:
|
||||
|
||||
@ -38,14 +38,19 @@ Describe "plain-text twins" {
|
||||
Test-Path -LiteralPath $script:Generator | Should -BeTrue
|
||||
}
|
||||
|
||||
It "every .md that owes a twin has one" {
|
||||
$pairs = Select-String -Path $script:Generator -Pattern '^\s*\("([^"]+\.md)",\s*"([^"]+\.txt)"\),' |
|
||||
ForEach-Object { [pscustomobject]@{ Md = $_.Matches[0].Groups[1].Value; Txt = $_.Matches[0].Groups[2].Value } }
|
||||
# Asks the REPOSITORY what markdown it has, not the generator what it was
|
||||
# told about. Scraping the generator's own list could only ever prove the
|
||||
# list was self-consistent - a document nobody added to it was invisible to
|
||||
# the check, which is how ELEVATOR_PITCH.md and TOKEN_SAVINGS.md sat here
|
||||
# with no twin while this test passed.
|
||||
It "every .md at the repository root has a twin" {
|
||||
$mds = Get-ChildItem -LiteralPath $script:RepoRoot -Filter *.md -File
|
||||
|
||||
$pairs.Count | Should -BeGreaterThan 0 -Because "PAIRS in plaintext_twins.py is what this suite checks"
|
||||
foreach ($p in $pairs) {
|
||||
Test-Path -LiteralPath (Join-Path $script:RepoRoot $p.Md) | Should -BeTrue -Because "$($p.Md) is listed in PAIRS"
|
||||
Test-Path -LiteralPath (Join-Path $script:RepoRoot $p.Txt) | Should -BeTrue -Because "$($p.Md) owes a twin at $($p.Txt)"
|
||||
$mds.Count | Should -BeGreaterThan 0 -Because "the repo documents itself in markdown"
|
||||
foreach ($md in $mds) {
|
||||
$txt = [IO.Path]::ChangeExtension($md.FullName, ".txt")
|
||||
Test-Path -LiteralPath $txt |
|
||||
Should -BeTrue -Because "$($md.Name) owes a twin at $(Split-Path -Leaf $txt)"
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
Loading…
Reference in New Issue
Block a user