docs: reconcile headline figures with measured data; price ingest at input rates

- One headline number everywhere: ~26,500 measured tokens/day (README said 3,000-7,000; ELEVATOR_PITCH said 157,000-540,000)
- Measurement note: measured output volume, estimated tokenization; /3.5 is conservative for code-heavy output
- Raw-orchestration column explicitly labeled an upper bound, not a prediction
- Measured dollar column priced at input rates (blended kept for est-raw only); note on re-sent context tokens
- Untrack md/Z-ScriptTokenData.md (superseded internal notes; md/ gitignored)
This commit is contained in:
KellyMichels 2026-07-16 11:33:57 -05:00
parent 3546a12564
commit a4c557bc87
5 changed files with 35 additions and 187 deletions

3
.gitignore vendored
View File

@ -13,3 +13,6 @@ Thumbs.db
desktop.ini desktop.ini
.idea/ .idea/
.vscode/ .vscode/
# Internal working notes (superseded by TOKEN_SAVINGS.md)
md/

View File

@ -12,11 +12,11 @@ Licensed under the MIT License. See LICENSE.
## The 30-second version ## The 30-second version
Every time you ask an AI coding agent to deploy your app, it burns 15,000–35,000 tokens reading Docker build output, SSH logs, and health-check responses — before writing a single line of code. Do that 10–15 times a day on a rapid dev cycle and you've spent $1.41–$16 just on infrastructure chatter (Sonnet 5 to Fable 5), plus context window space that should go to your actual problem. Every time your AI coding agent runs your infrastructure for you — a deploy, a restart, a health check — the full output lands in its context window: Docker layers, SSH banners, health-check chatter. We measured it: an active dev day pushes **~26,500 tokens of pure script output** through the agent, and a single full Docker rebuild adds ~35,000 more. The dollars are small; the context is not — every line of infrastructure noise crowds out the code your agent is supposed to be reasoning about.
Token Savers collapses the infrastructure side into short, one-word commands you run yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. Describe each project once in `zconfig.json` — where it lives, what kind it is, where it deploys — and every command just knows. You run the deploy; your agent edits the code. You run the health check; your agent reads the result and fixes whatever's wrong. Token Savers collapses the infrastructure side into short, one-word commands you run yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. Describe each project once in `zconfig.json` — where it lives, what kind it is, where it deploys — and every command just knows. You run the deploy; your agent edits the code. You run the health check; your agent reads the result and fixes whatever's wrong.
**Estimated savings: 157,000–540,000 tokens per active development day.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script breakdown and measurement methodology. **Measured: ~26,500 tokens of script output per active development day** — kept out of your agent's context entirely when you run the commands yourself. See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script measurements and method.
## Why it's different ## Why it's different

View File

@ -8,7 +8,7 @@ Licensed under the MIT License. See LICENSE.
When an AI coding agent orchestrates your infrastructure — starting dev servers, deploying to EC2, diagnosing 502s — it spends hundreds to thousands of tokens per operation on SSH plumbing, Docker output, and retry logic. Those tokens should go to code. When an AI coding agent orchestrates your infrastructure — starting dev servers, deploying to EC2, diagnosing 502s — it spends hundreds to thousands of tokens per operation on SSH plumbing, Docker output, and retry logic. Those tokens should go to code.
Token Savers gives you short, one-word commands to run those parts yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. You handle the deterministic infrastructure; your agent handles code. **Running these scripts manually instead of asking your agent to orchestrate them saves an estimated 3,000–7,000 tokens per active development day.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script breakdown. Token Savers gives you short, one-word commands to run those parts yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. You handle the deterministic infrastructure; your agent handles code. **Running these scripts manually instead of asking your agent to orchestrate them keeps a measured ~26,500 tokens of infrastructure output per active development day out of your agent's context window.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script measurements and method.
Every command is a tiny PowerShell script driven by a single JSON config file. The project key you define in that config **is** the command argument — add `myapp` to the config and `zstart myapp`, `zdeploy myapp`, `zbackup myapp` all just work, no script edits needed. Every command is a tiny PowerShell script driven by a single JSON config file. The project key you define in that config **is** the command argument — add `myapp` to the config and `zstart myapp`, `zdeploy myapp`, `zbackup myapp` all just work, no script edits needed.

View File

@ -29,9 +29,12 @@ Dollar equivalents use a blended input/output rate: **Sonnet 5 ≈ $9/1M** | **O
> **Measurement note:** "Measured" figures come from `token-count.ps1`, which runs > **Measurement note:** "Measured" figures come from `token-count.ps1`, which runs
> each script under `Start-Transcript` and counts output characters ÷ 3.5 > each script under `Start-Transcript` and counts output characters ÷ 3.5
> chars/token. Captured in **Claude Code (Sonnet 4.6)** against the `sp` project. > chars/token. Captured in **Claude Code (Sonnet 4.6)** against the `sp` project.
> "Estimated (raw)" figures are *not* measured — they approximate manual > Strictly speaking that's *measured output volume with estimated tokenization*:
> orchestration and are marked *est.* throughout. Other models/interfaces tokenize > ÷3.5 is a prose heuristic, and code-heavy output (paths, JSON, container IDs)
> differently. > fragments into **more** tokens per character under a real BPE tokenizer — so the
> token figures here are likely conservative. "Estimated (raw)" figures are *not*
> measured — they approximate manual orchestration and are marked *est.*
> throughout. Other models/interfaces tokenize differently.
--- ---
@ -233,22 +236,33 @@ The **~26,500 tokens/day measured** is the honest, reproducible savings from run
these yourself during an active tool-development day (mostly cached deploys). The these yourself during an active tool-development day (mostly cached deploys). The
**~115k–295k est.** upper figure is what it would cost to have Claude drive the raw **~115k–295k est.** upper figure is what it would cost to have Claude drive the raw
`ssh`/`docker` sequences instead — dominated by per-step reasoning on `zdeploy` and `ssh`/`docker` sequences instead — dominated by per-step reasoning on `zdeploy` and
`zrestart`, not by output volume. A day with several full-rebuild deploys pushes the `zrestart`, not by output volume. Treat that column as an **upper bound, not a
measured figure higher too, since each rebuild streams ~34,600 tokens. prediction**: a capable agent asked to deploy might well write its own wrapper
script and ingest very little — the counterfactual depends entirely on how the
agent chooses to work. A day with several full-rebuild deploys pushes the measured
figure higher too, since each rebuild streams ~34,600 tokens.
**Daily dollar savings during active tool development:** **Daily dollar savings during active tool development:**
| Model | Blended rate | Measured/day | Est. raw/day (no scripts) | Script output the agent ingests is billed at **input** rates, so the measured column
|-------|-------------|-------------:|--------------------------:| uses input pricing. The est.-raw column keeps the **blended** rate, because raw
| **Sonnet 5** | $9/1M | ~$0.24 | ~$1.04–$2.66 | orchestration also generates agent *output* (reasoning and tool calls between steps).
| **Opus 4.8** | $15/1M | ~$0.40 | ~$1.73–$4.43 |
| **Fable 5** | $30/1M | ~$0.80 | ~$3.45–$8.85 |
Over a ~22-day working month, the measured savings run **~$5–$17/mo** (Sonnet → | Model | Measured/day @ input rate | Est. raw/day @ blended rate |
Fable); the raw-orchestration estimate runs **~$23–$195/mo**. Either way, the |-------|--------------------------:|----------------------------:|
token-budget point stands: every token saved on infrastructure is a token your agent | **Sonnet 5** | ~$0.08 ($3/1M) | ~$1.04–$2.66 ($9/1M) |
keeps for the actual problem — and that context-window quality is worth more than the | **Opus 4.8** | ~$0.13 ($5/1M) | ~$1.73–$4.43 ($15/1M) |
raw dollar figure suggests. | **Fable 5** | ~$0.27 ($10/1M) | ~$3.45–$8.85 ($30/1M) |
One-time ingest slightly understates the true cost: tokens that enter the context are
re-sent on every later turn of the session (at cheaper cache-read rates when prompt
caching applies), so the cumulative figure is somewhat higher than a single ingest.
Over a ~22-day working month, the measured savings run **~$2–$6/mo** (Sonnet →
Fable); the raw-orchestration estimate runs **~$23–$195/mo**. The honest dollar
figure is small — the real currency is **context**: every infrastructure token kept
out of the window is context your agent keeps for the actual problem, and that's
worth more than the dollars suggest.
--- ---

View File

@ -1,169 +0,0 @@
# Z-Scripts Summary: Manual Automation to Reduce Token Usage
Running these scripts **manually** instead of asking Claude Code to orchestrate them saves significant token usage because you skip the overhead of Claude reasoning about deployment/build/testing orchestration.
---
## **Local Development Control**
### `zstart.ps1` — Start dev servers
**What:** Starts local dev servers for any project defined in `zconfig.json`
```powershell
zstart viteapp # Start Vite dev server
zstart pyapp -Port 3000 # Start Python app on custom port
zstart nextapp -Detached # Start in background
```
**Why run manually:** Eliminates Claude's need to track startup, wait for health checks, or validate ports. You start the server once and Claude just edits files—tokens saved on orchestration, ~150-300 tokens per use.
---
### `zkill.ps1` — Stop dev servers
**What:** Kills Node/Python/Docker processes on specified ports
```powershell
zkill viteapp # Kill Vite dev server
zkill pyapp nextapp # Kill multiple projects
```
**Why run manually:** You control when to stop iteration cycles. Saves Claude from having to reason about process cleanup, ~100-200 tokens.
---
### `zrestart.ps1` — Restart in one command
**What:** Calls `zkill` + `zstart` atomically; useful after major changes
```powershell
zrestart viteapp # Kill + restart Vite app
```
**Why run manually:** Cleaner than telling Claude "stop the server and start it again"—one script handles the sequence. Saves ~200-300 tokens on orchestration logic.
---
## **Build & Deployment**
### `zdeploy.ps1` — Deploy to EC2
**What:** Zips source, SCP to EC2, runs docker compose up, verifies build version
```powershell
zdeploy pyapp -Note "Fix nav alignment"
zdeploy edge # Just reload edge nginx config
zdeploy nextapp # Deploy Next.js app + db
zdeploy all # Deploy all projects (edge first)
```
**Why run manually:** Deployment is deterministic once code is ready. You test locally, then run the deploy script—Claude never needs to understand EC2 SSH, zip compression, docker compose, or deployment verification. Saves ~800-1200 tokens that would otherwise go to deployment orchestration.
---
### `zstart_docker.ps1` — Ensure Docker daemon is running
**What:** Starts Docker Desktop if not running (Windows convenience)
**Why run manually:** One-time setup; doesn't need Claude involvement.
---
## **Backup & Sync**
### `zbackup.ps1` — Backup projects locally
**What:** Zips any project defined in `zconfig.json` to the local backups folder with a timestamp
```powershell
zbackup # Backup everything + this scripts folder
zbackup viteapp nextapp # Backup just those two projects
zbackup pyapp -Tag "pre-refactor"
```
**Why run manually:** You decide when to snapshot. Running this yourself before risky changes means Claude never needs to reason about backup strategy or file compression. Saves ~300-500 tokens per session.
---
### `zsync.ps1` — Sync backups to OneDrive
**What:** Robocopy new files from the local backups folder to OneDrive (incremental)
```powershell
zsync # Sync all new backup files
zsync viteapp # Build + mirror vite dist to $env:ZSYNC_DEST
```
**Why run manually:** You manage backup cadence independently. Claude doesn't need to reason about incremental sync logic or file enumeration. Saves ~250-400 tokens.
---
### `zbackup_ec2.ps1` — Remote backups on EC2
**What:** SSH to EC2, tar application and database data, pull to local backups folder
**Why run manually:** Separates database/app backup concerns from code changes. Claude focuses on code; you manage infrastructure snapshots.
---
## **Diagnostics & Troubleshooting**
### `zec2.ps1` — Check EC2 reachability
**What:** TCP + HTTP connectivity tests + live build version for all projects with a `domain`
```powershell
zec2 viteapp # Check if Vite app is up on EC2
zec2 # Test all projects
```
**Why run manually:** When a deploy fails, you run this to verify EC2 is reachable before asking Claude to debug. Eliminates Claude doing network diagnostics blind. Saves ~400-600 tokens of troubleshooting overhead.
---
### `zec2online.ps1` — Deep EC2 health check
**What:** Full health check; auto-starts downed stacks, streams diagnostics
**Why run manually:** Quick check before starting work; Claude doesn't need to validate infrastructure state.
---
### `zrepair.ps1` — Audit & repair container routing
**What:** SSH to EC2, verify nginx proxy routes, check docker compose health, run smoke tests
```powershell
zrepair viteapp # Audit proxy path + smoke test
zrepair all # Audit all projects
```
**Why run manually:** When pages 502, you run this first to isolate whether it's routing, DNS, or app logic. Saves ~1000-1500 tokens of "try this, check logs, try that" debugging.
---
### `zsetup_mail.ps1` — Email account provisioning
**What:** Automates creation of email accounts on EC2; displays Route 53 DNS requirements
**Why run manually:** One-time setup task; doesn't benefit from Claude guidance.
---
## **Token Usage Impact by Script**
| Script | Manual Run Saves | Without Script (Claude orchestrates) |
|--------|-----------------|--------------------------------------|
| **zstart** | ~150-300 tokens | Claude tracks startup, validates ports, polls health |
| **zkill** | ~100-200 tokens | Claude enumerates processes, checks exit codes |
| **zrestart** | ~200-300 tokens | Claude chains stop→wait→start with error handling |
| **zdeploy** | **~800-1200 tokens** | Claude manages zip, SSH, SCP, compose, verification |
| **zbackup** | ~300-500 tokens | Claude enumerates, compresses, manages timestamps |
| **zsync** | ~250-400 tokens | Claude tracks file diffs, runs robocopy, verifies copy |
| **zec2** | ~400-600 tokens | Claude does TCP/HTTP tests, parses output |
| **zrepair** | **~1000-1500 tokens** | Claude SSH, grep logs, run smoke tests, interpret failures |
**Total potential savings per day of active development: 3,000–7,000 tokens** if you run these manually vs. asking Claude to orchestrate.
---
## **Claude Model Token Costs** *(Approximate, May 2026)*
| Model | Input Cost | Output Cost | Use Case |
|-------|-----------|-----------|----------|
| **Haiku 4.5** | ~$0.80/1M | ~$4/1M | Quick code edits, small changes |
| **Sonnet 4.6** | ~$3/1M | ~$15/1M | Daily coding, medium complexity |
| **Opus 4.8** | ~$15/1M | ~$60/1M | Complex reasoning, multi-file refactors |
**Token savings example:**
- Running `zdeploy` manually: ~1000 tokens saved × $15/1M (Sonnet input) = ~$0.015 saved
- Running `zrepair` manually: ~1500 tokens saved × $15/1M = ~$0.0225 saved
- Running all scripts daily: ~5000 tokens × $0.015 = ~$0.075 saved per day, ~$22.50/month
**More importantly:** Manual scripts let Claude focus on *code logic* instead of *infrastructure orchestration*—where Claude adds real value.
---
## **TL;DR**
Run these scripts manually when:
- ✅ You know the exact deployment/test/backup action needed
- ✅ The script is deterministic (same input = same result)
- ✅ You want to parallelize (run zstart while asking Claude for code)
- ✅ You're troubleshooting and need fast feedback loops
Ask Claude to *invoke* them only when:
- ❌ You need complex conditional logic (e.g., "if this test fails, try X")
- ❌ You're chaining many operations that depend on each other's output
- ❌ You want Claude to interpret script output and decide next steps
**Bottom line:** Your z-scripts are optimized for **you** to run directly. Use them. Save tokens. Let Claude focus on coding.