From a4c557bc870689da1a21da59f92c09a500f0ad5e Mon Sep 17 00:00:00 2001 From: KellyMichels Date: Thu, 16 Jul 2026 11:33:57 -0500 Subject: [PATCH] docs: reconcile headline figures with measured data; price ingest at input rates - One headline number everywhere: ~26,500 measured tokens/day (README said 3,000-7,000; ELEVATOR_PITCH said 157,000-540,000) - Measurement note: measured output volume, estimated tokenization; /3.5 is conservative for code-heavy output - Raw-orchestration column explicitly labeled an upper bound, not a prediction - Measured dollar column priced at input rates (blended kept for est-raw only); note on re-sent context tokens - Untrack md/Z-ScriptTokenData.md (superseded internal notes; md/ gitignored) --- .gitignore | 3 + ELEVATOR_PITCH.md | 4 +- README.md | 2 +- TOKEN_SAVINGS.md | 44 +++++++---- md/Z-ScriptTokenData.md | 169 ---------------------------------------- 5 files changed, 35 insertions(+), 187 deletions(-) delete mode 100644 md/Z-ScriptTokenData.md diff --git a/.gitignore b/.gitignore index 30ce77c..bc7a0e2 100644 --- a/.gitignore +++ b/.gitignore @@ -13,3 +13,6 @@ Thumbs.db desktop.ini .idea/ .vscode/ + +# Internal working notes (superseded by TOKEN_SAVINGS.md) +md/ diff --git a/ELEVATOR_PITCH.md b/ELEVATOR_PITCH.md index 8c94a3b..884ab1a 100644 --- a/ELEVATOR_PITCH.md +++ b/ELEVATOR_PITCH.md @@ -12,11 +12,11 @@ Licensed under the MIT License. See LICENSE. ## The 30-second version -Every time you ask an AI coding agent to deploy your app, it burns 15,000–35,000 tokens reading Docker build output, SSH logs, and health-check responses — before writing a single line of code. Do that 10–15 times a day on a rapid dev cycle and you've spent $1.41–$16 just on infrastructure chatter (Sonnet 5 to Fable 5), plus context window space that should go to your actual problem. +Every time your AI coding agent runs your infrastructure for you — a deploy, a restart, a health check — the full output lands in its context window: Docker layers, SSH banners, health-check chatter. We measured it: an active dev day pushes **~26,500 tokens of pure script output** through the agent, and a single full Docker rebuild adds ~35,000 more. The dollars are small; the context is not — every line of infrastructure noise crowds out the code your agent is supposed to be reasoning about. Token Savers collapses the infrastructure side into short, one-word commands you run yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. Describe each project once in `zconfig.json` — where it lives, what kind it is, where it deploys — and every command just knows. You run the deploy; your agent edits the code. You run the health check; your agent reads the result and fixes whatever's wrong. -**Estimated savings: 157,000–540,000 tokens per active development day.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script breakdown and measurement methodology. +**Measured: ~26,500 tokens of script output per active development day** — kept out of your agent's context entirely when you run the commands yourself. See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script measurements and method. ## Why it's different diff --git a/README.md b/README.md index d153867..db08884 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ Licensed under the MIT License. See LICENSE. When an AI coding agent orchestrates your infrastructure — starting dev servers, deploying to EC2, diagnosing 502s — it spends hundreds to thousands of tokens per operation on SSH plumbing, Docker output, and retry logic. Those tokens should go to code. -Token Savers gives you short, one-word commands to run those parts yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. You handle the deterministic infrastructure; your agent handles code. **Running these scripts manually instead of asking your agent to orchestrate them saves an estimated 3,000–7,000 tokens per active development day.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script breakdown. +Token Savers gives you short, one-word commands to run those parts yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. You handle the deterministic infrastructure; your agent handles code. **Running these scripts manually instead of asking your agent to orchestrate them keeps a measured ~26,500 tokens of infrastructure output per active development day out of your agent's context window.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script measurements and method. Every command is a tiny PowerShell script driven by a single JSON config file. The project key you define in that config **is** the command argument — add `myapp` to the config and `zstart myapp`, `zdeploy myapp`, `zbackup myapp` all just work, no script edits needed. diff --git a/TOKEN_SAVINGS.md b/TOKEN_SAVINGS.md index bf4be18..d7a95c6 100644 --- a/TOKEN_SAVINGS.md +++ b/TOKEN_SAVINGS.md @@ -29,9 +29,12 @@ Dollar equivalents use a blended input/output rate: **Sonnet 5 ≈ $9/1M** | **O > **Measurement note:** "Measured" figures come from `token-count.ps1`, which runs > each script under `Start-Transcript` and counts output characters ÷ 3.5 > chars/token. Captured in **Claude Code (Sonnet 4.6)** against the `sp` project. -> "Estimated (raw)" figures are *not* measured — they approximate manual -> orchestration and are marked *est.* throughout. Other models/interfaces tokenize -> differently. +> Strictly speaking that's *measured output volume with estimated tokenization*: +> ÷3.5 is a prose heuristic, and code-heavy output (paths, JSON, container IDs) +> fragments into **more** tokens per character under a real BPE tokenizer — so the +> token figures here are likely conservative. "Estimated (raw)" figures are *not* +> measured — they approximate manual orchestration and are marked *est.* +> throughout. Other models/interfaces tokenize differently. --- @@ -233,22 +236,33 @@ The **~26,500 tokens/day measured** is the honest, reproducible savings from run these yourself during an active tool-development day (mostly cached deploys). The **~115k–295k est.** upper figure is what it would cost to have Claude drive the raw `ssh`/`docker` sequences instead — dominated by per-step reasoning on `zdeploy` and -`zrestart`, not by output volume. A day with several full-rebuild deploys pushes the -measured figure higher too, since each rebuild streams ~34,600 tokens. +`zrestart`, not by output volume. Treat that column as an **upper bound, not a +prediction**: a capable agent asked to deploy might well write its own wrapper +script and ingest very little — the counterfactual depends entirely on how the +agent chooses to work. A day with several full-rebuild deploys pushes the measured +figure higher too, since each rebuild streams ~34,600 tokens. **Daily dollar savings during active tool development:** -| Model | Blended rate | Measured/day | Est. raw/day (no scripts) | -|-------|-------------|-------------:|--------------------------:| -| **Sonnet 5** | $9/1M | ~$0.24 | ~$1.04–$2.66 | -| **Opus 4.8** | $15/1M | ~$0.40 | ~$1.73–$4.43 | -| **Fable 5** | $30/1M | ~$0.80 | ~$3.45–$8.85 | +Script output the agent ingests is billed at **input** rates, so the measured column +uses input pricing. The est.-raw column keeps the **blended** rate, because raw +orchestration also generates agent *output* (reasoning and tool calls between steps). -Over a ~22-day working month, the measured savings run **~$5–$17/mo** (Sonnet → -Fable); the raw-orchestration estimate runs **~$23–$195/mo**. Either way, the -token-budget point stands: every token saved on infrastructure is a token your agent -keeps for the actual problem — and that context-window quality is worth more than the -raw dollar figure suggests. +| Model | Measured/day @ input rate | Est. raw/day @ blended rate | +|-------|--------------------------:|----------------------------:| +| **Sonnet 5** | ~$0.08 ($3/1M) | ~$1.04–$2.66 ($9/1M) | +| **Opus 4.8** | ~$0.13 ($5/1M) | ~$1.73–$4.43 ($15/1M) | +| **Fable 5** | ~$0.27 ($10/1M) | ~$3.45–$8.85 ($30/1M) | + +One-time ingest slightly understates the true cost: tokens that enter the context are +re-sent on every later turn of the session (at cheaper cache-read rates when prompt +caching applies), so the cumulative figure is somewhat higher than a single ingest. + +Over a ~22-day working month, the measured savings run **~$2–$6/mo** (Sonnet → +Fable); the raw-orchestration estimate runs **~$23–$195/mo**. The honest dollar +figure is small — the real currency is **context**: every infrastructure token kept +out of the window is context your agent keeps for the actual problem, and that's +worth more than the dollars suggest. --- diff --git a/md/Z-ScriptTokenData.md b/md/Z-ScriptTokenData.md deleted file mode 100644 index 3a4c3eb..0000000 --- a/md/Z-ScriptTokenData.md +++ /dev/null @@ -1,169 +0,0 @@ -# Z-Scripts Summary: Manual Automation to Reduce Token Usage - -Running these scripts **manually** instead of asking Claude Code to orchestrate them saves significant token usage because you skip the overhead of Claude reasoning about deployment/build/testing orchestration. - ---- - -## **Local Development Control** - -### `zstart.ps1` — Start dev servers -**What:** Starts local dev servers for any project defined in `zconfig.json` -```powershell -zstart viteapp # Start Vite dev server -zstart pyapp -Port 3000 # Start Python app on custom port -zstart nextapp -Detached # Start in background -``` -**Why run manually:** Eliminates Claude's need to track startup, wait for health checks, or validate ports. You start the server once and Claude just edits files—tokens saved on orchestration, ~150-300 tokens per use. - ---- - -### `zkill.ps1` — Stop dev servers -**What:** Kills Node/Python/Docker processes on specified ports -```powershell -zkill viteapp # Kill Vite dev server -zkill pyapp nextapp # Kill multiple projects -``` -**Why run manually:** You control when to stop iteration cycles. Saves Claude from having to reason about process cleanup, ~100-200 tokens. - ---- - -### `zrestart.ps1` — Restart in one command -**What:** Calls `zkill` + `zstart` atomically; useful after major changes -```powershell -zrestart viteapp # Kill + restart Vite app -``` -**Why run manually:** Cleaner than telling Claude "stop the server and start it again"—one script handles the sequence. Saves ~200-300 tokens on orchestration logic. - ---- - -## **Build & Deployment** - -### `zdeploy.ps1` — Deploy to EC2 -**What:** Zips source, SCP to EC2, runs docker compose up, verifies build version -```powershell -zdeploy pyapp -Note "Fix nav alignment" -zdeploy edge # Just reload edge nginx config -zdeploy nextapp # Deploy Next.js app + db -zdeploy all # Deploy all projects (edge first) -``` -**Why run manually:** Deployment is deterministic once code is ready. You test locally, then run the deploy script—Claude never needs to understand EC2 SSH, zip compression, docker compose, or deployment verification. Saves ~800-1200 tokens that would otherwise go to deployment orchestration. - ---- - -### `zstart_docker.ps1` — Ensure Docker daemon is running -**What:** Starts Docker Desktop if not running (Windows convenience) -**Why run manually:** One-time setup; doesn't need Claude involvement. - ---- - -## **Backup & Sync** - -### `zbackup.ps1` — Backup projects locally -**What:** Zips any project defined in `zconfig.json` to the local backups folder with a timestamp -```powershell -zbackup # Backup everything + this scripts folder -zbackup viteapp nextapp # Backup just those two projects -zbackup pyapp -Tag "pre-refactor" -``` -**Why run manually:** You decide when to snapshot. Running this yourself before risky changes means Claude never needs to reason about backup strategy or file compression. Saves ~300-500 tokens per session. - ---- - -### `zsync.ps1` — Sync backups to OneDrive -**What:** Robocopy new files from the local backups folder to OneDrive (incremental) -```powershell -zsync # Sync all new backup files -zsync viteapp # Build + mirror vite dist to $env:ZSYNC_DEST -``` -**Why run manually:** You manage backup cadence independently. Claude doesn't need to reason about incremental sync logic or file enumeration. Saves ~250-400 tokens. - ---- - -### `zbackup_ec2.ps1` — Remote backups on EC2 -**What:** SSH to EC2, tar application and database data, pull to local backups folder -**Why run manually:** Separates database/app backup concerns from code changes. Claude focuses on code; you manage infrastructure snapshots. - ---- - -## **Diagnostics & Troubleshooting** - -### `zec2.ps1` — Check EC2 reachability -**What:** TCP + HTTP connectivity tests + live build version for all projects with a `domain` -```powershell -zec2 viteapp # Check if Vite app is up on EC2 -zec2 # Test all projects -``` -**Why run manually:** When a deploy fails, you run this to verify EC2 is reachable before asking Claude to debug. Eliminates Claude doing network diagnostics blind. Saves ~400-600 tokens of troubleshooting overhead. - ---- - -### `zec2online.ps1` — Deep EC2 health check -**What:** Full health check; auto-starts downed stacks, streams diagnostics -**Why run manually:** Quick check before starting work; Claude doesn't need to validate infrastructure state. - ---- - -### `zrepair.ps1` — Audit & repair container routing -**What:** SSH to EC2, verify nginx proxy routes, check docker compose health, run smoke tests -```powershell -zrepair viteapp # Audit proxy path + smoke test -zrepair all # Audit all projects -``` -**Why run manually:** When pages 502, you run this first to isolate whether it's routing, DNS, or app logic. Saves ~1000-1500 tokens of "try this, check logs, try that" debugging. - ---- - -### `zsetup_mail.ps1` — Email account provisioning -**What:** Automates creation of email accounts on EC2; displays Route 53 DNS requirements -**Why run manually:** One-time setup task; doesn't benefit from Claude guidance. - ---- - -## **Token Usage Impact by Script** - -| Script | Manual Run Saves | Without Script (Claude orchestrates) | -|--------|-----------------|--------------------------------------| -| **zstart** | ~150-300 tokens | Claude tracks startup, validates ports, polls health | -| **zkill** | ~100-200 tokens | Claude enumerates processes, checks exit codes | -| **zrestart** | ~200-300 tokens | Claude chains stop→wait→start with error handling | -| **zdeploy** | **~800-1200 tokens** | Claude manages zip, SSH, SCP, compose, verification | -| **zbackup** | ~300-500 tokens | Claude enumerates, compresses, manages timestamps | -| **zsync** | ~250-400 tokens | Claude tracks file diffs, runs robocopy, verifies copy | -| **zec2** | ~400-600 tokens | Claude does TCP/HTTP tests, parses output | -| **zrepair** | **~1000-1500 tokens** | Claude SSH, grep logs, run smoke tests, interpret failures | - -**Total potential savings per day of active development: 3,000–7,000 tokens** if you run these manually vs. asking Claude to orchestrate. - ---- - -## **Claude Model Token Costs** *(Approximate, May 2026)* - -| Model | Input Cost | Output Cost | Use Case | -|-------|-----------|-----------|----------| -| **Haiku 4.5** | ~$0.80/1M | ~$4/1M | Quick code edits, small changes | -| **Sonnet 4.6** | ~$3/1M | ~$15/1M | Daily coding, medium complexity | -| **Opus 4.8** | ~$15/1M | ~$60/1M | Complex reasoning, multi-file refactors | - -**Token savings example:** -- Running `zdeploy` manually: ~1000 tokens saved × $15/1M (Sonnet input) = ~$0.015 saved -- Running `zrepair` manually: ~1500 tokens saved × $15/1M = ~$0.0225 saved -- Running all scripts daily: ~5000 tokens × $0.015 = ~$0.075 saved per day, ~$22.50/month - -**More importantly:** Manual scripts let Claude focus on *code logic* instead of *infrastructure orchestration*—where Claude adds real value. - ---- - -## **TL;DR** - -Run these scripts manually when: -- ✅ You know the exact deployment/test/backup action needed -- ✅ The script is deterministic (same input = same result) -- ✅ You want to parallelize (run zstart while asking Claude for code) -- ✅ You're troubleshooting and need fast feedback loops - -Ask Claude to *invoke* them only when: -- ❌ You need complex conditional logic (e.g., "if this test fails, try X") -- ❌ You're chaining many operations that depend on each other's output -- ❌ You want Claude to interpret script output and decide next steps - -**Bottom line:** Your z-scripts are optimized for **you** to run directly. Use them. Save tokens. Let Claude focus on coding.