mirror of
https://github.com/kellymichels/zscripts-token-savers
synced 2026-10-07 07:18:18 +00:00
docs: reconcile headline figures with measured data; price ingest at input rates
- One headline number everywhere: ~26,500 measured tokens/day (README said 3,000-7,000; ELEVATOR_PITCH said 157,000-540,000) - Measurement note: measured output volume, estimated tokenization; /3.5 is conservative for code-heavy output - Raw-orchestration column explicitly labeled an upper bound, not a prediction - Measured dollar column priced at input rates (blended kept for est-raw only); note on re-sent context tokens - Untrack md/Z-ScriptTokenData.md (superseded internal notes; md/ gitignored)
This commit is contained in:
parent
3546a12564
commit
a4c557bc87
3
.gitignore
vendored
3
.gitignore
vendored
@ -13,3 +13,6 @@ Thumbs.db
|
|||||||
desktop.ini
|
desktop.ini
|
||||||
.idea/
|
.idea/
|
||||||
.vscode/
|
.vscode/
|
||||||
|
|
||||||
|
# Internal working notes (superseded by TOKEN_SAVINGS.md)
|
||||||
|
md/
|
||||||
|
|||||||
@ -12,11 +12,11 @@ Licensed under the MIT License. See LICENSE.
|
|||||||
|
|
||||||
## The 30-second version
|
## The 30-second version
|
||||||
|
|
||||||
Every time you ask an AI coding agent to deploy your app, it burns 15,000–35,000 tokens reading Docker build output, SSH logs, and health-check responses — before writing a single line of code. Do that 10–15 times a day on a rapid dev cycle and you've spent $1.41–$16 just on infrastructure chatter (Sonnet 5 to Fable 5), plus context window space that should go to your actual problem.
|
Every time your AI coding agent runs your infrastructure for you — a deploy, a restart, a health check — the full output lands in its context window: Docker layers, SSH banners, health-check chatter. We measured it: an active dev day pushes **~26,500 tokens of pure script output** through the agent, and a single full Docker rebuild adds ~35,000 more. The dollars are small; the context is not — every line of infrastructure noise crowds out the code your agent is supposed to be reasoning about.
|
||||||
|
|
||||||
Token Savers collapses the infrastructure side into short, one-word commands you run yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. Describe each project once in `zconfig.json` — where it lives, what kind it is, where it deploys — and every command just knows. You run the deploy; your agent edits the code. You run the health check; your agent reads the result and fixes whatever's wrong.
|
Token Savers collapses the infrastructure side into short, one-word commands you run yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. Describe each project once in `zconfig.json` — where it lives, what kind it is, where it deploys — and every command just knows. You run the deploy; your agent edits the code. You run the health check; your agent reads the result and fixes whatever's wrong.
|
||||||
|
|
||||||
**Estimated savings: 157,000–540,000 tokens per active development day.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script breakdown and measurement methodology.
|
**Measured: ~26,500 tokens of script output per active development day** — kept out of your agent's context entirely when you run the commands yourself. See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script measurements and method.
|
||||||
|
|
||||||
## Why it's different
|
## Why it's different
|
||||||
|
|
||||||
|
|||||||
@ -8,7 +8,7 @@ Licensed under the MIT License. See LICENSE.
|
|||||||
|
|
||||||
When an AI coding agent orchestrates your infrastructure — starting dev servers, deploying to EC2, diagnosing 502s — it spends hundreds to thousands of tokens per operation on SSH plumbing, Docker output, and retry logic. Those tokens should go to code.
|
When an AI coding agent orchestrates your infrastructure — starting dev servers, deploying to EC2, diagnosing 502s — it spends hundreds to thousands of tokens per operation on SSH plumbing, Docker output, and retry logic. Those tokens should go to code.
|
||||||
|
|
||||||
Token Savers gives you short, one-word commands to run those parts yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. You handle the deterministic infrastructure; your agent handles code. **Running these scripts manually instead of asking your agent to orchestrate them saves an estimated 3,000–7,000 tokens per active development day.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script breakdown.
|
Token Savers gives you short, one-word commands to run those parts yourself: `zdeploy myapp`, `zrepair myapp`, `zstart myapp`. You handle the deterministic infrastructure; your agent handles code. **Running these scripts manually instead of asking your agent to orchestrate them keeps a measured ~26,500 tokens of infrastructure output per active development day out of your agent's context window.** See [TOKEN_SAVINGS.md](TOKEN_SAVINGS.md) for the per-script measurements and method.
|
||||||
|
|
||||||
Every command is a tiny PowerShell script driven by a single JSON config file. The project key you define in that config **is** the command argument — add `myapp` to the config and `zstart myapp`, `zdeploy myapp`, `zbackup myapp` all just work, no script edits needed.
|
Every command is a tiny PowerShell script driven by a single JSON config file. The project key you define in that config **is** the command argument — add `myapp` to the config and `zstart myapp`, `zdeploy myapp`, `zbackup myapp` all just work, no script edits needed.
|
||||||
|
|
||||||
|
|||||||
@ -29,9 +29,12 @@ Dollar equivalents use a blended input/output rate: **Sonnet 5 ≈ $9/1M** | **O
|
|||||||
> **Measurement note:** "Measured" figures come from `token-count.ps1`, which runs
|
> **Measurement note:** "Measured" figures come from `token-count.ps1`, which runs
|
||||||
> each script under `Start-Transcript` and counts output characters ÷ 3.5
|
> each script under `Start-Transcript` and counts output characters ÷ 3.5
|
||||||
> chars/token. Captured in **Claude Code (Sonnet 4.6)** against the `sp` project.
|
> chars/token. Captured in **Claude Code (Sonnet 4.6)** against the `sp` project.
|
||||||
> "Estimated (raw)" figures are *not* measured — they approximate manual
|
> Strictly speaking that's *measured output volume with estimated tokenization*:
|
||||||
> orchestration and are marked *est.* throughout. Other models/interfaces tokenize
|
> ÷3.5 is a prose heuristic, and code-heavy output (paths, JSON, container IDs)
|
||||||
> differently.
|
> fragments into **more** tokens per character under a real BPE tokenizer — so the
|
||||||
|
> token figures here are likely conservative. "Estimated (raw)" figures are *not*
|
||||||
|
> measured — they approximate manual orchestration and are marked *est.*
|
||||||
|
> throughout. Other models/interfaces tokenize differently.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@ -233,22 +236,33 @@ The **~26,500 tokens/day measured** is the honest, reproducible savings from run
|
|||||||
these yourself during an active tool-development day (mostly cached deploys). The
|
these yourself during an active tool-development day (mostly cached deploys). The
|
||||||
**~115k–295k est.** upper figure is what it would cost to have Claude drive the raw
|
**~115k–295k est.** upper figure is what it would cost to have Claude drive the raw
|
||||||
`ssh`/`docker` sequences instead — dominated by per-step reasoning on `zdeploy` and
|
`ssh`/`docker` sequences instead — dominated by per-step reasoning on `zdeploy` and
|
||||||
`zrestart`, not by output volume. A day with several full-rebuild deploys pushes the
|
`zrestart`, not by output volume. Treat that column as an **upper bound, not a
|
||||||
measured figure higher too, since each rebuild streams ~34,600 tokens.
|
prediction**: a capable agent asked to deploy might well write its own wrapper
|
||||||
|
script and ingest very little — the counterfactual depends entirely on how the
|
||||||
|
agent chooses to work. A day with several full-rebuild deploys pushes the measured
|
||||||
|
figure higher too, since each rebuild streams ~34,600 tokens.
|
||||||
|
|
||||||
**Daily dollar savings during active tool development:**
|
**Daily dollar savings during active tool development:**
|
||||||
|
|
||||||
| Model | Blended rate | Measured/day | Est. raw/day (no scripts) |
|
Script output the agent ingests is billed at **input** rates, so the measured column
|
||||||
|-------|-------------|-------------:|--------------------------:|
|
uses input pricing. The est.-raw column keeps the **blended** rate, because raw
|
||||||
| **Sonnet 5** | $9/1M | ~$0.24 | ~$1.04–$2.66 |
|
orchestration also generates agent *output* (reasoning and tool calls between steps).
|
||||||
| **Opus 4.8** | $15/1M | ~$0.40 | ~$1.73–$4.43 |
|
|
||||||
| **Fable 5** | $30/1M | ~$0.80 | ~$3.45–$8.85 |
|
|
||||||
|
|
||||||
Over a ~22-day working month, the measured savings run **~$5–$17/mo** (Sonnet →
|
| Model | Measured/day @ input rate | Est. raw/day @ blended rate |
|
||||||
Fable); the raw-orchestration estimate runs **~$23–$195/mo**. Either way, the
|
|-------|--------------------------:|----------------------------:|
|
||||||
token-budget point stands: every token saved on infrastructure is a token your agent
|
| **Sonnet 5** | ~$0.08 ($3/1M) | ~$1.04–$2.66 ($9/1M) |
|
||||||
keeps for the actual problem — and that context-window quality is worth more than the
|
| **Opus 4.8** | ~$0.13 ($5/1M) | ~$1.73–$4.43 ($15/1M) |
|
||||||
raw dollar figure suggests.
|
| **Fable 5** | ~$0.27 ($10/1M) | ~$3.45–$8.85 ($30/1M) |
|
||||||
|
|
||||||
|
One-time ingest slightly understates the true cost: tokens that enter the context are
|
||||||
|
re-sent on every later turn of the session (at cheaper cache-read rates when prompt
|
||||||
|
caching applies), so the cumulative figure is somewhat higher than a single ingest.
|
||||||
|
|
||||||
|
Over a ~22-day working month, the measured savings run **~$2–$6/mo** (Sonnet →
|
||||||
|
Fable); the raw-orchestration estimate runs **~$23–$195/mo**. The honest dollar
|
||||||
|
figure is small — the real currency is **context**: every infrastructure token kept
|
||||||
|
out of the window is context your agent keeps for the actual problem, and that's
|
||||||
|
worth more than the dollars suggest.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@ -1,169 +0,0 @@
|
|||||||
# Z-Scripts Summary: Manual Automation to Reduce Token Usage
|
|
||||||
|
|
||||||
Running these scripts **manually** instead of asking Claude Code to orchestrate them saves significant token usage because you skip the overhead of Claude reasoning about deployment/build/testing orchestration.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **Local Development Control**
|
|
||||||
|
|
||||||
### `zstart.ps1` — Start dev servers
|
|
||||||
**What:** Starts local dev servers for any project defined in `zconfig.json`
|
|
||||||
```powershell
|
|
||||||
zstart viteapp # Start Vite dev server
|
|
||||||
zstart pyapp -Port 3000 # Start Python app on custom port
|
|
||||||
zstart nextapp -Detached # Start in background
|
|
||||||
```
|
|
||||||
**Why run manually:** Eliminates Claude's need to track startup, wait for health checks, or validate ports. You start the server once and Claude just edits files—tokens saved on orchestration, ~150-300 tokens per use.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zkill.ps1` — Stop dev servers
|
|
||||||
**What:** Kills Node/Python/Docker processes on specified ports
|
|
||||||
```powershell
|
|
||||||
zkill viteapp # Kill Vite dev server
|
|
||||||
zkill pyapp nextapp # Kill multiple projects
|
|
||||||
```
|
|
||||||
**Why run manually:** You control when to stop iteration cycles. Saves Claude from having to reason about process cleanup, ~100-200 tokens.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zrestart.ps1` — Restart in one command
|
|
||||||
**What:** Calls `zkill` + `zstart` atomically; useful after major changes
|
|
||||||
```powershell
|
|
||||||
zrestart viteapp # Kill + restart Vite app
|
|
||||||
```
|
|
||||||
**Why run manually:** Cleaner than telling Claude "stop the server and start it again"—one script handles the sequence. Saves ~200-300 tokens on orchestration logic.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **Build & Deployment**
|
|
||||||
|
|
||||||
### `zdeploy.ps1` — Deploy to EC2
|
|
||||||
**What:** Zips source, SCP to EC2, runs docker compose up, verifies build version
|
|
||||||
```powershell
|
|
||||||
zdeploy pyapp -Note "Fix nav alignment"
|
|
||||||
zdeploy edge # Just reload edge nginx config
|
|
||||||
zdeploy nextapp # Deploy Next.js app + db
|
|
||||||
zdeploy all # Deploy all projects (edge first)
|
|
||||||
```
|
|
||||||
**Why run manually:** Deployment is deterministic once code is ready. You test locally, then run the deploy script—Claude never needs to understand EC2 SSH, zip compression, docker compose, or deployment verification. Saves ~800-1200 tokens that would otherwise go to deployment orchestration.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zstart_docker.ps1` — Ensure Docker daemon is running
|
|
||||||
**What:** Starts Docker Desktop if not running (Windows convenience)
|
|
||||||
**Why run manually:** One-time setup; doesn't need Claude involvement.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **Backup & Sync**
|
|
||||||
|
|
||||||
### `zbackup.ps1` — Backup projects locally
|
|
||||||
**What:** Zips any project defined in `zconfig.json` to the local backups folder with a timestamp
|
|
||||||
```powershell
|
|
||||||
zbackup # Backup everything + this scripts folder
|
|
||||||
zbackup viteapp nextapp # Backup just those two projects
|
|
||||||
zbackup pyapp -Tag "pre-refactor"
|
|
||||||
```
|
|
||||||
**Why run manually:** You decide when to snapshot. Running this yourself before risky changes means Claude never needs to reason about backup strategy or file compression. Saves ~300-500 tokens per session.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zsync.ps1` — Sync backups to OneDrive
|
|
||||||
**What:** Robocopy new files from the local backups folder to OneDrive (incremental)
|
|
||||||
```powershell
|
|
||||||
zsync # Sync all new backup files
|
|
||||||
zsync viteapp # Build + mirror vite dist to $env:ZSYNC_DEST
|
|
||||||
```
|
|
||||||
**Why run manually:** You manage backup cadence independently. Claude doesn't need to reason about incremental sync logic or file enumeration. Saves ~250-400 tokens.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zbackup_ec2.ps1` — Remote backups on EC2
|
|
||||||
**What:** SSH to EC2, tar application and database data, pull to local backups folder
|
|
||||||
**Why run manually:** Separates database/app backup concerns from code changes. Claude focuses on code; you manage infrastructure snapshots.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **Diagnostics & Troubleshooting**
|
|
||||||
|
|
||||||
### `zec2.ps1` — Check EC2 reachability
|
|
||||||
**What:** TCP + HTTP connectivity tests + live build version for all projects with a `domain`
|
|
||||||
```powershell
|
|
||||||
zec2 viteapp # Check if Vite app is up on EC2
|
|
||||||
zec2 # Test all projects
|
|
||||||
```
|
|
||||||
**Why run manually:** When a deploy fails, you run this to verify EC2 is reachable before asking Claude to debug. Eliminates Claude doing network diagnostics blind. Saves ~400-600 tokens of troubleshooting overhead.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zec2online.ps1` — Deep EC2 health check
|
|
||||||
**What:** Full health check; auto-starts downed stacks, streams diagnostics
|
|
||||||
**Why run manually:** Quick check before starting work; Claude doesn't need to validate infrastructure state.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zrepair.ps1` — Audit & repair container routing
|
|
||||||
**What:** SSH to EC2, verify nginx proxy routes, check docker compose health, run smoke tests
|
|
||||||
```powershell
|
|
||||||
zrepair viteapp # Audit proxy path + smoke test
|
|
||||||
zrepair all # Audit all projects
|
|
||||||
```
|
|
||||||
**Why run manually:** When pages 502, you run this first to isolate whether it's routing, DNS, or app logic. Saves ~1000-1500 tokens of "try this, check logs, try that" debugging.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### `zsetup_mail.ps1` — Email account provisioning
|
|
||||||
**What:** Automates creation of email accounts on EC2; displays Route 53 DNS requirements
|
|
||||||
**Why run manually:** One-time setup task; doesn't benefit from Claude guidance.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **Token Usage Impact by Script**
|
|
||||||
|
|
||||||
| Script | Manual Run Saves | Without Script (Claude orchestrates) |
|
|
||||||
|--------|-----------------|--------------------------------------|
|
|
||||||
| **zstart** | ~150-300 tokens | Claude tracks startup, validates ports, polls health |
|
|
||||||
| **zkill** | ~100-200 tokens | Claude enumerates processes, checks exit codes |
|
|
||||||
| **zrestart** | ~200-300 tokens | Claude chains stop→wait→start with error handling |
|
|
||||||
| **zdeploy** | **~800-1200 tokens** | Claude manages zip, SSH, SCP, compose, verification |
|
|
||||||
| **zbackup** | ~300-500 tokens | Claude enumerates, compresses, manages timestamps |
|
|
||||||
| **zsync** | ~250-400 tokens | Claude tracks file diffs, runs robocopy, verifies copy |
|
|
||||||
| **zec2** | ~400-600 tokens | Claude does TCP/HTTP tests, parses output |
|
|
||||||
| **zrepair** | **~1000-1500 tokens** | Claude SSH, grep logs, run smoke tests, interpret failures |
|
|
||||||
|
|
||||||
**Total potential savings per day of active development: 3,000–7,000 tokens** if you run these manually vs. asking Claude to orchestrate.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **Claude Model Token Costs** *(Approximate, May 2026)*
|
|
||||||
|
|
||||||
| Model | Input Cost | Output Cost | Use Case |
|
|
||||||
|-------|-----------|-----------|----------|
|
|
||||||
| **Haiku 4.5** | ~$0.80/1M | ~$4/1M | Quick code edits, small changes |
|
|
||||||
| **Sonnet 4.6** | ~$3/1M | ~$15/1M | Daily coding, medium complexity |
|
|
||||||
| **Opus 4.8** | ~$15/1M | ~$60/1M | Complex reasoning, multi-file refactors |
|
|
||||||
|
|
||||||
**Token savings example:**
|
|
||||||
- Running `zdeploy` manually: ~1000 tokens saved × $15/1M (Sonnet input) = ~$0.015 saved
|
|
||||||
- Running `zrepair` manually: ~1500 tokens saved × $15/1M = ~$0.0225 saved
|
|
||||||
- Running all scripts daily: ~5000 tokens × $0.015 = ~$0.075 saved per day, ~$22.50/month
|
|
||||||
|
|
||||||
**More importantly:** Manual scripts let Claude focus on *code logic* instead of *infrastructure orchestration*—where Claude adds real value.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## **TL;DR**
|
|
||||||
|
|
||||||
Run these scripts manually when:
|
|
||||||
- ✅ You know the exact deployment/test/backup action needed
|
|
||||||
- ✅ The script is deterministic (same input = same result)
|
|
||||||
- ✅ You want to parallelize (run zstart while asking Claude for code)
|
|
||||||
- ✅ You're troubleshooting and need fast feedback loops
|
|
||||||
|
|
||||||
Ask Claude to *invoke* them only when:
|
|
||||||
- ❌ You need complex conditional logic (e.g., "if this test fails, try X")
|
|
||||||
- ❌ You're chaining many operations that depend on each other's output
|
|
||||||
- ❌ You want Claude to interpret script output and decide next steps
|
|
||||||
|
|
||||||
**Bottom line:** Your z-scripts are optimized for **you** to run directly. Use them. Save tokens. Let Claude focus on coding.
|
|
||||||
Loading…
Reference in New Issue
Block a user