Anthropic
Claude
Deep reasoning for architecture, planning, and the Wave 3 readiness gate — the model the pipeline was first built around.

βάθος — Greek for “depth, the abyss.”
v0.4.0 · Rust engine · six model providers, chosen wave by wave
The team is model-agnostic. The same 17 roles, 7-wave pipeline, and hard gates now run on your choice of frontier model — six providers, no rewrite, no lock-in. Assign the model that fits each role and each wave.
Anthropic
Deep reasoning for architecture, planning, and the Wave 3 readiness gate — the model the pipeline was first built around.
OpenAI
Sharp at implementation and refactoring. It runs as a delegated subprocess, so it can share a wave with any other runtime.
Zhipu · Z.ai
The one swapped-in backend with a live-verified connection in this repo. Capable and cost-efficient for high-volume roles.
Moonshot AI
A million-token flagship plus dedicated coding tiers, reached over Moonshot's Anthropic-compatible endpoint.
DeepSeek
Flagship and fast tiers at low cost. The catalog records its quirks too, including the fields its endpoint quietly ignores.
Alibaba Cloud
Flagship, balanced, and flash tiers through Model Studio, whose endpoint is per-workspace rather than one fixed URL.
Pick a model per role, per wave, or leave it alone entirely — the deterministic Rust engine, the gates, and the audit chain stay exactly the same. Live API connections are verified for Claude, Codex, and GLM; Kimi, DeepSeek, and Qwen are wired and documented from their providers' own specs, not yet proven against a live key.
Build the Rust engine for your platform, point BATHOS_BIN at it, then drive the pipeline from Claude Code with your model of choice — any of the six supported providers.
1 — Get the code & build the engine
2 — Link BATHOS into your project
Follows the platform you picked in step 1. Use this repo as your working directory, or adopt βάθος into your own project.
# Option A — use this repo as your working directory
# .claude/ (commands, agents, hooks) is already wired — just open Claude Code here.
# Option B — adopt βάθος into your own project:
./install.sh --into /path/to/your/project # --force overwrites an existing .claude/
export BATHOS_BIN="/path/to/bathos/core/target/release/bathos"3 — Drive the pipeline
Open Claude Code in your project, choose your model with /model-config, and run the wave commands in order.
/team-kickoff
/route /abs/path/to/project
/wave1-discovery /abs/path
/wave2-design /abs/path
/wave3-story-gate /abs/path
/wave5-implement /abs/path
/wave6-verify-report /abs/path
/team-confirmRequires Claude Code v2.1.32+ with Agent Teams enabled — on any of the six supported providers — plus the Rust toolchain (and jq for the bash hooks).
A single long LLM conversation drifts: context is lost between design and build, checks get skipped, and the same model both writes and approves its own work. BATHOS replaces that with structure.
You mostly type slash commands. A single Rust binary is the deterministic core those commands call underneath.
Markdown under .claude/
Slash commands, roles, and hooks that run the waves and spawn, review, and retire teammates, driven by the lead: your main coding session, running on the provider you picked for that wave.
A single static Rust binary
Computes and enforces state, gates, wave transitions, routing, story freshness, and plugins. The hooks and commands invoke it automatically.
Work moves from discovery to a verified release through explicit waves. The mainline is W0 → W1 → W2 → W3 → W5 → W6; IP & research (W4) is an optional plug-in off the critical path.
W0
Caleb
Optional pre-brief: brainstorm, forge the idea, product brief.
W1
John · Caleb
Reverse-engineering and market research; sharpen the USP.
W2
Joshua → James · Jonnathan
Planning gates the wave, then architecture and UX in parallel.
W3
Matthew + Thomas · Matthias
Condense to self-contained story files; the implementation-readiness gate.
W4
Mark · Nathanael
Optional plug-in, off the critical path. Patents and papers.
W5
Phillip · Andrew · Stephen
Backend, frontend, and ML build against the story files.
W6
Thomas · Michael · Hananiah · Martin
Review, security audit, behavior-preserving refactor, and the release report.
Wave 3 · the heart
Wave 3 closes the design-to-build gap: the story engineer condenses upstream work into a self-contained story file where every technical claim is tagged to its source, independent reviewers sign off, and on FAIL a hook physically blocks entry into implementation.
Discovery, architecture, and implementation are not the same job — so they no longer have to run on the same model. Every wave can name its own provider, and the engine holds you to it.
$ bathos model set --wave W2 --runtime claude
$ bathos model set --wave W5 --runtime kimi
$ bathos model set phillip --runtime codex
# a role assignment beats its wave
$ bathos model validate --wave W5
→ exit 2 · session is claude, W5 wants kimi
save → set env → restart → /cold-startDeclare the provider per wave; validate stops the run before a wave starts on the wrong backend.
A usage limit is where most agent runs die — the teammate goes quiet, and everything it was holding in its head goes with it. Here it is a pause. Every handoff already lives on disk, so you re-point the team at a provider that still has budget from the command line, and the wave picks up where it stopped.
1 · Recognize
A teammate that stops producing output is almost never a crash — nine times out of ten it is the account's session limit, and its own transcript says so. Don't kill it. bathos model show prints the live session backend in its header, so you can see which provider ran dry and which roles were standing on it.
2 · Re-point
bathos model set --wave W5 --runtime glm --model glm-5.3 writes that wave's new provider into model-plan.json; naming a role instead overrides only that role, and bathos model unset drops it back to the fallback. Just the entry you changed is written — the rest of the plan is left alone.
3 · Resume
GLM, Kimi, DeepSeek and Qwen are reached by swapping ANTHROPIC_BASE_URL, and that variable is process-global — so the switch costs a save, an env change and a fresh session, never a silent hot-swap behind your back. Codex, a delegated subprocess, costs neither. /cold-start restores the session and the wave re-runs from the artifacts on disk.
# 1 — which backend am I on, and what just ran dry?
$ bathos model detect # ANTHROPIC_BASE_URL → claude|glm|kimi|deepseek|qwen
$ bathos model show # effective runtime/model per role, and its source
# 2 — re-point the wave (or one role) at a provider that still has budget
$ bathos model set --wave W5 --runtime glm --model glm-5.3
$ bathos model set stephen-ml-engineer --runtime codex
# a role assignment beats its wave · unset returns it to the fallback
$ bathos model validate --wave W5
→ exit 2 · session is claude, W5 wants glm
1) take the whole batch to glm 2) move the role to claude / codex
3) split the wave: finish this batch, shut down, restart on glm
# 3 — env-swap runtimes only: save, swap, restart, resume
/save
$ export ANTHROPIC_BASE_URL=https://api.z.ai/api/anthropic
$ export ANTHROPIC_AUTH_TOKEN=… # DeepSeek reads ANTHROPIC_API_KEY instead
# restart Claude Code → /cold-start → re-run /wave5-implementvalidate refuses to start a wave on the wrong backend, and prints the three ways out instead of a stack trace.
Changing provider changes nothing else: the same 17 roles, the same PASS / CONCERNS / FAIL vocabulary, the same audit chain, the same wave order. Model IDs are free-form rather than an allowlist, so a cheaper or older model on the same account works too — and stepping down a tier is often the fastest way past a limit, no second vendor required.
BATHOS activates only the waves a task needs. The router recommends a level from four stakes axes; you confirm it.
| Level | Work type | Active waves |
|---|---|---|
| Lv0 | Bug fix / trivial | W5 (+ minimal W6) |
| Lv1 | Small feature / local refactor | Light W2 + W3 (slim) + W5 + light W6 |
| Lv2 | Standard feature / module | W1 + W2 + W3 + W5 + W6 |
| Lv3 | New product / large | W0–W6 (W4 optional) |
| Lv4 | Enterprise / deep-tech / regulated | Full W0–W6 + W4 |
The lead is your main session and is never spawned. Specialists are spawned per wave, with concurrency capped at three.
Every wave gate speaks one vocabulary, and the critical one is enforced in code. No evidence-free auto-pass.
PASS
Criteria met, no blockers
Proceed to the next wave
CONCERNS
Conditional pass, non-blocking risk
Log the risk, then proceed
FAIL
Blocking defect
Entry blocked; remediate and re-gate
Every state change is written to a keyed audit hash-chain (HMAC-SHA256). A single command, bathos audit verify, proves the trail has not been altered.
Left to itself, an AI answers every task with new code — a new abstraction, a new dependency, one more file. Wave 5 implementers carry an explicit discipline instead: before writing anything they climb a seven-rung ladder and stop at the first rung that holds. Lazy about the solution, never about the reading.
Speculative need is skipped, and said out loud in one line. YAGNI.
A helper, type, or pattern that already lives here gets reused. Re-implementing what sits a few files over is the most common slop.
Then the standard library does it.
A date input over a picker library, CSS over JS, a database constraint over application code.
Use it. Never add a new dependency for what a few lines can do.
Then it is one line.
The shortest working diff wins — but only once you actually understand what the change has to touch.
Never simplified away
The ladder governs what gets built. How completely the settled scope gets built is a different axis, and it belongs to Boil the Ocean. Neither is ever traded for the other.
// ponytail: single global lock, split per-wave
// if profiling shows contention
# ponytail: fixed backoff, go exponential once
# the API starts rate-limiting
$ /bathos-debt # CONCERNS: docs + ponytail: src
→ 2 open · 1 no-triggerEvery deliberate shortcut leaves its ceiling and its upgrade trigger in the code, and /bathos-debt collects them — with the markers that name no trigger flagged, because those are the ones that rot in silence.
One switch sets how hard it bites: /bathos intensity lite · full · ultra · off.
Engineering principles adapted from ponytail by Dietrich Gebert (MIT). The persona and branding are not adopted — BATHOS is an orchestration product, not a character.
The critical invariants (state, gates, routing, story freshness) live in a single static Rust binary, so they are computed and enforced the same way every time.
$ echo '{"scope":"feature","novelty":true}' \
| bathos --state-dir _state route decide
→ {"recommended_level":2,"requires_confirmation":true}Recommend a level from the stakes. The engine proposes, you confirm.
Each work session ends with a self-contained HTML report — written automatically when the session closes, or on demand with /taskreport. It is assembled from disk and git facts only; nothing is invented.
result_report/
├─ task_report_20260716_104150_session_no12.html
└─ … # one file per session, auto-numbered
# on /exit → SessionEnd hook · or run /taskreportOne report per session, numbered in sequence — the paper trail of the whole build.
Assign a model and runtime to each role before a wave, then follow the run in a live panel — without ever taking the keyboard from the lead.
$ bathos model set james --runtime claude --model opus
$ bathos model set phillip --runtime codex
$ bathos model validate --wave W5 # mixed-batch guard
→ PASS
$ bathos panes --mode tui --wave W5 # live · read-onlyAssign models before the wave, then watch it run — read-only, with an inbox for your proposals.
Read the source, or explore more on the web.