BATHOS home

βάθος — Greek for “depth, the abyss.”

Give your Multi Modal AI the depth of a full product team.

BATHOS turns a single AI coding session — on Claude, Codex, GLM, Kimi, DeepSeek, or Qwen — into a disciplined team:17 specialist roles across a 7-wave pipeline, a model chosen wave by wave, scale-adaptive routing, and hard quality gates, all backed by a deterministic Rust engine.

v0.4.0 · Rust engine · six model providers, chosen wave by wave

Multi Modal based Agent Team as a Full Product Team

The team is model-agnostic. The same 17 roles, 7-wave pipeline, and hard gates now run on your choice of frontier model — six providers, no rewrite, no lock-in. Assign the model that fits each role and each wave.

Anthropic

Claude

Deep reasoning for architecture, planning, and the Wave 3 readiness gate — the model the pipeline was first built around.

OpenAI

Codex

Sharp at implementation and refactoring. It runs as a delegated subprocess, so it can share a wave with any other runtime.

Zhipu · Z.ai

GLM

The one swapped-in backend with a live-verified connection in this repo. Capable and cost-efficient for high-volume roles.

Moonshot AI

Kimi

A million-token flagship plus dedicated coding tiers, reached over Moonshot's Anthropic-compatible endpoint.

DeepSeek

DeepSeek

Flagship and fast tiers at low cost. The catalog records its quirks too, including the fields its endpoint quietly ignores.

Alibaba Cloud

Qwen

Flagship, balanced, and flash tiers through Model Studio, whose endpoint is per-workspace rather than one fixed URL.

Pick a model per role, per wave, or leave it alone entirely — the deterministic Rust engine, the gates, and the audit chain stay exactly the same. Live API connections are verified for Claude, Codex, and GLM; Kimi, DeepSeek, and Qwen are wired and documented from their providers' own specs, not yet proven against a live key.

Install

Build the Rust engine for your platform, point BATHOS_BIN at it, then drive the pipeline from Claude Code with your model of choice — any of the six supported providers.

1 — Get the code & build the engine

macOS / Linuxbash

POSIX shell with bash hooks. Build the static binary and export BATHOS_BIN.

git clone <your-fork-url> bathos && cd bathos

# Build the single static engine binary (~5.9 MB)
cd core
cargo build --release        # → core/target/release/bathos
cargo test                   # 510 tests, all green (optional)
cd ..

# Make the engine discoverable by hooks/commands:
export BATHOS_BIN="$(pwd)/core/target/release/bathos"

3 — Drive the pipeline

Run the waves on the provider you picked

Open Claude Code in your project, choose your model with /model-config, and run the wave commands in order.

/team-kickoff
/route              /abs/path/to/project
/wave1-discovery    /abs/path
/wave2-design       /abs/path
/wave3-story-gate   /abs/path
/wave5-implement    /abs/path
/wave6-verify-report /abs/path
/team-confirm

Requires Claude Code v2.1.32+ with Agent Teams enabled — on any of the six supported providers — plus the Rust toolchain (and jq for the bash hooks).

Structure where a chat has none

A single long LLM conversation drifts: context is lost between design and build, checks get skipped, and the same model both writes and approves its own work. BATHOS replaces that with structure.

Plain LLM chat
BATHOS
One conversation, growing context drift
17 specialist roles across a 7-wave pipeline
Context lost between design and implementation
Zero-context-loss, self-contained story files (Wave 3)
Implicit, one-size-fits-all effort
A scale-adaptive router with explicit Lv0–4
Implementation starts whenever
A hard readiness gate physically blocks the build on FAIL
The author also “verifies” its own work
Independent reviewers and a tamper-evident audit chain
Advice that quietly overrides you
User Sovereignty: the AI proposes, you decide

Two planes, one runtime

You mostly type slash commands. A single Rust binary is the deterministic core those commands call underneath.

Markdown under .claude/

Orchestration

Slash commands, roles, and hooks that run the waves and spawn, review, and retire teammates, driven by the lead: your main coding session, running on the provider you picked for that wave.

A single static Rust binary

Engine

Computes and enforces state, gates, wave transitions, routing, story freshness, and plugins. The hooks and commands invoke it automatically.

A seven-wave delivery pipeline

Work moves from discovery to a verified release through explicit waves. The mainline is W0 → W1 → W2 → W3 → W5 → W6; IP & research (W4) is an optional plug-in off the critical path.

W0

Analysis

Caleb

Optional pre-brief: brainstorm, forge the idea, product brief.

W1

Discovery & market

John · Caleb

Reverse-engineering and market research; sharpen the USP.

W2

Plan · architecture · design

Joshua → James · Jonnathan

Planning gates the wave, then architecture and UX in parallel.

W3

Story engineering

Matthew + Thomas · Matthias

Condense to self-contained story files; the implementation-readiness gate.

W4

IP & research

Mark · Nathanael

Optional plug-in, off the critical path. Patents and papers.

W5

Implementation

Phillip · Andrew · Stephen

Backend, frontend, and ML build against the story files.

W6

Verify · docs · report

Thomas · Michael · Hananiah · Martin

Review, security audit, behavior-preserving refactor, and the release report.

Wave 3 · the heart

Wave 3 closes the design-to-build gap: the story engineer condenses upstream work into a self-contained story file where every technical claim is tagged to its source, independent reviewers sign off, and on FAIL a hook physically blocks entry into implementation.

Pick the model wave by wave

Discovery, architecture, and implementation are not the same job — so they no longer have to run on the same model. Every wave can name its own provider, and the engine holds you to it.

  • Claudenative
  • Codexsubprocess
  • GLMenv-swap
  • Kimienv-swap
  • DeepSeekenv-swap
  • Qwenenv-swap
A plan per wave, and per role
bathos model set --wave W5 --runtime kimi records the wave's choice in model-plan.json. A role assignment beats its wave, and resolution falls through five steps — role → wave → defaults → agent frontmatter → runtime default — with every value showing which step it came from.
Six providers, one interface
Claude runs natively, Codex as a delegated subprocess, and GLM, Kimi, DeepSeek and Qwen over their Anthropic-compatible endpoints. A single predicate decides which runtimes can share a batch, so adding a seventh provider means teaching that one predicate — not rewriting the rules.
A gate, not a promise
Environment variables are process-global, so two swapped-in providers cannot run in one session. bathos model validate --wave stops that with exit 2 and prints the switch procedure — instead of pretending the models hot-swap underneath you.
The catalog is a menu, not a fence
Thirty catalogued models across the six providers, each marked for whether it was confirmed at the provider's own docs, with retired IDs kept separate. Model IDs stay free-form, so an older or cheaper model still works — and a new release is a JSON edit, not an engine rebuild.
$ bathos model set --wave W2 --runtime claude
$ bathos model set --wave W5 --runtime kimi
$ bathos model set phillip --runtime codex
    # a role assignment beats its wave

$ bathos model validate --wave W5
  → exit 2 · session is claude, W5 wants kimi
    save → set env → restart → /cold-start

Declare the provider per wave; validate stops the run before a wave starts on the wrong backend.

Out of tokens? Change the model, not the plan.

A usage limit is where most agent runs die — the teammate goes quiet, and everything it was holding in its head goes with it. Here it is a pause. Every handoff already lives on disk, so you re-point the team at a provider that still has budget from the command line, and the wave picks up where it stopped.

1 · Recognize

A silent teammate is usually quota

A teammate that stops producing output is almost never a crash — nine times out of ten it is the account's session limit, and its own transcript says so. Don't kill it. bathos model show prints the live session backend in its header, so you can see which provider ran dry and which roles were standing on it.

2 · Re-point

One command, per wave or per role

bathos model set --wave W5 --runtime glm --model glm-5.3 writes that wave's new provider into model-plan.json; naming a role instead overrides only that role, and bathos model unset drops it back to the fallback. Just the entry you changed is written — the rest of the plan is left alone.

3 · Resume

Restart once, lose nothing

GLM, Kimi, DeepSeek and Qwen are reached by swapping ANTHROPIC_BASE_URL, and that variable is process-global — so the switch costs a save, an env change and a fresh session, never a silent hot-swap behind your back. Codex, a delegated subprocess, costs neither. /cold-start restores the session and the wave re-runs from the artifacts on disk.

# 1 — which backend am I on, and what just ran dry?
$ bathos model detect        # ANTHROPIC_BASE_URL → claude|glm|kimi|deepseek|qwen
$ bathos model show          # effective runtime/model per role, and its source

# 2 — re-point the wave (or one role) at a provider that still has budget
$ bathos model set --wave W5 --runtime glm --model glm-5.3
$ bathos model set stephen-ml-engineer --runtime codex
    # a role assignment beats its wave · unset returns it to the fallback

$ bathos model validate --wave W5
  → exit 2 · session is claude, W5 wants glm
    1) take the whole batch to glm    2) move the role to claude / codex
    3) split the wave: finish this batch, shut down, restart on glm

# 3 — env-swap runtimes only: save, swap, restart, resume
/save
$ export ANTHROPIC_BASE_URL=https://api.z.ai/api/anthropic
$ export ANTHROPIC_AUTH_TOKEN=…      # DeepSeek reads ANTHROPIC_API_KEY instead
    # restart Claude Code → /cold-start → re-run /wave5-implement

validate refuses to start a wave on the wrong backend, and prints the three ways out instead of a stack trace.

Changing provider changes nothing else: the same 17 roles, the same PASS / CONCERNS / FAIL vocabulary, the same audit chain, the same wave order. Model IDs are free-form rather than an allowlist, so a cheaper or older model on the same account works too — and stepping down a tier is often the fastest way past a limit, no second vendor required.

Scale-adaptive routing

BATHOS activates only the waves a task needs. The router recommends a level from four stakes axes; you confirm it.

LevelWork typeActive waves
Lv0Bug fix / trivialW5 (+ minimal W6)
Lv1Small feature / local refactorLight W2 + W3 (slim) + W5 + light W6
Lv2Standard feature / moduleW1 + W2 + W3 + W5 + W6
Lv3New product / largeW0–W6 (W4 optional)
Lv4Enterprise / deep-tech / regulatedFull W0–W6 + W4

Seventeen specialists, one lead

The lead is your main session and is never spawned. Specialists are spawned per wave, with concurrency capped at three.

00PaulLead / final confirmall
01JohnReverse specialistW1
02CalebMarket analysis / USPW1
03JoshuaService planningW2
04JamesSW / cloud architectW2
05MarkIP specialist (patents)W4
06NathanaelResearch writerW4
07JonnathanChief designer (UX/UI)W2
08PhillipBackend & data leadW5
09AndrewFrontend & mobile leadW5
10StephenAI / ML leadW5
11TimothyDev-definition docsW6
12ThomasCode reviewerW6
13MichaelSecurity specialistW6
14HananiahRefactoring specialistW6
15MatthiasQA / validationW6
16MartinMonitoring / reportW6
17MatthewScrum master / story engineerW3

Gates that actually gate

Every wave gate speaks one vocabulary, and the critical one is enforced in code. No evidence-free auto-pass.

PASS

Criteria met, no blockers

Proceed to the next wave

CONCERNS

Conditional pass, non-blocking risk

Log the risk, then proceed

FAIL

Blocking defect

Entry blocked; remediate and re-gate

Tamper-evident by design

Every state change is written to a keyed audit hash-chain (HMAC-SHA256). A single command, bathos audit verify, proves the trail has not been altered.

It builds like a well-trained engineer, not like a demo

Left to itself, an AI answers every task with new code — a new abstraction, a new dependency, one more file. Wave 5 implementers carry an explicit discipline instead: before writing anything they climb a seven-rung ladder and stop at the first rung that holds. Lazy about the solution, never about the reading.

  1. Does this need to exist?

    Speculative need is skipped, and said out loud in one line. YAGNI.

  2. Is it already in this codebase?

    A helper, type, or pattern that already lives here gets reused. Re-implementing what sits a few files over is the most common slop.

  3. Does the standard library do it?

    Then the standard library does it.

  4. Is it native to the platform?

    A date input over a picker library, CSS over JS, a database constraint over application code.

  5. Does an installed dependency cover it?

    Use it. Never add a new dependency for what a few lines can do.

  6. Can it be one line?

    Then it is one line.

  7. Only then: the minimum that works.

    The shortest working diff wins — but only once you actually understand what the change has to touch.

Never simplified away

  • Input validation at trust boundaries
  • Error handling that prevents data loss
  • Security measures
  • Accessibility basics
  • Anything you were explicitly asked for
  • Understanding the problem — the ladder shortens the solution, never the reading

The ladder governs what gets built. How completely the settled scope gets built is a different axis, and it belongs to Boil the Ocean. Neither is ever traded for the other.

// ponytail: single global lock, split per-wave
//   if profiling shows contention
# ponytail: fixed backoff, go exponential once
#   the API starts rate-limiting

$ /bathos-debt    # CONCERNS: docs + ponytail: src
  → 2 open · 1 no-trigger

Every deliberate shortcut leaves its ceiling and its upgrade trigger in the code, and /bathos-debt collects them — with the markers that name no trigger flagged, because those are the ones that rot in silence.

One switch sets how hard it bites: /bathos intensity lite · full · ultra · off.

Engineering principles adapted from ponytail by Dietrich Gebert (MIT). The persona and branding are not adopted — BATHOS is an orchestration product, not a character.

A deterministic core, not vibes

The critical invariants (state, gates, routing, story freshness) live in a single static Rust binary, so they are computed and enforced the same way every time.

Single static binary
One compact bathos executable, no runtime dependencies.
Schema-validated SSOT
manifest.json is the source of truth: schema-validated, atomic writes.
Tamper-evident audit
A keyed HMAC hash-chain with a one-command verify.
Proven
510 Rust tests + 86 hook determinism checks, all green; clippy clean.
$ echo '{"scope":"feature","novelty":true}' \
    | bathos --state-dir _state route decide

→ {"recommended_level":2,"requires_confirmation":true}

Recommend a level from the stakes. The engine proposes, you confirm.

A report for every session

Each work session ends with a self-contained HTML report — written automatically when the session closes, or on demand with /taskreport. It is assembled from disk and git facts only; nothing is invented.

Auto-generated
A SessionEnd hook writes the report on exit; /taskreport produces the same file on demand, mid-session.
Six fixed sections
Start, end, and total duration; the session's key work; notable issues; and the full git commit / push / PR / merge history.
Facts, not fabrication
The narrative comes from session state, the git trail from the repository. Anything unrecorded is marked “(not recorded).”
Self-contained & numbered
One HTML file with no external assets, named by timestamp and an auto-incrementing session number.
result_report/
├─ task_report_20260716_104150_session_no12.html
└─ …                     # one file per session, auto-numbered

# on /exit → SessionEnd hook   ·   or run /taskreport

One report per session, numbered in sequence — the paper trail of the whole build.

Choose per role, watch it live

Assign a model and runtime to each role before a wave, then follow the run in a live panel — without ever taking the keyboard from the lead.

Per-role model plan
bathos model set / show / validate records a role → runtime/model plan in model-plan.json, the source of truth every wave resolves against. No plan means each role falls back to its own default, so a project that never touches this is unchanged.
Six runtimes, one plan
Assign Claude, Codex, GLM, Kimi, DeepSeek, or Qwen across roles; a mixed-batch guard blocks incompatible combinations before anyone is spawned.
Live wave panes
bathos panes opens a read-only tmux or built-in TUI panel over the same inspect data — control, status, and an inbox, side by side.
Proposals, not overrides
Panel confirm and feedback land as files in an inbox; the final gate still belongs to the lead. User Sovereignty holds.
$ bathos model set james   --runtime claude --model opus
$ bathos model set phillip --runtime codex
$ bathos model validate --wave W5      # mixed-batch guard
  → PASS

$ bathos panes --mode tui --wave W5    # live · read-only

Assign models before the wave, then watch it run — read-only, with an inbox for your proposals.

Bring depth to your own build.

Read the source, or explore more on the web.