---
name: bworlds-cli
description: Operate Builds with the public bworlds CLI from a local code agent or automated runner, including safe context checks, Audits, Findings, Dossier, telemetry, and repository work.
---

# BWorlds CLI

Use the installed `bworlds` binary and the public documentation at `https://docs.bworlds.co/llms.txt`. Read the relevant linked page before an unfamiliar operation; use `https://docs.bworlds.co/command-reference.json` for the exact versioned flags. Do not assume access to BWorlds source code or internal scripts.

An operator is a Builder acting through the CLI, directly or through an agent. It is not a separate role. An active Workspace Membership as Owner or Member grants the same operational access to every Build that Workspace owns.

Before operating on a Build, verify the personal identity, Workspace, and Build:

```console
bworlds --version
bworlds auth status --json
bworlds build info BUILD_SLUG --json
```

For investigation patterns, read `https://docs.bworlds.co/docs/agent-workflows`. Start with current server state through `audit show`, `findings list`, and the telemetry commands before proposing a code change.

`auth login` uses Device Approval. Relay its one-time code and approval URL when the current machine has no browser. Any Builder already signed in to the web app can approve it; approval requires that user session, not a CLI credential. The resulting Operator Session is stored by the CLI. Use `bworlds auth token` when another local process needs the active credential, treat its output as a secret, and never read the credential file.

Use `BWORLDS_TOKEN` for unattended execution. Never print it, put it in a repository file, or enable shell or Git credential tracing around it. A token represents its creator; the server reloads that Builder's current Memberships on every request and creates no agent role.

Prefer `--json` for agent consumption. Read successful data from standard output and diagnostics from standard error. Interpret documented exit codes before deciding to retry. A `not_found` result can mean the identifier is outside accessible Workspaces; do not probe other identifiers.

Every JSON result carries `schemaVersion` and names its fields in camelCase at every depth. Select `.startedAt`, not `.started_at`: the snake_case name ships beside it until the next `schemaVersion` removes it. The keys inside a `metadata` or `payload` value are data, not fields, and never change.

Treat `audit run`, `audit reevaluate`, and `control run` as potentially paid mutations. Inspect the target Build first, obtain the user's authorization for the cost, pass `--confirm`, and retain the request ID. Reuse that ID after a lost response. The server determines authorization, billing, and attribution from the authenticated identity. Follow a detached run with `bworlds audit status AUDIT_RUN_ID --json`.

Use `bworlds control run BUILD_SLUG CONTROL_ID --confirm --json` to check one Control of any available Audit after a targeted fix, without rerunning its Audit. It waits for the final ControlEvaluation, up to `--timeout` (default 10m), and prints it as JSON. The Workspace pays in tokens only when BWorlds settles the evaluation on `pass` or `fail`: 20 for a deterministic Control, 100 for a live agentic Control, 150 for a repository agentic Control. A question Control cannot run this way; only the Builder answers it. `audit reevaluate` follows its evaluation the same way. After a timeout or a pending Builder answer, read the evaluation with `bworlds control status BUILD_SLUG EVALUATION_ID --json`, which never charges and exits 0 whatever the state (read `executionStatus`), or rerun with the same `--request-id` to wait again without a second check. Treat a result with `evidenceBasis: declared` as the Builder's statement, not verified evidence.

Require explicit authorization immediately before repository push, override changes, Dossier push, Finding mutations, and credential creation or revocation. Inspect current server state before each mutation, then pass `--confirm`. This flag is a CLI safety check, not a server authorization boundary; the server still decides access and attribution.

Repository credentials are short-lived and repository-scoped; use `bworlds repo` commands instead of copying a credential into Git configuration. If refresh or push reports an origin mismatch, re-clone through the CLI instead of changing the remote to bypass the check. Only a new commit created by `repo push` receives BWorlds actor trailers. Retrying an already-created commit preserves its existing metadata, and direct Git use is outside the CLI's attribution guarantees.

For a live incident, read the same period across each source:

```console
bworlds telemetry uptime BUILD_SLUG --json
bworlds telemetry errors BUILD_SLUG --window 7 --json
bworlds telemetry sessions BUILD_SLUG --has-errors --json
bworlds telemetry sessions BUILD_SLUG SESSION_ID --failed-only --json
bworlds build notifications BUILD_SLUG --window 7 --json
```

A session detail returns a summary: identity, pages, captured requests, signals, and recorded errors and console lines. Read it first. Use `--depth full` only when the summary lacks what the investigation needs; it returns the whole server record, including every recorded interaction and the signed replay media links.

Telemetry withholds the email address of the Builder's end users and reports `"userEmailWithheld": true` in its place. Do not pass `--include-personal-data` to put that address into your own context. Ask for it only when the user needs to contact that person, and keep it out of Findings, commits, and messages.

`build notifications` reads only notifications associated with the named Build. Use `findings edit ... --confirm` to change content and `findings history` to verify attributed activity. Both Finding commands take the Build slug and Finding ID; never reuse a Finding ID from another Build.

A Finding's origin decides whether the Builder's Findings list shows it, and `findings edit` never changes that origin. A Finding's own link opens whatever its origin, so an open link proves nothing about that list. Before telling a Builder where to find a Finding, read its `origin` in `findings list --json` and compare it with the list rules in `https://docs.bworlds.co/docs/concepts/product-model`. When the list may hide it, put the fix instructions in the message itself.

## Writing Findings

A Finding is how work becomes visible in the app instead of dying in a report. Any skill that produces Findings follows the steps below; the command, its flags and its output are in `https://docs.bworlds.co/docs/guides/production` and `https://docs.bworlds.co/docs/guides/shared-memory`.

### 1. Inventory what already exists

```bash
bworlds findings list <slug> --json
```

Skip creating a Finding when an `open`, `needs_user` or `in_progress` Finding already tracks the risk, whichever produced it. Match on the `--source-ref` risk slug, or on title plus area. Use `bworlds findings update-status <slug> <finding-id> <status> --reason "..." --confirm` when its lifecycle changed.

When a run sharpens a tracked risk with a diagnosed cause, a concrete prescription, a different severity or a different service line, edit that Finding instead of creating a second one:

```bash
bworlds findings history <slug> <finding-id>
bworlds findings edit <slug> <finding-id> --rationale "..." --severity high --confirm
```

The Finding keeps its id, its status and its Builder-facing link, so the Builder keeps one item. Content ownership moves to whoever edited it: later Audit projections refresh its evidence and status without rewriting those words.

### 2. Propose one Finding per outcome

| Outcome | Action |
|---------|--------|
| Open risk (critical, high or medium) | `findings create ... --status open` |
| Risk needing a decision, input or a deploy from the Builder | `findings create ... --status needs_user` |
| Real concern verified handled, including fixed during the run | `findings create ... --status resolved` |
| Investigated and ruled out | `findings create ... --status dismissed` |
| Action considered and deliberately declined | `findings create ... --status dismissed --disposition restraint` |
| Tracked risk this run sharpened | `findings edit <finding-id> ...` |

- **One Finding per architectural theme.** Instances, whether endpoints, files or queries, are evidence inside the Finding, never separate Findings.
- **Concrete prescription.** The description states the move, for example "precompute look status with a trigger". If you cannot state the prescription, the diagnosis is not finished: investigate further or drop it.

### 3. Tag every Finding

- `--severity`: `critical`, `high` or `medium`.
- `--area`: the area the risk belongs to.
- `--service-line`: `launch` blocks safe sharing or selling, `run` is production reliability and incidents, `improve` is operational drag and code health.
- `--source-ref {producer}:{YYYY-MM-DD}:{risk-slug}`, where `{producer}` names the skill that wrote it. Lowercase kebab plus colons, 100 characters or fewer, unique per Finding. Never share one ref across two Findings of the same run.

### 4. Write for the Builder

The Builder reads these. Plain language, framed against their milestone, their critical workflow or their growth plans. Inline `file:line` evidence. 2000 characters or fewer. No Control ids, no raw table names, no internal jargon.

Never put a secret value in a title, a description or metadata. Name the location and the key type only, for example "a live Stripe secret key is hardcoded at src/config.js:12", and never the value, not even partially.

### 5. Review gate

Present the whole proposal as one table before running any `findings create`, `findings edit` or `update-status`:

| Action | Title | Status (+disposition) | Severity | Area | Service line | Source ref |
|--------|-------|----------------------|----------|------|--------------|------------|

Include the planned edits and status changes, naming the fields each edit rewrites, and the planned skips with their reasons. Then offer three choices:

- **Write all N Findings**, when the plan mirrors the report you presented.
- **Adjust first**: drop Findings or change status, severity or wording, then re-present the amended plan.
- **Skip writes this run**: print the summary line with `0 created` and stop.

Execute only what was approved.

### 6. A failed write never blocks the run

- stderr says `already tracked` (a 409 duplicate) → print `Already tracked: "<title>" — skipped` and count it as skipped. The risk is already written.
- any other non-zero exit → print `WARNING: findings write failed for "<title>" — continuing` and count it as failed.

Move to the next Finding either way. Never abort the rest of the run because a write failed.

If two consecutive writes fail with the same authentication, unknown-flag or connection error, stop, apply the remedy the error names, then resume. If the remedy fails, count every remaining Finding as failed.

If the step 1 inventory itself fails, print the same warning and continue with the creates. The 409 backstop only dedupes within a run, because refs embed the run date, so cross-run duplicates are possible until the inventory call works again.

Close with one line: `Findings written: N created, N edited, N updated, N skipped, N failed`. Print it even when there was nothing to write, because that line is how anyone knows this step ran.

Finding writes cost no tokens. Remedies for a missing binary, a rejected flag or a refused credential are in `https://docs.bworlds.co/docs/reference/troubleshooting`.
