Skip to content

Run a swarm

A swarm starts from the Swarms tab or one tool call and returns at once. You watch it in Keelson or ClickClack and collect the result when it ends.

The task is what the lead receives and what every agent sees in its system prompt. A good task names the question, the scope, and what a finished answer looks like:

Find why `bun test` on main takes four minutes when it took forty seconds a
month ago. Report the slowest suites with timings, the commit range where the
regression landed, and the most likely cause. Do not propose a fix.

The task is capped at 8,000 characters. Anything agents must read that is not in the checkout goes in task context, not in the task.

In the Swarms tab, type the task into Start a swarm, then press Start swarm. No project · chat only is selected by default, so agents read nothing on disk. Pick a project to grant read access; each option shows its name and path, with the home directory shortened to ~.

Choose a plan under Size. Scout is selected by default. Each card shows its agents, cost ceiling, minutes and the effective provider’s models, and, once that plan has run here, its median and highest cost. Providers without pins use the matching class model.

Plan For Agents Turns Minutes
Scout A narrow question, or a first pass before a bigger run. 3 20 15
Crew Most tasks: investigate, debate, and decide. 5 40 30
Fleet Wide or hard problems that are worth the spend. 8 80 60

Lead model and Workers model sit under the cards, beside Project, and start on the plan’s lead and the plan’s workers. Provider groups contain each provider’s default model, class models and pinned models, without duplicates within a group. Other… on Lead model accepts a model name and uses the effective default provider. Each picker replaces its own row on the selected card and keeps the card selected; picking only a lead keeps the plan’s workers, and picking only workers keeps the plan’s lead. Lead and workers must come from one provider. Picking a card clears both picks and restores that plan.

Start swarm sits at the end of the form with one sentence beside it: N agents for up to N min or $N, then the picked models, then where they work, such as 3 agents for up to 15 min or $1, chat only. Untouched Scout sends no size, power or model overrides. Crew records medium/balanced; Fleet records large/deep. A named model records size, model and provider, with no power.

New project… appears last in Project only when the host exposes optional createProject, even with no registered projects. Older hosts omit it and refuse crafted creation requests. Enter a required Name and an optional Folder. The placeholder ~/keelson/<name> is illustrative: a blank Folder uses the host’s workspace root plus the name, not a universal home-directory path. The host expands a leading ~; the rib passes it unchanged.

Start asks the host to create and register the project before admitting the swarm. The host initializes a missing or empty folder with git and a first empty “Initialize project” commit. An existing git repository or a nonempty non-git folder is registered untouched; the rib does not repair it. Missing git identity can cause a host initialization error. Host refusal messages appear unchanged in a toast, with no swarm started. If creation succeeds but swarm admission fails, the registered project remains available for retry.

For New project…, Write is on and locked on (checked and disabled). The swarm uses the returned registered project ID and starts with write access. Writers can work locally without origin after the project has a branch and a first commit. Each writer uses a branch-isolated worktree. An existing repository supplied as Folder follows the engine’s remote or local write rules. The sentence ends creating <name>, followed by , then <workflows> only when present, then , with beads when tracker intent is on.

Use the tracker defaults on when beads_init is reachable, even with no reachable lead tools. You can switch it off. After creation, the rib rechecks initialization reachability and awaits callTool("beads", "beads_init", { project: created.name }) before admitting the write swarm. Only successful initialization enables the requested, currently reachable tracker lead tools, rechecked after initialization. The launcher-only beads_init call is not a lead tool. Tracker off, unreachable initialization or a missing reachability hook skips initialization and starts without tracker tools. Without reachable initialization, the switch is off and disabled with “no tracker yet in a new project”; the row is hidden when no tracker lead tool is reachable either. A failed initialization (including a missing cross-rib caller) still starts a write swarm without tracker tools and reports beads_init with the original error in a toast. An explicit reachability-probe error or admission refusal still refuses Start. There is no automatic retry or project rollback.

Grant initialization and the six lead tools in config.json:

{
"crossRibGrants": {
"swarm": {
"beads": [
"beads_init",
"beads_ready",
"beads_show",
"beads_create",
"beads_update",
"beads_close",
"beads_dep"
]
}
}
}

The switch and initialization create no grants. The rib never runs bd init itself; the host-owned beads tool initializes and refreshes the tracker. Existing-project starts, Retry and Go deeper do not initialize a tracker.

With no project selected, the ALSO ALLOW group is absent. Selecting an existing project reveals Write, Run workflows and Use the tracker, all off.

Write permits code changes. The lead spawns writers, and only those writers receive their own worktrees and branches. Turning it off restores read access.

Run workflows shows removable workflow chips. Enter or comma adds names; paste whitespace/comma-separated names to add a batch. Duplicates are ignored. Remove a chip with its remove button. The limit is 10 distinct workflows. Valid pending text is added on Start; invalid or over-limit text stays in the input and blocks Start. Names become { name, isolated: true } grants and still need the operator’s ribWorkflowGrants. A host without workflow dispatch support disables the switch and explains why. Answer approvals in Workflows when the host has not granted automatic responses; remembered approval refusals append a note to that row. The separate ribApprovalGrants policy still applies to lead responses.

Use the tracker lists beads_ready, beads_show, beads_create, beads_update, beads_close and beads_dep, in that order. Only host-reported reachable tools are sent. Muted chips say needs your grant: crossRibGrants. The switch does not create host grants. Without a reachability hook it is disabled: This host does not say which tools a lead may hold. For existing projects, a supported host reporting no reachable tools leaves it usable, with all chips muted. The rib rechecks lead-tool reachability on Start, Retry and Go deeper.

Turning switches off omits their grants; workflow chips stay for that project. Changing or clearing the project resets all switches and chips, not the task. Returning to New project… reapplies its defaults and keeps local Name and Folder edits. The sentence follows your choices: reading <name>, then and writing on a branch when Write is on, then , then <workflows> when workflows are named, then , with beads when the tracker is on. The beads suffix records switch intent, not a promise that every tracker tool was granted. Ordinary refreshes preserve your draft, switches and chips. Project-list, provider, capability, dispatch or remembered-refusal configuration changes can replace the page. On Keelson v0.119.0 or later, replacement documents restore the task verbatim, plan and model choices, project, switches and chips, pending field text, and expanded presentation. A restored expanded or multiline draft opens the full controls instead of compact defaults.

New project… selection, Name and Folder restore verbatim while creation remains available, even after the project list grows. Losing creation capability restores chat-only with elevated access and workflow chips cleared. Replacement documents force Write on for creation and preserve explicit tracker opt-out. Losing initialization capability clears tracker consent. Projects restore by ID for existing projects; a removed or hidden project becomes chat-only and clears its switches and workflow chips. Current capability restrictions still apply. A changed project root clears elevated consent until you opt in again. Named models keep their selected provider; an unavailable provider requires choosing a model or plan again. Workflow chips restore exactly as typed, in order; the host refuses unknown workflows at Start. The launcher does not detect removed workflows.

For a GitHub issue or PR, put its link in the task: Start reads it with the gh CLI and attaches it as a context item (title, body and comments, with its URL and retrieval time), and the task hint says Will attach issue #N from owner/repo. At most 5 links are read. Any other URL or a bare #N is refused with a host toast; paste the text instead. Starting… is a two-second duplicate-click guard, not a completion signal; the task stays in place and a successful start opens the swarm on the index. Each Start dispatch clears the saved draft before sending the action, even if the host refuses it. Local validation failures do not clear it. The current document keeps its fields for retry; the next edit saves a fresh draft.

The bridge keeps state only in browser-tab memory, not durable storage. A browser-page reload loses saved drafts. The host retains at most 64 view keys and caps each JSON snapshot at 65,536 UTF-8 bytes. An oversized save leaves the last accepted snapshot intact without truncating visible text. Older hosts without the state bridge can discard local edits when the launcher page is replaced. See the launcher for draft retention and retry behavior.

For advanced inputs, call chat_swarm_start from chat or any MCP client:

{
"tool": "chat_swarm_start",
"input": {
"task": "Find why `bun test` on main takes four minutes ...",
"project": "keelson"
}
}

project is a registered Keelson project, by id or name. With it, agents get Read, Grep, and Glob confined to that project’s root. Leave it out and the swarm is chat only.

The result carries the swarm id and a run id:

swarm s3fk started in #swarm-s3fk (run 211bdfbe-...). Poll chat_swarm_status("s3fk") or run_status("211bdfbe-...").

For a larger or smaller job, pass size: small (3 agents, 20 turns, 15 minutes), medium (the default), or large (8 agents, 80 turns, 4 turns at once, 60 minutes). To set single ceilings on top of that, pass max_agents (up to 12), max_turns (up to 200), max_turns_per_agent (up to 100), turn_timeout_s (up to 1,800), and max_minutes (up to 240).

To choose how much model the agents get, pass power: fast, balanced (the default) or deep. Each provider maps a power to one of its models, and the host’s modelClasses setting can change that map. To name a model instead, pass provider and model. model runs every agent, or the lead alone when worker_model is also set, and workers then run worker_model. An agent with a named model ignores power for its model.

power also sets the reasoning effort every turn asks for: fast asks for low, balanced for medium, deep for high. To set it apart from the power, pass effort (none, low, medium, high or xhigh); it applies to every agent, named model or not. A provider without effort support ignores it. chat_swarm_status reports the effort asked for as effort.

Some models take no effort at all: on Copilot, claude-haiku-4.5 refuses one. When a model refuses the power’s effort, the turn is retried once without it, every later turn on that model goes without, and the activity log says so. An effort you name is not dropped: a model that refuses it fails the turn, and the swarm ends once the lead’s turns have failed three times.

The chat-swarm workflow wraps start, wait, and report. It takes the task as its argument, holds until the swarm ends, and reports the status, the conclusion, who took part, and the turn cost:

{ "tool": "workflow_run", "input": { "name": "chat-swarm", "arguments": "Find why ..." } }

The workflow is written to pass only a task. For a project, limits, a model, or task context, call chat_swarm_start directly. The model pinned on the workflow’s nodes runs only its start, wait, and report steps, not the swarm’s agents.

Open the swarm-<id> channel in ClickClack. The first message is the task. You will see the lead delegate with mentions, workers answer in threads, and findings land on the board. The rib also streams a progress line per turn to the run, which run_events returns.

To read the channel from Keelson instead, call chat_swarm_transcript with the swarm id. It returns every message in order, thread replies included, for a running or ended swarm, and pages long transcripts by offset.

chat_swarm_status with the swarm id returns the summary:

{
"id": "s3fk",
"status": "done",
"channelName": "swarm-s3fk",
"turnsUsed": 9,
"agents": [{ "handle": "s3fk-lead", "role": "Lead: owns the outcome", "turns": 3 }],
"conclusion": "The regression landed in ..."
}

Only done means the lead concluded. For any other status, error holds the reason and the channel holds whatever was found. When the lead’s conclusion was refused as too long and none landed, draftConclusion holds its last draft. chat_swarm_wait blocks until the swarm ends or its timeout passes, which is what a workflow wants; from chat, poll chat_swarm_status.

The run completes only when the swarm concluded or was stopped. Any other ending fails it with the status and reason, and the summary is its last progress frame.

A swarm that did not finish offers Retry, with a Retry with model picker seeded with its effective model/provider. A finished Scout or Crew offers Go deeper, which starts the same launch on the next plan up with that plan’s models; a finished Fleet offers neither. Both keep the same task, project, workflows and context. The hint names the evidence and when it was captured; context is not refreshed.

Retry reuses the saved size and plan power unless a model is named. Accepting an unchanged plan-derived lead keeps the pair, not one model for everyone. Deliberately saved model/worker overrides are retained unchanged. Choosing another model drops the old worker override; clearing the model restores saved power or its omitted default.

A swarm’s answer varies from run to run, so the first thing to measure is agreement: run the same tasks several times and see how often they pass the same checks. Until that pass rate is steady, a prompt change cannot be told from noise. Measure before you tune.

The rib ships a case set at evals/chat-swarm.eval.yaml: three tasks, each run three times, graded by a judge against claims a reader could check in the workflow’s report. It runs through the chat-swarm workflow, so it needs a running Keelson with this rib and a ClickClack it can reach, and every case starts a real swarm that spends model turns.

Terminal window
keelson eval run evals/chat-swarm.eval.yaml

Read the test split’s pass rate and its interval first. A NOISE warning means the interval is too wide to judge a change: add reps or cases before touching a prompt. Once it is narrow, change one thing, run the set again, and let keelson eval compare before.json after.json say whether the change stayed within noise. Keelson’s Evaluating workflows guide covers the case file and the verdict.