Site navigation

Watch a Coding Mission run

CODING TASKS

Stop babysitting your coding agents.

Throw a real job — Claude Max is GA on your Mac; Codex, Cursor, and Grok are local previews. HiveBase selects an eligible route, independently verifies on a different family, and only pulls you in for judgment. Cloud subscription runners are on the roadmap.

Claude Max · GA on Mac · local previews

00THE BRIEF

Throw a real job — not a blank agent session.

The ask arrives carrying company context, the judgment calls are answered, and the shape of the work is already drawn.

Ship the enterprise tier before the Acme renewal
M-12
JB
You

Ship the enterprise tier before the Acme renewal — they're on Business now and the renewal is in 18 days.

+6 more
Claude Code
Claude Code

I’ll run this as a program. Only these judgment calls need you:

Q1Audit logging — every write, or just admin actions?
Just admin actions
Q2The Enterprise price — should I set it, or propose options?
Set $499/mo, but hold it for my sign-off
Explored your codebase
Here’s the plan· 3 phases · 4 sessions
Phase 1 · parallel
SSO / SAML
Audit log
Phase 2
Seat management
Phase 3
Pricing & billing
Auto
Message…

Tap a beat · same Mission surface as desktop

SUPERVISOR

Leave the specifics to the Supervisor.

One right hand over the mission — a single coding session or a multi-agent wave. It closes what company context already knows, escalates only what still needs you, and when you redirect mid-flight it lands at the program level with a visible receipt. You don’t run each chat separately.

Claude Code
Claude CodeOpus

Plan is grounded. Still open: audit scope, OAuth, enterprise price, and seats vs Acme renewal.

Auto
Message…
SEATS YOU PLUG IN

Plug the heavy metal. We run the fleet.

Claude Max is GA on your Mac. Codex, Cursor, and Grok are local previews of the same plug-in. Lineup is a domain × effort Auto map you can pin — one primary harness per mission. Independent verification is mandatory and cross-family. Cloud subscription runners are on the roadmap. Signed-in and BYOK sessions count as tasks with no HiveBase token charge.

  • Claude · GA · Mac
  • GPT Pro · Local preview
  • SuperGrok · Local preview
  • Cursor · Local preview

No seats plugged.

auth-refactor
  • Scope auth surface
  • Implement token path
  • Review PR #412

How subscription economics work →

RIGHT MODEL FOR THE WORK

Stop picking a model for every subtask.

Your lineup is domain × effort — Auto maps each cell from certified routes you can pin. One primary harness per mission. Cursor has no public quota API, so HiveBase cannot pre-check Cursor remaining; exhaustion is reactive at execution. Exhausted clock windows queue; exhausted monthly pools route or Hold — never a silent bill.

How subscription economics work →

SCALES WITH THE ASK

Same Code Task. Simple or multi-session.

Throw a quick fix or a broad mission — the product adjusts. A simple ask stays one session; complex work becomes multi-session under the same Supervisor. No mode switch, no “group it later,” no second tool.

One intake, size-adaptive. A quick task is a program of one; larger work adds waves on the same surface.

01Program of one

Add SAML before renewal

  • 1 work item
  • 1 agent path
  • status + hold

Add SAML before renewal

1 hold
CodeReview
JB
Youkickoff
Add SAML SSO before the Acme renewal — it's on Business now.
enriched with your contextlaunch-saml.md
Claude Code
Claude CodeSonnet

Read the codebase and your context. Plan: add a SAML verifier, route OAuth through PKCE, update Business pricing copy.

Analyzed your codebase · 12 files
Claude Code
Claude CodeSonnet

Implemented across 3 files and opened a pull request.

Writesrc/auth/saml.ts
Editsrc/auth/oauth.ts
Editapp/pricing/page.tsx
Auto
Message…

One intake. Size-adaptive. No mode switch — a quick task is a program of one.

HOLDS

Soft where you touch it. Hard where trust is minted.

Product judgment and money paths wait for a human — the Supervisor doesn’t guess. Cap exhaustion never silently bills: spill-to-on-demand is a Hold. Experience parks with a typed reason; trust fails closed.

M-12 · decision hold · spend hold

See the Engineering Squad
Decision · your call
#250·Pricing & billing

Enterprise price — your call

Code and checkout are verified in preview. The price itself is the one customer-facing call I held for you.

Recommended · Approve $499/mo (from enterprise-gtm.md)

not offered$499 / mo
your rule · hold customer-facing pricing
Spend · fail closed
Included pool·monthly included pool

Monthly included pool is spent

Continuing on this leg would spill to pay-as-you-go at API rates. This is a trust Hold, not a silent continue.

Recommended · Route remaining work to another eligible seat · or approve on-demand spend

Clock window

Claude · 5h window exhausted

Queued until reset — refill on the clock

Monthly pool

Monthly included pool exhausted

Hold or route — queueing is useless for weeks

REVIEW THE OUTCOME

Re-enter on a brief — not a text dump.

One session or a whole mission: see what happened in a visual brief, not a wall of logs and diffs. What changed and why · proven vs reviewed vs unverified · what the Supervisor decided · callouts for what still needs you. Open any PR when you want depth.

Independent proof binds to a SHA with criterion↔evidence. Criteria without evidence stay labeled — never a naked green check. Experience fails soft; trust fails closed.

Mission briefM-12
CloudLinkShare
Verifiedsha 8f31c2a

Mission review

Enterprise tier before the Acme renewal

Acme renews in 18 days on Business. You asked for enterprise readiness as one shippable outcome — not four half-finished tickets.

Architecture

Rendering diagram…

Acceptance

CriterionEvidenceStatus
SAML enterprise loginExecuted in previewmet
OAuth path preservedCross-session guardrailmet
Seat count → StripeMeter increments by 1met
Enterprise price liveHeld — your ruleheld
Needs you$499 / mo enterprise price

Code and checkout are verified in preview. Your rule held customer-facing pricing — recommend $499/mo from enterprise-gtm.md, waiting on your sign-off before it goes live.

Self-correction · What I got wrong: Phase 1 audit missed one mutation path on first review — fixed and re-verified without pulling you in.

Depth#247 · #248 · #249 · #250

Session or mission · visual brief · human callouts — not PR archaeology every time.

COMING SOON

After missions fly reliably

Shipping first: multi-agent missions, Supervisor over the program, and honest proof. These deepen the flywheel — not claimed as live today.

  • Soon

    Memory flywheel

    Crown-admitted claims from proven missions feed the next run — decisions, dead ends, and constraints with evidence pointers, not agent-asserted notes.

  • Soon

    Isolation → cloud venue

    When ports, local DBs, or .env collide, isolation can become a routing signal toward a single-tenant cloud runner on the same subscription — on the roadmap, not shipped. We rent the sandbox; we don’t become one.

  • Soon

    Open-weight lane

    Mechanical refactors, tests, and boilerplate on competitive open-weight models once reliability clears our bar — task class, not benchmark theater.

QUESTIONS, ANSWERED PLAINLY

What a Coding Mission actually does.

A thin control plane over seats you plug in — Claude Max is GA on your Mac; Codex, Cursor, and Grok are local previews. We select an eligible route, supervise, and prove. Not another coding agent, runtime, planner, or bare chat with a model.

What is an AI coding agent supervisor?

Your right hand over coding agents you plug in — not another chat, runtime, or planner. One Supervisor covers a single session or a multi-agent mission: company and repo context let it inject constraints, cover criteria, double-check mechanical work, and escalate only when the decision is still yours. It also selects an eligible route, lands mid-flight steers at the program level with a visible receipt, and proves results independently. Most of the loop never reaches you; the ones that do are the calls you would want to make anyway.

How is this different from just running a coding agent yourself?

Running an agent alone still makes you the router, the context-loader, the babysitter, and the QA function: you re-explain taste and decisions every session, unstick mid-flight, and review every diff cold. HiveBase’s Supervisor answers what it can for you from company context so you don’t babysit each chat in a mission — mechanical orchestration recedes, and only clarifying judgment, money paths, and other calls that stay yours pull you in.

Is a quick fix a different product from a big mission?

No. One intake, size-adaptive. A quick task is a program of one — same Supervisor, same surface, no mode switch, no “group it later.” Multi-agent waves only appear when the work actually needs them.

What happens when I interrupt mid-run?

You steer the Supervisor, not each agent chat. Every interjection enters at the program level — from task chat, Linear, a hold answer, or a scoped hint in an open session. The Supervisor triages the meaning, fences it to the plan revision it was judged against, and returns a visible receipt so the redirect doesn’t feel like it vanished even when it worked.

Do I need API keys or BYOK?

Claude Max is GA on your Mac. Codex, Cursor, and SuperGrok are local previews of the same plug-in. Cloud subscription runners are on the roadmap. Metered HiveBase Cloud (API key or credits) is a separate preview — consumer subscription credentials stay off that venue. API keys and BYOK are optional overflow. When a pool would spill into unexpected on-demand billing, HiveBase Holds for your approval instead of continuing silently.

What happens when a subscription hits its cap?

Usage is not one number: some legs refill on a clock (queue and resume), others are monthly pools (route elsewhere or Hold). Cursor has no public quota API, so HiveBase cannot pre-check Cursor remaining — exhaustion is reactive at execution. If continuing would spill into surprise on-demand billing, HiveBase stops and asks you first — it never burns your card quietly.

What does “done” mean here?

Every criterion carries a proof tier: verified with independent evidence, reviewed only, or unverified — stated plainly. Agent self-checks are not verification. Producer and grader are different families when proof is claimed. Experience surfaces fail soft with honest labels; trust surfaces fail closed.

What happens when a Mission hits a wall?

It parks with a typed reason and a next action — never bricks. It ships what already passed verification, surfaces the blocked work with the real error, and never silently half-merges. Safe partial delivery is part of the contract.

CODING TASKS

Most of the run never reaches you. What does is yours to decide.

HiveBase selects an eligible route, supervises the waves, and proves the result. The two things that still stop for a human are product judgment and money paths.