Beehive AI Labs

Beehive AI Labs is a studio. Turncoat is the first paid offer.

We audit whether the authorization controls you already shipped still run on every agent tool path.

Turncoat

A white-box authorization-enforcement audit of whether your agent policy still runs on auto, YOLO, headless, subagent, and MCP paths.

We require a sibling path that still enforces, and a reproduction. No finding without both.

Who it is for

AI-agent harness owners. The champion is Staff or Head of Product Security, or the engineer who owns the harness.

You fit if all of these are true:

  1. You own the TypeScript, JavaScript, or Python source for the harness, CLI, IDE, SDK, MCP host, or internal wrapper. A SaaS seat is not enough.
  2. The product can run tools (shell, file write, network, MCP), not only chat.
  3. You already ship a policy layer: deny rules, allowlists, approval, sandbox, or workspace trust.
  4. You also ship a second path that skips or weakens the human prompt: auto, YOLO, headless, subagent, or MCP.

Typical buyers: vendors of coding agents; MCP hosts and agent control planes with their own approval layer; product companies with an internal TypeScript, JavaScript, or Python harness that has more than one dispatcher or more than one approval mode.

You are not the buyer if you only use Cursor, Claude Code, or Copilot as a seat; if you want prompt-injection or jailbreak testing; if you want a runtime guardrail or Copilot dashboard; if you cannot share source under NDA; or if you want a public CVE.

What you get

A time-boxed white-box authorization-enforcement audit of one owned repository, pinned to one commit, in TypeScript, JavaScript, or Python.

Every engagement includes:

Focused (16 hours, 5 business days): one dispatcher family, one authoritative control, up to two confirmed-gap reproductions, 30-minute readout.

Standard (32 hours, 10 business days, default): at least two named modes, private gap matrix for those paths, 45-minute champion readout.

Deep (48 hours, 15 business days): the mode families you actually ship (minimum three), written exec summary, champion readout plus signer readout. Optional architecture-only reachability note. Still one repo. Not a pentest.

Patch-verify after you ship a fix is a separate engagement. Never bundled.

A clean "no confirmed gap in the hour cap" is a valid, paid outcome.

What it costs

Fixed-fee USD. 50% on kickoff, 50% on delivery. No time-and-materials.

SKUListHoursCalendar
Focused$6,000165 business days
Standard (default)$8,5003210 business days
Deep$13,5004815 business days
Patch-verify (add-on, never bundled)$2,5008after you ship a SHA

Clock starts when the deposit clears and the repo is in hand.

How to start

  1. Reply to the email that brought you here. You get a one-page scope back: SKU, hours, calendar, what we will do, what we will not do, list price. If you do not pick a SKU, Standard is the default.
  2. Sign the NDA. No repo before that.
  3. Intake: pinned commit SHA, language, named modes, which control is supposed to bind, and a security champion with access.
  4. Agree the statement of work (one sentence, hours, deliverables, out of scope, list price).
  5. Pay the 50% deposit. Kickoff pins the SHA. Work starts.

If intake is still incomplete after five business days, the engagement is cancelled or re-scoped. There is no unpaid look at the repo.

What we will not claim

Out of scope for every SKU: prompt-injection and jailbreak testing; pentest of product, cloud, or tenants; SAST of AI-generated application code; runtime guardrails; retainers; remediation except as Patch-verify.