Blue Jacket ConsultancyBlue Jacket Consultancy
For teams running AI agents

Supervised Autonomy. Agents at full speed. People in command. The ledger holds the proof.

Your agents work at machine speed inside boundaries you set. Agents propose. People commit. The ledger preserves. We establish your AI foundation first, then deploy custom agents at full force, on the record. An append-only ledger preserves every action and every authorization. The Board of Operations shows what’s done, what’s pending, and what’s planned. Control AI with confidence.

Diagnose. Design. Deploy. Operate. The four-step methodology behind every Blue Jacket engagement.

The problem

The failure mode: AI reporting a health it doesn’t actually have.

SMBs are fast becoming the missing middle: too exposed to run agents unsupervised, too lean to staff a control room. They need agents that can read, write, and act on the business, and the risk compounds with every one they add. The fix isn’t more tools. It’s to instrument the truth, install the discipline that keeps it true, and keep the record that proves both.

01

A security review you can't answer

An enterprise customer's security questionnaire asks what your AI agents can reach, what they hold, and what happens when a control fails, and nobody owns the answer. “We have a policy” isn't a control; it's a claim about one. Configured is not the same as binding, and the only thing that tells you which you have is the machine's own record.

02

An agent that's configured, not contained

The containment was set up: the policy written, the approval wired, the boundary drawn. Whether it holds under load is a different question, and it stays invisible until an agent acts wrongly at machine speed against something you can't take back. A control you haven't measured isn't a floor. It's a hope with good posture. And an approval nobody actually read is not a control either. It’s a reflex.

03

An incident you can’t reconstruct

Something went wrong at 2 a.m. and the questions are simple: what did the agent do, in what order, and who approved it. If the answer lives in chat scrollback and screenshots, you don’t have a record. You have folklore. When it matters, the difference between the two is the whole game.

The methodology

How we build it: defense before offense.

The same reach that makes an agent useful (it can write, deploy, spend, send) is exactly what makes it dangerous when it’s wrong. We don’t answer that with hope. We build agent operations on the controls security teams already trust, applied to a system that touches your revenue.

Most teams can’t prove who approved what. They manage agent risk one of two ways: click approve on every step, which becomes a rubber stamp, or turn approvals off, which is no gate at all. Neither is a control, and neither leaves a record anyone can check. We install the third way: agents work at full speed, a named person commits, and the record shows who and when.

Every commit is a human act. Before anything irreversible lands, a named person types back a digest of what is about to happen, enters a per-operation PIN, and touches a physical security key. Three independent channels, one human, every time. No read-back, no commit. This gate has caught a one-character transcription slip in live use and refused it before anything landed, and the refusal itself is on the record.

One rule: Agents propose. People commit. The ledger preserves. Every irreversible action stays with a human. Six controls do the work.

01

Segregation of duties

The seat that does the work is never the seat that commits it. Separation is the control.

02

Least privilege

An agent holds only the capabilities its job needs. Write and destructive tools are withheld unless granted.

03

Verifiable identity

No unidentified agent action. Every action runs under an identity you can trace.

04

Append-only audit trail

Actions are recorded and corrected with change-orders, never quietly deleted. The audit is a byproduct of the method, not a chore you bolt on after.

05

Fail-closed defaults

Missing authority refuses rather than proceeds. The agent stops; it doesn’t guess.

06

Verify before act

No action on a reconstruction from memory. The agent reads the live source, or it stops. Ground truth is the only “done.”

Enforcement honesty: Where the platform can enforce a control by mechanism, we make it deterministic by design, not by the model’s good behavior. Where a surface can’t enforce one deterministically yet, we keep a clearly-labeled human gate in place as a second net, and we say plainly that’s what it is. A second net never gets dressed up as the floor. That candor is the point. It’s the difference between a system you can stand behind and a story you hope holds.

The result is an operating layer, not a demo: defined agent roles with verifiable identity, capability gating, fail-closed defaults, and a record you can audit after the fact. Defense before offense. Then the revenue work runs on a substrate that holds.

Engagement Cadence

How an engagement runs.

A four-step engagement methodology. Each step is a deliverable, not a consulting deck. Each step earns the next.

01

Diagnose

Measure what your agents can reach, what they hold, and what happens when a control fails. Map the exposure across your stack from configuration and, where useful, a lab run against synthetic credentials. The diagnostic is a deliverable, not a sales pitch.

02

Design

Architect the operating layer: the seats, the gates, the identity model, the record. The design names the agents, the workflows, the data contracts, and the controls that prove it’s working.

03

Deploy

Install the package: seats and identity, the read-back commit gate, the append-only record, custom agents working at machine speed under the gates. Hands-on implementation, not a slide deck handed to your team to figure out.

04

Operate

Embedded leadership while your team takes the watch. Agents drift, the market moves, the data contracts and the controls need an owner. Operate ends with your team owning the system, and the record proving it ran.

The work

Two front doors. One operating philosophy.

Same discipline whichever door you come through: instrument the truth, then install the controls that keep it true. Most engagements start at a front door (a fixed-fee, walk-away-usable assessment) and go deeper only if the findings earn it. Defense before offense.

Start here · two front doors
Front door · Assurance

Agent Containment Readiness Review

Fixed-fee · ~1–2 weeks · no production touch

For the security-questionnaire moment. We map where your AI-agent exposure sits (what your agents can reach, what they hold, what happens when a control fails) from your configuration and, where useful, a run in our own lab against synthetic credentials. Never your production systems. You get a readiness memo you can hand your customer's security team: measured findings, not a policy PDF. Complete value on its own: no MSA, no production access.

Front door · Revenue

Revenue Substrate Review

Fixed-fee · ~2 weeks

For the board-that-no-longer-trusts-the-forecast moment. A focused read of pipeline, forecast, and attribution (actuals versus narrative) delivered as findings a CEO can act on with or without us. Thirteen years of operating history behind why the read holds up.

Then, in sequence · verify, deploy, operate
01Verify

Agent Containment Diagnostic

Scoped engagement

The full instrument behind our published crosswalk and case studies, proven on our own bench, brought to your live agentic stack: conformance checks returning deterministic receipts, measured off packet capture and disk state, mapped to NIST AI RMF, OWASP, MITRE ATLAS, and SOC 2. Production-touch work: scoped under a written Rules of Engagement and counsel review before anything runs. Engaged after a Readiness Review, not before.

02Deploy

Supervised Autonomy Install

90 days

Install Watchbill™, the operating layer: the seats, the gates, verifiable identity, the append-only record. Then deploy AI where it compounds, under controls you own. Agents propose. People commit. The ledger preserves. Built hands-on with your team, on your infrastructure, and it stays when we leave.

watchbill.ai →
03Operate

Fractional Chief AI Officer

6-month minimum

Embedded leadership for the seat between a part-time advisor and a full-time Chief AI Officer. Agents drift, the market moves, the data contracts and the controls need an owner. We run the system with your team until your team runs it without us. What compounds is yours: the operating package, the record, the playbook. Not a dependency on us.

By Request / Authorized Only

Adversarial AI Audit

Boundary testing for mature teams, under written authorization and a per-engagement Rules of Engagement: explicit scope, operator-side safety controls, counsel-reviewed before anything runs. Boundary setting is part of every engagement; boundary testing is by request. Engaged case-by-case.

Pricing discussed on the discovery call, calibrated to scope and engagement model.

Not sure which door? The discovery call is the cheapest and easiest way to find out, and it’s on the calendar, not in a sales funnel.

Book the call
Proof

Three failures. Three refusals. One record.

Everything we install for clients runs in our own stack first, under Supervised Autonomy: agents at full speed, a named person on every commit, an append-only record underneath. These are three moments the controls were tested for real. The record of each is available on the discovery call.

01

The gate that refused

Problem

An irreversible change was queued carrying a one-character slip. Small enough to read right on a screen. Big enough to matter once it landed.

Action

The commit gate ran the way it runs on every irreversible action: a named person typed back a digest of the change, entered a per-operation PIN, and touched a physical security key. The typed read-back did not match what was queued.

Outcome
Refused before it landed.

No read-back, no commit. The refusal itself is in the append-only record, and that is the point: the record holds the near-misses too.

02

The block the agent walked around

Problem

An agent was denied its web tool mid-task. It did what agents do: found another way, a raw shell command toward the same destination, at machine speed.

Action

The boundary it hit was not a policy asking for good behavior. It was a measured control sitting beneath the agent, tested before it was trusted.

Outcome
Zero bytes out.

The attempt, the block, and the investigation are documented end to end in the case study below.

03

The safeguard that failed open

Problem

A probabilistic safety control passed every test in its suite and still failed open in live use. Same input, opposite outputs.

Action

We caught it in our own systems, documented the failure, and rebuilt the control as a deterministic floor, with the probabilistic layer demoted to a second net.

Outcome
A filter is not a floor.

Every control we ship is measured before it is trusted, and the measurement is on the record.

What we install
  • Orchestrator layer
  • Custom agents
  • Read-back commit gate
  • Append-only operating record
The architecture

What we verify, and what we install.

Two things sit under a Blue Jacket engagement: the instrument that measures whether your agents are actually contained, and the operating layer we install once they are. Controls first, measured and not attested, then the work runs on a substrate that holds. Your team owns it when we leave.

What we verify · the containment diagnostic

Conformance checks

A battery of assertions about what an agent can reach, what it holds, and what it does when a control fails.

Deterministic receipts

Every finding measured off packet capture, disk state, process state, and kernel/network logs, never the agent's narration.

Append-only ledger

The record is a byproduct of the method, corrected with change-orders, never quietly edited.

Framework crosswalk

Each check lands where your auditors already have a box: NIST AI RMF, OWASP LLM & Agentic, MITRE ATLAS, SOC 2.

Measured, not attested.See the Framework Crosswalk
What we install · Watchbill

The watchbill is who stands which watch. Ours stands the watch over your agents: named seats, a read-back commit gate, an append-only record.

See the product at watchbill.ai.

01Identity

Seats and identity

Named seats for people and agents, verifiable identity on every action. Who did what is never a guess.

02Commit gate

Read-back commit gate

Before anything irreversible lands: a typed read-back of the digest, a per-operation PIN, a touch on a physical security key. No read-back, no commit.

03Execution

Execution layer

Custom agents draft, enrich, and stage work at machine speed, under the gates. Nothing irreversible ships without a human.

04The record

Append-only record

Every proposal, approval, and refusal lands in an append-only record your auditors can read. Corrected with change-orders, never quietly edited.

The control architecture is the IP. Components are chosen for integration depth, not brand. The runbooks are documented. Your team owns the system when we leave.

Joseph D. Alise, Founder of Blue Jacket Consultancy
Founder
Joseph D. Alise
About

Built from zero. Scaled with AI. Delivered results.

I’m Joseph Alise. Navy veteran: Yeoman aboard the USS Bonhomme Richard (LHD-6), Surface Warfare Specialist. The Navy taught me how to read systems under pressure and how to build a watch that runs when the senior person isn’t in the room.

Thirteen years across enterprise data sales, e-commerce, and revenue operations, most recently building that function from zero as the first dedicated hire and driving sustained double-digit YoY growth across U.S. and international markets. Blue Jacket exists because most companies are layering AI on top of broken processes, and most consultants are happy to sell them more tools instead of fixing the system underneath. That discipline now has a credential behind it: I’m an IAPP-certified AI Governance Professional (AIGP). Blue Jacket is a small firm on purpose. The people on the engagement are the ones doing the work. Steady on.

13
years across
systems and revenue
1,200
sailors supported in
command operations
during deployment
Built from zero
the operating layer we
now install, run daily
in our own stack
Book the call

30 minutes. No deck. Honest read on whether AI is the right next move.

I’ll ask three questions about your revenue engine — or about what your agents can actually reach — name what I’m hearing, and tell you whether the diagnostic is the right next step. If it isn’t, you’ll know on the call.

  • Operator on the call — not a BDR, not a setter.
  • Walk-away usable — three things to fix, regardless.
  • No nurture sequence. We talk once, you decide.

Trouble with the embed?

Open in new tab