Founded 2026 · Private beta open

Your AI agents get least privilege. Nobody has to write the policy.

Codify watches the tasks your team keeps asking agents to do, distils how they're actually done, and locks each one to exactly what its runs touched: hosts, paths, secrets. It enforces that scope at the container boundary.

“Nobody wrote this policy. It is what these runs already did.”

codify / governed-tasks

Auto-promoted · v1

Audit dependencies → reports/audit.md

Learned from 50 runs · 37 different wordings

Enforced

Network

registry.npmjs.org

Filesystem

./reports RW

./repo RO

Secrets

None injected

Live boundary log

  • budgetrun admission — 429 over ceiling
  • allowregistry.npmjs.org — 200 via broker
  • pathwrite finance/q3.xlsx — EROFS

Runs on the stack you already have

◆ Docker◆ Volcengine Ark◆ Codex CLI◆ Node.js 22◆ Fastify◆ React◆ OpenTelemetry-style traces◆ Allowlisting egress broker◆ Docker◆ Volcengine Ark◆ Codex CLI◆ Node.js 22◆ Fastify◆ React◆ OpenTelemetry-style traces◆ Allowlisting egress broker

The problem

Nobody runs agents with least privilege, because nobody can write the policy.

Enterprises are putting autonomous agents on real repos, real data and real credentials. The permission model hasn't caught up.

01

Agents ship with the union of everything

You can't specify what a non-deterministic actor may touch before you've seen what it touches. So every agent gets every host, every path and every token anyone might need.

02

Fifty people, fifty wordings, fifty answers

The same task gets re-invented in every chat. Quality varies run to run, and nobody captures how it's actually done well.

03

Controls people route around are worthless

Security that slows people down gets bypassed. Least privilege only sticks if the governed path is also the better one.

How Codify works

Policy that's measured, not imagined

Codify is middleware for agent platforms. It plugs in at three seams and enforces at one boundary, with no new data layer and no rewrite.

◎

Step 1 · Observe

Watch the work

Every run is traced: the hosts it reached, the paths it wrote, the credentials it used. Prompts are redacted before anything is stored.

≋

Step 2 · Recognise

Spot the recurring task

Three matching channels (lexical fingerprint, containment and embeddings) recognise the same task across rewordings, padding and translation.

✦

Step 3 · Promote

Distil a specialist

Codify writes a brief from how the task is really done and promotes a specialist agent. Person 51 gets the distilled version, not a blank page.

⬡

Step 4 · Enforce

Lock it at the boundary

The specialist runs on an internal network behind an allowlisting broker, with a read-only workspace apart from the paths it actually writes.

It keeps learning after promotion.

When several people give a specialist the same correction, such as “use more colour in the headings”, it stops being a preference and becomes a proposed rule. The contract is versioned and the brief is rewritten. Scopes can only narrow automatically, never widen.

v1v2v3

Enforcement, observed live

Four refusals. All watched happening.

With the agent's own sandbox switched off, every one of these was blocked by Codify at the container boundary. The provider key never enters the agent container.

Egress⊘

Unlisted host blocked

A second rates API the contract didn't name was refused. The run finished on the allowed host.

$ 403 broker: host not in contract
Path⊘

Write outside scope fails

A write to finance/ failed in the kernel. The directory was unchanged.

$ EROFS: read-only file system
Budget⊘

Over-spend stopped at the door

The token ceiling was enforced at admission, before the run existed.

$ HTTP 429 at admission
Secret⊘

Credentials not auto-granted

A GITHUB_TOKEN seen in past runs was withheld from auto-promotion and needs a human decision.

$ secret withheld: GITHUB_TOKEN

And because the prompt comes from the caller, the scope never depends on recognising it. A specialist runs under its own contract's scope, whatever it is asked. Defeating the matcher costs the brief and gains no capability.

Measured results

We publish the numbers, including the ones that broke us.

Our first matcher design failed open on real users. We measured it, fixed it, and kept the evidence.

99.4%

task recognition recall

at the calibrated threshold, with zero cross-contract confusion

0

false matches on 2,788 real prompts

WildChat, BigCodeBench and SWE-bench background sets

1 vs 4

output structures

governed agent across 3 runs vs ungoverned control across 4

283

automated tests

plus typecheck and both builds, on every change

Why three channels: recall under adversarial rewording

ChannelsRecallFails openRecall bar
Fingerprint only1.7%354
+ Containment8.1%330
+ Semantic embedding100.0%0
All three channels100.0%0

They fail in opposite directions

  • Padding attack: the fingerprint drops, but containment holds at 1.000.
  • Reword or translate: both lexical channels read 0.00, but the embedding holds at 0.78.
  • Route if any channel clears. An attacker has to beat all three at once, and even then gains no capability.

The governance console

Every scope has a receipt.

Security teams see each promoted specialist, the evidence its scope came from, and a one-click revoke. Approvals are refused by the control plane, not just hidden in the UI.

app.blueberryjam.site/governance
Codify governance console showing three auto-promoted specialists, each with network, filesystem and secret scopes

Roadmap

Day one was 2026.

We're early, and we're building in the open with the teams who need this most.

  1. Q1 2026Shipped

    Blueberry Jam founded

    Four engineers, one thesis: nobody runs agents with least privilege because nobody can write the policy.

  2. Q3 2026Shipped

    Codify proof of concept

    Observe → distil → promote → enforce, running end to end on Docker and Volcengine Ark. 283 tests passing.

  3. Q4 2026In progress

    Private beta with design partners

    Onboarding a small cohort of teams who run internal agent platforms. Contracts, briefs and refusals on their real workloads.

  4. H1 2027Next

    Public launch

    Self-hosted and managed editions, SSO, policy export to OPA and Kubernetes NetworkPolicy, and SOC 2 Type I.

The team

Small team. Sharp focus.

Security, infrastructure and ML engineers who got tired of shipping agents with admin-shaped permissions.

AT

Aria Tan

Co-founder & CEO

Ran platform security for an internal AI rollout and watched every agent ship with admin-shaped permissions.

DO

Daniel Okafor

Co-founder & CTO

Container and networking engineer. Built the broker that keeps provider keys out of the agent sandbox.

PR

Priya Raman

Founding Engineer, ML

Built the three-channel matcher and the benchmark that proved the original design wrong.

LH

Leo Hartmann

Founding Engineer, Product

Designs the governance console, so security teams can read the evidence and act on it.

FAQ

Questions we get asked

Do I have to write any policy?+

No. Codify watches recurring tasks and derives the scope from what those runs did: hosts reached, paths written, credentials used. You can review, narrow or revoke anything, but nobody has to author the first draft.

What if someone rephrases a prompt to dodge the matcher?+

They lose the brief and gain no capability. A promoted specialist runs under its own contract's scope whatever it is asked, so defeating recognition never widens access.

Where is the policy enforced?+

At the container boundary, not in the prompt. The specialist runs on an internal network with no route off-host, reaching out only through an allowlisting broker. The workspace is read-only except the paths the task really writes.

Does the agent ever see my API keys?+

No. The broker holds the real provider key, and raw prompt text is redacted on the control-plane side. Neither enters the agent container.

What does it run on today?+

The proof of concept runs Codex CLI agents in Docker on Volcengine Ark. Support for more runtimes and model providers is on the 2027 roadmap.

Is it production ready?+

Not yet. We are a 2026 startup in private beta. Design partners get early access, direct support from the founders, and a say in the roadmap.

Private beta · Limited cohort

Spread least privilege like jam.

We're onboarding a small group of design partners who run agents on internal platforms. You get early access, direct founder support, and a real say in what we build next.

  • ✓Self-hosted, so your data stays in your environment
  • ✓Hands-on onboarding to your agent runtime
  • ✓Free during the beta

This opens your email app, so nothing is stored on this site.