AuthBox documentation — User guide

Documentation embedded in this build.

Who decides, and where AuthBox goes

Two questions come up in every first conversation about this product, and until now neither had a page. Where does it go? — one of these in front of everything, or one beside every service? And who decides what? — because the platform team, the tenant and the application team each think the answer is them, and in a well-run organisation all three are partly right.

How AuthBox works, on one page: the entity, the mission, the marking, the door and the receipt, and the one rule that joins them.
How AuthBox works, on one page: the entity, the mission, the marking, the door and the receipt, and the one rule that joins them.
One box: every capability in one sealed instance — the directory and its policy decision point, the certificate authority, the console and LDAPS, the signed audit chain — with no network required.
One box: every capability in one sealed instance — the directory and its policy decision point, the certificate authority, the console and LDAPS, the signed audit chain — with no network required.
Chained: HQ publishes a signed bundle, a region adopts and republishes it, an enclave adopts it; a partner's trust anchor is appended without its data being believed; origin survives every hop and freshness is budgeted.
Chained: HQ publishes a signed bundle, a region adopts and republishes it, an enclave adopts it; a partner's trust anchor is appended without its data being believed; origin survives every hop and freshness is budgeted.

This page answers both with one rule and its consequences, and then works the rule through a Kubernetes cluster, which is where it is asked most often.

Two words are used throughout and both are glossed here rather than assumed.

A door is an AuthBox process that stands in the path of a request, checks the certificate the caller presented, decides whether that caller may have what they asked for, and only then passes the request on. The program is called authboxproxy. People also call it a front door, because a deployment usually has one in front of everything a browser or an API client can reach.

A mission is a name for a need-to-know — "the people who may see the depot relocation plan" — that a person either holds or does not. Rules are written in terms of missions, groups, and facts about a person, and never in terms of individuals.

The rule

A door per authority, never a door per process.

An authority is somebody who may decide: a party that owns something and therefore holds the right to write the rules for it. The platform team owns the cluster. The tenant owns their namespace. The organisation owns its project and the people in it.

Two doors are right when two different parties each have a decision to make. Two doors are redundant when one party would be deciding the same question twice. That is the whole test, and it is a question about your organisation chart rather than about your topology.

It is why a cluster-wide door and a per-tenant door are both correct — the platform team and the tenant are different parties with different rules to write — and why a door beside every workload is not. Fifty workloads owned by one team is one authority fifty times, and fifty copies of one decision is fifty places for it to drift.

Five things follow from that sentence. They are the substance of the page.

1. What scales out is authorship, not enforcement

A service mesh takes a decision point and puts a copy of it next to every process, so that every hop between services is checked. That is a coherent design and it answers a real question. It is not this one.

AuthBox scales the other half. What gets distributed here is the right to write the rules — pushed out to whoever owns the thing being governed — while the deciding stays at boundaries, where there are few enough of them to name, watch and audit.

The practical difference: adding a team to AuthBox usually adds no new process anywhere. It adds a group that may author its own rules. Adding a team to a mesh adds as many sidecars as they have workloads.

2. A team can own its rules three ways, and only the third needs a second deployment

This is the question behind most "do we need our own instance?" conversations, and the answer is usually no. There are three rungs, and you climb only as far as the trust boundary actually requires.

Rung one — the team owns its rules inside your deployment. A project's own admins author the project's own missions and publish its own agreements, from the web portal or the command line, without filing a ticket with anybody. The identifier of a project mission carries the project's name, derived from the project's own directory name and never typed, so one project cannot write into another's half of the namespace. The same admins mint identities for their own services, so a web server gets a certificate of its own whose name derives from the project's, with the key born on the machine that uses it. See Your project's own missions and A certificate for a service.

For most teams this is the whole answer. They own their rules; you still own the deployment.

Rung two — another deployment owns the project outright. A second AuthBox, run by somebody else, holds the project and its records. Yours adopts it as a named upstream: a configured authority with a freshness budget attached. The authority is a name in your configuration and never a URL, because a URL is a location and nothing signs a location. If that upstream's word goes stale past its budget, your deployment stops authorizing that authority's projects rather than failing everything.

This is the rung for a genuinely separate organisation — a partner, a subsidiary, a classified enclave — not for a team that merely wants autonomy.

Rung three — another system asserts facts, and you still decide. Sometimes the other party owns a fact rather than a project: a training system knows who is current, a badging system knows whose badge expires when. That system publishes signed statements about one attribute name, your deployment ingests them, and every decision is still yours. The fact's owner is not an AuthBox at all in most cases, and need not be. See Where an org fact lives.

Only rung two is a second deployment. Rung one is a permission and rung three is a feed.

3. A boundary that does not enforce itself is advice

A per-tenant door is a real boundary only if something stops traffic going round it.

Say that plainly for Kubernetes, because this is the most common mistake on the page. A namespace is an authorization and scoping boundary for the Kubernetes API. It is not a kernel isolation boundary. It decides who may create, read and delete objects, and it gives names a place to live. It does not, by itself, stop one pod talking to another: pods in different namespaces share nodes and can reach each other's addresses directly, and nothing about the namespace refuses that connection.

So a per-tenant door in front of a namespace is advice until a default-deny network policy stands in front of it — a rule that refuses all pod-to-pod traffic except what is named, so that the door is the only way in. This repository states that as a prerequisite for the shape rather than as hardening to do later, because the shape does not do what it appears to do without it. Check your cluster's networking plugin enforces those policies: some popular ones do not, and a policy nothing enforces is a second piece of advice on top of the first.

The general form of the rule, beyond Kubernetes: draw the boundary, then find the thing that makes going round it impossible. If there isn't one, you have documentation.

4. Plain HTTP to an application is permitted, and graded

The last hop — from the door to the application behind it — may be plain HTTP. This surprises people, so here is the reasoning.

The risk that matters on that hop is not eavesdropping. It is bypass: somebody reaching the application directly and never passing the door at all. An encrypted network fabric fixes eavesdropping and does nothing about bypass, which is why mutual TLS is the default here — the point is not that the traffic is encrypted, it is that the application refuses everyone except the door.

Plain HTTP is therefore permitted and named rather than forbidden or silent. The deployment reports it as a weakening you can read off the door's own status, and the grade depends on what travels:

Two further facts about that hop. The decision receipt keeps its own integrity whatever the transport: it is a signed document the application verifies by itself, offline, and forging one takes the signing key rather than a position on the network. But a plain hop cannot carry the chain — the record of who called through which door — because the chain's trust rests on the mutual TLS binding that a plain hop does not have. If your inner service needs to know what happened in front of it, the hop has to be mutually authenticated.

5. An application gets its identity from here, and its own team mints it

The service behind the door needs a certificate of its own. The failure everybody has lived through is that a person claims it, the key is generated on that person's laptop, and the certificate says "web server" while the thing holding its key is whoever was on shift.

Here, the project's admin mints a service identity whose name derives from the project's own, and the workload enrols with a single-use token that expires, generating its key where it will be used and renewing itself from then on. Nothing hands a key around, and there is no download option anywhere in any shipped interface. Teams who would rather declare a certificate the way they declare everything else can use the standard certificate-management path instead; the identity is the same either way.

Worked example: a Kubernetes cluster

Apply the rule to the case it is asked about most.

One door for the cluster. The platform team owns the cluster, so they own one door. Run as a Kubernetes ingress, it discovers what it fronts from a single annotation on a Service — no static list of applications to maintain, no restart when one arrives. This repository stands that up on a real API server and proves four things about it: a route appears while the door is running and starts answering; a caller who does not meet the route's rule is refused through the identical route; an annotation somebody breaks leaves the previous table serving rather than taking the application off the air; and deleting the Service removes the route.

One door per tenant namespace, when the tenant is a different authority. If each namespace belongs to a team that writes its own rules, each gets a door, and each team authors through rung one above. If the namespaces are just a filing convention used by one platform team, they do not — that is one authority with several folders, and the cluster door already decides for it.

The namespace door is advice until the network policy is there. Everything in §3, in the place it actually bites. Default-deny first, then the door, then the tenant's rules.

What the inner services do: nothing. They verify the receipt the door hands them, or they simply refuse every client certificate except the door's. They do not each decide again, and they do not each get a door.

Read on