Kafka topic authorization belongs in your policy set, not in a per cluster ACL list

AAlex OlivierSeptember 18, 20269 min read
Kafka topic authorization belongs in your policy set, not in a per cluster ACL list

Ask a platform team which services can write to a given Kafka topic and the honest answer is usually that somebody will have to go and look. The answer lives in an ACL list on the cluster, expressed as bindings between a principal, an operation and a resource pattern, and it has been accumulating for as long as the cluster has been running. Every binding was correct on the day it was added. Almost none of them were ever removed.

That list also has no relationship to how access is described anywhere else in the organization. The application knows about teams, services and data classifications. The ACL list knows about User:payments-writer and a literal topic name. Translating between them is a job that lives in somebody's head and does not get handed over when they move on. Add a second cluster and there are two lists, drifting apart on their own schedule.

The general shape of this is older than Kafka. NIST's guide to ABAC draws the contrast directly, noting that access control lists depend on static membership that is re-evaluated only when something notifies the system. Static models assume stable roles, coarse permissions and a decision made once. A long lived cluster's ACL list encodes that assumption and then outlives it.

So this is not a failure of Kafka's authorization model. It is what happens when access rules live in a system that has no idea what your organization looks like. Kafka never insisted on owning them, which makes this a wiring problem rather than a policy authoring one, and wiring problems are what external authorization is for.

Kafka's authorizer is a plugin interface, not a fixed implementation

The broker decides authorization through a class named in authorizer.class.name. Kafka ships an implementation backed by its own ACL store, and that is the one most clusters run, but the Authorizer interface is the actual contract. The broker starts the configured authorizer, waits for it to report ready before it accepts connections on a listener, then calls it on the request thread for every operation a client attempts.

There is an established plugin for that interface which does not answer from a local store at all. It delegates each decision over HTTP using the request and response shapes established by the Open Policy Agent project, and reads the verdict out of the response body. Once a broker is configured that way, where the rules live stops being a Kafka question. The endpoint at the other end can be anything that speaks the protocol, including a Cerbos deployment that already governs the rest of the stack.

That division of labour is what reference architectures deliberately design towards. Germany's federal API authorization blueprint, published open source by Saxony-Anhalt with FITKO, is built around a strict split between the component that enforces and the component that decides, with the enforcing side kept deliberately thin. A broker authorizer plugin is close to a textbook instance, since posing the question and acting on the answer is all it knows how to do.

What a Kafka authorization request contains, principal, operation and resource

Three things, and no more. The principal, as authenticated by the listener, which is a SASL username or the subject of a client certificate. The operation, drawn from a fixed set that includes READ, WRITE, CREATE, DELETE, ALTER, DESCRIBE, ALTER_CONFIGS and several more. And the resource, which is a TOPIC, GROUP, CLUSTER, TRANSACTIONAL_ID, DELEGATION_TOKEN or USER, carrying a name and whether that name is a literal or a prefix.

That shape maps onto a Cerbos check almost directly. The resource type becomes the resource kind, so each Kafka resource type gets its own resource policy. The operation becomes the action, so rules address WRITE and ALTER_CONFIGS by their Kafka names rather than through a translation table somebody has to keep current. The principal name becomes the principal ID, and the connection metadata the broker already sends becomes principal attributes.

Cerbos Synapse serves that endpoint with a route extension, which claims a path, reshapes the raw payload into a check and formats the decision back into the body the plugin expects. That reshaping is a declarative mapping of expressions rather than a compiled adapter, so there is no plugin code to build and no second policy language in the stack.

Kafka brokers, service meshes and Kubernetes ask the same authorization question

The reason this generalises is that the interface has already been settled. The OpenID Foundation's Authorization API reached Final Specification in January 2026 and standardises how an enforcement point asks a decision point a question and reads the answer back. Its information model is a subject, an action, a resource and a context, the same four part shape Kafka's authorizer contract already produces. Decisions default to closed, and the spec is strict that a transport error and a policy denial are different things.

What that settles is the protocol rather than the architecture. The specification is explicit that the policy language and the state management of a decision point sit outside its scope. Where the attributes come from, what is cached and what happens when the source is unavailable is left to whoever builds it, and that gap is most of the work in a real deployment, whether the enforcement point is a broker, a query engine or an API.

Which is why the pattern repeats across the estate instead of being a Kafka trick. Companion pieces cover row filtering in Trino, mesh enforcement with Istio and admission control in Kubernetes, each answering on the protocol the component already speaks.

Kafka topic naming conventions become authorization rules

The consequence worth dwelling on is what the rules can now say. An ACL binding is an enumeration, and prefixed patterns aside, a binding still has to exist for every combination somebody wants to permit.

A policy evaluates a condition instead. The topic name and the principal name both arrive as ordinary strings, so the relationship between them is something the policy computes rather than something an operator enumerates. If a team owns the topics carrying its prefix, that convention becomes one rule covering every topic that exists and every topic that will exist.

In practice that reads as a set of derived roles turning the shape of a principal name into a capability, so a name ending in -consumer gains a consumer role, alongside a local variable computing the prefix that principal owns. A rule then grants READ on a topic to the consumer role when the topic name starts with that prefix. Six resource policies cover the six Kafka resource types, and a topic created tomorrow inherits its rules with nothing to provision.

A naming convention is only the cheapest available attribute, not the only one. The same condition could compare a topic against ownership recorded in a service catalogue, a data classification tag, or group membership resolved from the identity provider at decision time. Assembling those attributes before evaluation is a named role in the same NIST model rather than an implementation detail, the context handler, and it is the position Synapse occupies here.

NIST also names the payoff. So long as a new subject carries the attributes a rule needs, neither the rule nor the object has to change to accommodate it. In Kafka terms, a team onboarded next quarter and every topic it creates are covered already. That is where this stops being purely role based and starts describing intent.

Kafka topic access across clusters from one policy set and one decision log

The rules are policy files, reviewed in the same repository and the same process as the rules governing the API. A second cluster becomes another enforcement point reading the same policy rather than another list to maintain, and promoting a change is a policy version rather than a run of the ACL CLI against a different bootstrap server. Where an estate already expresses tenant isolation somewhere, the broker honours the same definition of a tenant as everything else.

The record changes too. Every broker decision lands in the same decision log as every application check, carrying the principal, the resource, the action and the rule that decided. Answering who could write to a settlements topic last quarter becomes a query against one log rather than a reconstruction from ACL change history and application logs that identify people differently. The audit trail is where the per cluster model fails hardest, because ACL bindings record their current state and not the decisions they produced. That matters more as event streams become where automated activity is recorded and replayed, turning who may read which topic into a governance question.

Caching, fail closed and the request thread, where external Kafka authorization gets hard

The Kafka javadoc is direct about the cost. authorize() is a synchronous call made on the request thread, and implementations "should avoid time-consuming remote communication that may block request threads". A network call is exactly what that warning is about.

Caching is therefore part of the design rather than a tuning exercise. The plugin caches decisions with a configurable expiry, and that expiry is the most consequential number in the integration. It sets how much authorization traffic the brokers generate and bounds how long a policy change takes to reach them. A one second expiry makes changes feel immediate and puts the decision service on the hot path of a busy cluster. A five minute expiry does the reverse. Size against the cache miss rate rather than the request rate, and treat the expiry as the staleness window you are choosing.

The failure mode is the other question, and it has to be settled before the first deployment, not during the first incident. The plugin exposes a setting for what happens when the endpoint is unreachable. Deny, and an unavailable decision service halts production traffic. Allow, and it quietly removes access control from the cluster. This is the first objection any Kafka operator will raise.

The strongest answer is not better uptime, it is fewer live dependencies. The same German work is designed around decision points that carry on deciding when the components feeding them are unavailable, holding policy and attribute state locally and synchronizing by polling rather than depending on a live call. Applied to a broker, the decision service answers from local state, with no synchronous call to the identity provider or the policy control plane on the request path. Failing closed is tolerable once there is no runtime dependency left to lose.

A third cost only appears at scale. Policy quality depends on attributes being complete, accurate and available at decision time, so centralising every attribute can move the bottleneck out of authorization and into the data layer. Where the information lives has to be designed alongside where the decision point runs, not after it.

Two smaller boundaries are worth naming. The superuser list stays in broker configuration and short circuits the check before any policy is consulted, so it remains part of your access model whether or not you think of it that way. And the scope here is operations on topics, groups and the cluster, not the contents of messages. A principal permitted to read a topic reads all of it. Field level control over payloads is a different problem, solved at the producer or in a schema aware processing layer.

What is left on the cluster

Kafka did not need a new authorization model. It has a hook designed from the start to be answered outside the broker, and the reason most clusters answer it from a local ACL list is that a local ACL list was the only thing on offer.

Moving the answer out leaves one plugin configuration on the broker and a decision service to reach. Everything past that point is policy the organization already runs, reviewed by the people who already review it, recorded where the rest of the record lives. Which team can write to which topic gets an answer that does not require anyone to go and look.

Try Cerbos to see how this works in practice, or book a call to talk through your architecture with the team.

Go deeper:

FAQ

Free policy workshop

Get your first Cerbos policy written by our team.

Book a session to talk through your requirements and walk away with a working policy.

Book a session