---
title: "Kafka topic authorization belongs in your policy set, not in a per cluster ACL list"
description: "Kafka ACLs accumulate per cluster. This guide covers how Kafka's pluggable authorizer hands topic access decisions to an external policy engine, how ACL bindings map onto resource policies, caching and fail closed tradeoffs on the broker hot path, and what moves topic access into one audited policy set."
author: "Alex Olivier"
date: "2026-09-18T13:27:00.000Z"
canonical: "https://www.cerbos.dev/blog/kafka-topic-authorization-belongs-in-policy-set"
image: "https://stylish-appliance-1c1cc1c30d.media.strapiapp.com/Kafka_topic_authorization_belongs_in_your_policy_set_not_in_a_per_cluster_ACL_list_64fa4cfb72.png"
tags: ["engineering","guide"]
source: "https://www.cerbos.dev/blog/kafka-topic-authorization-belongs-in-policy-set"
---

# Kafka topic authorization belongs in your policy set, not in a per cluster ACL list

Ask a platform team which services can write to a given Kafka topic and the honest answer is usually that somebody will have to go and look. The answer lives in an ACL list on the cluster, expressed as bindings between a principal, an operation and a resource pattern, and it has been accumulating for as long as the cluster has been running. Every binding was correct on the day it was added. Almost none of them were ever removed.

That list also has no relationship to how access is described anywhere else in the organization. The application knows about teams, services and data classifications. The ACL list knows about `User:payments-writer` and a literal topic name. Translating between them is a job that lives in somebody's head and does not get handed over when they move on. Add a second cluster and there are two lists, drifting apart on their own schedule.

The general shape of this is older than Kafka. NIST's [guide to ABAC](https://csrc.nist.gov/pubs/sp/800/162/upd2/final) draws the contrast directly, noting that access control lists depend on static membership that is re-evaluated only when something notifies the system. Static models assume stable roles, coarse permissions and a decision made once. A long lived cluster's ACL list encodes that assumption and then outlives it.

So this is not a failure of Kafka's authorization model. It is what happens when access rules live in a system that has no idea what your organization looks like. Kafka never insisted on owning them, which makes this a wiring problem rather than a policy authoring one, and wiring problems are what [external authorization](https://www.cerbos.dev/blog/why-external-authorization) is for.

## Kafka's authorizer is a plugin interface, not a fixed implementation

The broker decides authorization through a class named in `authorizer.class.name`. Kafka ships an implementation backed by its own ACL store, and that is the one most clusters run, but the [Authorizer interface](https://kafka.apache.org/39/javadoc/org/apache/kafka/server/authorizer/Authorizer.html) is the actual contract. The broker starts the configured authorizer, waits for it to report ready before it accepts connections on a listener, then calls it on the request thread for every operation a client attempts.

There is an established plugin for that interface which does not answer from a local store at all. It delegates each decision over HTTP using the request and response shapes established by the [Open Policy Agent](https://www.openpolicyagent.org/) project, and reads the verdict out of the response body. Once a broker is configured that way, where the rules live stops being a Kafka question. The endpoint at the other end can be anything that speaks the protocol, including a [Cerbos deployment](https://www.cerbos.dev/ecosystem/cerbos-kafka) that already governs the rest of the stack.

That division of labour is what reference architectures deliberately design towards. Germany's federal API authorization [blueprint](https://gitlab.opencode.de/sachsen-anhalt/mid/foederale-api-autorisierungsinfrastruktur), published open source by Saxony-Anhalt with FITKO, is built around a strict split between the component that enforces and the component that decides, with the enforcing side kept deliberately thin. A broker authorizer plugin is close to a textbook instance, since posing the question and acting on the answer is all it knows how to do.

## What a Kafka authorization request contains, principal, operation and resource

Three things, and no more. The principal, as authenticated by the listener, which is a SASL username or the subject of a client certificate. The operation, drawn from a fixed set that includes `READ`, `WRITE`, `CREATE`, `DELETE`, `ALTER`, `DESCRIBE`, `ALTER_CONFIGS` and [several more](https://kafka.apache.org/40/javadoc/org/apache/kafka/common/acl/AclOperation.html). And the resource, which is a `TOPIC`, `GROUP`, `CLUSTER`, `TRANSACTIONAL_ID`, `DELEGATION_TOKEN` or `USER`, carrying a name and whether that name is a literal or a prefix.

That shape maps onto a Cerbos check almost directly. The resource type becomes the resource kind, so each Kafka resource type gets its own [resource policy](https://docs.cerbos.dev/cerbos/latest/policies/resource_policies.html). The operation becomes the action, so rules address `WRITE` and `ALTER_CONFIGS` by their Kafka names rather than through a translation table somebody has to keep current. The principal name becomes the principal ID, and the connection metadata the broker already sends becomes principal attributes.

[Cerbos Synapse](https://www.cerbos.dev/blog/introducing-cerbos-synapse-unified-authorization-context-and-enforcement-across-your-stack) serves that endpoint with a [route extension](https://docs.cerbos.dev/synapse/latest/extensions/route-extensions.html), which claims a path, reshapes the raw payload into a check and formats the decision back into the body the plugin expects. That reshaping is a declarative mapping of expressions rather than a compiled adapter, so there is no plugin code to build and no second policy language in the stack.

## Kafka brokers, service meshes and Kubernetes ask the same authorization question

The reason this generalises is that the interface has already been settled. The OpenID Foundation's [Authorization API](https://openid.net/specs/authorization-api-1_0.html) reached Final Specification in January 2026 and standardises how an enforcement point asks a decision point a question and reads the answer back. Its information model is a subject, an action, a resource and a context, the same four part shape Kafka's authorizer contract already produces. Decisions default to closed, and the spec is strict that a transport error and a policy denial are different things.

What that settles is the protocol rather than the architecture. The specification is explicit that the policy language and the state management of a decision point sit outside its scope. Where the attributes come from, what is cached and what happens when the source is unavailable is left to whoever builds it, and that gap is most of the work in a real deployment, whether the enforcement point is a broker, a query engine or an API.

Which is why the pattern repeats across the estate instead of being a Kafka trick. Companion pieces cover [row filtering](https://www.cerbos.dev/blog/row-level-security-for-apache-trino) in Trino, [mesh enforcement](https://www.cerbos.dev/blog/istio-authorization-stops-at-identity) with Istio and [admission control](https://www.cerbos.dev/blog/kubernetes-admission-control-does-not-need-a-second-policy-language) in Kubernetes, each answering on the protocol the component already speaks.

## Kafka topic naming conventions become authorization rules

The consequence worth dwelling on is what the rules can now say. An ACL binding is an enumeration, and prefixed patterns aside, a binding still has to exist for every combination somebody wants to permit.

A policy evaluates a condition instead. The topic name and the principal name both arrive as ordinary strings, so the relationship between them is something the policy computes rather than something an operator enumerates. If a team owns the topics carrying its prefix, that convention becomes one rule covering every topic that exists and every topic that will exist.

In practice that reads as a set of [derived roles](https://docs.cerbos.dev/cerbos/latest/policies/derived_roles.html) turning the shape of a principal name into a capability, so a name ending in `-consumer` gains a consumer role, alongside a [local variable](https://docs.cerbos.dev/cerbos/latest/policies/conditions.html) computing the prefix that principal owns. A rule then grants `READ` on a topic to the consumer role when the topic name starts with that prefix. Six resource policies cover the six Kafka resource types, and a topic created tomorrow inherits its rules with nothing to provision.

A naming convention is only the cheapest available attribute, not the only one. The same condition could compare a topic against ownership recorded in a service catalogue, a data classification tag, or group membership resolved from the identity provider at decision time. Assembling those attributes before evaluation is a named role in the same NIST model rather than an implementation detail, the context handler, and it is the position Synapse occupies here.

NIST also names the payoff. So long as a new subject carries the attributes a rule needs, neither the rule nor the object has to change to accommodate it. In Kafka terms, a team onboarded next quarter and every topic it creates are covered already. That is where this stops being purely [role based](https://www.cerbos.dev/blog/rbac-vs-abac) and starts describing intent.

## Kafka topic access across clusters from one policy set and one decision log

The rules are policy files, reviewed in the same repository and the same process as the rules governing the API. A second cluster becomes another enforcement point reading the same policy rather than another list to maintain, and promoting a change is a policy version rather than a run of the ACL CLI against a different bootstrap server. Where an estate already expresses [tenant isolation](https://www.cerbos.dev/blog/authorization-challenges-in-a-multitenant-system) somewhere, the broker honours the same definition of a tenant as everything else.

The record changes too. Every broker decision lands in the same [decision log](https://docs.cerbos.dev/cerbos/latest/configuration/audit.html) as every application check, carrying the principal, the resource, the action and the rule that decided. Answering who could write to a settlements topic last quarter becomes a query against one log rather than a reconstruction from ACL change history and application logs that identify people differently. The [audit trail](https://www.cerbos.dev/blog/why-audit-logs-are-important) is where the per cluster model fails hardest, because ACL bindings record their current state and not the decisions they produced. That matters more as event streams become where automated activity is recorded and replayed, turning who may read which topic into a governance question.

## Caching, fail closed and the request thread, where external Kafka authorization gets hard

The Kafka javadoc is direct about the cost. `authorize()` is a synchronous call made on the request thread, and implementations "should avoid time-consuming remote communication that may block request threads". A network call is exactly what that warning is about.

Caching is therefore part of the design rather than a tuning exercise. The plugin caches decisions with a configurable expiry, and that expiry is the most consequential number in the integration. It sets how much authorization traffic the brokers generate and bounds how long a policy change takes to reach them. A one second expiry makes changes feel immediate and puts the decision service on the hot path of a busy cluster. A five minute expiry does the reverse. Size against the cache miss rate rather than the request rate, and treat the expiry as the staleness window you are choosing.

The failure mode is the other question, and it has to be settled before the first deployment, not during the first incident. The plugin exposes a setting for what happens when the endpoint is unreachable. Deny, and an unavailable decision service halts production traffic. Allow, and it quietly removes access control from the cluster. This is the first objection any Kafka operator will raise.

The strongest answer is not better uptime, it is fewer live dependencies. The same German work is designed around decision points that carry on deciding when the components feeding them are unavailable, holding policy and attribute state locally and synchronizing by polling rather than depending on a live call. Applied to a broker, the decision service answers from local state, with no synchronous call to the identity provider or the policy control plane on the request path. Failing closed is tolerable once there is no runtime dependency left to lose.

A third cost only appears at scale. Policy quality depends on attributes being complete, accurate and available at decision time, so centralising every attribute can move the bottleneck out of authorization and into the data layer. Where the information lives has to be designed alongside where the decision point runs, not after it.

Two smaller boundaries are worth naming. The superuser list stays in broker configuration and short circuits the check before any policy is consulted, so it remains part of your access model whether or not you think of it that way. And the scope here is operations on topics, groups and the cluster, not the contents of messages. A principal permitted to read a topic reads all of it. Field level control over payloads is a different problem, solved at the producer or in a schema aware processing layer.

## What is left on the cluster

Kafka did not need a new authorization model. It has a hook designed from the start to be answered outside the broker, and the reason most clusters answer it from a local ACL list is that a local ACL list was the only thing on offer.

Moving the answer out leaves one plugin configuration on the broker and a [decision service](https://www.cerbos.dev/product-cerbos-synapse) to reach. Everything past that point is policy the organization already runs, reviewed by the people who already review it, recorded where the rest of the record lives. Which team can write to which topic gets an answer that does not require anyone to go and look.

[**Try Cerbos**](https://hub.cerbos.cloud) to see how this works in practice, or [**book a call**](https://www.cerbos.dev/workshop) to talk through your architecture with the team.

**Go deeper:**

- [Building a scalable authorization system, a step-by-step blueprint](https://solutions.cerbos.dev/building-a-scalable-authorization-system) (eBook) for a structured approach to designing enforcement points across services and infrastructure rather than one component at a time

## FAQ

### Can Apache Kafka delegate authorization decisions to an external service?

Apache Kafka can delegate authorization decisions to an external service because the broker's authorizer is a plugin interface set in broker configuration, not a fixed implementation. Kafka ships an authorizer backed by its own ACL store, but any class implementing the Authorizer interface can answer instead, including a plugin that sends each decision over HTTP to a policy decision point. The Cerbos authorization management platform serves that endpoint through [Cerbos Synapse](https://www.cerbos.dev/ecosystem/cerbos-kafka), so topic, group and cluster operations are decided by the same policies that govern the rest of the stack.

### What is the alternative to Kafka ACLs for controlling topic access?

The alternative to Kafka ACLs for controlling topic access is a policy that evaluates a condition instead of enumerating bindings. An ACL list needs an entry for every principal, operation and resource pattern someone wants to permit, and those entries are rarely removed. A policy receives the principal name and topic name as [attributes](https://www.cerbos.dev/blog/rbac-vs-abac) and computes the relationship between them, so a rule saying a team may read the topics carrying its prefix covers every topic that exists and every topic created later.

### How do Kafka ACL concepts map onto an authorization policy?

Kafka ACL concepts map onto an authorization policy almost directly. The Kafka resource type, such as TOPIC, GROUP or CLUSTER, becomes the resource kind, the operation, such as READ, WRITE or ALTER\_CONFIGS, becomes the action, and the authenticated principal name becomes the principal ID. In Cerbos that means one [resource policy](https://docs.cerbos.dev/cerbos/latest/policies/resource_policies.html) per Kafka resource type, with rules that address operations by their Kafka names, so there is no translation table to maintain alongside the cluster.

### Does external authorization add latency to Kafka brokers?

External authorization adds latency to Kafka brokers unless decisions are cached, because the broker calls the authorizer synchronously on the request thread for every operation a client attempts. Authorizer plugins that call out over HTTP cache decisions with a configurable expiry, and that expiry sets both how much authorization traffic the brokers generate and how long a policy change takes to reach them. Size the decision service against the cache miss rate rather than the raw request rate, and treat the expiry as the staleness window you are choosing.

### What happens to a Kafka cluster if the external authorization service is unreachable?

What happens to a Kafka cluster when the external authorization service is unreachable depends on a plugin setting that has to be decided before the first deployment. Failing closed halts production traffic until the service is back, while failing open removes access control from the cluster. The stronger answer is fewer live dependencies rather than better uptime, so the decision service should hold policy and attribute state locally and answer without a synchronous call to an identity provider or control plane on the request path.

### How do you audit who had access to a Kafka topic?

Auditing who had access to a Kafka topic is hard with ACLs alone, because ACL bindings record their current state and not the decisions they produced, so the answer has to be rebuilt from ACL change history and application logs that identify people differently. When broker decisions are made by an external policy engine, every decision lands in the same [decision log](https://docs.cerbos.dev/cerbos/latest/configuration/audit.html) as every application check, carrying the principal, the resource, the action and the rule that decided. Answering who could write to a settlements topic last quarter becomes a query against one [audit trail](https://www.cerbos.dev/blog/why-audit-logs-are-important).
