Skip to content

Best Kafka management tools for energy and utility companies

Comparisons
Chad Harris·October 1, 2026·24 min read·Updated

The best Kafka management tool for a grid operator or utility is one that the operator can run inside its own segmented network without adding anything to the path that carries market, dispatch and metering data: one component with no external database, no vendor proxy and no need for the internet, an audit trail that names the person behind every action and data query, production access granted for a task and then removed, and one tool across self-managed clusters on Kubernetes, on-prem data centres and the cloud, signed in through the operator’s directory. Kpow, Kafbat UI, AKHQ, StreamsHub Console, Lenses and Conduktor each cover part of that. Scored on the six weighted criteria explained below the rankings, Kpow ranks first with 89 out of 100, ahead of Kafbat UI at 67 and AKHQ at 61.

Tools compared

Kafka management tools for energy and utility companies scored against this page’s rubric (read 1 October 2026). Total is the weighted score out of 100, with the criteria in order of weight; the weights are explained under how these tools were scored. Conduktor is listed last whatever its total; on its total of 59 it would place fourth.
Rank Tool Total (out of 100) Out of the data path Audit trail per person Production access on request On-prem and cloud together Runs on Kubernetes Directory and Kafka sign-in Cost a year (modelled)
1 Kpow 89 One container, state in your Kafka, not a proxy, runs offline Every action with the IdP user, data queries included Time-boxed temporary policies via API, staged approvals Any distribution, 12 clusters per instance Helm chart, one pod, liveness and readiness endpoints SAML, OIDC, LDAP; mTLS, SCRAM, Kerberos or OAuth to brokers $20,880
2 Kafbat UI 67 One stateless container Optional, reads at level ALL, no view Per-resource RBAC, no approvals, no expiring grant Confluent Cloud broke in v1.4.x and v1.5.0 Helm chart, clusters in static configuration OAuth2, OIDC, LDAP; no SAML $11,520
3 AKHQ 61 One stateless container Opt-in, no reads, no view Regex groups, UI-only if JWT secret unset Named connections Helm chart and YAML, memory-growth reports LDAP, OIDC; no SAML $16,320
4 StreamsHub Console 56 Operator and console, no database None documented OIDC roles, no approvals or masking Any Kafka by properties, Kubernetes only Own operator and Console resource OIDC only $8,640
5 Lenses 54 HQ on PostgreSQL, agent and database per cluster In-product audit log from Team tier Strict global masking, no approvals Any Kafka API, agent per cluster Helm charts, a PostgreSQL for HQ and each agent SSO incl. Keycloak, Entra ID and Okta $2,880 plus quoted licence
6 Conduktor 59 Console on PostgreSQL; Gateway proxy in the data path 70+ event types with user, in the UI Per-viewer masking, cross-team access requests Confluent Cloud, Aiven, MSK, Cloudera Kubernetes deployments, plus PostgreSQL LDAP, OIDC $122,880; $212,880 with Gateway Core and Protect

No tool meets every column, so a grid operator or utility pairs a management tool for people with broker ACLs or an authorizer for services, and with network segmentation that keeps the tool on the same side as the clusters it manages.

The tools, ranked for energy and utility companies

Rank 1

89 out of 100 Total

Try Kpow in the live demo No signup needed.

Cost a year
$18,000 licence for 4 clusters plus $2,880 operator time, so $20,880 (modelled)
Deployment
One container, JAR or Helm chart, no external database, runs fully offline
Sign-in
SAML, OIDC, LDAP; mTLS, SCRAM, Kerberos or OAuth to brokers
Out of the data path ×3 weight, this criterion counts 3 times toward the total
9 out of 10
Audit trail per person ×2 weight, this criterion counts 2 times toward the total
9 out of 10
Production access on request ×2 weight, this criterion counts 2 times toward the total
9 out of 10
On-prem and cloud together
8 out of 10
Runs on Kubernetes
9 out of 10
Directory and Kafka sign-in
9 out of 10
Why these scores for Kpow
Out of the data path 9 out of 10
It is one container or JAR whose state lives in Kafka topics on your own cluster, and it connects as an ordinary Kafka client, so nothing sits between your applications and the brokers.
Audit trail per person 9 out of 10
Every action is recorded with the user from the identity provider and the policy that allowed it, including data inspect queries, with a seven-day view in the product, the record written to an audit topic on your own cluster, and webhooks that send it to a SIEM for long-term retention.
Production access on request 9 out of 10
Temporary policies grant time-boxed access that an admin or a change system calling the Kpow API can create, staged mutations hold any action for approval, and data policies mask fields in inspection, though masking is per resource rather than per viewer.
On-prem and cloud together 8 out of 10
One deployment manages self-managed Apache Kafka, Confluent Platform, Confluent Cloud and MSK together, capped at 12 clusters per instance before you run another. Its documentation asks for it to run close to its clusters and does not officially support multi-region installations.
Runs on Kubernetes 9 out of 10
Factor House publishes Helm charts for Kpow and Community Edition at charts.factorhouse.io, it runs as one pod with no database, and its /healthy endpoint serves as a liveness probe and /up as a readiness probe.
Directory and Kafka sign-in 9 out of 10
People sign in with SAML, OpenID or LDAP through Jetty JAAS, and Kpow connects to brokers with any SASL mechanism, GSSAPI by default, or SSL.

For a grid operator or utility. Kpow runs inside the operator’s own network and gives every engineer a governed way into Kafka without adding anything to the path that carries operational data. It is one container or JAR whose snapshots, metrics and audit log live in topics on the operator’s own cluster, it installs on Kubernetes from the Helm charts, and Factor House states on the Kpow product page that it runs fully offline, the same way in an air-gapped network as anywhere else. Temporary policies grant production access for a fixed time, staged mutations hold a change until a second person approves it, and the audit log records each action and data query with the person who took it.

Where it falls short. Kpow governs people working through Kpow. Applications still authenticate to the brokers with their own principals, so broker ACLs or an authorizer remain the control for services. Kpow is designed to run as a single instance, restarted by the platform through its liveness endpoint; for failover its documentation describes an active/passive pair, in which the passive instance takes over without the active instance’s last hour of history. Its documentation also asks for it to run close to its clusters and does not officially support multi-region installations, so sites in different regions each get their own instance. The in-app audit view covers seven days, and Kpow’s audit topic keeps one week of records by default, so a longer record needs the topic’s retention raised or a webhook into the operator’s SIEM. Single sign-on, RBAC, temporary policies, staged mutations, the audit log, webhooks and Prometheus endpoints are Enterprise features; Community Edition is free for 3 clusters and 10 users.

Cost a year. $20,880 on this page’s model of a grid operator or utility running 4 clusters (development, test, pre-production and production) for 100 engineers. Kpow Enterprise is published at $4,500 per cluster per year with 100 users included, so the licence is $18,000, and the model adds 2 engineer-hours a month at $120 an hour, $2,880, to run one container and keep it current. With 100 users included per cluster, the bill does not grow with headcount at this size. Kpow is also sold on AWS Marketplace as Kpow for Apache Kafka (Annual), for a utility that runs part of its Kafka on AWS.

Rank 2

67 out of 100 Total

Cost a year
$0 licence, about $11,520 in operator time (modelled)
Sign-in
OAuth2, OIDC, LDAP or Active Directory; no SAML
Deployment
One stateless container
Out of the data path ×3 weight, this criterion counts 3 times toward the total
9 out of 10
Audit trail per person ×2 weight, this criterion counts 2 times toward the total
6 out of 10
Production access on request ×2 weight, this criterion counts 2 times toward the total
4 out of 10
On-prem and cloud together
6 out of 10
Runs on Kubernetes
7 out of 10
Directory and Kafka sign-in
7 out of 10
Why these scores for Kafbat UI
Out of the data path 9 out of 10
It is one stateless container with no database and no proxy, the same pass as Kpow.
Audit trail per person 6 out of 10
Its audit log names the logged-in user and records reads when the level is set to ALL, but it writes to a topic or the console with no view in the product, so reading the trail is something you build.
Production access on request 4 out of 10
RBAC grants actions per resource and a cluster can be set read-only, but there is no approval step, no time-boxed grant, and its masking applies the same way to every viewer.
On-prem and cloud together 6 out of 10
It covers self-managed Kafka, MSK and other managed services, but Confluent Cloud connectivity broke in v1.4.x and v1.5.0.
Runs on Kubernetes 7 out of 10
Its Helm chart is maintained and documented, but the Kafbat UI review records that adding clusters through the UI on Kubernetes returns 400 Bad Request, so cluster connections stay in static configuration.
Directory and Kafka sign-in 7 out of 10
It supports OAuth2 and OIDC, including Microsoft Entra ID, and LDAP or Active Directory, and its documentation does not list SAML.

For a grid operator or utility. Kafbat UI is the maintained open-source fork of the original kafka-ui, Apache 2.0, with free RBAC, server-side remove, replace and mask policies, and an optional audit log. It stays out of the data path the same way Kpow does, needs nothing outside the operator’s network, and for a small team on one set of clusters it covers day-to-day inspection and topic work at no licence cost.

Where it falls short. There is no way to grant production access for a task and have it expire, and no approval before a change runs. Reading its audit trail means building a consumer first, which is a gap when an incident review asks who read a topic. Its last release, v1.5.0, shipped in April 2026. There is no SLA, and paid help is a professional services engagement from the maintainers, quoted rather than listed.

Cost a year. $11,520 on this page’s estimate, with no licence fee. Running, securing and upgrading it is 6 engineer-hours a month at $120 an hour, $8,640, and a utility that signs people in with SAML also runs a proxy such as oauth2-proxy in front of it, 2 hours a month, $2,880.

Rank 3

AKHQ

akhq.io

61 out of 100 Total

Cost a year
$0 licence, about $16,320 in operator time and review (modelled)
Sign-in
LDAP, OIDC, header auth; no SAML
Deployment
One stateless container
Out of the data path ×3 weight, this criterion counts 3 times toward the total
9 out of 10
Audit trail per person ×2 weight, this criterion counts 2 times toward the total
4 out of 10
Production access on request ×2 weight, this criterion counts 2 times toward the total
3 out of 10
On-prem and cloud together
7 out of 10
Runs on Kubernetes
7 out of 10
Directory and Kafka sign-in
6 out of 10
Why these scores for AKHQ
Out of the data path 9 out of 10
It is one stateless container with no database and no proxy, the same pass as Kpow.
Audit trail per person 4 out of 10
Audit events are opt-in to a Kafka topic, reads are not recorded, and there is no view for the trail.
Production access on request 3 out of 10
Groups bind actions to resources by regex, but there is no approval step or time-boxed grant, masking is global, and without the JWT signing secret the restriction is in the UI only.
On-prem and cloud together 7 out of 10
Each cluster is a named connection, with Confluent Cloud and MSK IAM examples in its documentation.
Runs on Kubernetes 7 out of 10
Its configuration is YAML deployed through a Helm chart, which fits a GitOps workflow, docked for the memory-growth reports at heaps up to 14 GB that the AKHQ review records with no published fix.
Directory and Kafka sign-in 6 out of 10
It supports LDAP, OIDC and header authentication from a proxy, does not list SAML, and ships with security disabled until you enable it.

For a grid operator or utility. AKHQ is free under Apache 2.0, configured in YAML that fits a GitOps review, with roles that combine resource types and cluster patterns, and it browses topic data well. It needs no database and nothing outside the operator’s network. Its latest release, 0.28.0, shipped in August 2026.

Where it falls short. Its documentation warns that if the JWT signing secret is not set, the API will not enforce the group role, so a misconfiguration turns access control into a UI restriction. Audit is opt-in and reads are not in it, so it cannot show who looked at a production topic during an incident, and there is no approval step or expiring grant.

Cost a year. $16,320 on this page’s estimate, with no licence fee. Running, securing and upgrading it is 6 engineer-hours a month at $120 an hour, $8,640; a utility that signs people in with SAML runs oauth2-proxy in front of it, 2 hours a month, $2,880; and the model adds one access review a year, 40 hours or $4,800, because of the JWT secret behaviour above.

Rank 4

StreamsHub Console

github.com/streamshub/console

56 out of 100 Total

Cost a year
$0 licence, about $8,640 in operator time (modelled)
Sign-in
OIDC provider such as Keycloak or Dex
Deployment
Operator plus a Console resource, no database; release 0.14.1
Out of the data path ×3 weight, this criterion counts 3 times toward the total
9 out of 10
Audit trail per person ×2 weight, this criterion counts 2 times toward the total
2 out of 10
Production access on request ×2 weight, this criterion counts 2 times toward the total
3 out of 10
On-prem and cloud together
5 out of 10
Runs on Kubernetes
9 out of 10
Directory and Kafka sign-in
5 out of 10
Why these scores for StreamsHub Console
Out of the data path 9 out of 10
It runs as a REST API and a UI managed by its own operator, with no database and no proxy, and Prometheus is needed only for its metrics charts.
Audit trail per person 2 out of 10
Its documentation describes no audit trail of what each person read or changed.
Production access on request 3 out of 10
OIDC sign-in maps users and groups to roles that allow listed actions on topics, records, groups and rebalances per cluster, but there is no approval step, no time-boxed grant and no masking.
On-prem and cloud together 5 out of 10
It connects to any Kafka cluster given connection properties in its Console resource, with extra features for clusters the Strimzi operator manages, but its documented deployments are all on Kubernetes and it gives no connection examples for managed cloud services.
Runs on Kubernetes 9 out of 10
It installs through its own operator, from Operator Lifecycle Manager or plain manifests, and is configured with a Console custom resource that can live in Git beside the Kafka resources.
Directory and Kafka sign-in 5 out of 10
Its documentation describes sign-in through an OIDC provider such as Keycloak or Dex only, with roles assigned from group claims, so an LDAP or SAML directory sits behind that provider; it connects to brokers as a Strimzi KafkaUser or with its own connection properties.

For a grid operator or utility. StreamsHub Console is an Apache 2.0 web console for Kafka, built from a Quarkus REST API, a Next.js UI and a Kubernetes operator, per its GitHub repository, and it adds the most when the clusters are managed by the Strimzi Cluster Operator. It shows topics and their messages, Kafka nodes including KRaft, and consumer groups with offset changes, and it installs through its own operator, from Operator Lifecycle Manager or plain Kubernetes resources. People sign in through an OIDC provider such as Keycloak or Dex. The Strimzi-specific detail is on best Kafka UI tools for Strimzi.

Where it falls short. It is pre-1.0 software, at release 0.14.1 in September 2026, and its Console resource is still a v1alpha1 API. There is no approval step, no time-boxed access, no masking and no documented audit trail, so it cannot answer who read or changed a production topic. Outside Kubernetes it has no documented production deployment.

Cost a year. $8,640 on this page’s estimate, with no licence fee: 6 engineer-hours a month at $120 an hour to run the operator and the console, secure them and keep them current.

Rank 5

Lenses

lenses.io

54 out of 100 Total

Cost a year
$2,880 operator time, plus a licence quoted above 15 users (modelled)
Sign-in
SSO with Okta, Keycloak, OneLogin, Google, Entra ID
Deployment
HQ on PostgreSQL, an agent and database per cluster
Out of the data path ×3 weight, this criterion counts 3 times toward the total
4 out of 10
Audit trail per person ×2 weight, this criterion counts 2 times toward the total
7 out of 10
Production access on request ×2 weight, this criterion counts 2 times toward the total
4 out of 10
On-prem and cloud together
7 out of 10
Runs on Kubernetes
6 out of 10
Directory and Kafka sign-in
7 out of 10
Why these scores for Lenses
Out of the data path 4 out of 10
It runs a central HQ on PostgreSQL plus an agent and an agent database beside every cluster, and HQ has no high-availability option.
Audit trail per person 7 out of 10
Audit logs can be read in the product, with no need to build a consumer first.
Production access on request 4 out of 10
Its masking is the strictest view-time model, global with no escape even for admins, but no approval step or time-boxed grant is described.
On-prem and cloud together 7 out of 10
It connects to any provider exposing a Kafka-compatible API, one agent per cluster.
Runs on Kubernetes 6 out of 10
Helm charts are published for HQ and for the agent, but each needs a PostgreSQL database you provide, which is more to run inside the cluster than a single pod.
Directory and Kafka sign-in 7 out of 10
SSO spans Okta, Keycloak, OneLogin, Google and Entra ID, with basic authentication only on Community.

For a grid operator or utility. Lenses brings vendor-backed RBAC, SSO, in-product audit logs and SQL Studio for querying topics, which is the reason to choose it if analysts need SQL over Kafka. Its data policies redact by field name across Kafka topics, Postgres tables and Elasticsearch indices.

Where it falls short. Every cluster adds an agent and a database to deploy, patch and clear through a segmented network, and HQ is a single node that every cluster depends on, so a cluster at a site has to reach HQ through the operator’s network controls. Its policies apply to Lenses interfaces only, and no approval step or expiring grant is described.

Cost a year. $2,880 of operator time on this page’s estimate, 2 hours a month at $120 an hour, plus a licence that is not published. The published Team Edition is $4,000 a year for up to 15 users on one cluster, so 100 engineers across 4 clusters is Multi-Kafka Enterprise at a custom quote.

Rank 6

Conduktor

conduktor.io

59 out of 100 Total

Cost a year
100 seats at $1,200 plus $2,880 operator time, so $122,880; Gateway Core adds $60,000 and Gateway Protect, which carries encryption and masking, a further $30,000 (modelled)
Sign-in
LDAP, OIDC; no SAML described
Deployment
Console on PostgreSQL 13+; data-level controls through Gateway, a proxy
Out of the data path ×3 weight, this criterion counts 3 times toward the total
3 out of 10
Audit trail per person ×2 weight, this criterion counts 2 times toward the total
8 out of 10
Production access on request ×2 weight, this criterion counts 2 times toward the total
6 out of 10
On-prem and cloud together
8 out of 10
Runs on Kubernetes
7 out of 10
Directory and Kafka sign-in
7 out of 10
Why these scores for Conduktor
Out of the data path 3 out of 10
Console needs PostgreSQL 13 or later, and its encryption, data-level masking and Virtual Clusters only work when client traffic goes through Gateway, a proxy in the data path.
Audit trail per person 8 out of 10
Console logs produce, consume and admin requests across more than 70 event types with user, IP and timestamp, browsable in the UI and exported as CloudEvents.
Production access on request 6 out of 10
Masking can exempt users or groups, which beats every other tool here on who sees unmasked data, and cross-team access requests are approved by the owning team, but no expiring grant is described and topic creation that passes policy is a direct API call.
On-prem and cloud together 8 out of 10
Its cluster configuration covers Confluent Cloud, Aiven, Amazon MSK and Cloudera, and Console works across clusters.
Runs on Kubernetes 7 out of 10
Console and Gateway each have documented Kubernetes deployments, and Console needs a PostgreSQL database alongside it.
Directory and Kafka sign-in 7 out of 10
Its SSO configuration covers LDAP and OIDC, with guides for Okta, Entra ID and Keycloak, and does not describe SAML.

For a grid operator or utility. Conduktor pairs Console, a web UI, with Gateway, a Kafka protocol proxy. Conduktor’s Gateway documentation describes it as “a Kafka-compliant middle layer between clients and Kafka clusters”, and that is where field encryption, masking of the data itself, policy enforcement on client traffic and Virtual Clusters are applied. Console alone connects to clusters directly and masks in its UI, with exemptions per user or group. The full picture is in the Conduktor review.

Where it falls short. Its data-level controls apply only to the applications that connect through Gateway, so using them puts a vendor’s proxy between those applications and the brokers. A proxy adds a network hop that can slightly increase end-to-end latency, and it is a tier the operator has to size, keep available and clear through its critical-infrastructure review. Console also needs its own PostgreSQL. Per-seat pricing grows with every engineer who needs access.

Cost a year. $122,880 on this page’s model of 100 engineers. Conduktor’s published Team Edition price is $1,200 a seat a year, $120,000, and the model adds 2 engineer-hours a month at $120 an hour, $2,880. On AWS Marketplace, Conduktor Enterprise lists Gateway Core, which carries Virtual Clusters and policy enforcement, at $60,000 a year and Gateway Protect, the add-on for encryption and masking, at a further $30,000, so the data-level controls take the total to $212,880. Conduktor prices Gateway per cluster with a 3-cluster minimum, and the listing does not say how many clusters that figure covers.

What energy and utility companies need from a Kafka management tool

This page is about companies that run the power system and its networks: transmission grid owners and system operators, distribution network operators, generators and retail utilities, and the gas and water utilities that run Kafka under similar rules. Energy trading desks share the market data but answer to market regulators rather than critical-infrastructure regimes, so they are covered in best Kafka management tools for energy trading firms. For Kafka run on Kubernetes by the Strimzi operator, best Kafka UI tools for Strimzi goes deeper, and for self-managed clusters in general see best Kafka tools for self-managed Apache Kafka. For banks, which run Kafka as shared infrastructure under a financial regulator, see best Kafka management tools for banks.

Kafka sits close to the operational core in this sector. At Transpower, which owns New Zealand’s national grid and runs its wholesale electricity market, Red Hat’s case study describes AMQ Streams, the Apache Kafka component of Red Hat AMQ, running on OpenShift and feeding the Market System applications that produce real-time electricity prices. Kai Waehner’s survey of Kafka for smart grid, utilities and energy production describes the common shape: clusters at the edge beside SCADA systems, PLCs and sensors, often disconnected from the cloud or remote data centres, with replication to cloud clusters that aggregate and monitor the assets of the grid. So a utility’s Kafka is often self-managed, sometimes on its own Kubernetes platform, and spread across sites, data centres and the cloud.

The first test for a management tool is the operator’s critical-infrastructure review. In North America, the NERC CIP standards set mandatory cyber security requirements for the bulk electric system, and one of them, CIP-013, requires supply chain risk management controls for BES Cyber Systems. In the EU, the NIS2 Directive, Directive (EU) 2022/2555, covers energy and water among the sectors it treats as vital to the economy and society and introduces risk management measures for them, as Ireland’s National Cyber Security Centre summarises it. In Australia, the Security of Critical Infrastructure Act 2018 applies to 11 sectors, energy among them. None of these names Kafka or any tool, but each makes every component that can reach operational data part of the operator’s security and supplier review. A management tool that runs as one component inside the operator’s own network, with no database of its own, no vendor proxy in front of the brokers and no need to reach the internet, gives the shortest answer. A tool that routes application traffic through a vendor’s proxy, or needs its own database, adds to the review and to what has to be kept running inside a segmented network.

Whether a given Kafka cluster falls inside those rules is the operator’s own categorisation. CIP-002-5.1a describes BES Cyber Assets as those that, “if rendered unavailable, degraded, or misused, would adversely impact the reliable operation of the BES within 15 minutes”, and it adds that redundant systems must not be counted when the test is applied, because “redundancy does not mitigate cyber security vulnerabilities.” A cluster that feeds dispatch or the market system can meet that test, a cluster that carries meter readings for billing may not, and many operators run both kinds. The same test explains why a proxy weighs more heavily here than in most industries. A component that every dispatch message passes through can degrade the feed when it fails, so it is a candidate for the same 15-minute assessment as the cluster it sits in front of, and running it as a highly available pair does not change that assessment. A tool that connects as one more Kafka client can stop without any application noticing.

CIP-013-2 turns the supplier question into a list. Its requirement R1.2 has the operator’s procurement process address six things: notification by the vendor of incidents, coordination of the response to them, notification when remote or onsite access for vendor staff should end, “disclosure by vendors of known vulnerabilities”, “verification of software integrity and authenticity of all software and patches”, and “coordination of controls for vendor-initiated remote access.” Where the software runs decides how many of the six have anything behind them. With a tool that the operator runs itself and the vendor cannot reach, the two access items have nothing to coordinate, and the review comes down to the items about the vendor’s own software: incident notification and response, vulnerability disclosure and software integrity. For Kpow the last two are public: Factor House’s security documentation describes daily scans of each release against the National Vulnerability Database and a dependency report published for every release, and its security policy commits to acknowledging vulnerability reports within 2 to 3 business days.

CIP-005-7 makes the access half of that list concrete. It requires methods for “determining active vendor remote access sessions (including Interactive Remote Access and system-to-system remote access)” and methods for disabling them. A console hosted by its vendor, or an agent that reports to a vendor’s cloud, is system-to-system access of that kind, and the operator has to be able to see it and cut it. A self-hosted tool needs neither method, provided it makes no calls out, which is worth checking for any tool. For Kpow the answer carries one qualification: the server runs fully offline, but the web UI sends product usage analytics from the user’s browser unless a paid licence switches them off, as its data collection documentation describes.

Each component a tool brings is also a line in the operator’s configuration records. CIP-010-4 requires a baseline configuration for each applicable system that lists its operating system, its installed commercial or open-source software with versions, its “logical network accessible ports” and its security patches, and every change that deviates from the baseline has to be authorised and documented, with the baseline updated within 30 calendar days. CIP-007-6 adds a patch cycle, with a tracked source of security patches for each asset and an evaluation of new patches “at least once every 35 calendar days.” At the perimeter, CIP-005-7 requires “inbound and outbound access permissions, including the reason for granting access”, with all other access denied by default. A tool that is one container has one entry in each of those records. A tool with a database of its own adds the database engine, its patch source and its port, a tool with an agent beside every cluster repeats that at every site, and a proxy adds listeners that the firewall rule of every client application has to name. None of this appears in a feature comparison, although all of it is recurring work for the team that keeps the compliance evidence.

How people reach the tool matters as much as how the tool reaches Kafka. Desktop Kafka clients and command-line tools run on each engineer’s workstation or on a jump box and connect to the brokers directly, so every workstation becomes a Kafka client that the perimeter has to admit. CIP-005-7 requires all Interactive Remote Access to go through an Intermediate System “such that the Cyber Asset initiating Interactive Remote Access does not directly access an applicable Cyber Asset”, with multi-factor authentication on every session. A web tool that runs inside the perimeter fits that model, because engineers reach it with a browser through the intermediate system, and only the tool’s own Kafka principal talks to the brokers, with whatever access the cluster’s ACLs give that principal. Kpow’s minimum ACL permissions are documented, and the same page states that Kpow does not read from or write to topics other than its own internal ones as part of normal operation, so topic data is read only when a person runs a query.

The second is a record of who did what. Grid operators and utilities run their systems around the clock, and after an incident the first questions are who could read or change a production topic and who actually did. When people work through a shared tool, the broker only sees the tool’s service account, so only the tool’s own log can name the person, and it needs to cover data reads as well as changes. Production access belongs with it: an engineer diagnosing a stalled pricing or metering feed needs to look at a production topic, so access has to be granted for the task, expire on its own and leave a record, and risky changes such as an offset reset should wait for a second person.

The service account that a shared tool uses is what the CIP standards treat as a shared account once people know its credentials. CIP-007-6 requires operators to “identify individuals who have authorized access to shared accounts”, and on high impact systems CIP-004-7 requires the passwords of shared accounts known to a person to be changed within 30 calendar days of that person leaving or changing role. Where engineers run command-line tools with a team credential, both requirements apply at every departure. Where each engineer signs in to the tool under their own name and the tool alone holds the Kafka credential, nobody outside the platform team knows a shared secret, the list of people with access is a directory group, and a departure is a directory change.

Retention is where a default most often falls short. For the security event logs it covers, CIP-007-6 requires retention “for at least the last 90 consecutive calendar days” where technically feasible, and on high impact systems a review of logged events “at intervals no greater than 15 calendar days.” An operator that holds its Kafka audit trail to the same figures cannot rely on a view that covers a week, so the audit topic’s retention has to be confirmed and set deliberately, or the records forwarded to the operator’s SIEM from the first day. Reporting deadlines make the same point from the other direction. Article 23 of NIS2 requires an early warning “within 24 hours of becoming aware of the significant incident”, an incident notification within 72 hours and a final report no later than one month after that. Australia’s SOCI Act gives 12 hours for an incident with a significant impact and 72 hours for one with a relevant impact, as Ashurst’s summary of the reporting regime sets out. The first report is due before any forensic work is finished, so the question of who read or changed a topic has to be answerable by a search on the same day. A trail that exists only as records on a topic, with a consumer still to be written, uses hours that the deadline does not allow.

Access itself is governed by CIP-004-7, and its wording favours grants that end. The standard requires a process to “authorize based on need”, a check “at least once each calendar quarter that individuals with active electronic access” have authorisation records, and a review at least once every 15 calendar months that accounts, groups and role privileges are still correct. It also requires an individual’s Interactive Remote Access to be removed within 24 hours of a termination, and electronic access that is no longer needed after a transfer to be revoked by the end of the next calendar day. Standing read access to production topics has to be justified again at each of those reviews, for every person who holds it. Access that is requested through the change system, granted for the hours a task needs and removed automatically produces the authorisation record at the moment of the request, and leaves nothing to find at the quarterly check. Roles that follow directory groups meet the 24-hour rule with one directory change, while a tool with local accounts is a second place to remember.

Around that sit the platform constraints. The tool should install on the Kubernetes or OpenShift platform the Kafka already runs on, as one pod with health endpoints the platform can act on, connect with the mutual TLS or SCRAM credentials the clusters already use, sign people in through the operator’s directory, and reach every cluster, at a site, in a data centre or in the cloud, from one deployment.

Segmentation decides how many deployments there are. NIST’s Guide to Operational Technology Security lists among its security objectives a DMZ with firewalls “to prevent network traffic from passing directly between the corporate and OT networks”, and it describes unidirectional gateways, also called data diodes, which let traffic leave the operational network and let none enter. Kai Waehner’s overview of Kafka cluster deployment strategies notes that some industries run edge clusters in safety-critical environments behind exactly that kind of gateway. A management tool speaks the Kafka protocol, which needs traffic in both directions, so a tool on the corporate side cannot manage a cluster behind a diode, and a tool inside a zone should not be given a path out of it. The number of instances therefore follows the zone map, with one instance inside each zone that holds clusters. Kpow’s documented limits point the same way: its system requirements ask for it to run close to its clusters and do not officially support multi-region installations, so an instance for each site or zone is the supported layout as well as the one the network imposes.

Three details then matter in a topology like that. In an instance that manages several clusters, the first configured cluster is the primary cluster that holds Kpow’s internal topics, the audit log included, so it should be a cluster whose replication, retention and access controls suit an audit record, and a small cluster at a remote site is the wrong choice. Older brokers are common in long-lived installations: Factor House’s account of a breaking change in the Kafka producer notes teams still running Kafka 1.0.0, and Kpow works with every version from 1.0.0 onwards, so a site that is upgraded only at planned outages is not left without tooling. And at a site that is sometimes cut off from the centre, topic retention is a deadline. The Kafka documentation says retention.ms “represents an SLA on how soon consumers must read their data”, so retention at the site has to outlast the longest loss of the link that the operator plans for, or readings are deleted before they have been copied to the centre, and a size limit set with retention.bytes on a small disk can remove them sooner than the time limit suggests.

On Kubernetes, an isolated network changes what installing a tool involves. Nodes with no route to the internet pull their images from a private registry, so every image a tool needs has to be mirrored, scanned and approved before it can run, and every chart value reviewed. Kpow adds one image and one chart to that list. The Helm chart sets resource requests equal to limits by default, which places the pod in the Guaranteed class that Kubernetes evicts last when a node runs short of resources, and it can take its configuration from a ConfigMap and a Secret, so a change to the tool is a reviewed change to two manifests, which is the kind of record CIP-010 asks for. The chart also has a preset for running with a read-only root filesystem, a common policy on hardened clusters, and Kpow is published as a Red Hat Universal Base Image build as well as the standard image, for platforms such as OpenShift where UBI-based images are the standard. The chart is listed on Artifact Hub under a verified publisher, where each version carries a security report from a scan of its container image that a reviewer can read before the image is mirrored.

Sign-in has a requirement of its own in an operational network. The same NIST guide recommends “separate authentication mechanisms and credentials for users of the OT network and the corporate network”, so the instance inside an operational zone often cannot use the corporate identity provider, and a cloud identity service is unreachable from a segment with no route out. What is available there is whatever directory or identity provider the operator runs inside the zone, which makes LDAP or LDAPS and self-hosted SAML or OIDC providers the options that count, and a tool that supports only one protocol narrows the operator’s choice of zone design. On the corporate side, large directories bring a different failure. Microsoft documents that Entra ID stops listing a user’s groups in the token once the user belongs to more than 150 groups for SAML or 200 for JWT and adds an overage claim that points at Microsoft Graph instead, so a tool that maps groups to roles without handling that claim can leave such users without their roles; Kpow’s changelog records group overage support for Entra. The multi-factor authentication that CIP-005 requires for remote sessions is enforced at the intermediate system and the identity provider, so the test for the tool is that it has no local accounts that sidestep them.

None of these controls holds if the brokers themselves accept any client. The Kafka documentation states that security is optional and that non-secured clusters are supported, and on a cluster that was deployed years ago without authentication, anyone with a network path to the brokers can bypass every role defined in a management tool. The access control in any tool on this page is therefore only as strong as the brokers’ own authentication and ACLs.

What energy and utility companies use Kpow for

Two current Kpow customers from the energy sector are named here: Transpower, New Zealand’s national grid owner and system operator, and EDF Trading, the wholesale energy trading business of EDF Group. Neither has published what it uses Kpow for, so each card describes the company and, for Transpower, the Kafka platform its own technology partner has described in public. Both run on real-time data: Red Hat describes Transpower’s streaming platform feeding real-time electricity prices, and EDF Trading’s own site names global real time data as key to keeping up with the pace of trading.

  • Transpower

    National grid owner and system operator, New Zealand

    • Grid operator
    • Real-time electricity pricing
    • Kafka on OpenShift

    Public context about the company. How it uses the product has not been published.

    Transpower designs, builds and maintains New Zealand’s national electricity grid, and it also operates the power system and manages the wholesale electricity market in real time. Red Hat’s case study describes 174 substations, 25,000 transmission towers and more than 6,800 miles of lines, and a Market System used in the control centres to keep the power system secure and dispatch prices in real time. Its applications are being modernised as Java microservices on Red Hat OpenShift, and AMQ Streams, the Apache Kafka component of Red Hat AMQ, running on OpenShift provides the streaming data behind the real-time pricing Transpower launched in 2022. Factor House counts Transpower among its Kpow customers.

    Source: Red Hat, Transpower New Zealand success story

  • EDF Trading

    Wholesale energy trader, part of EDF Group, United Kingdom

    • Wholesale energy markets
    • Part of EDF Group

    Public context about the company. How it uses the product has not been published.

    EDF Trading is part of the EDF Group and describes itself as a specialist in the wholesale energy markets, with around 800 employees who connect the EDF Group and third-party customers to the world’s energy markets. It trades wholesale power, natural gas, oil, LPG and environmental products, and LNG through JERA Global Markets, and its own site says that advanced analytics and global real time data will be key to keeping up with the pace of trading. Factor House counts EDF Trading among its Kpow customers.

    Source: EDF Trading, about EDF Trading

How a grid operator or utility runs its Kafka with Kpow

The workflows below are how a grid operator’s or utility’s platform team puts Kpow to work across its Kafka clusters. Each one is built from documented Kpow features.

Install it inside the operational network. Kpow runs as one Docker container or Java JAR, or on Kubernetes with the Helm charts, inside the operator’s own network. It needs no external database, because its snapshots, metrics and audit log live in topics on the operator’s own clusters, and the Kpow product page states that it runs fully offline, so it can run in a network segment that has no route out. Applications keep talking to the brokers directly. Where policy forbids plaintext passwords in configuration, Kpow accepts encrypted configuration values through the open source Shroud library, whose own documentation says encrypted configuration is “not a replacement for secret managers”, so it is the fallback for environments without one.

Let the platform keep it healthy. Kpow serves a liveness endpoint at /healthy, which returns 503 when the instance is unhealthy and needs no user authentication, and a readiness endpoint at /up, so Kubernetes or OpenShift can restart the pod on its own if Kpow loses its connection to a cluster for longer than the timeout, instead of waiting for someone to notice. A healthy pod says nothing about the Kafka cluster, however. A Kubernetes probe tests the container, and in an incident Claritev describes in its case study, Kubernetes reported every pod healthy and Prometheus raised no alert while one broker was not running, which the team saw only in the Kafka-level view. Kpow’s Prometheus endpoints share the UI’s port and are not secured by default, so basic authentication is configured before that port is opened to the monitoring system:

PROMETHEUS_USERNAME and PROMETHEUS_PASSWORD

Keep a passive instance ready. Kpow’s deployment notes recommend one instance, and describe an active/passive model for teams that need a second one. The active instance writes Kpow’s internal topics, and a passive instance runs beside it without writing them, set with:

PERSISTENCE_MODE="audit" (keeps the audit log) or PERSISTENCE_MODE="none"

If the active instance becomes unavailable, the operator fails over to the passive one, which starts without the active instance’s last hour of metric history. In audit mode the Kpow Streams Agent integration is disabled, according to the environment variable reference, so an operator that relies on the Streams Agent views plans for their absence on the passive instance.

Connect with the clusters’ own security. Kpow takes the standard Kafka client security settings, and for clusters run by the Strimzi operator its documentation covers mutual TLS, SCRAM-SHA-512 and OAuth with the certificates and users Strimzi already creates, using the Strimzi build of the Kpow image for OAuth. One instance then reaches self-managed Apache Kafka, Confluent Platform, Confluent Cloud and Amazon MSK together, up to 12 clusters according to the Kpow multi-cluster page, so clusters in the data centre and in the cloud sit in the same view, and a site in another region, or one cut off from the network, runs its own instance.

Sign people in through the directory. People sign in with SAML, including Keycloak and Microsoft Entra ID, with OpenID Connect, or with LDAP, and directory groups map to roles through RBAC, which sets Allow, Deny or Stage per action and resource, so production can stay read-only for everyone outside the operations team. With LDAP, roles are read from directory groups, the connection can use LDAPS and the bind password can be stored encrypted, which suits a directory run inside an operational zone. With SAML, the interval before a user has to authenticate again is one hour by default and is set with:

SAML_SESSION_S

Grant production access for one task. An engineer who needs to inspect a production topic raises a request in the operator’s change system. Once it is approved, that system calls the Kpow API to create a temporary policy that grants inspect access on the named topic for the time the task needs. Temporary policies joined the admin API endpoints in release 94.1, and Self-service Kafka governance with ServiceNow walks through the request and approval flow. The policy expires on its own, is capped at seven days by default, and is recorded in the audit log. Each grant then exists as an approved request with a start and an end, which is the authorisation record a quarterly access review looks for.

Hold risky changes for a second person. With staged mutations, a role can be set to Stage on actions such as resetting a consumer group’s offsets or deleting a topic, so the request waits in Kpow until an administrator approves or denies it, and the full history lands in the audit log. A team that will not give any console write access to an operational cluster can run Kpow read-only on that cluster by granting no mutating actions, because any action without a matching policy is implicitly denied.

Keep the record of who did what. The audit log records each action, data inspect queries included, with the user from the operator’s directory and the policy that allowed it. Kpow shows the last seven days in the product and writes the record to the __oprtr_audit_log topic on the operator’s own cluster, where Kpow’s topics default to one week of retention, so for a longer record webhooks send mutations, queries or both to Slack, Microsoft Teams or any endpoint the operator chooses, such as its SIEM. An operator that applies the 90 days CIP-007 sets for security event logs to this trail treats the SIEM copy, or the topic with its retention confirmed to cover the period, as the record, and the seven-day view as a convenience. The audit record holds the content of each request, search filters included, so a search for a customer’s meter or account number writes that number into the audit topic, and read access to that topic is best limited to the people who review it.

Find the message behind a bad price or reading. Data Inspect searches by key, value or header across one topic or several at once, with kJQ filters run on the server, so an engineer can find the messages for one node, meter or interval without writing a consumer against production. When a reading is missing altogether, the producer’s acknowledgement setting is part of the answer. With acks=0 the producer does not wait for the broker and no guarantee can be made that the record arrived, which is a reasonable trade for high-volume sensor telemetry and the wrong one for dispatch instructions or readings used in settlement.

Watch Kafka Streams applications and lag. Applications built on Kafka Streams can register with the Kpow Streams Agent, which shows their topologies and metrics in Kpow and exposes them for alerting. Kpow publishes consumer group offsets and lag, with broker, topic and connector metrics, on Prometheus endpoints for Grafana, AlertManager or the operator’s own monitoring, so a consumer that falls behind on a pricing or telemetry feed raises an alert. The agent is the only Factor House code that runs inside the operator’s own applications, and it is open source under the Apache 2.0 licence, so the operator’s software review can read and build it. For broker health, Kpow counts under-replicated partitions by walking every topic partition instead of asking each broker, as its article on the change explains, because Kafka’s own UnderReplicatedPartitions metric is reported per broker and a broker that is offline reports nothing.

Govern Flink jobs the same way. Operators that run Apache Flink beside Kafka can put the same directory sign-in and Allow, Deny or Stage policies over their Flink jobs with Flex.

Keep a separate control for applications. Kpow governs people working through Kpow and does not govern services connecting to brokers, so broker ACLs or an authorizer remain the control for market, dispatch and metering applications. This is the main difference from Conduktor for a grid operator or utility. Conduktor Console connects to clusters directly and masks data in its own UI, but it needs an external PostgreSQL database, and Conduktor’s encryption, masking of the data itself, policy enforcement on client traffic and Virtual Cluster multi-tenancy all run through Conduktor Gateway, which Conduktor’s own documentation describes as a Kafka proxy between client applications and brokers. Kpow gives people RBAC, temporary access, staged approvals and an audit log without putting anything in front of the brokers. The trade runs in both directions, because Gateway can enforce policy on applications, which Kpow does not attempt, at the cost of a component that operational data passes through.

Know the limits before the review asks. Kpow is one Docker container or JAR, and everything it needs to operate, snapshots, metrics and the audit log, is held in topics on the operator’s own cluster, so beyond Kafka it has no dependencies, and it connects to brokers as an ordinary Kafka client with the same security settings as any other client. Several consequences belong in the operator’s design. Kpow snapshots each cluster every minute, which suits investigation, trend views and alerts on lag, and does not replace the control room’s own real-time alarms. It is designed to run as a single instance restarted by the platform, with an active/passive pair as the documented route to failover, and it should sit close to its clusters, so sites in different regions each run their own instance, and one instance manages up to 12 clusters. Every governance feature on this page is an Enterprise feature, and Community Edition holds none of them.

To see these screens before installing anything, the live Kpow demo needs no signup. It shows what a utility’s engineers do every day, inspecting topic data, following consumer groups and lag and managing topics across two MSK clusters, and the __oprtr_audit_log topic on MSK Secondary shows what the audit trail records. The demo runs on MSK and has no SSO configured, so the Kubernetes install, the health endpoints and sign-in are the things to test in the operator’s own environment. To run the workflows against the operator’s own clusters, install Kpow from its container image, JAR or Helm chart; single sign-on, RBAC, temporary policies, staged mutations, the audit log, webhooks, Prometheus endpoints and the Streams Agent are Kpow Enterprise features, and Community Edition is free for 3 clusters and 10 users.

Kpow live demo

Open Kpow the way a grid operator's engineers would

The live Kpow demo needs no signup. Browse clusters, consumer groups and topic data across two clusters, then read the audit trail on the __oprtr_audit_log topic of the MSK Secondary cluster.

For platform teams running self-managed Kafka under market, dispatch and metering systems.

Try the Kpow demo

FAQ

What is the best Kafka management tool for a utility or grid operator?

On this page’s rubric, Kpow: it runs as one container inside the operator’s own network with no external database, no vendor proxy and no need for the internet, records every action and data query with the person from the operator’s directory, grants time-boxed production access through its API, manages self-managed, on-prem and cloud clusters from one deployment, and installs on Kubernetes from a Helm chart with liveness and readiness endpoints. Kafbat UI and AKHQ are the strongest free options and stay out of the data path too, but neither has expiring access grants or a readable audit trail. Kpow also has a free Community Edition for up to 3 clusters and 10 users, without the governance features.

How do energy and utility companies use Kafka?

Red Hat’s case study describes Transpower, New Zealand’s grid owner and system operator, using AMQ Streams on OpenShift to give its Market System applications the streaming data behind real-time electricity prices. Kai Waehner’s survey of Kafka for smart grid, utilities and energy production describes Kafka at the edge beside SCADA systems, PLCs and sensors, often disconnected from the cloud, with smart meter data preprocessed at the edge and the assets of the smart grid aggregated and monitored in the cloud.

Can Kpow run in an air-gapped network?

Yes. Kpow runs fully offline with no data leaving the operator’s environment, and works the same way in an air-gapped network as anywhere else. Its snapshots, metrics and audit log are held in topics on the operator’s own cluster, so beyond Kafka it has no further dependencies: no external database and no vendor-hosted service.

Do NERC CIP, NIS2 or the SOCI Act apply to Kafka tooling?

None of them names Kafka or any particular tool. They put cyber security and risk management obligations on the operator of the critical infrastructure, and NERC CIP-013 adds supply chain risk management for the bulk electric system, so Kafka clusters that carry operational data, and any tool that can read them, fall inside the operator’s own risk management. A tool that runs inside the operator’s network, keeps no data outside its clusters and names the person behind each action makes that part of the review shorter.

Does Kpow work with Kafka on Kubernetes and Strimzi?

Yes. Kpow installs from Factor House’s Helm charts as one pod with no database, its /healthy and /up endpoints serve as liveness and readiness probes, and its documentation covers connecting to Strimzi clusters with mutual TLS, SCRAM-SHA-512 or OAuth. Best Kafka UI tools for Strimzi compares the options for that platform in detail.

Does a Kafka management tool need to be a proxy to govern access?

No. Governing what people can see and do in a tool, with RBAC, temporary access and an audit log, works from a tool that connects as an ordinary Kafka client. A proxy is only needed to enforce policy on application traffic, and it puts a component between the brokers and every application that routes through it.

Can one Kafka tool manage clusters at the edge, in the data centre and in the cloud?

Yes, if the tool is vendor-agnostic and can reach each cluster. Kpow manages self-managed Apache Kafka, Confluent Platform, Confluent Cloud and Amazon MSK from one deployment, up to 12 clusters per instance. Its documentation asks for it to run close to its clusters, so sites in another region run their own instance.

How these tools were scored

Six criteria, each taken from a situation grid operators and utilities face under critical-infrastructure rules or in day-to-day operations, score every option from 0 to 10. They are listed here in order of weight.

1. Out of the data path. A management tool that runs in the operator’s own network, connects as an ordinary Kafka client and keeps no data outside the operator’s clusters gives the shortest answer when a critical-infrastructure review asks which suppliers’ components can reach operational data. Scored lower: tools that need an external database, and tools whose controls work only when application traffic passes through a vendor’s proxy. A self-hosted container with no external database and no proxy scores 9, a tool with a database of its own 6, one with several databases or a component on the brokers 4, and one that needs both a database and a proxy for its controls 3; 10 is kept for an option with nothing to deploy at all. This criterion is scored the same way on every Factor House page that uses it, and only its weight changes with the reader.

2. Audit trail per person. When people work through a shared tool, the broker only sees the tool’s service account, so only the tool’s own log can name the person. A grid operator or utility needs that log to include data reads as well as changes, so an incident review can see who looked at a topic as well as who changed it, and to be readable without building a consumer first. Best tools for Kafka audit logging compares the layers in detail.

3. Production access on request. Can an engineer be given read access to a production topic for a task, approved and time-boxed, without a standing grant? Can production changes such as offset resets be held for a second person’s approval? And is sensitive data masked for the people who do get in? The wider set of controls over deletes and offset resets is compared in best tools to control destructive Kafka operations.

4. On-prem and cloud together. Utilities rarely run one distribution. Clusters at sites and in data centres run beside newer clusters in the cloud, and the same tool should reach all of them, from as few deployments as the network allows; the general comparison is best tools to manage multiple Kafka clusters from one place.

5. Runs on Kubernetes. A utility that runs Kafka on its own Kubernetes or OpenShift platform needs the tool to install from a maintained Helm chart or operator, run as one pod and need nothing else in the cluster. Most tools here pass; tools with recorded Kubernetes problems or a database of their own score lower. The scores are the same as on best Kafka UI tools for Strimzi.

6. Directory and Kafka sign-in. People should sign in through the operator’s directory, whether that is LDAP directly or SAML and OIDC in front of Active Directory, Entra ID or Keycloak, with directory groups mapped to roles. On the cluster side the tool has to connect the way the clusters already authenticate clients, whether that is mutual TLS, SCRAM, Kerberos or OAuth. A tool that supports only OIDC scores 5. Best tools for Kafka SSO integration covers the protocol detail.

The cost figures model a grid operator or utility running 4 clusters (development, test, pre-production and production) for 100 engineers at $120 an engineer-hour, and each card prints its own assumptions. The general listicle view, without the energy and utilities weighting, is in best Kafka management tools.

F1 What a grid operator or utility runs into, and what the Kafka tool has to do about it
What happens at a grid operator or utility What the Kafka tool has to do
Critical-infrastructure review Security and risk review every component that can reach market, dispatch or metering data Run inside the operator's own network as a Kafka client, with no external database and no vendor proxy
Isolated networks Operational networks are segmented, and some have no route to the internet Run fully offline, with no vendor-hosted service
Who did what After an incident, the operator has to show who read or changed a production topic Log each action and data query with the person from the directory, on the operator's own infrastructure
Reading production data An engineer needs to look at a production topic to diagnose a stalled pricing or metering feed Grant read access on request, time-boxed, and hold changes for a second person's approval
Self-managed Kafka on Kubernetes Kafka runs on the operator's own Kubernetes or OpenShift platform, managed by an operator such as Strimzi Install from a Helm chart as one pod, with liveness and readiness probes the platform can act on
Edge, data centre and cloud Clusters at sites and in data centres run beside newer clusters in the cloud Manage clusters of any distribution from one tool, with an instance per region where sites are far apart
Each row is a situation grid operators and utilities face under critical-infrastructure rules or in day-to-day operations, followed by the behaviour a management tool needs in order to handle it without a workaround.

Every option is scored from 0 to 10 on each criterion, from the evidence and sources this page cites, and the reason for each score is on its card. The criteria are weighted: Out of the data path counts three times, Audit trail per person counts twice, Production access on request counts twice, On-prem and cloud together counts once, Runs on Kubernetes counts once and Directory and Kafka sign-in counts once, for a total out of 100. Out of the data path counts three times because grid operators and utilities are critical infrastructure under NERC CIP in North America, NIS2 in the EU and the SOCI Act in Australia, and every component that can reach operational data is a supply-chain and security question: a tool that sits between applications and brokers, or keeps its state in a database of its own, is more to review, more to run inside a segmented network, and one more thing that can fail on the path that carries pricing and dispatch data. The per-person audit trail and production access count twice, because after an incident the first questions are who could read or change production, and who actually did. On-prem and cloud together, running on Kubernetes and directory sign-in count once. This page is published by Factor House, which makes Kpow. Every option is scored on the same rubric and the same sources: Kpow's per-criterion scores are set the same way as every other option's and are not adjusted, and the weights apply to every option alike. Kpow ranks first on its total of 89 out of 100. The other options follow by total. Conduktor is listed last whatever its total; on its total of 59 it would place fourth.

Related reading