Chad Harris, Head of Product at Factor House, shows how Kpow delivers governed, agentic Kafka operations using the Factor CLI with Claude. The same agentic skills are also available with Copilot, Codex, and other LLMs.
The skills work through the Factor CLI, so every query runs under your existing Kpow single sign-on, RBAC, and tenancy configuration. They give the assistant the right context from Kpow for each type of question, so it queries only the data it needs. In the demo, Claude finds that one of twelve connectors has failed, pulls the underlying exception, and identifies it as a database connectivity problem rather than a Kafka Connect one. It then proves that a service reporting no data is in fact consuming from all four of its topics, with production rates, consumer group state, lag, and last commit times as evidence, and finishes by reporting the most lagging topic.
You can install the skills today with fh skills install, which places them wherever your agent (Claude, Copilot, Codex, or another LLM-based assistant) normally looks for skills.
Full transcript
The big one now is our agentic skills, so let's have a look at what they can do for us. Here is my Claude session. The first thing I want to know is what was happening with that connector, because the Kafka Connect cluster thought it was healthy, but the connector itself wasn't working when we looked at it in the terminal UI.
So I'll ask: "Are my connectors healthy?" For those who haven't used Claude before, this is the Claude terminal interface. I'm literally just asking it that question, and it takes about thirty seconds to a minute for Claude to work through the data and come back with an answer.
The most important thing with agentic workflows is the context you give them. I think we're all starting to learn that agents perform quite poorly if you give them the wrong context or poor context. If you give them good data and good instructions, they perform well, and that's what we're going for with our skills. Kpow has a rich and accurate set of data across all of your clusters, so it has all of the context your agents need. By packaging that up in these skills, we give Claude the correct context to make the right decision. Over time, as we refine them, we'll give it better and better data so it can reach these decisions quicker and use fewer tokens. If you give an agent heaps of data it uses more tokens; the more accurate the context, the fewer tokens it needs.
So what does it come back with? It's telling us that eleven of our twelve connectors are healthy. One is down: the Debezium airline Postgres connector. The airline Postgres connector isn't moving any data, and Claude has even pulled the exception out of the connector for us. It can't obtain the encoding for the database: it's getting a connect exception because the connection is refused. The database is refusing TCP connections, so this is an RDS issue, not a Kafka Connect issue. The connector reads as running only because the worker is up. That's the important part: the task can be up and Connect will advertise it as running, even though it failed to connect to the database. Claude also tells us that restarting the task will just fail again, so don't bother restarting the connector; you need to actually fix the database. Otherwise it gives us a quick summary that everything else is healthy.
What I'm curious about next is a question that, as a platform engineer, I got asked almost every day of my career: a team asking why their service isn't getting any data. So let's ask why the airline analytics service isn't getting any data. Thankfully the AI will interpret my very poor typing.
That question always took time to research. At least fifty percent of the time the service was getting data, and something internal to the service was going wrong, but it was always hard to gather the data to prove that a service was consuming. On the flip side, when the service really wasn't getting data, it was a lot of extra work to figure out why. Is there no producer? Is it a poison pill event, where it's actually getting data but tripping up on one particular event? It always took us a long time to answer. Claude can answer it because we give it all of the information it needs. When we write these skills, we tell it: when somebody asks why their service isn't getting data, these are the data sets you need to ask Kpow for, and Kpow gives it that data.
Here's the verdict: in this case, the consumer is getting data. On a good day, that would have taken me ten minutes to prove at my previous job, so that's ten minutes saved. It's told us that all four topics are having data produced to them and we're consuming them all. We're producing at around 0.7 messages a second across all of the partitions, all of them are in sync, and the last write was less than a minute ago. For each of the four topics it gives us the message count, confirms the consumer groups are stable, and shows the lag, the consume rate, and the last commit. So we can see that everything is actually fine.
It's given us a bit more insight as well. It noticed that one of the nodes is restarting fairly regularly, and a continually restarting node can cause some lag from time to time, depending on your configuration. I know why that's happening: this is our test environment, and every twenty minutes an automated task removes one of the nodes and replaces it. It's quite cool that the agent was able to make that observation, because it could be something you'd want to look at. It's also given us the CLI command, so we can run it ourselves and monitor this directly.
I'll ask it one more question: "What is my biggest lagging topic?"
While Claude works on that: in summary, we're really excited about the CLI, the terminal UI, and the set of agentic skills, and we'll talk a bit more at the end about what's coming next for the skills. I suspect most of us grew up in engineering waiting for the build or waiting for the tests to run. Engineers now spend their days waiting for Claude instead.
Here we go. Our Iceberg control topic has the most lag, but it's only 114 messages, which is probably fine for a busy topic. It's told us what our biggest lagging topics are, and there's nothing there to worry about.
This is all available to download today. The CLI works today, and the agent skills work today. There are a couple of ways to install the agentic skills. The easiest is the CLI: run fh skills install. If you work with specific agents, the CLI installs the skills wherever that agent normally keeps them, whether that's Copilot, Codex, or Claude. You can also use the plugin mechanisms in Claude and Codex by adding our marketplace and installing the skills that way, but I think it's easier to use the CLI.
Speaker
Head of Product, Factor House
Chad Harris is Head of Product at Factor House, bringing 18 years of experience across software engineering, application architecture, and engineering leadership. He has deep hands-on expertise with Apache Kafka, high-volume transactional systems, and PCI-compliant architectures, most recently as an Engineering Leader at Block (formerly Square).
Try Kpow for Apache Kafka
The Kafka management console built for platform and data engineers.
Learn more