for self-managed kafka
Your Kafka
on-call engineer, in software.
Cairn watches every broker, topic, and Schema Registry event on your fleet twenty-four hours a day. When consumer lag drifts, schema compat breaks, or a broker cascades, the right playbook fires before most teams finish paging in — and the incident report lands in Slack the way you wish you had time to write it.
Per-cluster pricing · one Helm chart · runs in your VPC · data stays in your VPC.
cairn · watches 24/7
orders-eu-west-1 · topic payments.events.v3
Consumer-lag drift on payments.events.v3
- 02:47:11lag probe · consumer-group=checkout-svc · p99 = 14,820 msgs
- 02:47:14drift > 3σ over 90s window · escalation tier 2
- 02:47:18rebalance: add 6 partitions to consumer-group checkout-svc
- 02:47:23lag probe · p99 = 1,094 msgs · recovering
- 02:47:31resolved · 20s end-to-end · paged 0 humans
what cairn closes
Observability shows the problem. Cairn ends it.
Datadog, Lenses, Conduktor, Kpow and AKHQ will all tell you a consumer is lagging or a schema just broke. Cairn is the agent that takes the next step — and writes the report.
lag · p99
Consumer-lag rebalancing
Detects per-partition drift in seconds, allocates warm consumers smoothly, and replays poisoned messages via the dead-letter queue before a single page fires.
compat · BACKWARD
Schema-evolution safety net
Watches every register on the Schema Registry and pins BACKWARD-incompatible versions before the first producer hits publish.
isr · leaders
Broker cascade response
Notices ISR shrinkage, reassigns leaders, and pre-emptively shrinks/re-expands replication factors when a single broker wobbles.
partitions · retention
Topic & partition tuning
Auto-tunes partition counts, retention, and compaction against actual traffic patterns — not a quarterly guess.
report → slack/pagerduty
Incident reports you actually read
Every remediation ships with a plain-English report pushed to Slack and PagerDuty — written the way on-call engineers wish they had time to write theirs.
getting started
One chart. Three minutes. The pager goes quiet.
Cairn is delivered as a single Helm chart that runs alongside your brokers. No code changes, no producer rewrites, no per-topic wiring.
- 01
Detect before a human would notice
cairn-agent probes every topic, consumer group, broker ISR, and Schema Registry subject on your fleet. Anomaly baselines are calibrated against your own traffic — not a vendor out-of-the-box threshold — so lag drift, incompat publishing, and ISR shrinkage surface within tens of seconds.
WATCH topics=* brokers=* registry=* - 02
Auto-remediate via the right playbook
Each signal runs through the playbook that closes it: consumer-lag rebalancing with DLQ replay for the lag spike, schema rollback-and-pin for the breaking register, ISR shrink-and-reassign for the broker wobble. Guardrails (replication floor, DLQ cap, severity ceiling) keep remediation inside the boundaries you set.
FIX dlq-replay | schema-rollback | broker-rebalance - 03pager stays quiet
Plain-English report → Slack / PagerDuty
Every remediation ships with an incident report written the way on-call engineers wish they had time to write. Slack and PagerDuty get it the moment the loop closes — Cairn only pages a human when a real decision is genuinely required.
REPORT → slack,pagerduty (escalate only if needed)
playbooks in the box
Drawn from real incidents, not theoretical ones.
Cairn is built by a senior engineer who maintains parts of Confluent Platform and contributes upstream across Apache Kafka, Netty, Spark, Thrift, and Shapeless. The default playbooks are calibrated against the failures that actually cost teams SLA penalties — like the $240,000 cascading-broker crash a Factor House post-mortem documented.
rebalance-and-replay
Detects consumer-lag drift across a topic and replays dropped messages via the DLQ.
rollback-and-pin
Pins a Schema Registry subject to the last compatible version and blocks promotion of breaking changes.
isr-shrink-and-reassign
Shrinks ISR before partitions go under-replicated and reassigns leaders to the healthiest brokers.
partition-and-tune
Recommends or applies partition splits, retention sweeps, and compaction tuning based on observed traffic.
pricing
One plan. Priced per cluster.
No per-seat counting. No per-event metering. No surprise overage lines. The whole product — every playbook, every report, every Slack and PagerDuty integration — is included on every cluster you connect.
- Unlimited playbooks per cluster
- Slack + PagerDuty reports
- Helm chart installs in minutes
- All Kafka distributions supported
per cluster
Sized to you
Pricing for your fleet depends on broker count, topic count, and region mix. We reply to every request inside one business day.
- Contact
- cairn-t0i77x@polsia.app
- Reply window
- ≤ 1 business day
- Trial
- 30 days · Helm chart only
faq
The questions we get most.
ready when you are
Stop paging on every Kafka anomaly.
Tell us how many clusters and which distribution — we will get back to you with a sized quote and trial install instructions the same day.
join the private beta
Get the Helm chart before the public launch.
One email — we will share the install steps, the playbook list, and a date for the first live walkthrough. No drip campaign.