for self-managed kafka

Your Kafka
on-call engineer, in software.

Cairn watches every broker, topic, and Schema Registry event on your fleet twenty-four hours a day. When consumer lag drifts, schema compat breaks, or a broker cascades, the right playbook fires before most teams finish paging in — and the incident report lands in Slack the way you wish you had time to write it.

Per-cluster pricing · one Helm chart · runs in your VPC · data stays in your VPC.

cairn · watches 24/7

live

orders-eu-west-1 · topic payments.events.v3

Consumer-lag drift on payments.events.v3

warn
  • 02:47:11lag probe · consumer-group=checkout-svc · p99 = 14,820 msgs
  • 02:47:14drift > 3σ over 90s window · escalation tier 2
  • 02:47:18rebalance: add 6 partitions to consumer-group checkout-svc
  • 02:47:23lag probe · p99 = 1,094 msgs · recovering
  • 02:47:31resolved · 20s end-to-end · paged 0 humans
// playbook/rebalance-and-replay · v4.2cairn.report→slack,pagerduty
24/7watch on every broker, topic, Schema Registry event
<60smedian time from anomaly to remediation
14playbooks shipped by default — drawn from real incidents
0humans paged when a playbook closes the loop cleanly

what cairn closes

Observability shows the problem. Cairn ends it.

Datadog, Lenses, Conduktor, Kpow and AKHQ will all tell you a consumer is lagging or a schema just broke. Cairn is the agent that takes the next step — and writes the report.

lag · p99

Consumer-lag rebalancing

Detects per-partition drift in seconds, allocates warm consumers smoothly, and replays poisoned messages via the dead-letter queue before a single page fires.

primary playbook

compat · BACKWARD

Schema-evolution safety net

Watches every register on the Schema Registry and pins BACKWARD-incompatible versions before the first producer hits publish.

isr · leaders

Broker cascade response

Notices ISR shrinkage, reassigns leaders, and pre-emptively shrinks/re-expands replication factors when a single broker wobbles.

partitions · retention

Topic & partition tuning

Auto-tunes partition counts, retention, and compaction against actual traffic patterns — not a quarterly guess.

report → slack/pagerduty

Incident reports you actually read

Every remediation ships with a plain-English report pushed to Slack and PagerDuty — written the way on-call engineers wish they had time to write theirs.

getting started

One chart. Three minutes. The pager goes quiet.

Cairn is delivered as a single Helm chart that runs alongside your brokers. No code changes, no producer rewrites, no per-topic wiring.

  1. 01

    Detect before a human would notice

    cairn-agent probes every topic, consumer group, broker ISR, and Schema Registry subject on your fleet. Anomaly baselines are calibrated against your own traffic — not a vendor out-of-the-box threshold — so lag drift, incompat publishing, and ISR shrinkage surface within tens of seconds.

    WATCH topics=* brokers=* registry=*
  2. 02

    Auto-remediate via the right playbook

    Each signal runs through the playbook that closes it: consumer-lag rebalancing with DLQ replay for the lag spike, schema rollback-and-pin for the breaking register, ISR shrink-and-reassign for the broker wobble. Guardrails (replication floor, DLQ cap, severity ceiling) keep remediation inside the boundaries you set.

    FIX dlq-replay | schema-rollback | broker-rebalance
  3. 03
    pager stays quiet

    Plain-English report → Slack / PagerDuty

    Every remediation ships with an incident report written the way on-call engineers wish they had time to write. Slack and PagerDuty get it the moment the loop closes — Cairn only pages a human when a real decision is genuinely required.

    REPORT → slack,pagerduty (escalate only if needed)

playbooks in the box

Drawn from real incidents, not theoretical ones.

Cairn is built by a senior engineer who maintains parts of Confluent Platform and contributes upstream across Apache Kafka, Netty, Spark, Thrift, and Shapeless. The default playbooks are calibrated against the failures that actually cost teams SLA penalties — like the $240,000 cascading-broker crash a Factor House post-mortem documented.

01

rebalance-and-replay

Detects consumer-lag drift across a topic and replays dropped messages via the DLQ.

lag
02

rollback-and-pin

Pins a Schema Registry subject to the last compatible version and blocks promotion of breaking changes.

schema
03

isr-shrink-and-reassign

Shrinks ISR before partitions go under-replicated and reassigns leaders to the healthiest brokers.

brokers
04

partition-and-tune

Recommends or applies partition splits, retention sweeps, and compaction tuning based on observed traffic.

topics

pricing

One plan. Priced per cluster.

No per-seat counting. No per-event metering. No surprise overage lines. The whole product — every playbook, every report, every Slack and PagerDuty integration — is included on every cluster you connect.

  • Unlimited playbooks per cluster
  • Slack + PagerDuty reports
  • Helm chart installs in minutes
  • All Kafka distributions supported

per cluster

annual

Sized to you

Pricing for your fleet depends on broker count, topic count, and region mix. We reply to every request inside one business day.

Reply window
≤ 1 business day
Trial
30 days · Helm chart only

faq

The questions we get most.

ready when you are

Stop paging on every Kafka anomaly.

Tell us how many clusters and which distribution — we will get back to you with a sized quote and trial install instructions the same day.

join the private beta

Get the Helm chart before the public launch.

One email — we will share the install steps, the playbook list, and a date for the first live walkthrough. No drip campaign.