🧮
AI Kafka Cluster Sizing & ADR Generator
Principal Enterprise Architect Tool🌱 Learn how Senior Architects calculate the exact number of servers, disks, and memory needed for Kafka.
50 MB/s (~400 Mbps)
5 MB/s (Startup)100 MB/s (Enterprise)500 MB/s (Hyper-Scale)
RF = 3
7 Days
1 Day (Buffer)7 Days (Standard)30 Days (Extended)
3 Downstream Apps
Tiered Storage Architecture
Store older log segments in S3/GCS object storage
Recommended Cluster Hardware TopologyKRaft Quorum
Brokers
5 Nodes
vCPU / Node
8 Cores
RAM / Node
32 GB
NVMe / Node
3.21 TB
Total Network Bandwidth (In + Egress + ISR):2.34 Gbps
Estimated Monthly Infrastructure Cost:
Self-Managed EC2
$4,901
KRaft + EBS gp3
Confluent Cloud
$5,230
Fully Managed CKU
Amazon MSK
$6,616
m5.2xlarge clusters
Automated ADR-001 Architecture Decision Document
# Architectural Decision Record (ADR-001)
## Topic: Kafka Cluster Capacity Planning & Hardware Sizing
* **Status:** APPROVED
* **Date:** 2026-10-05
* **Author:** Principal Enterprise Architect
* **Target Workload:** Real-Time Enterprise Streaming Platform
---
### 1. Context & Business Requirements
The enterprise requires a resilient, high-throughput Apache Kafka cluster capable of sustaining continuous message ingestion with zero data loss ($RPO=0$).
* **Sustained Ingress Throughput:** 50 MB/s (~400 Mbps)
* **Peak Burst Ratio:** 2.5x (125 MB/s peak burst)
* **Replication Factor ($RF$):** 3 (Min In-Sync Replicas = 2)
* **Retention Window:** 7 days
* **Consumer Group Fanout:** 3 independent consumer applications
* **Tiered Storage Architecture:** ENABLED (Local NVMe + Cloud Object Storage)
---
### 2. Sizing Calculations & Resource Envelope
| Dimension | Calculated Value | Operational Rationale |
| :--- | :--- | :--- |
| **Total Daily Ingestion** | 4219 GB/day | Base raw event payload |
| **Total Retention Storage** | 112.47 TB | Factoring $RF=3$ and 30% OS/Index headroom |
| **Local NVMe Fast Tier** | 16.07 TB | Sub-millisecond read cache |
| **Cold S3/GCS Object Tier** | 96.41 TB | Long-term compacted historical events |
| **Inter-Broker Replication Traffic** | 100.0 MB/s | Synchronous ISR log replication |
| **Consumer Egress Traffic** | 150.0 MB/s | 3 concurrent consumer groups |
| **Total Sustained Network I/O** | 2.34 Gbps | Ingress + Egress + ISR catchup |
---
### 3. Recommended Infrastructure Topology
* **Cluster Topology:** 5 KRaft Broker Nodes (Dedicated Quorum)
* **Compute per Node:** 8 vCPU / 32 GB RAM (minimum 50% reserved for Linux OS PageCache)
* **Storage per Node:** 3.21 TB NVMe block device
* **Filesystem:** XFS with `noatime` mount options and Linux kernel `dirty_ratio=20`
```mermaid
flowchart LR
P["Producers (50 MB/s)"] -->|Ingress| K["KRaft Quorum (5 Nodes)"]
K -->|Replication (100 MB/s)| K
K -->|Fanout (150 MB/s)| C["3 Consumer Groups"]
K -->|Tiered Segments| S3["Object Storage (S3 / GCS)"]
```
---
### 4. Enterprise Commercial TCO Comparison (Monthly Est.)
* **AWS EC2 Self-Managed (KRaft):** ~$4,901/mo
* **Confluent Cloud Serverless / Dedicated:** ~$5,230/mo
* **Amazon MSK Managed:** ~$6,616/mo
### 5. Architectural Consequences & Tradeoffs
* Enabling Tiered Storage reduces NVMe storage footprint by up to 86%, drastically accelerating broker rebalance and rolling restart recovery times.
* All producer clients must enforce `acks=all` and `enable.idempotence=true` to guarantee strictly once processing under network retries.