Autumn of Learning Sale is live now!Claim Up to 50% Off
🧮

AI Kafka Cluster Sizing & ADR Generator

Principal Enterprise Architect Tool

🌱 Learn how Senior Architects calculate the exact number of servers, disks, and memory needed for Kafka.

50 MB/s (~400 Mbps)
5 MB/s (Startup)100 MB/s (Enterprise)500 MB/s (Hyper-Scale)
RF = 3
7 Days
1 Day (Buffer)7 Days (Standard)30 Days (Extended)
3 Downstream Apps
Tiered Storage Architecture
Store older log segments in S3/GCS object storage
Recommended Cluster Hardware TopologyKRaft Quorum
Brokers
5 Nodes
vCPU / Node
8 Cores
RAM / Node
32 GB
NVMe / Node
3.21 TB
Total Network Bandwidth (In + Egress + ISR):2.34 Gbps
Estimated Monthly Infrastructure Cost:
Self-Managed EC2
$4,901
KRaft + EBS gp3
Confluent Cloud
$5,230
Fully Managed CKU
Amazon MSK
$6,616
m5.2xlarge clusters

Automated ADR-001 Architecture Decision Document

# Architectural Decision Record (ADR-001) ## Topic: Kafka Cluster Capacity Planning & Hardware Sizing * **Status:** APPROVED * **Date:** 2026-10-05 * **Author:** Principal Enterprise Architect * **Target Workload:** Real-Time Enterprise Streaming Platform --- ### 1. Context & Business Requirements The enterprise requires a resilient, high-throughput Apache Kafka cluster capable of sustaining continuous message ingestion with zero data loss ($RPO=0$). * **Sustained Ingress Throughput:** 50 MB/s (~400 Mbps) * **Peak Burst Ratio:** 2.5x (125 MB/s peak burst) * **Replication Factor ($RF$):** 3 (Min In-Sync Replicas = 2) * **Retention Window:** 7 days * **Consumer Group Fanout:** 3 independent consumer applications * **Tiered Storage Architecture:** ENABLED (Local NVMe + Cloud Object Storage) --- ### 2. Sizing Calculations & Resource Envelope | Dimension | Calculated Value | Operational Rationale | | :--- | :--- | :--- | | **Total Daily Ingestion** | 4219 GB/day | Base raw event payload | | **Total Retention Storage** | 112.47 TB | Factoring $RF=3$ and 30% OS/Index headroom | | **Local NVMe Fast Tier** | 16.07 TB | Sub-millisecond read cache | | **Cold S3/GCS Object Tier** | 96.41 TB | Long-term compacted historical events | | **Inter-Broker Replication Traffic** | 100.0 MB/s | Synchronous ISR log replication | | **Consumer Egress Traffic** | 150.0 MB/s | 3 concurrent consumer groups | | **Total Sustained Network I/O** | 2.34 Gbps | Ingress + Egress + ISR catchup | --- ### 3. Recommended Infrastructure Topology * **Cluster Topology:** 5 KRaft Broker Nodes (Dedicated Quorum) * **Compute per Node:** 8 vCPU / 32 GB RAM (minimum 50% reserved for Linux OS PageCache) * **Storage per Node:** 3.21 TB NVMe block device * **Filesystem:** XFS with `noatime` mount options and Linux kernel `dirty_ratio=20` ```mermaid flowchart LR P["Producers (50 MB/s)"] -->|Ingress| K["KRaft Quorum (5 Nodes)"] K -->|Replication (100 MB/s)| K K -->|Fanout (150 MB/s)| C["3 Consumer Groups"] K -->|Tiered Segments| S3["Object Storage (S3 / GCS)"] ``` --- ### 4. Enterprise Commercial TCO Comparison (Monthly Est.) * **AWS EC2 Self-Managed (KRaft):** ~$4,901/mo * **Confluent Cloud Serverless / Dedicated:** ~$5,230/mo * **Amazon MSK Managed:** ~$6,616/mo ### 5. Architectural Consequences & Tradeoffs * Enabling Tiered Storage reduces NVMe storage footprint by up to 86%, drastically accelerating broker rebalance and rolling restart recovery times. * All producer clients must enforce `acks=all` and `enable.idempotence=true` to guarantee strictly once processing under network retries.