🔬
Kafka Log Segment & PageCache X-Ray
Binary Storage Engine🌱 Visual anatomy of how Kafka saves data on disk without slowing down under millions of messages.
🌱 ELI5 Book Index AnalogyLog Segment: 00000000000000000000.log
Why is Kafka so fast? Imagine a 1,000-page book. Instead of reading from page 1 to find chapter 680, you glance at the index at the back of the book, jump straight to byte 114,400 on page 500, and read just a few lines. That is exactly how Kafka's .index and .log files work!
Interactive Sparse Offset Binary Search
Target Offset Requested: 680
Offset 0 (Segment Base)Offset 1,000 (Mid-Segment)Offset 1,999 (High Watermark HW)
Sparse Index (.index) Entries (Binary Search O(log N)):Jumped to Entry: Offset 500 @ Byte 84,300
Offset
0
@0B
Offset
250
@42,150B
Offset
500
@84,300B
Offset
750
@126,450B
Offset
1000
@168,600B
Offset
1250
@210,750B
Offset
1500
@252,900B
Offset
1750
@295,050B
Calculated Log Segment Byte Seek: 114,540 bytes
Only scanned 180 records from index anchor!
Linux Zero-Copy (sendfile) vs Standard JVM Transfer
⚡ Zero-Copy DMA Pipeline (2 Context Switches, 0 CPU Copies):100% Throughput Efficiency
1. NVMe Block Disk
Commit Log
2. Direct DMA Channel
Linux OS PageCache
3. NIC Socket Ring
Consumer Network
Data moves directly from OS page cache into the network card buffer via DMA (Direct Memory Access). The JVM user space memory and Garbage Collector are completely bypassed, avoiding GC pauses.
Partition Log Directory Structure on Disk
| File Name | File Type | Internal Structure | Purpose & Lifecycle |
|---|---|---|---|
| 00000000000000000000.log | Data Commit Log | RecordBatch: MagicByte, CRC32, PID, Key, Value | Raw sequential messages appended until segment.bytes (default 1GB). |
| 00000000000000000000.index | Offset Sparse Index | 8 bytes/entry: 4B relative offset + 4B byte position | Maps logical message offset to physical byte position for rapid seeks. |
| 00000000000000000000.timeindex | Timestamp Index | 12 bytes/entry: 8B epoch timestamp + 4B relative offset | Supports timestamp queries (offsetsForTimes) and retention cleanup. |
| leader-epoch-checkpoint | Epoch Metadata | Epoch ID + Start Offset mapping | Prevents log divergence and truncation during KRaft quorum failovers. |