What Is a Kafka Partition?
A Kafka partition is an ordered, append-only sequence of records inside a Kafka topic. Topics are divided into partitions so Kafka can distribute data across brokers, process records in parallel, and scale producers and consumers beyond a single machine.
Partitions are one of the most important concepts in Apache Kafka. They determine where records are stored, what ordering Kafka can guarantee, how consumer groups parallelize processing, and how a Kafka cluster distributes load.
Table of Contents
- Why Kafka Needs Partitions
- How a Kafka Partition Works
- Topics and Partitions
- How Producers Choose a Partition
- Partition Ordering
- Partitions and Consumer Groups
- Partitions and Parallelism
- Partition Replication
- Partition Leaders and Followers
- Choosing the Number of Partitions
- Partition Hotspots
- Adding Partitions
- Production Design Example
- Common Kafka Partition Mistakes
- Production Checklist
- Frequently Asked Questions
- Conclusion
Why Kafka Needs Partitions
Imagine a Kafka topic implemented as one large log on one machine:
orders
Broker A
┌──────────────────────────────────────────────┐
│ e1 │ e2 │ e3 │ e4 │ e5 │ e6 │ e7 │ ... │
└──────────────────────────────────────────────┘
Every producer writes through that broker, the complete log is stored there, and consumer parallelism is constrained by the single stream.
As traffic grows, that machine eventually becomes a bottleneck.
Kafka partitions the topic into multiple independent logs:
orders
Partition 0: e1 → e4 → e7 → e10
Partition 1: e2 → e5 → e8 → e11
Partition 2: e3 → e6 → e9 → e12
Those partitions can be distributed across different brokers:
Broker A Broker B Broker C
Partition 0 Partition 1 Partition 2
e1 e2 e3
e4 e5 e6
e7 e8 e9
e10 e11 e12
Now storage, writes, reads, and consumer processing can be distributed across the cluster.
This is the core reason Kafka uses partitions: a partition is the unit that allows a topic to scale horizontally.
How a Kafka Partition Works
Each partition is an ordered sequence of records. New records are appended to the end rather than inserted into arbitrary positions.
Partition 0
┌────────┬────────┬────────┬────────┬────────┐
│ Event A│ Event B│ Event C│ Event D│ Event E│
└────────┴────────┴────────┴────────┴────────┘
0 1 2 3 4
Offset
The partition provides the ordering boundary. Kafka preserves the position of records within that partition.
Records and Offsets
Every record in a partition receives an offset.
orders-0
Offset 0 → order.created
Offset 1 → order.paid
Offset 2 → order.shipped
Offset 3 → order.delivered
An offset identifies a record's position within a particular partition.
It is not a global identifier across the topic:
Partition 0 Partition 1
Offset 0 → A Offset 0 → X
Offset 1 → B Offset 1 → Y
Offset 2 → C Offset 2 → Z
Therefore, a record position is effectively identified by the combination of topic, partition, and offset:
Topic: orders
Partition: 2
Offset: 91827
Consumers track their progress through these offsets.
Append-Only Log
A Kafka partition behaves as an append-only log.
Existing records
A → B → C → D
New record E
A → B → C → D → E
↑
append
This sequential model is important for Kafka's throughput. Brokers can write records efficiently while consumers read sequential ranges of the log.
Records are not normally removed when a consumer reads them. Kafka retains them according to topic retention or compaction policies, which allows multiple consumers to read the same data independently and can enable replay.
Kafka Topics, Partitions, and Offsets Explained covers the relationship between these three concepts in more detail.
Topics and Partitions
A Kafka topic is the logical stream of records. A partition is one physical subdivision of that stream.
For example, an orders topic might contain six partitions:
orders
│
├── Partition 0
├── Partition 1
├── Partition 2
├── Partition 3
├── Partition 4
└── Partition 5
Producers publish records to the topic, but every record ultimately belongs to exactly one partition.
Producer
│
↓
orders
│
├── P0 ← some records
├── P1 ← some records
└── P2 ← some records
Kafka can distribute these partitions across brokers, allowing the topic to exceed the storage and throughput capacity of one server.
How Producers Choose a Partition
A producer must determine which partition receives each record. The decision is important because it affects ordering and load distribution.
Keyed Records
Records can include a key:
{
"key": "customer-481",
"event": {
"type": "order.created",
"order_id": "ORD-9182"
}
}
A partitioner uses the key to select a partition. Conceptually:
partition = hash(key) → partition selection
Records with the same key are normally routed consistently to the same partition as long as the relevant partitioning conditions remain unchanged.
customer-481
│
├── order.created
├── order.paid
└── order.shipped
│
↓
Partition 2
This is useful when events for the same business entity must preserve relative ordering.
Another customer can be assigned elsewhere:
customer-481 ──→ Partition 2
customer-732 ──→ Partition 0
customer-915 ──→ Partition 4
The choice of key therefore becomes part of the application's data-distribution and ordering design.
Records Without Keys
When records do not have keys, the producer can distribute them across available partitions according to the producer's partitioning behavior.
Record A ──→ P0
Record B ──→ P1
Record C ──→ P2
Record D ──→ P0
Record E ──→ P1
This can provide balanced throughput when there is no requirement to keep related records together.
The exact producer behavior depends on the Kafka client and configuration, so production designs should not depend on an assumed partitioning algorithm without verifying the client being used.
Kafka Producers Explained: Partitioning, Batching, and Delivery Guarantees covers producer partitioning, batching, acknowledgements, and delivery behavior.
Partition Ordering
Kafka provides ordering within a partition, not one global ordering across every partition in a topic.
Consider two partitions:
Partition 0
A1 → A2 → A3
Partition 1
B1 → B2 → B3
Kafka preserves:
A1 before A2 before A3
and
B1 before B2 before B3
But there is no inherent total ordering such as:
A1 → B1 → A2 → B2 → A3 → B3
because the partitions are independent logs and may be consumed concurrently.
This is why partition keys are important when business operations require ordering.
For example, using order_id as the key can keep all events for one order together:
ORD-100
│
├── created
├── paid
├── packed
└── shipped
│
↓
Partition 3
Events for ORD-200 can simultaneously be processed through another partition.
Partition 3 Partition 7
ORD-100 created ORD-200 created
ORD-100 paid ORD-200 paid
ORD-100 packed ORD-200 cancelled
ORD-100 shipped
The result is ordering where it matters and parallelism where it does not.
Kafka Ordering Guarantees and Message Deduplication covers ordering boundaries and duplicate-processing considerations in greater detail.
Partitions and Consumer Groups
Partitions also determine how Kafka consumer groups divide work.
Within one consumer group, each partition is assigned to at most one consumer at a time.
Suppose a topic has four partitions and a consumer group has two consumers:
Topic: orders
P0 ──┐
├──→ Consumer A
P1 ──┘
P2 ──┐
├──→ Consumer B
P3 ──┘
Each consumer processes multiple partitions.
Add two more consumers:
P0 ──→ Consumer A
P1 ──→ Consumer B
P2 ──→ Consumer C
P3 ──→ Consumer D
The group can now process the four partitions with greater parallelism.
But adding a fifth consumer does not create a fifth active partition assignment:
P0 ──→ Consumer A
P1 ──→ Consumer B
P2 ──→ Consumer C
P3 ──→ Consumer D
Consumer E → idle
There are only four partitions available to divide among consumers in that group.
This creates one of Kafka's most important capacity relationships:
Maximum active consumers for partition processing
within one consumer group
≈
number of partitions
Different consumer groups are independent. An analytics group and a fulfillment group can each consume all partitions of the same topic.
orders
│
├──→ Analytics Group
│ P0 P1 P2 P3
│
└──→ Fulfillment Group
P0 P1 P2 P3
Kafka Consumers and Consumer Groups Explained covers assignments, rebalancing, offsets, and consumer scaling in detail.
Partitions and Parallelism
Partition count creates an upper bound on useful consumer parallelism for one consumer group.
Consider a topic with one partition:
Partition 0
│
↓
Consumer A
Consumer B → idle
Consumer C → idle
Adding consumers does not split one partition among several consumers in the same group.
With six partitions:
P0 ──→ C1
P1 ──→ C2
P2 ──→ C3
P3 ──→ C4
P4 ──→ C5
P5 ──→ C6
up to six consumers can simultaneously own partitions.
This makes partition count a capacity-planning decision, not merely a storage setting.
More partitions can enable:
- higher producer throughput;
- higher consumer parallelism;
- better distribution across brokers;
- larger aggregate topic capacity.
But more partitions are not free. They increase metadata, files, replication work, partition leadership, rebalancing complexity, and operational overhead.
Partition Replication
Partitions provide scalability, but a partition stored on only one broker would still create a single point of failure.
Kafka can maintain multiple replicas of each partition.
Partition 0
Broker A → Leader
Broker B → Replica
Broker C → Replica
With a replication factor of three, the partition has three copies distributed across brokers.
A topic with several partitions might look conceptually like:
Broker A Broker B Broker C
Partition 0 Leader Replica Replica
Partition 1 Replica Leader Replica
Partition 2 Replica Replica Leader
Leadership is distributed so one broker does not necessarily handle every partition's client traffic.
Replication provides redundancy if a broker becomes unavailable.
Kafka Replication and Fault Tolerance Explained covers replication factors, leaders, followers, in-sync replicas, and broker failure behavior.
Partition Leaders and Followers
For each partition, one replica acts as the leader. Other replicas act as followers.
Producer
│
↓
Partition 4 Leader
│
├──→ Follower
└──→ Follower
Producers send records through the partition leader. Followers replicate the partition's data.
If the leader's broker fails, Kafka can elect an eligible replica as the new leader:
Before
Broker A: P4 Leader
Broker B: P4 Follower
Broker C: P4 Follower
Broker A fails
After Election
Broker B: P4 Leader
Broker C: P4 Follower
This is how partitioning and replication work together: partitions provide distribution and parallelism, while replicas provide fault tolerance.
Choosing the Number of Partitions
There is no universal correct partition count.
The decision should consider expected throughput, consumer parallelism, partition-key distribution, future growth, broker capacity, and operational overhead.
Suppose one consumer instance safely processes approximately 2,000 events per second and the target workload is 20,000 events per second.
Required processing rate: 20,000 events/sec
Per-consumer capacity: 2,000 events/sec
20,000 / 2,000 = 10 consumers
The topic needs enough partitions to allow approximately that level of consumer parallelism:
10 consumers
│
↓
at least 10 useful partition assignments
That does not automatically mean ten is the correct final partition count. Growth margin and producer throughput may justify more.
At the same time, creating hundreds or thousands of partitions for a tiny workload can add complexity without useful capacity.
Partition count should therefore be treated as an architectural capacity parameter rather than a number chosen arbitrarily.
Partition Hotspots
Having many partitions does not guarantee balanced traffic. The partition key must also distribute records appropriately.
Consider a topic with eight partitions where one customer produces 70% of all events:
Customer A → P3 → 70% traffic
Everyone else:
P0 P1 P2 P4 P5 P6 P7 → 30% traffic
Partition 3 becomes a hot partition.
P0 ██
P1 ███
P2 ██
P3 █████████████████████████
P4 ██
P5 ███
P6 ██
P7 ██
The topic may have plenty of aggregate cluster capacity, but one partition becomes the bottleneck because records belonging to the hot key cannot simply be spread across partitions without changing the ordering or partitioning model.
Potential strategies include:
- choosing a higher-cardinality partition key;
- splitting exceptionally hot entities into controlled subkeys;
- removing unnecessary ordering requirements;
- separating unusually heavy workloads into another topic;
- increasing capacity where the hot partition is processed.
Changing the key should not be done casually. Distribution improvements can alter ordering guarantees and downstream assumptions.
Adding Partitions
Kafka allows the partition count of a topic to be increased.
For example:
Before
orders:
P0 P1 P2 P3
After
orders:
P0 P1 P2 P3 P4 P5 P6 P7
This provides more potential parallelism for future records and consumers.
However, increasing partition count can affect keyed record placement. If the partitioning calculation depends on the number of partitions, the same key may map differently after the count changes.
Before:
customer-481 → P2
Partition count increases
After:
customer-481 → P6
Old records remain in their existing partitions while new records for the same key may be written elsewhere.
This can break assumptions about per-key ordering across the partition-count change.
Increasing partitions should therefore be planned carefully when applications depend on stable key-to-partition assignment or strict per-key ordering.
Production Design Example
Consider an e-commerce platform publishing order lifecycle events.
The topic receives approximately 30,000 events per second during peak periods:
order.created
order.paid
order.packed
order.shipped
order.delivered
order.cancelled
Events for one order must be processed in order, but different orders can be processed independently.
Using order_id as the partition key provides that boundary:
ORD-1001 ──→ P2
ORD-1002 ──→ P7
ORD-1003 ──→ P1
ORD-1004 ──→ P5
Every event for ORD-1001 goes to the same partition:
Partition 2
ORD-1001 created
↓
ORD-1001 paid
↓
ORD-1001 packed
↓
ORD-1001 shipped
Suppose a fulfillment consumer instance safely processes about 2,500 events per second.
Peak workload: 30,000 events/sec
Consumer capacity: 2,500 events/sec
30,000 / 2,500 = 12 consumers
The topic therefore needs enough partitions to support at least that required processing parallelism, with additional capacity considered for growth and uneven distribution.
Assume the design chooses 18 partitions:
orders topic
18 partitions
replication factor = 3
Fulfillment Group:
12 consumers
Analytics Group:
18 consumers
Notification Group:
6 consumers
Each consumer group uses the same topic differently.
The analytics group can use all 18 partitions simultaneously. Fulfillment uses 12 consumers, so some consumers own multiple partitions. Notifications uses only six consumers because its processing rate is lower.
The topic is distributed across six Kafka brokers:
Broker A Broker B Broker C Broker D Broker E Broker F
│ │ │ │ │ │
leaders and replicas distributed across the cluster
The production system should monitor partition-level metrics rather than only topic-wide averages.
Partition Events/sec Consumer Lag
P0 1,610 21
P1 1,720 18
P2 1,580 24
P3 1,690 19
P4 1,640 22
...
P11 5,900 84,210 ← problem
The topic's average throughput might look acceptable while P11 is overloaded.
Investigation could reveal that a small number of extremely active merchants dominate the partition key distribution. The solution depends on whether strict merchant-level or order-level ordering is actually required.
Consumer lag is especially important operationally. A partition can continue receiving records faster than its assigned consumer can process them, causing the consumer's position to fall further behind.
What Is Kafka Consumer Lag? explains how lag develops, how to interpret it, and what to monitor in production.
Common Kafka Partition Mistakes
- Assuming ordering exists across the entire topic. Kafka ordering is fundamentally partition-scoped.
- Choosing a low-cardinality partition key. A few keys can concentrate most traffic into a few partitions.
- Choosing a highly skewed key. One large tenant or entity can create a hot partition.
- Using random keys when per-entity ordering is required. Related records may be distributed across different partitions.
- Creating too few partitions. Consumer groups may be unable to reach the required parallelism.
- Creating huge numbers of partitions without need. Partitions have broker and operational overhead.
- Adding consumers beyond the partition count. Extra consumers in the same group may remain idle.
- Increasing partition count without considering keyed ordering. Future records for a key can map to a different partition.
- Monitoring only topic-wide throughput. A single overloaded partition can be hidden by healthy averages.
- Confusing partitions with replicas. Partitions provide distribution; replicas provide redundant copies.
- Ignoring downstream capacity. More partitions and consumers can simply move the bottleneck to a database or external API.
Production Checklist
- Define the required ordering scope.
- Choose the partition key from business ordering requirements.
- Check partition-key cardinality.
- Check expected key skew and hot-key behavior.
- Estimate peak producer throughput.
- Estimate per-consumer processing capacity.
- Choose enough partitions for required consumer parallelism.
- Include realistic growth margin.
- Avoid excessive partition counts without a capacity reason.
- Configure an appropriate replication factor.
- Distribute partition leaders across brokers.
- Monitor partition-level throughput.
- Monitor consumer lag by partition.
- Monitor partition size and storage growth.
- Alert on under-replicated partitions.
- Test broker failures and leader elections.
- Test consumer rebalancing.
- Load-test skewed partition keys.
- Plan partition-count increases before they become necessary.
- Document the consequences of changing partition count.
- Review downstream capacity before increasing consumer parallelism.
Frequently Asked Questions
Kafka partitions affect storage, ordering, consumer parallelism, replication, and throughput at the same time. Understanding the boundaries of a partition prevents many common Kafka architecture mistakes.
What Is the Difference Between a Topic and a Partition?
A topic is the logical stream to which producers publish records. A partition is one ordered subdivision of that topic.
Topic
│
├── Partition 0
├── Partition 1
└── Partition 2
Every record in the topic belongs to one partition.
Is a Kafka Partition the Same as a Broker?
No. A broker is a Kafka server. A partition is a portion of a topic stored on Kafka brokers.
One broker usually stores many partitions, and replicas of one partition can exist on multiple brokers.
Can Multiple Consumers Read the Same Partition?
Consumers in different consumer groups can independently read the same partition.
Within one consumer group, however, a partition is assigned to at most one consumer at a time. This preserves the group's partition processing model and creates the relationship between partition count and maximum active consumers.
Do More Partitions Break Ordering?
More partitions do not break ordering inside each individual partition. They do mean there is no single total order across all records in the topic.
If related records need ordering, they should be partitioned according to an appropriate stable key.
Can the Number of Partitions Be Changed?
The partition count can be increased, but the change should be planned carefully.
For keyed records, changing the number of partitions can change key-to-partition mapping for future records. Existing records are not automatically reorganized to recreate the previous per-key sequence, so ordering assumptions can be affected.
Conclusion
A Kafka partition is an ordered, append-only portion of a topic. Kafka divides topics into partitions to distribute storage and traffic across brokers and to allow consumer groups to process data in parallel.
Partitions also define Kafka's ordering boundary. Records inside one partition have a defined sequence, while records across different partitions do not have a single global order. Choosing the right partition key therefore determines both data distribution and business-event ordering.
Key takeaway: Kafka partition design is a balance between ordering, parallelism, throughput, distribution, and future growth. The partition count and partition key should be deliberate architecture decisions rather than default configuration values.
Comments (0)