Category: Messaging & Event Streaming Tags: kafka-consumer-groups apache-kafka consumer-scaling

What Is a Kafka Consumer Group?

By Oleksandr Andrushchenko — Published on
0 Likes
0 Dislikes
What Is a Kafka Consumer Group?
What Is a Kafka Consumer Group?

A Kafka consumer group is a set of consumers that cooperate to process records from Kafka topics. Kafka distributes topic partitions among consumers in the group so that each partition is assigned to at most one consumer in that group at a time.

Consumer groups are how Kafka combines parallel processing with independent consumption. More consumers can divide a topic's partitions, while different consumer groups can independently process the same records for completely different purposes.

Table of Contents

Why Kafka Needs Consumer Groups

Suppose an orders topic receives 20,000 events per second. A single consumer can process only 5,000 events per second.

Kafka Topic
20,000 events/sec
       │
       ↓
Single Consumer
5,000 events/sec

Backlog: +15,000 events/sec

The consumer cannot keep up, so consumer lag continuously increases.

Kafka needs a way to distribute processing across multiple consumer instances. A consumer group provides that mechanism.

orders
  │
  ├── Partition 0 ──→ Consumer A
  ├── Partition 1 ──→ Consumer B
  ├── Partition 2 ──→ Consumer C
  └── Partition 3 ──→ Consumer D

Each consumer processes a different subset of partitions. Together, the consumers behave as one logical application processing the topic.

This provides horizontal scaling without having every consumer process every record.

How a Consumer Group Works

Consumers identify themselves as members of a group through a group ID.

Conceptually:

group.id = fulfillment-service

Suppose the orders topic has six partitions:

orders

P0
P1
P2
P3
P4
P5

Three instances of the fulfillment service join the same consumer group:

Consumer Group: fulfillment-service

Consumer A
Consumer B
Consumer C

Kafka distributes the partitions among them:

P0 ──→ Consumer A
P1 ──→ Consumer A

P2 ──→ Consumer B
P3 ──→ Consumer B

P4 ──→ Consumer C
P5 ──→ Consumer C

Every partition is owned by one consumer in this group, while each consumer can own multiple partitions.

If another application uses a different group ID, it receives its own independent view of the topic.

Partitions and Consumer Assignment

Kafka consumer groups operate around partition ownership.

Within one consumer group, a partition is assigned to at most one consumer at a time:

Consumer Group A

P0 ──→ Consumer 1
P1 ──→ Consumer 2
P2 ──→ Consumer 3

Kafka does not normally split records from P0 between Consumer 1 and Consumer 2 while both consumers belong to the same group and assignment generation.

This matters because a Kafka partition is an ordered log. Keeping one partition assigned to one group member preserves sequential partition consumption while still allowing different partitions to be processed concurrently.

The relationship between partitions, offsets, and ordering is covered in Kafka Topics, Partitions, and Offsets Explained.

Consumer Groups and Parallelism

The number of partitions limits useful consumer parallelism within one consumer group.

Consider a topic with four partitions and one consumer:

P0 ──┐
P1 ──┤
P2 ──┼──→ Consumer A
P3 ──┘

One consumer handles all four partitions.

With two consumers:

P0 ──┐
     ├──→ Consumer A
P1 ──┘

P2 ──┐
     ├──→ Consumer B
P3 ──┘

With four consumers:

P0 ──→ Consumer A
P1 ──→ Consumer B
P2 ──→ Consumer C
P3 ──→ Consumer D

Now add a fifth consumer:

P0 ──→ Consumer A
P1 ──→ Consumer B
P2 ──→ Consumer C
P3 ──→ Consumer D

Consumer E → idle

There is no fifth partition to assign.

The practical relationship is:

Maximum consumers actively owning partitions
within one group
=
number of partitions

This is why partition count is also a consumer-scaling decision. Adding application instances cannot create more partition-level parallelism after every partition already has its own consumer.

Multiple Consumer Groups

The same Kafka topic can be consumed by many independent consumer groups.

Suppose an order-events topic is needed by fulfillment, analytics, and notifications:

                         ┌──→ Fulfillment Group
                         │
order-events ────────────┼──→ Analytics Group
                         │
                         └──→ Notification Group

Each group independently reads the topic.

For a topic with three partitions:

order-events

P0 ──→ Fulfillment Consumer A
P1 ──→ Fulfillment Consumer B
P2 ──→ Fulfillment Consumer C

P0 ──→ Analytics Consumer A
P1 ──→ Analytics Consumer B
P2 ──→ Analytics Consumer C

P0 ──→ Notification Consumer A
P1 ──→ Notification Consumer A
P2 ──→ Notification Consumer A

Each group has its own partition assignments and consumption progress.

This gives Kafka both major messaging behaviors:

  • work sharing inside a consumer group;
  • independent event consumption across consumer groups.

This model is closely related to the fan-out behavior described in Event-Driven Architecture in Distributed Systems.

Consumer Group Offsets

A consumer group needs to remember how far it has processed each partition.

Kafka uses offsets for this purpose.

Partition 0

Offset:
100   101   102   103   104   105
 │     │     │     │     │     │
 A     B     C     D     E     F

             ↑
      committed progress

The consumer group's committed offset records its processing position for the partition.

Different groups can have completely different positions:

Partition 0
0 ───────────────────────────────→ 10,000

Fulfillment Group       ↑ 9,980
Analytics Group         ↑ 8,450
Notification Group      ↑ 9,995

Fulfillment is nearly caught up, analytics is substantially behind, and notifications is almost current.

The groups do not interfere with each other's offsets.

Committed Offsets

A committed offset allows a consumer group to resume processing after a restart or partition reassignment.

Suppose Consumer A processes records through offset 500 and commits its progress:

Processed:
496
497
498
499
500

Committed progress
      │
      ↓
next processing position

If Consumer A later crashes, another group member can receive the partition and continue from the group's committed position rather than starting the entire partition from the beginning.

Offset management is therefore closely connected to processing reliability.

Automatic vs Manual Offset Commits

Consumers can use automatic offset management or explicitly control when processing progress is committed.

The critical question is:

When is a record considered successfully processed?

If an offset is committed before the business operation finishes:

1. Receive record
2. Commit offset      ✓
3. Update database    ✗ crash

the consumer may restart after that record even though the database update never completed.

If the business operation completes before the offset is committed:

1. Receive record
2. Update database    ✓
3. Crash
4. Commit offset      never happened

Kafka can redeliver the record after recovery, potentially causing duplicate processing.

This is one reason idempotent consumers are important in at-least-once processing designs.

Kafka Delivery Semantics: At-Most-Once, At-Least-Once, and Exactly-Once covers how offset handling interacts with processing guarantees.

Consumer Group Rebalancing

Consumer group membership changes over time. Consumers start, stop, fail, and scale.

Kafka must redistribute partitions when the group's membership or subscribed partition set changes. This process is called a rebalance.

Consider four partitions assigned to two consumers:

Before

P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer B
P3 ──→ Consumer B

A third consumer joins:

Consumer C joins
       │
       ↓
   Rebalance

The group can receive a new assignment:

After

P0 ──→ Consumer A
P1 ──→ Consumer B
P2 ──→ Consumer C
P3 ──→ Consumer A

The exact assignment depends on the configured assignment strategy and group protocol behavior.

What Triggers a Rebalance?

Common causes include:

  • a consumer joining the group;
  • a consumer leaving normally;
  • a consumer crashing or becoming unresponsive;
  • consumer instances being added during autoscaling;
  • consumer instances being removed during scale-in;
  • topic partition changes affecting the subscription;
  • subscription changes.

Rebalancing is necessary because Kafka must maintain valid partition ownership as the group changes.

Why Rebalances Matter

Rebalances can temporarily affect processing while ownership moves between consumers.

A system with unstable consumers can enter a pattern like:

Consumer joins
     ↓
Rebalance
     ↓
Processing
     ↓
Consumer fails
     ↓
Rebalance
     ↓
Processing
     ↓
Consumer restarts
     ↓
Rebalance

Frequent rebalances can reduce effective throughput and increase lag.

Production systems should investigate repeated group membership changes instead of treating rebalances as harmless background activity.

What Happens When a Consumer Fails?

Suppose Consumer B owns P2 and P3:

P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer B
P3 ──→ Consumer B

Consumer B crashes.

Consumer B
    X

After Kafka determines that the consumer is no longer an active member, its partitions must be reassigned.

After reassignment

P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer A
P3 ──→ Consumer A

Consumer A can continue P2 and P3 from the group's committed offsets.

If another consumer starts later:

Consumer C joins

P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer C
P3 ──→ Consumer C

This ability to reassign partitions provides consumer-side fault recovery without requiring publishers to know which consumer instance is currently processing a partition.

Consumer Groups and Ordering

Kafka preserves ordering within a partition. Consumer groups preserve that boundary by assigning a partition to one group member at a time.

Partition 2

Offset 501 → order.created
Offset 502 → order.paid
Offset 503 → order.packed
Offset 504 → order.shipped

              │
              ↓
          Consumer B

However, different partitions are processed independently:

Partition 0 ──→ Consumer A
Partition 1 ──→ Consumer B
Partition 2 ──→ Consumer C

There is no global processing order across those three consumers.

Applications that require per-entity ordering should choose a partition key that keeps related records in the same partition.

Kafka Ordering Guarantees and Message Deduplication explains this ordering boundary and its production implications.

Consumer Groups and Consumer Lag

Consumer lag measures how far a consumer group's processing position is behind the latest available records.

Conceptually:

Latest offset:       12,500
Committed offset:    12,100

Lag:                    400

Lag should be considered per partition and per consumer group.

Group: fulfillment

Partition    Latest    Committed    Lag

P0           12,500       12,490     10
P1           11,900       11,895      5
P2           14,200       10,100  4,100
P3           13,600       13,590     10

The group-wide average can hide the fact that P2 has a serious problem.

A growing lag can indicate:

  • insufficient consumer capacity;
  • a slow downstream database or API;
  • a hot partition;
  • repeated processing failures;
  • frequent rebalances;
  • long processing times;
  • consumer pauses or application bugs.

Lag is not just a number of records. A lag of 100,000 records could represent seconds for a high-throughput stream or hours for a low-volume stream.

Record age and recovery rate are often important alongside the raw count.

What Is Kafka Consumer Lag? covers lag calculation, causes, monitoring, and remediation in detail.

Scaling Consumer Groups

A common response to increasing lag is to add more consumer instances.

This works only while partitions are available to assign.

Suppose a topic has twelve partitions and the group has four consumers:

12 partitions
4 consumers

≈ 3 partitions per consumer

Scaling to six consumers provides approximately:

12 partitions
6 consumers

≈ 2 partitions per consumer

Scaling to twelve consumers can allow:

12 partitions
12 consumers

≈ 1 partition per consumer

Scaling to twenty consumers does not create twenty-way partition processing:

12 partitions
20 consumers

12 consumers can own partitions
8 consumers have no partition assignment

Consumer scaling should therefore be planned together with topic partitioning.

More consumers also increase pressure on downstream systems. Doubling consumers can double concurrent database queries, HTTP requests, or writes.

The actual bottleneck may simply move:

Kafka
  │
  ↓
More Consumers
  │
  ↓
Database
  │
  X overloaded

Scaling should be based on end-to-end capacity rather than consumer count alone.

Kafka Performance and Scaling covers partition counts, producer throughput, consumer parallelism, and cluster-level capacity planning.

Consumer Groups and Delivery Semantics

Consumer groups distribute records, but they do not by themselves guarantee exactly-once business processing.

Consider a payment event:

payment.completed
       │
       ↓
Consumer
       │
       ├── Add loyalty points ✓
       │
       └── Commit offset      ✗ crash

After recovery, the consumer group can read the event again because the offset was not committed.

If the consumer blindly adds loyalty points again, the customer receives duplicate points.

An idempotent processing strategy can prevent that:

def handle_payment(event):
    event_id = event["event_id"]

    with database.transaction():
        if processed_events.exists(event_id):
            return

        loyalty.add_points(
            customer_id=event["customer_id"],
            amount=event["amount"],
        )

        processed_events.insert(event_id)

The consumer can safely encounter the same event again without repeating the business effect.

Offset strategy, idempotency, retries, Kafka transactions, and downstream transaction boundaries all contribute to the actual processing semantics.

Production Design Example

Consider an e-commerce platform with an order-events topic containing 24 partitions.

The topic carries:

order.created
order.paid
order.cancelled
order.packed
order.shipped
order.delivered

The platform has three independent applications consuming these events.

order-events
     │
     ├──→ fulfillment-service
     ├──→ analytics-service
     └──→ notification-service

Each application uses its own group ID:

fulfillment-service
analytics-service
notification-service

Fulfillment performs expensive warehouse operations and runs 12 consumers:

24 partitions
12 consumers

≈ 2 partitions per consumer

Analytics performs lightweight batch processing and runs six consumers:

24 partitions
6 consumers

≈ 4 partitions per consumer

Notifications runs eight consumers:

24 partitions
8 consumers

≈ 3 partitions per consumer

All three groups independently receive the complete logical event stream.

At peak traffic, fulfillment begins falling behind:

Consumer Group          Lag

fulfillment-service     185,000
analytics-service         1,200
notification-service        340

The Kafka topic itself is healthy. The problem is isolated to one consumer group.

Partition-level metrics reveal:

Fulfillment Group

Partition    Lag

P0            1,120
P1              980
P2            1,340
P3          142,000  ← hot partition
P4            1,010
...

Simply increasing fulfillment from 12 to 24 consumers may not solve the P3 problem if most of its lag comes from one unusually hot partition. One consumer still owns P3.

The investigation should include:

  • traffic distribution by partition key;
  • processing duration for P3 records;
  • downstream warehouse latency;
  • retry rates;
  • consumer errors;
  • partition assignment history;
  • recent rebalances.

If traffic is balanced and all partitions are moderately behind, scaling from 12 to 18 consumers may improve throughput.

After adding consumers, Kafka redistributes partitions:

12 consumers
      │
      ↓
scale out
      │
      ↓
18 consumers
      │
      ↓
rebalance
      │
      ↓
new partition assignments

Scaling should then be validated against downstream warehouse capacity. Increasing consumer concurrency without checking the downstream service could convert Kafka lag into warehouse overload.

A useful production dashboard should include at least:

Metric Why It Matters
Lag by partition Detects individual partitions falling behind
Oldest unprocessed event age Shows actual processing delay
Consumption rate Shows processing throughput
Processing duration Detects slow handlers
Rebalance frequency Detects unstable group membership
Retry and error rate Detects failed processing
Active consumers Confirms expected capacity
Downstream latency Identifies bottlenecks outside Kafka

Common Consumer Group Mistakes

  • Giving every consumer instance a different group ID. Each instance can receive the entire stream instead of sharing the workload.
  • Using the same group ID for independent applications. The applications divide partitions instead of each receiving all required records.
  • Adding more consumers than partitions. Extra consumers cannot increase partition-level parallelism.
  • Committing offsets before processing is safely complete. A crash can cause business operations to be skipped.
  • Assuming a successful business operation guarantees the offset was committed. Redelivery can still occur.
  • Ignoring idempotency. Duplicate processing can create duplicate payments, emails, inventory changes, or other side effects.
  • Monitoring only total lag. One hot partition can be hidden by healthy partitions.
  • Autoscaling aggressively without considering rebalances. Constant membership changes can reduce useful processing time.
  • Assuming more consumers always improve throughput. The bottleneck may be a database, external API, or one hot partition.
  • Performing very long processing without appropriate consumer configuration. Kafka may interpret an unhealthy or non-polling consumer as unavailable depending on the client behavior and configuration.
  • Ignoring shutdown behavior. Abrupt termination can create unnecessary rebalances and repeated processing.
  • Assuming consumer groups provide global ordering. Ordering remains partition-scoped.

Production Checklist

  • Use a stable group ID for each logical consuming application.
  • Use different group IDs for independent applications that each need the full stream.
  • Choose enough topic partitions for expected consumer parallelism.
  • Avoid running unnecessary consumers beyond the partition count.
  • Define when offsets are committed.
  • Understand failure behavior around business processing and offset commits.
  • Make important side effects idempotent.
  • Use stable event identifiers where deduplication is required.
  • Monitor lag by consumer group and partition.
  • Monitor oldest unprocessed event age.
  • Monitor consumer processing throughput.
  • Monitor processing latency.
  • Monitor rebalance frequency.
  • Monitor consumer membership changes.
  • Monitor retries and processing failures.
  • Test consumer crashes.
  • Test graceful shutdowns.
  • Test rolling deployments.
  • Test consumer scale-out and scale-in.
  • Verify downstream systems can handle increased concurrency.
  • Investigate hot partitions before blindly adding consumers.
  • Document partition keys and ordering requirements.
  • Define alert thresholds based on processing delay, not only raw lag.

Frequently Asked Questions

Consumer groups are straightforward once the distinction between a consumer instance, a partition, and a logical consuming application is clear.

Can Two Consumers in the Same Group Read the Same Partition?

A partition is assigned to at most one consumer within the same consumer group at a time. This is what allows Kafka to divide partition processing among group members.

Consumers in different groups can independently read that same partition.

What Happens If There Are More Consumers Than Partitions?

Some consumers will have no partition assignment.

If a topic has six partitions, at most six consumers in one group can simultaneously own those partitions. Adding a seventh consumer does not create additional partition-level parallelism.

Can Multiple Consumer Groups Read the Same Topic?

Yes. This is a core Kafka pattern.

Analytics, notifications, search indexing, and fulfillment can each use a different consumer group and independently consume the same topic while maintaining separate offsets.

What Is a Kafka Group ID?

The group ID identifies the logical consumer group to Kafka.

Consumers configured with the same group ID cooperate and divide partitions. Consumers with different group IDs belong to independent groups and maintain independent consumption progress.

Is a Consumer Group Like a Message Queue?

Inside one consumer group, Kafka has queue-like work-sharing behavior because partition ownership is distributed among consumers.

Across multiple groups, Kafka has Pub/Sub-like behavior because each group can independently consume the same topic.

This combination is one reason Kafka works well for event-driven architectures that need both scalable processing and multiple independent downstream applications.

Conclusion

A Kafka consumer group represents one logical application consuming Kafka data. Consumers inside the group cooperate by dividing topic partitions, while Kafka tracks the group's progress independently from other consumer groups.

Consumer groups make horizontal processing possible, but their capacity is closely tied to partition count. Offset management determines recovery behavior, rebalances redistribute work when membership changes, and consumer lag reveals whether processing is keeping pace with production.

Key takeaway: use one consumer group for consumers that should share the workload, and separate consumer groups for applications that should independently receive the same Kafka stream.

Comments (0)