Category: Messaging & Event Streaming Tags: kafka-offsets apache-kafka consumer-processing

What Is a Kafka Offset?

By Oleksandr Andrushchenko — Published on
0 Likes
0 Dislikes
What Is a Kafka Offset?
What Is a Kafka Offset?

A Kafka offset is a number that identifies the position of a record within a Kafka partition. As records are appended to a partition, Kafka assigns them increasing offsets such as 0, 1, 2, 3, and so on.

Offsets allow consumers to track how far they have processed a partition. They are fundamental to Kafka's ability to resume consumption after failures, replay old records, measure consumer lag, and let multiple consumer groups read the same topic independently.

Table of Contents

Why Kafka Needs Offsets

A Kafka partition is an ordered, append-only log. As records accumulate, consumers need a way to identify positions inside that log.

Partition 0

Record A → Record B → Record C → Record D → Record E
   0          1          2          3          4
                       Offset

Without offsets, a consumer would have no durable reference for questions such as:

  • Which records have already been processed?
  • Where should processing resume after a restart?
  • How far behind the producer is the consumer?
  • How can old records be replayed?

Offsets provide that position.

Suppose a consumer has successfully processed records through offset 1542:

Partition 0

1539 → processed
1540 → processed
1541 → processed
1542 → processed
1543 → next
1544
1545

The consumer can store its progress and later continue around that position instead of rereading the complete partition from the beginning.

How Kafka Offsets Work

Kafka assigns an offset to each record as it is appended to a partition.

orders-0

Offset 0 → order.created
Offset 1 → order.paid
Offset 2 → order.packed
Offset 3 → order.shipped
Offset 4 → order.delivered

The offset increases monotonically within that partition.

A record can therefore be located using three important values:

Topic:     orders
Partition: 0
Offset:    3

This identifies a position in the orders topic's partition 0.

The offset describes the record's position in Kafka. It does not indicate when the record was successfully processed by a particular application.

That distinction becomes important because different consumers can be at completely different positions in the same partition.

The relationship between topics, partitions, and offsets is covered in Kafka Topics, Partitions, and Offsets Explained.

Offsets Are Partition-Specific

Offsets are not unique across an entire Kafka topic.

Each partition has its own independent sequence:

Partition 0          Partition 1          Partition 2

0 → A                0 → D                0 → G
1 → B                1 → E                1 → H
2 → C                2 → F                2 → I

There are three different records at offset 1.

Therefore, saying:

offset = 1

is not enough to identify a Kafka record.

The partition is required:

orders / partition 2 / offset 1

This also means consumers track progress separately for every assigned partition.

Consumer Group: fulfillment

Partition 0 → position 18,420
Partition 1 → position 21,118
Partition 2 → position 17,991
Partition 3 → position 20,504

One partition can be significantly behind while the others are nearly current.

Consumer Position vs Committed Offset

Two offset-related concepts are especially important: the consumer's current position and its committed progress.

Imagine that a consumer has fetched and processed several records:

100 → processed
101 → processed
102 → processed
103 → processing
104
105

The consumer may currently be reading around offset 103 while its last durable committed position still represents earlier progress.

Partition

100   101   102   103   104   105
             ↑     ↑
         committed current
          progress position

The exact values represented by consumer APIs depend on the operation being discussed, but the architectural distinction is important:

  • current position describes where the active consumer is reading;
  • committed offset represents durable group progress used for recovery.

A consumer can therefore successfully process records that have not yet been reflected in the group's committed position.

If it crashes during that window, some records can be read again.

How Offset Commits Work

Consumers commit offsets to record durable progress for their consumer group.

Conceptually:

Fetch records
     │
     ↓
Process records
     │
     ↓
Commit progress
     │
     ↓
Continue

Offset commit timing directly affects failure behavior.

Automatic Offset Commits

Kafka clients can be configured to periodically commit offsets automatically.

This is convenient because the application does not explicitly manage every commit.

Consume
  │
  ├── records
  ├── records
  ├── records
  │
  ↓
automatic commit
  │
  ├── records
  └── ...

The trade-off is reduced control over the relationship between business processing and durable Kafka progress.

If an application performs important side effects, such as charging a customer or modifying inventory, the offset strategy must match the application's required processing semantics.

Manual Offset Commits

Manual commits allow the application to decide when progress is safe to record.

A common pattern is:

1. Receive record
2. Process business operation
3. Confirm successful processing
4. Commit offset

This provides more explicit control, but it does not automatically eliminate duplicates.

Consider:

1. Receive payment.completed
2. Add loyalty points      ✓
3. Application crashes     ✗
4. Commit offset           never happens

After restart, Kafka can deliver the record again because the consumer group's durable progress does not include that completed operation.

Manual commits therefore often work together with idempotent business processing.

Offsets and Consumer Groups

Offsets become particularly powerful when combined with consumer groups.

Different consumer groups maintain independent progress while reading the same topic.

orders / Partition 0

0 ───────────────────────────────────────→ 50,000

Fulfillment Group             ↑ 49,950
Analytics Group          ↑ 47,200
Search Index Group             ↑ 49,980

All three applications read the same partition, but each advances independently.

The analytics group can fall behind without changing the fulfillment group's progress.

This enables Kafka's Pub/Sub-like behavior:

orders
  │
  ├──→ Fulfillment Group
  ├──→ Analytics Group
  └──→ Search Group

Inside each group, consumers divide the partitions. Across groups, applications independently consume the stream.

Kafka Consumers and Consumer Groups Explained covers group membership, partition assignment, rebalancing, and consumer scaling in more detail.

Offsets and Consumer Failures

Committed offsets allow another consumer to continue processing after an instance fails.

Suppose Consumer A owns partition 2:

Partition 2
     │
     ↓
Consumer A

Committed progress: 8,500

Consumer A crashes.

Consumer A
    X

Kafka eventually reassigns the partition to another member of the group:

Partition 2
     │
     ↓
Consumer B

Resume from group's
committed progress

Consumer B does not need Consumer A's local memory. The group's durable offset state provides the recovery point.

This decouples processing progress from the lifetime of an individual consumer instance.

Offsets and Delivery Semantics

The order of business processing and offset commits affects whether records can be lost or processed more than once.

Consider committing first:

1. Receive record
2. Commit progress        ✓
3. Process business logic ✗ crash

After recovery, the group may continue beyond the record even though its business operation never completed.

Now process first:

1. Receive record
2. Process business logic ✓
3. Crash                  ✗
4. Commit progress        never happens

The record can be processed again after recovery.

This is the fundamental failure window behind many at-most-once and at-least-once designs.

Strategy Main Risk Typical Behavior
Commit before processing Lost processing At-most-once style
Process before commit Duplicate processing At-least-once style
Coordinated transactional processing More complexity and scope limitations Can provide stronger guarantees

Kafka's delivery semantics involve more than offsets alone. Producer configuration, retries, Kafka transactions, downstream storage, and consumer processing all matter.

Kafka Delivery Semantics: At-Most-Once, At-Least-Once, and Exactly-Once covers these trade-offs in detail.

Offsets and Consumer Lag

Offsets make consumer lag measurable.

Conceptually, lag compares the latest available partition position with the consumer group's progress.

Latest position:   150,000
Group position:    147,500

Lag:                 2,500 records

For multiple partitions:

Partition    Latest    Group Position    Lag

P0           150,000       149,980         20
P1           148,200       148,150         50
P2           172,400       160,000     12,400
P3           139,500       139,490         10

P2 is clearly falling behind even though the other partitions are healthy.

Lag can grow because:

  • producers are publishing faster than consumers can process;
  • a downstream database or API is slow;
  • one partition receives disproportionate traffic;
  • records repeatedly fail and retry;
  • consumer instances are unavailable;
  • frequent rebalances interrupt useful processing.

Raw lag should be interpreted alongside event age and processing rate. Ten thousand records might represent a few seconds for one system and several hours for another.

What Is Kafka Consumer Lag? covers lag monitoring and troubleshooting in more detail.

Replaying Kafka Records

Kafka retains records independently of whether a consumer has already processed them. This makes it possible to move consumption back to an earlier offset and process retained data again.

Suppose a bug affected analytics processing between offsets 40,000 and 50,000:

0 ────────── 40,000 ────────── 50,000 ────────── 70,000
                 │                 │
                 └── bad processing ──┘

After fixing the bug, the analytics application can reset its position and replay that data if the records are still retained.

Current position
      ↓
    70,000

Reset
      ↓
    40,000

Replay
40,000 ─────────────────────→ 70,000

This is very different from traditional messaging systems where successful consumption may permanently remove the message from the queue.

Replay is useful for:

  • rebuilding search indexes;
  • recalculating analytics;
  • recovering from consumer bugs;
  • rebuilding materialized projections;
  • running new consumers against historical events.

Replay must still be safe. Reprocessing events that trigger emails, payments, or other external side effects can produce duplicates unless those operations are designed accordingly.

Offset Reset Behavior

A consumer does not always have a valid committed offset.

This can happen when a consumer group is new or when previously committed offsets no longer point to retained data.

Kafka consumers therefore need a policy for determining where to begin when no usable offset exists.

Two common choices are conceptually:

earliest
   ↓
start from oldest available retained records


latest
   ↓
start near the current end of the partition

The choice can dramatically change application behavior.

An analytics application may want historical records:

earliest retained record
        │
        ↓
A → B → C → D → E → F → current

A real-time-only application may care only about newly arriving records:

A → B → C → D → E → F → current
                            │
                            ↓
                       start here

Offset reset configuration should therefore be treated as part of the application's recovery design rather than an arbitrary client default.

Offsets and Rebalancing

Consumer groups redistribute partitions when membership changes.

Before a rebalance:

P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer B
P3 ──→ Consumer B

Consumer B fails:

Consumer B
    X

After reassignment:

P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer C
P3 ──→ Consumer C

Consumer C needs to know where the group previously stopped processing P2 and P3. Committed offsets provide that durable handoff point.

If processing is substantially ahead of committed progress when ownership changes, the new consumer can repeat some work.

If progress was committed before corresponding business work safely completed, processing can instead be skipped.

Offset management and rebalance behavior therefore need to be designed together.

Offset Management Strategies

The correct offset strategy depends on the workload.

For a consumer performing idempotent database updates, a common design is:

Receive batch
     │
     ↓
Process records
     │
     ↓
Persist business changes
     │
     ↓
Commit Kafka progress

Duplicates remain possible around failures, but idempotency makes them safe.

For independent records, processing and commits can sometimes happen in batches to improve throughput.

Offsets 100-199
       │
       ↓
process batch
       │
       ↓
commit progress

Larger batches reduce commit overhead but increase the amount of potential repeated work after a crash.

Smaller commit intervals reduce the replay window but increase coordination overhead.

The design is therefore a trade-off among:

  • throughput;
  • duplicate-processing tolerance;
  • recovery time;
  • processing latency;
  • business transaction boundaries.

For important side effects, offsets should never be treated as a substitute for application-level idempotency.

Production Design Example

Consider an e-commerce platform with an order-events topic containing 12 partitions.

A fulfillment consumer group processes:

order.paid
order.cancelled
order.ready_for_fulfillment

The group runs six consumers:

12 partitions
6 consumers

≈ 2 partitions per consumer

Suppose Consumer C owns P4 and P5:

Consumer C

P4 → committed progress around 84,210
P5 → committed progress around 91,775

Consumer C receives an event from P4:

{
  "event_id": "evt-91827",
  "type": "order.ready_for_fulfillment",
  "order_id": "ORD-4812"
}

The consumer creates a fulfillment job in PostgreSQL.

def process_event(event):
    with database.transaction():
        if processed_events.exists(event["event_id"]):
            return

        fulfillment_jobs.create(
            order_id=event["order_id"],
        )

        processed_events.insert(event["event_id"])

Only after the application considers the processed records safe does it advance durable Kafka progress according to its chosen commit strategy.

Now consider a crash:

1. Read event at offset 84,211
2. Create fulfillment job       ✓
3. Store event as processed     ✓
4. Database commit              ✓
5. Consumer crashes             ✗
6. Kafka offset commit          not completed

After the group recovers, offset 84,211 can be delivered again.

The idempotency check sees the existing event_id and prevents a second fulfillment job.

Redelivery
    │
    ↓
event_id already processed?
    │
   yes
    │
    ↓
skip duplicate business effect

This design accepts at-least-once delivery while protecting the business operation from duplicate execution.

The group should also expose offset-related operational metrics:

Metric Purpose
Lag by partition Find partitions falling behind
Oldest unprocessed event age Measure real processing delay
Consumption rate Measure processing throughput
Commit failures Detect progress-tracking problems
Processing failures Detect business-handler failures
Rebalance frequency Detect unstable group membership

Suppose monitoring reports:

Partition    Latest    Group Position    Lag

P0           92,010       92,001           9
P1           88,421       88,415           6
P2           94,870       94,861           9
P3           90,114       90,110           4
P4           96,700       84,211      12,489
P5           91,800       91,775          25

P4 requires investigation. Looking only at a group-wide average could hide the affected partition.

Potential causes include a hot partition, unusually slow records, downstream database latency, repeated failures, or a consumer that cannot keep up.

Common Kafka Offset Mistakes

  • Thinking offsets are global across a topic. Every partition has its own offset sequence.
  • Treating an offset as a globally unique message ID. Topic and partition context are required, and business deduplication usually needs a separate event identifier.
  • Committing before business processing is safely complete. A failure can cause records to be skipped from the application's perspective.
  • Assuming committing after processing prevents duplicates. A crash between processing and the commit can cause redelivery.
  • Confusing consumer position with durable committed progress. An active consumer can be ahead of its recovery point.
  • Assuming commits delete records. Kafka retention is independent of one consumer group's progress.
  • Resetting offsets without understanding side effects. Replay can repeat emails, payments, notifications, or other operations.
  • Using raw lag as the only health metric. Event age and lag growth rate can provide more useful operational context.
  • Ignoring per-partition lag. One partition can be unhealthy while topic-wide averages appear normal.
  • Relying on offsets instead of idempotency. Offsets track consumption progress; they do not make external business operations exactly once.
  • Using an inappropriate offset reset policy. A new group can unexpectedly replay historical data or skip retained history.

Production Checklist

  • Understand that offsets are scoped to partitions.
  • Use stable consumer group IDs.
  • Define exactly when processing is considered successful.
  • Choose automatic or manual commits deliberately.
  • Document the failure window around processing and commits.
  • Make critical business operations idempotent.
  • Use stable event IDs where deduplication is required.
  • Define the offset reset policy explicitly.
  • Test consumer restarts.
  • Test crashes before an offset commit.
  • Test crashes after a business operation completes.
  • Test partition reassignment during rebalances.
  • Test offset resets before using them in production.
  • Protect replay workflows from unsafe side effects.
  • Monitor lag by partition and consumer group.
  • Monitor oldest unprocessed event age.
  • Monitor consumption and production rates.
  • Monitor commit errors.
  • Monitor processing failures and retries.
  • Alert on sustained lag growth rather than isolated spikes.
  • Keep retention long enough for required recovery and replay windows.

Frequently Asked Questions

Offsets are simple numbers, but several different Kafka concepts use those positions. Distinguishing record offsets, consumer positions, and committed group progress is important for understanding consumer behavior.

Is a Kafka Offset Global?

No. An offset is scoped to a partition.

Partition 0 can contain offset 100 while partition 1 also contains offset 100. A Kafka position therefore needs topic and partition context.

Is an Offset a Message ID?

An offset identifies a record's position within a partition, but it should generally not be treated as a business-level event identifier.

If an application needs deduplication across processing workflows, a stable event ID in the event itself is usually more appropriate.

Does Committing an Offset Delete the Record?

No. Kafka records are retained according to the topic's retention or compaction configuration.

One consumer group committing progress does not remove the record for other groups and does not normally remove it from the partition.

Can a Consumer Process the Same Offset Again?

Yes. A record can be processed again after a failure, an offset reset, or an intentional replay.

This is why important Kafka consumers are often designed to tolerate duplicate delivery.

Where Does Kafka Store Consumer Offsets?

Kafka stores committed consumer-group offsets in its internal offset-management infrastructure, commonly associated with the __consumer_offsets internal topic.

This allows group progress to survive the failure or replacement of an individual consumer instance.

Conclusion

A Kafka offset identifies a record's position inside a partition and provides the foundation for tracking consumer progress. Consumer groups use committed offsets to recover after failures, transfer partition ownership, measure lag, and independently consume the same Kafka data.

Offset management also defines important failure boundaries. Committing too early can skip unfinished work, while committing after processing can result in duplicate processing after a crash. Reliable applications therefore combine an appropriate commit strategy with idempotent business operations.

Key takeaway: an offset tells Kafka where a consumer group is in a partition; it does not by itself guarantee that a business operation happened exactly once.

Comments (0)