What Is a Kafka Offset?
A Kafka offset is a number that identifies the position of a record within a Kafka partition. As records are appended to a partition, Kafka assigns them increasing offsets such as 0, 1, 2, 3, and so on.
Offsets allow consumers to track how far they have processed a partition. They are fundamental to Kafka's ability to resume consumption after failures, replay old records, measure consumer lag, and let multiple consumer groups read the same topic independently.
Table of Contents
- Why Kafka Needs Offsets
- How Kafka Offsets Work
- Offsets Are Partition-Specific
- Consumer Position vs Committed Offset
- How Offset Commits Work
- Offsets and Consumer Groups
- Offsets and Consumer Failures
- Offsets and Delivery Semantics
- Offsets and Consumer Lag
- Replaying Kafka Records
- Offset Reset Behavior
- Offsets and Rebalancing
- Offset Management Strategies
- Production Design Example
- Common Kafka Offset Mistakes
- Production Checklist
- Frequently Asked Questions
- Conclusion
Why Kafka Needs Offsets
A Kafka partition is an ordered, append-only log. As records accumulate, consumers need a way to identify positions inside that log.
Partition 0
Record A → Record B → Record C → Record D → Record E
0 1 2 3 4
Offset
Without offsets, a consumer would have no durable reference for questions such as:
- Which records have already been processed?
- Where should processing resume after a restart?
- How far behind the producer is the consumer?
- How can old records be replayed?
Offsets provide that position.
Suppose a consumer has successfully processed records through offset 1542:
Partition 0
1539 → processed
1540 → processed
1541 → processed
1542 → processed
1543 → next
1544
1545
The consumer can store its progress and later continue around that position instead of rereading the complete partition from the beginning.
How Kafka Offsets Work
Kafka assigns an offset to each record as it is appended to a partition.
orders-0
Offset 0 → order.created
Offset 1 → order.paid
Offset 2 → order.packed
Offset 3 → order.shipped
Offset 4 → order.delivered
The offset increases monotonically within that partition.
A record can therefore be located using three important values:
Topic: orders
Partition: 0
Offset: 3
This identifies a position in the orders topic's partition 0.
The offset describes the record's position in Kafka. It does not indicate when the record was successfully processed by a particular application.
That distinction becomes important because different consumers can be at completely different positions in the same partition.
The relationship between topics, partitions, and offsets is covered in Kafka Topics, Partitions, and Offsets Explained.
Offsets Are Partition-Specific
Offsets are not unique across an entire Kafka topic.
Each partition has its own independent sequence:
Partition 0 Partition 1 Partition 2
0 → A 0 → D 0 → G
1 → B 1 → E 1 → H
2 → C 2 → F 2 → I
There are three different records at offset 1.
Therefore, saying:
offset = 1
is not enough to identify a Kafka record.
The partition is required:
orders / partition 2 / offset 1
This also means consumers track progress separately for every assigned partition.
Consumer Group: fulfillment
Partition 0 → position 18,420
Partition 1 → position 21,118
Partition 2 → position 17,991
Partition 3 → position 20,504
One partition can be significantly behind while the others are nearly current.
Consumer Position vs Committed Offset
Two offset-related concepts are especially important: the consumer's current position and its committed progress.
Imagine that a consumer has fetched and processed several records:
100 → processed
101 → processed
102 → processed
103 → processing
104
105
The consumer may currently be reading around offset 103 while its last durable committed position still represents earlier progress.
Partition
100 101 102 103 104 105
↑ ↑
committed current
progress position
The exact values represented by consumer APIs depend on the operation being discussed, but the architectural distinction is important:
- current position describes where the active consumer is reading;
- committed offset represents durable group progress used for recovery.
A consumer can therefore successfully process records that have not yet been reflected in the group's committed position.
If it crashes during that window, some records can be read again.
How Offset Commits Work
Consumers commit offsets to record durable progress for their consumer group.
Conceptually:
Fetch records
│
↓
Process records
│
↓
Commit progress
│
↓
Continue
Offset commit timing directly affects failure behavior.
Automatic Offset Commits
Kafka clients can be configured to periodically commit offsets automatically.
This is convenient because the application does not explicitly manage every commit.
Consume
│
├── records
├── records
├── records
│
↓
automatic commit
│
├── records
└── ...
The trade-off is reduced control over the relationship between business processing and durable Kafka progress.
If an application performs important side effects, such as charging a customer or modifying inventory, the offset strategy must match the application's required processing semantics.
Manual Offset Commits
Manual commits allow the application to decide when progress is safe to record.
A common pattern is:
1. Receive record
2. Process business operation
3. Confirm successful processing
4. Commit offset
This provides more explicit control, but it does not automatically eliminate duplicates.
Consider:
1. Receive payment.completed
2. Add loyalty points ✓
3. Application crashes ✗
4. Commit offset never happens
After restart, Kafka can deliver the record again because the consumer group's durable progress does not include that completed operation.
Manual commits therefore often work together with idempotent business processing.
Offsets and Consumer Groups
Offsets become particularly powerful when combined with consumer groups.
Different consumer groups maintain independent progress while reading the same topic.
orders / Partition 0
0 ───────────────────────────────────────→ 50,000
Fulfillment Group ↑ 49,950
Analytics Group ↑ 47,200
Search Index Group ↑ 49,980
All three applications read the same partition, but each advances independently.
The analytics group can fall behind without changing the fulfillment group's progress.
This enables Kafka's Pub/Sub-like behavior:
orders
│
├──→ Fulfillment Group
├──→ Analytics Group
└──→ Search Group
Inside each group, consumers divide the partitions. Across groups, applications independently consume the stream.
Kafka Consumers and Consumer Groups Explained covers group membership, partition assignment, rebalancing, and consumer scaling in more detail.
Offsets and Consumer Failures
Committed offsets allow another consumer to continue processing after an instance fails.
Suppose Consumer A owns partition 2:
Partition 2
│
↓
Consumer A
Committed progress: 8,500
Consumer A crashes.
Consumer A
X
Kafka eventually reassigns the partition to another member of the group:
Partition 2
│
↓
Consumer B
Resume from group's
committed progress
Consumer B does not need Consumer A's local memory. The group's durable offset state provides the recovery point.
This decouples processing progress from the lifetime of an individual consumer instance.
Offsets and Delivery Semantics
The order of business processing and offset commits affects whether records can be lost or processed more than once.
Consider committing first:
1. Receive record
2. Commit progress ✓
3. Process business logic ✗ crash
After recovery, the group may continue beyond the record even though its business operation never completed.
Now process first:
1. Receive record
2. Process business logic ✓
3. Crash ✗
4. Commit progress never happens
The record can be processed again after recovery.
This is the fundamental failure window behind many at-most-once and at-least-once designs.
| Strategy | Main Risk | Typical Behavior |
|---|---|---|
| Commit before processing | Lost processing | At-most-once style |
| Process before commit | Duplicate processing | At-least-once style |
| Coordinated transactional processing | More complexity and scope limitations | Can provide stronger guarantees |
Kafka's delivery semantics involve more than offsets alone. Producer configuration, retries, Kafka transactions, downstream storage, and consumer processing all matter.
Kafka Delivery Semantics: At-Most-Once, At-Least-Once, and Exactly-Once covers these trade-offs in detail.
Offsets and Consumer Lag
Offsets make consumer lag measurable.
Conceptually, lag compares the latest available partition position with the consumer group's progress.
Latest position: 150,000
Group position: 147,500
Lag: 2,500 records
For multiple partitions:
Partition Latest Group Position Lag
P0 150,000 149,980 20
P1 148,200 148,150 50
P2 172,400 160,000 12,400
P3 139,500 139,490 10
P2 is clearly falling behind even though the other partitions are healthy.
Lag can grow because:
- producers are publishing faster than consumers can process;
- a downstream database or API is slow;
- one partition receives disproportionate traffic;
- records repeatedly fail and retry;
- consumer instances are unavailable;
- frequent rebalances interrupt useful processing.
Raw lag should be interpreted alongside event age and processing rate. Ten thousand records might represent a few seconds for one system and several hours for another.
What Is Kafka Consumer Lag? covers lag monitoring and troubleshooting in more detail.
Replaying Kafka Records
Kafka retains records independently of whether a consumer has already processed them. This makes it possible to move consumption back to an earlier offset and process retained data again.
Suppose a bug affected analytics processing between offsets 40,000 and 50,000:
0 ────────── 40,000 ────────── 50,000 ────────── 70,000
│ │
└── bad processing ──┘
After fixing the bug, the analytics application can reset its position and replay that data if the records are still retained.
Current position
↓
70,000
Reset
↓
40,000
Replay
40,000 ─────────────────────→ 70,000
This is very different from traditional messaging systems where successful consumption may permanently remove the message from the queue.
Replay is useful for:
- rebuilding search indexes;
- recalculating analytics;
- recovering from consumer bugs;
- rebuilding materialized projections;
- running new consumers against historical events.
Replay must still be safe. Reprocessing events that trigger emails, payments, or other external side effects can produce duplicates unless those operations are designed accordingly.
Offset Reset Behavior
A consumer does not always have a valid committed offset.
This can happen when a consumer group is new or when previously committed offsets no longer point to retained data.
Kafka consumers therefore need a policy for determining where to begin when no usable offset exists.
Two common choices are conceptually:
earliest
↓
start from oldest available retained records
latest
↓
start near the current end of the partition
The choice can dramatically change application behavior.
An analytics application may want historical records:
earliest retained record
│
↓
A → B → C → D → E → F → current
A real-time-only application may care only about newly arriving records:
A → B → C → D → E → F → current
│
↓
start here
Offset reset configuration should therefore be treated as part of the application's recovery design rather than an arbitrary client default.
Offsets and Rebalancing
Consumer groups redistribute partitions when membership changes.
Before a rebalance:
P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer B
P3 ──→ Consumer B
Consumer B fails:
Consumer B
X
After reassignment:
P0 ──→ Consumer A
P1 ──→ Consumer A
P2 ──→ Consumer C
P3 ──→ Consumer C
Consumer C needs to know where the group previously stopped processing P2 and P3. Committed offsets provide that durable handoff point.
If processing is substantially ahead of committed progress when ownership changes, the new consumer can repeat some work.
If progress was committed before corresponding business work safely completed, processing can instead be skipped.
Offset management and rebalance behavior therefore need to be designed together.
Offset Management Strategies
The correct offset strategy depends on the workload.
For a consumer performing idempotent database updates, a common design is:
Receive batch
│
↓
Process records
│
↓
Persist business changes
│
↓
Commit Kafka progress
Duplicates remain possible around failures, but idempotency makes them safe.
For independent records, processing and commits can sometimes happen in batches to improve throughput.
Offsets 100-199
│
↓
process batch
│
↓
commit progress
Larger batches reduce commit overhead but increase the amount of potential repeated work after a crash.
Smaller commit intervals reduce the replay window but increase coordination overhead.
The design is therefore a trade-off among:
- throughput;
- duplicate-processing tolerance;
- recovery time;
- processing latency;
- business transaction boundaries.
For important side effects, offsets should never be treated as a substitute for application-level idempotency.
Production Design Example
Consider an e-commerce platform with an order-events topic containing 12 partitions.
A fulfillment consumer group processes:
order.paid
order.cancelled
order.ready_for_fulfillment
The group runs six consumers:
12 partitions
6 consumers
≈ 2 partitions per consumer
Suppose Consumer C owns P4 and P5:
Consumer C
P4 → committed progress around 84,210
P5 → committed progress around 91,775
Consumer C receives an event from P4:
{
"event_id": "evt-91827",
"type": "order.ready_for_fulfillment",
"order_id": "ORD-4812"
}
The consumer creates a fulfillment job in PostgreSQL.
def process_event(event):
with database.transaction():
if processed_events.exists(event["event_id"]):
return
fulfillment_jobs.create(
order_id=event["order_id"],
)
processed_events.insert(event["event_id"])
Only after the application considers the processed records safe does it advance durable Kafka progress according to its chosen commit strategy.
Now consider a crash:
1. Read event at offset 84,211
2. Create fulfillment job ✓
3. Store event as processed ✓
4. Database commit ✓
5. Consumer crashes ✗
6. Kafka offset commit not completed
After the group recovers, offset 84,211 can be delivered again.
The idempotency check sees the existing event_id and prevents a second fulfillment job.
Redelivery
│
↓
event_id already processed?
│
yes
│
↓
skip duplicate business effect
This design accepts at-least-once delivery while protecting the business operation from duplicate execution.
The group should also expose offset-related operational metrics:
| Metric | Purpose |
|---|---|
| Lag by partition | Find partitions falling behind |
| Oldest unprocessed event age | Measure real processing delay |
| Consumption rate | Measure processing throughput |
| Commit failures | Detect progress-tracking problems |
| Processing failures | Detect business-handler failures |
| Rebalance frequency | Detect unstable group membership |
Suppose monitoring reports:
Partition Latest Group Position Lag
P0 92,010 92,001 9
P1 88,421 88,415 6
P2 94,870 94,861 9
P3 90,114 90,110 4
P4 96,700 84,211 12,489
P5 91,800 91,775 25
P4 requires investigation. Looking only at a group-wide average could hide the affected partition.
Potential causes include a hot partition, unusually slow records, downstream database latency, repeated failures, or a consumer that cannot keep up.
Common Kafka Offset Mistakes
- Thinking offsets are global across a topic. Every partition has its own offset sequence.
- Treating an offset as a globally unique message ID. Topic and partition context are required, and business deduplication usually needs a separate event identifier.
- Committing before business processing is safely complete. A failure can cause records to be skipped from the application's perspective.
- Assuming committing after processing prevents duplicates. A crash between processing and the commit can cause redelivery.
- Confusing consumer position with durable committed progress. An active consumer can be ahead of its recovery point.
- Assuming commits delete records. Kafka retention is independent of one consumer group's progress.
- Resetting offsets without understanding side effects. Replay can repeat emails, payments, notifications, or other operations.
- Using raw lag as the only health metric. Event age and lag growth rate can provide more useful operational context.
- Ignoring per-partition lag. One partition can be unhealthy while topic-wide averages appear normal.
- Relying on offsets instead of idempotency. Offsets track consumption progress; they do not make external business operations exactly once.
- Using an inappropriate offset reset policy. A new group can unexpectedly replay historical data or skip retained history.
Production Checklist
- Understand that offsets are scoped to partitions.
- Use stable consumer group IDs.
- Define exactly when processing is considered successful.
- Choose automatic or manual commits deliberately.
- Document the failure window around processing and commits.
- Make critical business operations idempotent.
- Use stable event IDs where deduplication is required.
- Define the offset reset policy explicitly.
- Test consumer restarts.
- Test crashes before an offset commit.
- Test crashes after a business operation completes.
- Test partition reassignment during rebalances.
- Test offset resets before using them in production.
- Protect replay workflows from unsafe side effects.
- Monitor lag by partition and consumer group.
- Monitor oldest unprocessed event age.
- Monitor consumption and production rates.
- Monitor commit errors.
- Monitor processing failures and retries.
- Alert on sustained lag growth rather than isolated spikes.
- Keep retention long enough for required recovery and replay windows.
Frequently Asked Questions
Offsets are simple numbers, but several different Kafka concepts use those positions. Distinguishing record offsets, consumer positions, and committed group progress is important for understanding consumer behavior.
Is a Kafka Offset Global?
No. An offset is scoped to a partition.
Partition 0 can contain offset 100 while partition 1 also contains offset 100. A Kafka position therefore needs topic and partition context.
Is an Offset a Message ID?
An offset identifies a record's position within a partition, but it should generally not be treated as a business-level event identifier.
If an application needs deduplication across processing workflows, a stable event ID in the event itself is usually more appropriate.
Does Committing an Offset Delete the Record?
No. Kafka records are retained according to the topic's retention or compaction configuration.
One consumer group committing progress does not remove the record for other groups and does not normally remove it from the partition.
Can a Consumer Process the Same Offset Again?
Yes. A record can be processed again after a failure, an offset reset, or an intentional replay.
This is why important Kafka consumers are often designed to tolerate duplicate delivery.
Where Does Kafka Store Consumer Offsets?
Kafka stores committed consumer-group offsets in its internal offset-management infrastructure, commonly associated with the __consumer_offsets internal topic.
This allows group progress to survive the failure or replacement of an individual consumer instance.
Conclusion
A Kafka offset identifies a record's position inside a partition and provides the foundation for tracking consumer progress. Consumer groups use committed offsets to recover after failures, transfer partition ownership, measure lag, and independently consume the same Kafka data.
Offset management also defines important failure boundaries. Committing too early can skip unfinished work, while committing after processing can result in duplicate processing after a crash. Reliable applications therefore combine an appropriate commit strategy with idempotent business operations.
Key takeaway: an offset tells Kafka where a consumer group is in a partition; it does not by itself guarantee that a business operation happened exactly once.
Comments (0)