Category: Messaging & Event Streaming Tags: poison-messages dead-letter-queues message-processing

What Is a Poison Message?

By Oleksandr Andrushchenko — Published on
0 Likes
0 Dislikes

A poison message is a message that a consumer repeatedly fails to process successfully. The failure may come from invalid data, an incompatible schema, a permanent business-rule violation, corrupted content, or a bug triggered by that specific message.

Poison messages are dangerous because normal retry logic cannot fix a permanent failure. Without a retry limit or dead-letter strategy, one bad message can consume resources, increase queue lag, block ordered processing, and create an endless failure loop.

Poison Message Lifecycle Example
Poison Message Lifecycle Example

Table of Contents

How a Message Becomes Poison

Consider an order-processing consumer receiving messages from a queue:

Queue

Message A → Consumer → success
Message B → Consumer → success
Message C → Consumer → failure
Message D
Message E

A single failure does not necessarily mean Message C is poison. The database might be temporarily unavailable or an external API might have timed out.

The consumer retries:

Message C
   │
   ↓
Attempt 1 → failure
   │
   ↓
Attempt 2 → failure
   │
   ↓
Attempt 3 → failure
   │
   ↓
Attempt 4 → failure

If the message consistently fails because of something intrinsic to the message or the code path it triggers, retries are unlikely to help.

For example:

{
  "event_id": "evt-8192",
  "type": "order.created",
  "order_id": "ORD-4812",
  "total": "INVALID"
}

If the consumer expects total to be numeric, every processing attempt can fail in exactly the same way.

At that point, the message behaves as a poison message.

Transient vs Permanent Failures

The most important decision in message failure handling is determining whether another attempt has a reasonable chance of succeeding.

Failure Type Retry Usually Helps?
Database connection timeout Transient Yes
HTTP 503 from downstream service Transient Yes
Temporary rate limit Transient Yes, with delay
Invalid JSON Permanent No
Missing required field Usually permanent No
Unsupported schema version Usually permanent until code changes No immediate retry benefit
Application bug triggered by message Permanent until fixed No immediate retry benefit

A transient failure is often worth retrying:

Message
   │
   ↓
Database unavailable
   │
   ↓
wait
   │
   ↓
retry
   │
   ↓
Database recovered
   │
   ↓
success

A permanent failure behaves differently:

Invalid Message
   │
   ↓
retry
   │
   ↓
same error
   │
   ↓
retry
   │
   ↓
same error
   │
   ↓
...

The retry system must eventually stop treating a permanent failure like a temporary outage.

Why Poison Messages Are Dangerous

A poison message is rarely dangerous because of one failed processing attempt. The problem is what happens when failure handling has no boundary.

An unlimited retry loop can consume worker capacity:

Worker 1 → poison message → fail → retry
Worker 2 → poison message → fail → retry
Worker 3 → poison message → fail → retry

This can produce several secondary failures:

  • CPU and network resources are wasted on retries that cannot succeed;
  • healthy messages wait longer;
  • queue depth or consumer lag grows;
  • logs and alerts become noisy;
  • downstream dependencies receive repeated invalid requests;
  • ordered streams can stop progressing;
  • autoscaling can add workers that only perform more failed retries.

A single malformed message can therefore become a system-level reliability problem.

Common Causes of Poison Messages

Poison messages can originate from producers, consumers, deployment changes, or external data.

Common causes include:

  • malformed JSON or another invalid serialization format;
  • missing required fields;
  • incorrect field types;
  • unsupported event versions;
  • producer and consumer schema incompatibility;
  • unexpected enum values;
  • payloads exceeding application assumptions;
  • invalid business data;
  • references to resources that permanently do not exist;
  • consumer bugs triggered by particular values;
  • corrupted payloads;
  • old events being replayed into newer incompatible consumers.

For example, a producer might introduce a new event version:

{
  "event_type": "payment.completed",
  "schema_version": 3,
  "payment_id": "PAY-9182"
}

If a consumer only understands versions 1 and 2, version 3 may fail every time it is processed.

This is why schema evolution is part of messaging reliability, not merely an event-format concern.

Retrying Poison Messages

Retries are necessary in distributed systems because many failures are temporary. The problem is not retrying; the problem is retrying without limits or without distinguishing recoverable failures from permanent ones.

A safer flow is:

Message
   │
   ↓
Process
   │
   ├── success → acknowledge
   │
   └── failure
          │
          ↓
       retry?
       │    │
      yes   no
       │    │
       ↓    ↓
     delay  DLQ

Bounded Retries

A retry policy should have a maximum number of attempts.

MAX_ATTEMPTS = 5

def handle(message):
    try:
        process(message)
        acknowledge(message)
    except Exception:
        if message.attempts >= MAX_ATTEMPTS:
            send_to_dead_letter_queue(message)
            acknowledge(message)
        else:
            schedule_retry(message)

The exact number depends on the workload, but the important property is that failure eventually reaches a terminal path.

After the retry budget is exhausted, repeatedly executing the same failing operation usually provides no additional value.

Retry Backoff

Immediate retries can make temporary failures worse.

Downstream service fails
        │
        ↓
10,000 messages fail
        │
        ↓
10,000 immediate retries
        │
        ↓
service receives even more traffic

Delayed retries with exponential backoff reduce this pressure:

Attempt 1 → fail
wait 1s

Attempt 2 → fail
wait 2s

Attempt 3 → fail
wait 4s

Attempt 4 → fail
wait 8s

Backoff is especially useful for timeouts, temporary unavailability, throttling, and overloaded dependencies.

Retry policies should also consider jitter to prevent large numbers of consumers from retrying at exactly the same moment.

Dead-Letter Queues

A dead-letter queue or dead-letter topic provides a destination for messages that cannot be processed successfully after the configured failure policy is exhausted.

Main Queue
    │
    ↓
Consumer
    │
    ├── success ──→ done
    │
    └── failure
           │
           ↓
         retry
           │
           ↓
      retries exhausted
           │
           ↓
          DLQ

This removes the poison message from the normal processing path while preserving it for investigation or later recovery.

A useful dead-letter record should contain enough context to diagnose the failure:

{
  "original_message": {
    "event_id": "evt-8192",
    "type": "order.created",
    "order_id": "ORD-4812"
  },
  "failure": {
    "error_type": "ValidationError",
    "error_message": "total must be numeric",
    "attempts": 5
  }
}

Production systems often need additional metadata such as the original topic or queue, partition and offset when applicable, timestamps, consumer version, schema version, and correlation identifiers.

A DLQ should not become permanent unmonitored storage. What Is a Dead Letter Queue? covers the pattern and its operational responsibilities in more detail.

Poison Messages and Ordering

Poison messages become particularly difficult when strict ordering is required.

Consider an ordered sequence:

Offset 100 → order.created
Offset 101 → order.paid
Offset 102 → order.packed
Offset 103 → order.shipped

If offset 101 cannot be processed, simply skipping it and continuing with offset 102 may violate business assumptions.

100 → success
101 → poison
102 → should this run?
103 → should this run?

There is no universal answer. The correct policy depends on the ordering requirement.

Possible strategies include:

  • block the partition until the failed message is resolved;
  • move the message to a DLQ and continue processing;
  • pause only the affected partition or entity;
  • route the entity's subsequent events to a recovery workflow;
  • repair the message and replay it before dependent events continue.

Blocking preserves strict ordering but allows one poison message to stop progress. Skipping improves availability but may violate state transitions.

Kafka systems must consider this carefully because ordering is partition-scoped. Kafka Ordering Guarantees and Message Deduplication explains the partition ordering boundary in more detail.

Validation and Schema Evolution

Many poison messages can be prevented before they enter the messaging system.

Producer-side validation can reject malformed events:

def publish_order_created(event):
    if not event.get("order_id"):
        raise ValueError("order_id is required")

    if not isinstance(event.get("total"), (int, float)):
        raise ValueError("total must be numeric")

    broker.publish("order.created", event)

Consumer-side validation is still required because consumers can encounter historical records, events from multiple producers, and versions created before current validation existed.

Events should also evolve in backward-compatible ways whenever possible.

For example, adding an optional field is generally easier to support than changing the type of an existing field:

Safer evolution:

{
  "order_id": "ORD-1",
  "total": 42.50,
  "currency": "USD"
}


Riskier evolution:

"total": 42.50

becomes

"total": {
  "amount": 42.50,
  "currency": "USD"
}

Schema compatibility becomes especially important when producers and consumers are deployed independently.

For Kafka systems, Kafka Schema Evolution and Event Versioning covers these compatibility decisions in detail.

Idempotency During Retries

A message does not need to be poison to be processed multiple times. Any retry can repeat work that partially or completely succeeded before the failure was detected.

Consider:

1. Receive payment.completed
2. Add loyalty points       ✓
3. Acknowledgement fails    ✗
4. Message delivered again
5. Add loyalty points again ?

The retry itself can create a duplicate business effect.

A common protection is a stable event ID:

def process_payment(event):
    event_id = event["event_id"]

    with database.transaction():
        if processed_events.exists(event_id):
            return

        loyalty.add_points(
            customer_id=event["customer_id"],
            amount=event["amount"],
        )

        processed_events.insert(event_id)

Retries can then safely encounter an event that already completed.

This distinction is important:

  • retry policy determines whether processing is attempted again;
  • idempotency controls whether repeated attempts create duplicate effects;
  • dead-letter handling determines what happens when processing cannot succeed.

Idempotency and Deduplication in Distributed Systems covers duplicate-safe processing patterns in more detail.

Poison Message Handling Strategies

Different failures need different responses. A single retry policy for every exception is usually too simplistic.

Failure Typical Strategy
Network timeout Retry with backoff
HTTP 503 Retry with backoff and jitter
Rate limit Delayed retry
Invalid JSON Dead-letter immediately
Unsupported schema Dead-letter or hold for compatible consumer
Permanent business validation failure Dead-letter or reject
Unknown application exception Bounded retries, then dead-letter

A practical failure classifier might look like:

def handle(message):
    try:
        process(message)

    except InvalidPayloadError as exc:
        dead_letter(message, exc)

    except UnsupportedSchemaError as exc:
        dead_letter(message, exc)

    except TemporaryDependencyError:
        retry_with_backoff(message)

    except Exception as exc:
        if retry_budget_exhausted(message):
            dead_letter(message, exc)
        else:
            retry_with_backoff(message)

The exact exception hierarchy is application-specific, but separating clearly permanent failures from retryable failures prevents unnecessary work.

Production Design Example

Consider an e-commerce system processing payment.completed events.

Each event triggers three operations:

  • mark the order as paid;
  • create a fulfillment request;
  • record the processed event ID.

The normal flow is:

Payment Event
     │
     ↓
Main Queue
     │
     ↓
Consumer
     │
     ↓
Validate
     │
     ↓
Process
     │
     ↓
Acknowledge

Now a malformed event arrives:

{
  "event_id": "evt-78219",
  "type": "payment.completed",
  "order_id": null,
  "payment_id": "PAY-9918"
}

The consumer requires order_id, so the message cannot be processed.

The system classifies this as a permanent validation error:

Message
   │
   ↓
Validation
   │
   X missing order_id
   │
   ↓
Permanent failure
   │
   ↓
DLQ

There is no reason to make five identical attempts. The message can be dead-lettered immediately.

A different message fails because the fulfillment database times out:

Attempt 1 → DB timeout
              │
              ↓
           wait 1s

Attempt 2 → DB timeout
              │
              ↓
           wait 2s

Attempt 3 → success

This failure is transient, so retrying is useful.

Now consider an unexpected application bug:

Attempt 1 → unexpected exception
Attempt 2 → unexpected exception
Attempt 3 → unexpected exception
Attempt 4 → unexpected exception
Attempt 5 → unexpected exception
              │
              ↓
             DLQ

The message is preserved together with diagnostic metadata:

{
  "event_id": "evt-99201",
  "source": "payment-events",
  "failure_type": "UnexpectedProcessingError",
  "attempts": 5,
  "consumer_version": "2026.10.5",
  "failed_at": "2026-10-05T20:15:00Z"
}

An alert fires because the DLQ received messages.

After investigation, the bug is fixed and deployed. The affected messages are then replayed through a controlled recovery process:

DLQ
 │
 ↓
Review / Repair
 │
 ↓
Replay
 │
 ↓
Main Processing
 │
 ↓
Success

The recovery tool should preserve stable event IDs so that any partially completed earlier attempts remain safe through idempotent processing.

Production monitoring should include:

Metric Why It Matters
Processing failure rate Detects unhealthy consumers or messages
Retry rate Shows dependency or processing instability
Retries per message Identifies repeatedly failing messages
DLQ message count Detects messages requiring intervention
DLQ oldest message age Detects unresolved failures
Queue depth or consumer lag Shows whether failures affect throughput
Failure type Separates transient and permanent problems

Common Poison Message Mistakes

  • Retrying forever. Permanent failures never leave the normal processing path.
  • Retrying immediately. Consumers can overload an already unhealthy dependency.
  • Using the same policy for every error. Invalid payloads and network timeouts should not necessarily receive identical treatment.
  • Sending every first failure directly to a DLQ. Temporary failures that would recover after a retry become unnecessary manual work.
  • Creating a DLQ without monitoring it. Failed messages silently accumulate.
  • Storing only the original payload. Missing error and retry metadata makes investigation harder.
  • Logging sensitive payloads indiscriminately. Failure diagnostics can expose credentials or personal data.
  • Ignoring ordering requirements. Skipping one event may make later events invalid.
  • Replaying DLQ messages blindly. The underlying problem may still exist.
  • Ignoring idempotency during retries and replay. Recovery can duplicate business side effects.
  • Assuming every repeated failure is bad data. A deterministic consumer bug can poison otherwise valid messages.

Production Checklist

  • Classify failures as retryable or permanent where possible.
  • Set a maximum retry count.
  • Use backoff for transient failures.
  • Add jitter when many workers may retry together.
  • Dead-letter messages after the retry budget is exhausted.
  • Dead-letter clearly invalid payloads without unnecessary retries.
  • Preserve the original message for investigation.
  • Store failure type and error context.
  • Store retry count and failure timestamps.
  • Preserve correlation and event IDs.
  • Include source topic, queue, partition, or other routing context when useful.
  • Monitor DLQ depth and message age.
  • Alert on unusual retry and dead-letter rates.
  • Validate producer payloads.
  • Validate consumer inputs.
  • Define schema compatibility rules.
  • Make retried business operations idempotent.
  • Define how ordering behaves when a message cannot be processed.
  • Provide a controlled DLQ replay process.
  • Fix or classify the root cause before replaying failed messages.
  • Test malformed messages in non-production environments.
  • Test consumer bugs and repeated exceptions.
  • Test downstream outages and recovery.
  • Protect logs and DLQ payloads containing sensitive data.

Frequently Asked Questions

Poison messages are primarily a failure-handling concept. The exact broker mechanics differ, but the underlying problem is the same: a particular message cannot successfully move through the normal processing path.

Is Every Failed Message a Poison Message?

No. A message can fail because of a temporary database outage, network timeout, rate limit, or another transient problem.

A poison message is associated with a failure that persists across normal processing attempts and cannot be resolved by ordinary retry behavior.

Should a Poison Message Be Retried Forever?

No. Unlimited retries waste resources and can prevent healthy work from progressing.

Retries should normally have a defined budget. After that budget is exhausted, the message should move to an explicit failure path such as a dead-letter queue.

Should a Poison Message Be Deleted?

Immediately discarding the message loses information that may be required to diagnose the failure or recover the business operation.

Moving it to a dead-letter queue or another controlled failure store usually provides a safer operational path.

Does a Dead-Letter Queue Fix the Problem?

No. A DLQ isolates the failed message so normal processing can continue. It does not repair malformed data, incompatible schemas, consumer bugs, or invalid business state.

The underlying cause still needs investigation and remediation.

Can Poison Messages Be Reprocessed?

Yes, after the underlying cause has been fixed or the message has been safely repaired.

Replay should be controlled and observable, and business operations should tolerate duplicate delivery when an earlier attempt may have partially succeeded.

Conclusion

A poison message is a message that repeatedly fails normal consumer processing because retrying alone cannot resolve the underlying problem. Without proper handling, one bad message can waste worker capacity, increase lag, overload dependencies, or block ordered processing.

Reliable messaging systems distinguish transient failures from permanent ones, use bounded retries with backoff, isolate unrecoverable messages in a dead-letter path, monitor those failures, and provide a controlled recovery process.

Key takeaway: retries are for failures that may recover; poison messages need a terminal failure path that isolates them without losing the information required to diagnose and safely replay them.

Comments (0)