What Is Pub/Sub?
Publish/Subscribe (Pub/Sub) is a messaging pattern where producers publish messages to a topic without knowing which consumers will receive them. Consumers subscribe to topics and receive the messages that are relevant to them.
Pub/Sub is especially useful when one event needs to trigger multiple independent actions. An order.created event, for example, can independently trigger email delivery, analytics, inventory processing, and warehouse integration without the order service calling each system directly.
Table of Contents
- Why Pub/Sub Exists
- How Pub/Sub Works
- Fan-Out
- Pub/Sub vs Message Queue
- Subscriptions and Consumer Groups
- Delivery Guarantees
- Message Ordering
- Retries and Failed Messages
- Event Schema Design
- Scaling Pub/Sub
- Pub/Sub and Event-Driven Architecture
- Popular Pub/Sub Technologies
- When to Use Pub/Sub
- When Not to Use Pub/Sub
- Production Design Example
- Common Pub/Sub Mistakes
- Production Checklist
- Frequently Asked Questions
- Conclusion
Why Pub/Sub Exists
Consider an order service that must notify several systems whenever an order is created.
A direct synchronous design might look like this:
Order Service
│
├──→ Email Service
├──→ Analytics Service
├──→ Inventory Service
└──→ Warehouse Service
The order service now knows about every downstream consumer. Adding another consumer requires another integration, and failures in those dependencies can affect the order workflow.
The coupling increases as the system grows:
def create_order(order):
saved_order = database.save(order)
email_service.order_created(saved_order)
analytics_service.order_created(saved_order)
inventory_service.order_created(saved_order)
warehouse_service.order_created(saved_order)
return saved_order
Pub/Sub changes the relationship.
Order Service
│
│ publish
↓
order.created
│
├──→ Email
├──→ Analytics
├──→ Inventory
└──→ Warehouse
The order service publishes one event. Subscribers decide independently whether they are interested in that event.
A new fraud-detection service can later subscribe without requiring the order service to call it directly.
How Pub/Sub Works
A basic Pub/Sub system contains publishers, topics, and subscribers.
Publisher
│
│ publish
↓
┌───────────────┐
│ Topic │
└───────┬───────┘
│
├──→ Subscriber A
├──→ Subscriber B
└──→ Subscriber C
The publisher does not send a separate request to every subscriber. It publishes to the messaging infrastructure, which is responsible for making the message available to the appropriate subscriptions.
Publisher
A publisher produces an event or message.
For example:
{
"event_id": "evt-91827",
"type": "order.created",
"order_id": "ORD-8421",
"customer_id": "CUS-271",
"total": 149.95,
"currency": "USD",
"occurred_at": "2026-10-05T04:15:00Z",
"schema_version": 1
}
The publisher sends the event to a topic:
publisher.publish(
topic="order-events",
message={
"event_id": "evt-91827",
"type": "order.created",
"order_id": "ORD-8421",
"customer_id": "CUS-271",
},
)
The publisher does not need to know whether there are two subscribers or twenty.
Topic
A topic is a logical destination used to organize messages.
Topics
order-events
payment-events
user-events
inventory-events
Publishers send messages to topics, while subscribers create subscriptions for topics they need to consume.
Topic design should reflect meaningful event domains rather than individual consumers. A topic named order-events remains useful even when subscribers change. A topic named send-events-to-analytics couples the publisher architecture to a specific consumer.
Subscriber
A subscriber consumes messages from a subscription.
def handle_order_created(event):
order_id = event["order_id"]
send_order_confirmation(order_id)
Another subscriber can process the same event differently:
def handle_order_for_analytics(event):
record_order_metric(
order_id=event["order_id"],
total=event["total"],
)
Neither subscriber needs to know about the other.
Fan-Out
The defining Pub/Sub behavior is fan-out: one published message can be delivered to multiple independent subscribers.
┌──→ Email Subscription
│
order.created ──────┼──→ Analytics Subscription
│
├──→ Inventory Subscription
│
└──→ Fraud Subscription
Each subscription represents an independent processing path.
If 10,000 order events are published, each interested subscription can independently process those 10,000 events.
10,000 order events
│
├──→ Email 10,000
├──→ Analytics 10,000
└──→ Warehouse 10,000
This is different from placing 10,000 jobs into a work queue and having several workers compete to process them.
Fan-out is useful because new behavior can often be introduced without changing the original publisher.
Before:
order.created
├──→ Email
└──→ Analytics
Add Fraud Detection:
order.created
├──→ Email
├──→ Analytics
└──→ Fraud Detection
The order service continues publishing the same event.
Pub/Sub vs Message Queue
Pub/Sub and message queues both support asynchronous communication, but their core distribution models differ.
A traditional work queue distributes work among competing consumers:
┌──→ Worker A
Producer → Queue ────┼──→ Worker B
└──→ Worker C
One job → one successful worker
Pub/Sub distributes an event to independent subscriptions:
┌──→ Subscriber A
Publisher → Topic ─────┼──→ Subscriber B
└──→ Subscriber C
One event → multiple subscribers
| Characteristic | Message Queue | Pub/Sub |
|---|---|---|
| Primary purpose | Distribute work | Broadcast events |
| Typical delivery | One successful worker | Each independent subscription |
| Consumers | Often competing | Logically independent |
| Typical example | Resize an image | Order created |
| Adding a consumer | Usually increases worker capacity | Usually adds new behavior |
These models can also be combined. Each Pub/Sub subscription can itself have multiple competing workers.
Message Queues Explained: Producers, Consumers, and Brokers covers producers, consumers, brokers, acknowledgements, buffering, and queue-based worker processing in more detail.
Subscriptions and Consumer Groups
A common misunderstanding is that every individual process must receive every published event.
Usually, the important boundary is the logical subscriber, not the process.
Suppose an email service has four worker instances:
order-events
│
↓
Email Subscription
│
├──→ Email Worker 1
├──→ Email Worker 2
├──→ Email Worker 3
└──→ Email Worker 4
The email service should generally process each order event once as a logical service. Its four workers compete for messages within the email subscription.
Analytics can have a separate subscription:
┌──→ Email Subscription
│ ├── Worker 1
order-events ──────┤ └── Worker 2
│
└──→ Analytics Subscription
├── Worker 1
└── Worker 2
Both logical services receive the event, while workers inside each service divide the work.
Technologies use different names for this concept. Some use subscriptions, some use consumer groups, and others use queues bound to an exchange or topic.
Delivery Guarantees
Pub/Sub does not automatically mean that every event is delivered exactly once.
The messaging system must define what happens when delivery or acknowledgement fails.
Topic
│
↓
Subscriber
│
├── process event
│
└── acknowledgement
│
X connection failure
The broker may not know whether processing completed. Redelivering the event avoids silently losing it, but the subscriber may process it twice.
This is why many production event consumers are designed to be idempotent.
def process_event(event):
event_id = event["event_id"]
if processed_events.exists(event_id):
return
apply_business_change(event)
processed_events.save(event_id)
The business update and deduplication state need appropriate atomicity. Otherwise, a crash between the two operations can still produce inconsistent behavior.
At-most-once, at-least-once, and exactly-once semantics are discussed in Message Delivery Guarantees: At-Most-Once vs At-Least-Once vs Exactly-Once.
Message Ordering
Pub/Sub systems often process many messages concurrently, so global ordering should not be assumed.
Suppose the following customer events are published:
1. customer.created
2. customer.email_changed
3. customer.deleted
If they are processed concurrently, a subscriber might observe:
customer.created
customer.deleted
customer.email_changed
That can create incorrect state.
A common approach is to preserve ordering only for events sharing the same business key:
customer-101 ──→ Partition 1
customer-202 ──→ Partition 3
customer-303 ──→ Partition 2
Events for different customers can run concurrently while events for one customer remain ordered.
The ordering scope should be as narrow as the business requirement allows. Requiring global ordering can severely limit parallel processing.
Retries and Failed Messages
Subscribers will eventually fail to process some events. Network failures, database outages, rate limits, invalid data, and application bugs are unavoidable in production.
Transient failures can be retried:
Event
│
↓
Attempt 1 ──→ failure
│
↓
Backoff
│
↓
Attempt 2 ──→ failure
│
↓
Backoff
│
↓
Attempt 3 ──→ success
Retries should be bounded. A permanently invalid event should not continuously consume subscriber capacity.
After repeated failures, the event can be isolated:
Subscription
│
↓
Consumer
│
├── success ──→ ACK
│
└── repeated failure
│
↓
Dead-Letter Queue
A dead-letter workflow should include monitoring, investigation, and controlled replay rather than becoming permanent storage for forgotten failures.
What Is a Dead Letter Queue? explains this failure-isolation pattern in more detail.
Event Schema Design
Pub/Sub reduces runtime coupling between services, but publishers and subscribers are still coupled through the event contract.
An event should normally describe something that happened:
{
"event_id": "evt-1029",
"type": "payment.completed",
"payment_id": "PAY-741",
"order_id": "ORD-8421",
"amount": 149.95,
"currency": "USD",
"occurred_at": "2026-10-05T04:20:13Z",
"schema_version": 2
}
Stable event contracts are important because publishers and subscribers are deployed independently.
Suppose a publisher immediately renames:
customer_id → account_id
Existing subscribers expecting customer_id can fail.
Schema evolution should therefore favor backward-compatible changes where possible:
- add optional fields instead of immediately removing existing fields;
- use explicit schema versions when useful;
- avoid changing field meaning without changing the contract;
- test important consumers against new event versions;
- define how long old event formats remain supported.
Events should also contain enough information for consumers to perform their job without creating unnecessary synchronous dependencies back to the publisher.
Scaling Pub/Sub
Publishers and subscribers can generally scale independently.
Publishers
P1 P2 P3
\ | /
\ | /
Topic
│
├──→ Subscription A ──→ 4 workers
├──→ Subscription B ──→ 20 workers
└──→ Subscription C ──→ 2 workers
Analytics might need twenty workers while email delivery needs only four. Each subscriber can scale according to its own workload.
A slow subscriber should not necessarily slow other subscribers:
Topic
│
├── Email → healthy
├── Analytics → healthy
└── Warehouse → backlog growing
The warehouse backlog is isolated to its subscription if the messaging infrastructure and architecture provide independent delivery state.
This is a major advantage over synchronous fan-out, where one slow dependency can increase latency for the publisher.
Scaling still has limits. Adding hundreds of consumers can overload a shared database or downstream API. Consumer concurrency should reflect the capacity of the complete processing path.
Pub/Sub and Event-Driven Architecture
Pub/Sub is a common communication mechanism in event-driven systems.
Services publish facts about completed state changes:
Order Service
│
│ order.created
↓
Event Infrastructure
│
├──→ Billing
├──→ Inventory
├──→ Notifications
└──→ Analytics
The publisher owns the fact that an order was created. Subscribers decide what that fact means for their own domain.
This makes it possible to add new reactions without continuously modifying the producer.
However, asynchronous event-driven architecture introduces eventual consistency. The order may already exist while analytics, notifications, and inventory projections are still processing the event.
10:00:00.000 Order stored
10:00:00.020 Event published
10:00:00.080 Inventory updated
10:00:00.150 Email scheduled
10:00:01.400 Analytics updated
The architecture must tolerate these temporary differences.
Event-Driven Architecture in Distributed Systems covers these architectural trade-offs in more depth.
Popular Pub/Sub Technologies
Several messaging platforms support Pub/Sub-style communication, but their architectures and semantics differ.
| Technology | Typical Model | Common Use |
|---|---|---|
| Apache Kafka | Topics, partitions, consumer groups | Durable event streaming |
| Google Cloud Pub/Sub | Topics and subscriptions | Managed cloud messaging |
| Amazon SNS | Topics and subscribers | Fan-out and notifications |
| RabbitMQ | Exchanges, bindings, queues | Broker-based messaging and routing |
| NATS | Subjects and subscribers | Low-latency distributed messaging |
| Redis Pub/Sub | Channels and subscribers | Lightweight real-time messaging |
These technologies should not be selected solely because all can implement some form of Pub/Sub.
Important differences include durability, replay, ordering, message retention, acknowledgement behavior, throughput, routing capabilities, and what happens when subscribers are offline.
When to Use Pub/Sub
Pub/Sub is a strong fit when:
- one event needs to trigger multiple independent actions;
- publishers should not know every consumer;
- new consumers need to be added without modifying publishers;
- services should scale independently;
- event processing can happen asynchronously;
- temporary eventual consistency is acceptable;
- independent teams own different event consumers;
- business events need to feed operational and analytical systems.
Common examples include:
order.created
payment.completed
user.registered
shipment.delivered
document.uploaded
subscription.cancelled
inventory.changed
account.deleted
Each event can have several independent consumers.
When Not to Use Pub/Sub
Pub/Sub adds asynchronous behavior and distributed state. It should not replace simpler communication without a reason.
It may be a poor fit when:
- the caller requires an immediate response from the downstream operation;
- only one simple component needs to perform the work and a work queue is clearer;
- strict cross-service transactional behavior is required;
- the workflow cannot tolerate eventual consistency;
- debugging asynchronous workflows would add more complexity than value;
- the system is small enough that direct communication remains simpler and reliable.
For example, retrieving a user's current profile is naturally request/response:
Client → User API → Database → Response
Publishing profile.requested and waiting for another service to publish profile.loaded would usually add unnecessary complexity.
Production Design Example
Consider an e-commerce platform where successful payment should trigger several operations:
- mark the order as paid;
- send a receipt;
- update loyalty points;
- record analytics;
- start fulfillment;
- run post-payment fraud analysis.
Calling every service directly creates a long dependency chain:
Payment Service
│
├──→ Order Service
├──→ Email Service
├──→ Loyalty Service
├──→ Analytics Service
├──→ Fulfillment Service
└──→ Fraud Service
Instead, the payment service can publish one event:
{
"event_id": "evt-pay-9182",
"type": "payment.completed",
"payment_id": "PAY-9182",
"order_id": "ORD-5519",
"customer_id": "CUS-210",
"amount": 249.00,
"currency": "USD",
"occurred_at": "2026-10-05T04:25:00Z",
"schema_version": 1
}
The resulting architecture becomes:
Payment Service
│
│ payment.completed
↓
Payment Events
│
├──→ Order Subscription
│ └── mark order paid
│
├──→ Email Subscription
│ └── send receipt
│
├──→ Loyalty Subscription
│ └── award points
│
├──→ Analytics Subscription
│ └── record payment
│
├──→ Fulfillment Subscription
│ └── start fulfillment
│
└──→ Fraud Subscription
└── analyze payment
If analytics becomes unavailable, payment processing does not need to fail. Analytics messages can accumulate in that subscription while other subscribers continue normally.
Suppose the analytics consumer is unavailable for five minutes:
Payment Events
│
├──→ Order ✓
├──→ Email ✓
├──→ Loyalty ✓
├──→ Fulfillment ✓
└──→ Analytics ✗
│
↓
backlog grows
When analytics recovers, its consumers process the backlog.
Idempotency is important for business-critical consumers. The loyalty service, for example, must not award points twice if payment.completed is redelivered.
def handle_payment_completed(event):
event_id = event["event_id"]
if processed_events.exists(event_id):
return
with database.transaction():
loyalty.add_points(
customer_id=event["customer_id"],
amount=event["amount"],
)
processed_events.insert(event_id)
Another challenge exists on the publisher side. The payment service might commit the payment and then crash before publishing the event:
1. UPDATE payment SET status = 'completed' ✓
2. COMMIT ✓
3. Publish payment.completed ✗ crash
The database now says the payment completed, but subscribers never receive the event.
The Transactional Outbox pattern addresses this dual-write problem by storing the business change and outgoing event within the same database transaction.
Production monitoring should be performed independently for each subscription:
Subscription Backlog Oldest Event Failure Rate
orders 12 0.4s 0.1%
email 83 2.1s 0.3%
loyalty 7 0.2s 0.0%
analytics 8,412 94.0s 4.8%
fulfillment 21 0.8s 0.1%
A healthy topic does not imply every subscriber is healthy. Backlog age, consumption rate, retries, and dead-letter counts should be monitored per subscription.
Common Pub/Sub Mistakes
- Assuming Pub/Sub guarantees exactly-once business processing. Subscribers should understand the actual delivery semantics and handle duplicates when necessary.
- Publishing events before the business transaction commits. Consumers can react to a state change that later rolls back.
- Committing state and publishing independently without handling the dual-write problem. A crash between the operations can lose events.
- Designing topics around consumers. Topics should normally represent meaningful event domains rather than current service names.
- Putting business logic into the publisher for every subscriber. This recreates tight coupling.
- Assuming global ordering. Distributed processing frequently changes message order unless ordering is explicitly provided.
- Ignoring schema compatibility. Independently deployed subscribers may still consume older event formats.
- Retrying poison messages indefinitely. Permanent failures should be isolated.
- Ignoring slow subscribers. One subscription can accumulate a large backlog while the rest of the system appears healthy.
- Using events as hidden synchronous RPC. Request/reply behavior implemented through several topics can be harder to operate than a normal API.
- Publishing enormous event payloads. Large objects are often better stored externally with a stable reference in the event.
- Creating too many narrowly scoped topics without operational need. Topic sprawl makes governance and monitoring harder.
Production Checklist
- Define clear event names and semantics.
- Use stable event identifiers.
- Include event timestamps where useful.
- Version event schemas.
- Define backward-compatibility rules.
- Keep publishers independent of specific subscribers.
- Define publication reliability requirements.
- Handle the database-and-message dual-write problem.
- Know the broker's actual delivery guarantees.
- Make critical consumers idempotent.
- Define ordering requirements explicitly.
- Use bounded retries with backoff.
- Configure dead-letter handling for permanent failures.
- Define message retention requirements.
- Plan how failed events are replayed.
- Scale subscriptions independently.
- Limit consumers according to downstream capacity.
- Monitor backlog per subscription.
- Monitor oldest unprocessed event age.
- Monitor retry and dead-letter rates.
- Test subscriber outages.
- Test duplicate delivery.
- Test incompatible or malformed events.
- Document ownership for topics and event schemas.
Frequently Asked Questions
Pub/Sub is conceptually simple, but production behavior depends heavily on the messaging technology, delivery guarantees, and subscription model.
Is Pub/Sub Always Asynchronous?
Pub/Sub is generally used for asynchronous communication. Publishers send messages without waiting for every subscriber to complete its work.
This independence is one of the main reasons to use the pattern.
Can Multiple Subscribers Receive the Same Message?
Yes. That is the central fan-out property of Pub/Sub.
An order.created event can independently reach email, analytics, inventory, and fraud subscriptions. Multiple workers inside one subscription can then compete to process that subscription's workload.
What Happens If a Subscriber Is Offline?
It depends on the technology and configuration.
Durable Pub/Sub systems can retain messages for an offline subscription and deliver them after consumers recover. Ephemeral Pub/Sub implementations may discard messages when no active subscriber is listening.
This distinction is critical when choosing messaging infrastructure.
Is Pub/Sub the Same as an Event Bus?
Not exactly. Pub/Sub is a messaging pattern. An event bus is infrastructure that can implement Pub/Sub while also providing features such as event routing, filtering, transformation, integrations, and delivery policies.
Is Kafka Pub/Sub?
Kafka can implement Pub/Sub-style event distribution using topics and consumer groups, but Kafka is more specifically a durable distributed event-streaming platform.
Multiple consumer groups can independently read the same topic, while consumers inside one group divide partitions among themselves. Kafka also retains records independently of normal consumption, which enables replay and differs from many traditional messaging systems.
Apache Kafka Explained: How Kafka Works covers Kafka's topic, partition, offset, and consumer architecture.
Conclusion
Pub/Sub allows publishers to announce events without directly coordinating with every system that reacts to them. Topics provide the communication boundary, subscriptions create independent processing paths, and fan-out allows one event to trigger many different workflows.
The pattern improves decoupling and independent scalability, but it also introduces asynchronous failure handling, duplicate delivery, eventual consistency, schema evolution, ordering concerns, retries, and subscriber backlogs.
Key takeaway: Pub/Sub is most valuable when an event represents a fact that multiple independent components may need to react to, without requiring the publisher to know or synchronously call those components.
Comments (0)