What Is Service Discovery?

By Oleksandr Andrushchenko — Published on — Modified on
0 Likes
0 Dislikes

Service Discovery is the mechanism that allows applications to find the network locations of other services dynamically. Instead of hardcoding IP addresses and ports, a service asks a discovery system where another service is currently available.

What Is Service Discovery?
What Is Service Discovery?

Service discovery becomes especially important in microservices, containers, Kubernetes, autoscaling environments, and cloud infrastructure where service instances are continuously created, replaced, scaled, and moved.

Table of Contents

Why Service Discovery Exists

Consider an Order service that needs to call a Payment service.

In a simple environment, the Payment service could have a fixed address:

http://10.0.2.15:8080

The Order service could store that address in configuration and send every request there.

This becomes unreliable once the application starts running multiple dynamically managed instances.

Payment Service

10.0.2.15:8080
10.0.3.21:8080
10.0.4.37:8080

An autoscaler may create another instance:

10.0.5.12:8080

Later, 10.0.2.15 may be terminated and replaced with:

10.0.6.44:8080

Hardcoded addresses quickly become stale.

The caller instead needs a stable logical identity:

payment-service

Service discovery maps that logical service name to the currently available service instances.

How Service Discovery Works

The basic process has two sides: services become discoverable, and callers resolve those services before communicating with them.

Payment Instance
      ↓
Register
      ↓
Service Registry
      ↑
Lookup
      ↑
Order Service

The registry might contain:

payment-service
  ├── 10.0.3.21:8080
  ├── 10.0.4.37:8080
  └── 10.0.5.12:8080

The Order service does not need to know when those instances were created or where they are running.

It asks for payment-service and receives one or more reachable endpoints.

A complete discovery mechanism therefore needs to solve several problems:

  • how instances register;
  • how instances are removed;
  • how callers resolve service names;
  • how unhealthy instances are excluded;
  • how discovery data is cached;
  • how traffic is distributed across discovered instances.

Service Registry

A service registry maintains information about currently available service instances.

A simplified registry entry could look like:

{
  "service": "payment-service",
  "instances": [
    {
      "host": "10.0.3.21",
      "port": 8080,
      "status": "healthy"
    },
    {
      "host": "10.0.4.37",
      "port": 8080,
      "status": "healthy"
    }
  ]
}

The registry can also contain metadata such as:

  • service version;
  • availability zone;
  • region;
  • protocol;
  • health status;
  • deployment environment;
  • instance capabilities.

Callers can use this information to make smarter routing decisions.

For example, a service may prefer an instance in the same availability zone to reduce latency and cross-zone traffic.

Service Registration

Before an instance can be discovered, its location must become known to the discovery system.

There are two common registration models: self-registration and third-party registration.

Self-Registration

With self-registration, each application instance registers itself when it starts.

Service Starts
     ↓
Register Address
     ↓
Service Registry
     ↓
Send Heartbeats
     ↓
Service Stops
     ↓
Deregister

The service might conceptually send:

{
  "service": "payment-service",
  "host": "10.0.3.21",
  "port": 8080
}

The advantage is that the service knows when it is ready and can control its registration lifecycle.

The disadvantage is that discovery-specific logic becomes part of every service. Applications need registry clients, heartbeat logic, retries, authentication, and deregistration behavior.

Third-Party Registration

With third-party registration, infrastructure registers and removes service instances.

Service Instance
      ↓
Container / VM Platform
      ↓
Registration Controller
      ↓
Service Registry

The application itself does not need to understand the registry protocol.

This model is common in orchestrated environments because the platform already knows when containers, tasks, pods, or virtual machines are created and terminated.

Client-Side Service Discovery

With client-side discovery, the caller queries the service registry and selects an instance itself.

             Service Registry
                    ↑
                  Lookup
                    │
Order Service ──────┘
     ↓
Select Instance
     ↓
Payment Service

The registry may return:

10.0.3.21:8080
10.0.4.37:8080
10.0.5.12:8080

The client then applies a selection strategy such as:

  • round robin;
  • random selection;
  • least connections;
  • latency-aware routing;
  • zone-aware routing;
  • weighted routing.

This gives the client significant control over routing.

The trade-off is increased application complexity. Every client needs discovery and load-balancing behavior, which can be difficult when services are written in several programming languages.

Server-Side Service Discovery

With server-side discovery, clients send requests to a stable intermediary rather than selecting service instances themselves.

Order Service
     ↓
Load Balancer
     ↓
Service Discovery
     ↓
Payment Instance

The client only needs a stable endpoint:

https://payment.internal

The load balancer, reverse proxy, service mesh, or platform networking layer resolves the available backend instances.

This keeps discovery logic out of application code.

Server-side discovery is especially attractive when many services use different programming languages because the discovery implementation remains in infrastructure rather than being duplicated across client libraries.

DNS-Based Service Discovery

DNS can provide a simple form of service discovery by mapping a stable hostname to one or more service addresses.

payment.internal
      ↓
     DNS
      ↓
10.0.3.21
10.0.4.37
10.0.5.12

The application connects to the logical hostname instead of storing individual IP addresses.

DNS-based discovery is attractive because almost every application already knows how to resolve DNS names.

However, DNS introduces caching behavior.

If an instance disappears but an application or operating system still has its address cached, requests can continue reaching the stale endpoint until the cached record expires.

TTL configuration therefore becomes part of the discovery trade-off:

Long TTL
→ fewer DNS queries
→ slower reaction to topology changes

Short TTL
→ faster discovery updates
→ more DNS traffic

DNS architecture, caching, and resolution behavior are covered further in Domain Name System (DNS): Overview and Use Cases.

Health Checks and Instance Removal

Discovery data is useful only when it represents instances capable of receiving traffic.

Suppose the registry contains:

payment-service

10.0.3.21 → healthy
10.0.4.37 → unhealthy
10.0.5.12 → healthy

The unhealthy instance should stop receiving new requests.

This requires a mechanism for detecting failures.

Common approaches include:

  • periodic heartbeats from service instances;
  • active health checks from infrastructure;
  • container readiness information;
  • lease expiration;
  • connection and request failure feedback.

There is an important distinction between alive and ready.

A service process may be running while still unable to handle requests because it is warming caches, applying startup logic, waiting for dependencies, or draining during shutdown.

This distinction is explored in Health Checks, Readiness, and Liveness Probes.

Service Discovery and Load Balancing

Service discovery and load balancing solve related but different problems.

Service Discovery
→ Which instances exist?

Load Balancing
→ Which available instance should receive this request?

Discovery may return four healthy instances:

payment-service
  ├── Instance A
  ├── Instance B
  ├── Instance C
  └── Instance D

A load-balancing algorithm then selects one of them.

For example:

Request 1 → A
Request 2 → B
Request 3 → C
Request 4 → D
Request 5 → A

Discovery must therefore react quickly enough to topology changes that the load balancer does not continue routing traffic to terminated instances.

The routing side of this architecture is covered in Load Balancing Explained: Distributing Traffic at Scale.

Service Discovery in Kubernetes

Kubernetes provides built-in service discovery around Services and cluster DNS.

Pods are ephemeral. Their IP addresses can change whenever workloads are restarted, rescheduled, or scaled.

payment Pod A → 10.244.1.8
payment Pod B → 10.244.2.5

Pod A terminated

payment Pod C → 10.244.3.11

Calling Pod IP addresses directly would couple clients to this constantly changing topology.

A Kubernetes Service provides a stable logical identity:

payment-service
      ↓
Kubernetes Service
      ↓
Healthy Payment Pods

Another application can use a stable DNS name rather than individual Pod addresses.

http://payment-service

The Kubernetes networking layer then directs traffic toward the current endpoints behind that Service.

This allows Pods to be replaced or scaled without requiring callers to update their configuration.

The broader relationship between Services, application endpoints, and Kubernetes networking is covered in Services, Ingress, and Networking.

Caching Service Discovery Results

Querying a registry for every application request can create unnecessary latency and turn the registry into a high-throughput dependency.

Clients often cache discovery results.

Request
   ↓
Local Discovery Cache
   │
   ├── hit → use endpoints
   │
   └── expired → query registry

Caching improves performance and can allow applications to continue operating temporarily if the registry becomes unavailable.

But cached discovery information can become stale.

Suppose a client caches:

10.0.3.21
10.0.4.37
10.0.5.12

Then 10.0.4.37 is terminated.

Until the cache refreshes, the client may still attempt to use that address.

Discovery caching therefore needs a balance between:

  • lookup overhead;
  • topology-change speed;
  • registry availability;
  • stale endpoint tolerance.

Clients should usually combine cached discovery with timeouts, connection-error handling, and retries against a different healthy instance rather than assuming cached data is always current.

Failure Scenarios

Service discovery is part of the request path even when its lookup happens indirectly, so its failure behavior needs explicit design.

The registry becomes unavailable. Existing clients may continue using cached endpoints while new services or topology changes temporarily remain undiscoverable.

An instance crashes without deregistering. Health checks or lease expiration must eventually remove it.

A healthy instance is incorrectly marked unhealthy. Capacity decreases even though the application itself still works.

An unhealthy instance remains registered. Requests continue failing until discovery information is corrected.

DNS information is stale. Clients may continue connecting to an old address until the relevant cache expires.

The registry contains an instance before it is ready. Traffic reaches a process that has started but cannot yet serve requests.

An instance is removed too early during shutdown. In-flight requests may be interrupted if connection draining is not coordinated with deregistration.

These are distributed-system failures, so discovery should be designed together with timeouts, retries, health checks, load balancing, and graceful shutdown.

Production Design Example

Consider an e-commerce platform with Order, Payment, Inventory, and Notification services.

Each service runs multiple instances:

Order Service
  ├── O1
  ├── O2
  └── O3

Payment Service
  ├── P1
  ├── P2
  ├── P3
  └── P4

Inventory Service
  ├── I1
  └── I2

The number and location of instances change throughout the day as workloads scale.

The Order service should not contain configuration such as:

PAYMENT_HOST=10.0.4.37

Instead, it uses a logical service identity:

payment-service

The request path becomes:

Order Service
     ↓
Resolve payment-service
     ↓
Discovery / DNS
     ↓
Healthy Payment Endpoints
     ↓
Load Balancing
     ↓
Payment Instance

Suppose four Payment instances are healthy:

P1 → healthy
P2 → healthy
P3 → healthy
P4 → healthy

The platform starts two more instances because traffic increases:

P5 → starting
P6 → starting

They should not immediately receive production requests.

Only after readiness succeeds do they become eligible endpoints:

P5 → ready → discoverable
P6 → ready → discoverable

Later, P2 begins failing health checks:

P2
 ↓
Health Check Fails
 ↓
Removed from Healthy Endpoints
 ↓
New Requests Stop Reaching P2

The infrastructure can restart or replace P2 while the remaining instances continue serving requests.

During a deployment, old and new versions may briefly coexist:

payment-service

v1:
  P1
  P2

v2:
  P3
  P4

Discovery metadata can help infrastructure distinguish instance versions, but deployment traffic policy should remain explicit rather than depending accidentally on registration order.

Useful production metrics include:

  • registered instance count per service;
  • healthy instance count;
  • registration and deregistration rate;
  • health-check failures;
  • discovery lookup latency;
  • registry error rate;
  • DNS resolution failures;
  • stale endpoint failures;
  • endpoint changes per minute;
  • traffic distribution across instances.

One especially useful alert compares expected capacity with discoverable healthy capacity. Ten running instances are not useful if only three are registered and ready to receive traffic.

When Service Discovery Is Needed

Service discovery becomes valuable when service locations change frequently or when several instances represent the same logical service.

Common situations include:

  • microservice architectures;
  • container orchestration;
  • Kubernetes clusters;
  • autoscaling workloads;
  • ephemeral virtual machines or containers;
  • multi-zone deployments;
  • internal service-to-service communication;
  • dynamic worker or backend pools.

A small application with a few stable endpoints may not need a dedicated service registry. DNS plus a load balancer can often provide all necessary indirection.

The goal is not to add discovery infrastructure everywhere. The goal is to prevent callers from depending on network locations that are expected to change.

Common Service Discovery Mistakes

  • Hardcoding service IP addresses. Dynamic infrastructure eventually makes them stale.
  • Registering instances before they are ready. Startup traffic reaches applications that cannot yet serve requests.
  • Confusing liveness with readiness. A running process is not necessarily a healthy request destination.
  • Never expiring registrations. Crashed instances remain discoverable indefinitely.
  • Ignoring stale client caches. Removed instances may continue receiving connection attempts.
  • Querying the registry for every request. This adds latency and unnecessary registry load.
  • Caching discovery forever. Clients stop reacting to topology changes.
  • Assuming discovery replaces load balancing. Finding instances and selecting an instance are separate responsibilities.
  • Ignoring graceful shutdown. An instance can be terminated while clients are still sending requests to it.
  • Making the registry a single point of failure. Discovery infrastructure itself needs high availability.
  • Ignoring network partitions. Different clients may temporarily observe different registry state.
  • Monitoring only running instances. Running, registered, healthy, and ready are different states.

Frequently Asked Questions

Service discovery overlaps with DNS, load balancing, health checking, and container orchestration, so the boundaries between these components are easy to confuse.

Is Service Discovery the Same as DNS?

No. DNS is one mechanism that can be used to implement service discovery.

Service discovery is the broader problem of mapping a logical service identity to currently available instances. DNS can provide that mapping, while dedicated registries or orchestration platforms can provide richer metadata and health information.

Is Service Discovery the Same as Load Balancing?

No. Service discovery determines which instances are available. Load balancing determines which available instance should receive a particular request.

The two are commonly integrated, which is why the distinction may not be visible to application code.

Does Kubernetes Provide Service Discovery?

Yes. Kubernetes Services provide stable identities for groups of Pods, while cluster DNS allows workloads to resolve those Services by name.

This hides ephemeral Pod addresses from callers and allows the backend Pod set to change as workloads scale, restart, or move between nodes.

Does Every Microservice Need Service Discovery?

A microservice needs some way to locate its dependencies, but that does not mean every application needs to run its own discovery client or dedicated service registry.

In many environments, DNS, load balancers, Kubernetes Services, or platform networking provide discovery transparently. A dedicated registry is useful when the infrastructure requires richer registration, metadata, routing, or health-management capabilities.

Conclusion

Service discovery allows applications to communicate using stable logical service identities while the underlying instances, IP addresses, and infrastructure change dynamically.

A reliable discovery design needs more than a registry. Registration, health checks, readiness, caching, load balancing, deregistration, failure handling, and graceful shutdown all determine whether callers actually reach healthy service instances.

The core principle is: applications should depend on stable service identities, not on ephemeral network locations.

Author

Enjoyed this article?

Support Oleksandr Andrushchenko

Buy me a coffee

This helps Oleksandr Andrushchenko continue creating useful content

Related articles

Comments (0)