Category: Caching Tags: caching-strategies distributed-caching cache-consistency

Caching Strategies

5.0 out of 5 from 1 votes
By Team4Dev — Published on
1 Likes
0 Dislikes

Caching reduces latency and protects databases by keeping frequently accessed data closer to the application. The difficult part is not adding a cache—it is deciding where data is cached, how it gets there, when it is updated, and how stale data is prevented.

This guide covers the major caching strategies shown in the architecture above: distributed caching, write-through caching, cache-aside, write-back caching, and cache consistency. Each strategy moves complexity to a different part of the system and makes different trade-offs between latency, consistency, durability, and operational cost.

Caching Strategies
Caching Strategies

Table of Contents

Distributed Caching

A distributed cache is a cache shared by multiple application instances. Instead of every server maintaining an isolated copy of cached data, the application uses a cache service or cluster accessible over the network.

Distributed Caching
Distributed Caching

This architecture is common in horizontally scaled applications because requests can reach different application instances while still accessing the same logical cache.

Distributed Cache Architecture

Consider three application servers behind a load balancer:

             Load Balancer
            /      │      \
           ↓       ↓       ↓
        App A    App B    App C
           \       │       /
            \      │      /
             ↓     ↓     ↓
             Cache Cluster
                  │
                  ↓
               Database

If App A caches product:42, App B can access the same cached value on a later request.

This differs from an in-process cache:

In-process cache:
App A → local memory A
App B → local memory B

Distributed cache:
App A ─┐
App B ─┼→ shared cache
App C ─┘

The shared architecture improves cache reuse across application instances, but each cache operation now involves network communication and an external dependency.

Partitioning Cache Data

Large distributed caches commonly partition keys across multiple nodes.

user:123   → Node A
order:781  → Node B
product:42 → Node C

A good partitioning mechanism should distribute keys evenly and minimize unnecessary movement when nodes are added or removed.

Consistent hashing is one technique used to reduce key redistribution when cluster membership changes.

Partitioning introduces its own operational concerns: hot keys, uneven memory consumption, node failures, replication, and rebalancing.

Handling Cache Node Failures

A cache should normally improve application performance rather than become a new single point of failure.

If a cache node disappears, several things can happen depending on the architecture:

  • another replica serves the key;
  • the key becomes a cache miss;
  • the application falls back to the database;
  • the cache client remaps keys to healthy nodes.

The fallback path must be capacity-tested. If the database handles 5,000 queries per second normally but the cache absorbs another 40,000 reads per second, losing the cache can suddenly expose the database to nine times its normal read traffic.

Normal:

45k requests/s
   │
   ├── 40k → Cache
   └──  5k → Database


Cache outage:

45k requests/s
        ↓
     Database

A cache failure can therefore become a database outage if the origin cannot absorb the miss traffic.

Write-Through Caching

Write-through caching updates the cache as part of the write path while ensuring the authoritative database is also updated before the write is considered complete according to the implementation's rules.

Write-Through Caching
Write-Through Caching

The key property is that cache population is coupled to writes rather than waiting for a later read miss.

Write-Through Flow

Suppose a product price changes from $100 to $120.

UPDATE product:42 = $120
          │
          ↓
     Cache updated
          │
          ↓
    Database updated
          │
          ↓
       SUCCESS

A subsequent read can immediately find the new value in the cache:

GET product:42
       ↓
     Cache
       ↓
     $120

Unlike cache-aside, the first read after a successful write does not need to populate the key from the database.

Write-Through Trade-Offs

Write-through caching makes cached data available immediately after writes, but the write path becomes more expensive.

Property Write-Through
Read after write Cache is already populated
Write latency Higher because cache and persistent storage participate
Cold reads after writes Reduced
Cache pollution Can cache data that is written but rarely read
Failure handling Requires clear ordering and error semantics

The most important implementation question is what happens when one write succeeds and the other fails.

Cache write succeeds
        ↓
Database write fails
        ↓
What happens to cached value?

The system needs explicit failure semantics rather than assuming that two separate systems update atomically.

A deeper comparison is available in What Is Write-Through Caching?.

Cache-Aside

Cache-aside, also called lazy loading, keeps cache management in the application. The application checks the cache first and loads missing data from the database.

Cache-Aside
Cache-Aside

It is one of the most common caching strategies because the cache remains optional to the primary persistence path.

Cache-Aside Read Path

For a cache hit:

Application
     │
 GET product:42
     ↓
   Cache
     │
    HIT
     ↓
  Return value

For a cache miss:

Application
     │
     ↓
   Cache
     │
    MISS
     ↓
 Database
     │
     ↓
 Read value
     │
     ↓
Cache SET
     │
     ↓
Return value

A simplified implementation:

def get_product(product_id):
    key = f"product:{product_id}"

    cached = cache.get(key)
    if cached is not None:
        return cached

    product = database.get_product(product_id)

    if product is not None:
        cache.set(key, product, ttl=300)

    return product

Only data that is actually requested enters the cache, which makes cache-aside naturally demand-driven.

Cache-Aside Write Path

A common write strategy is:

UPDATE database
      ↓
COMMIT succeeds
      ↓
DELETE cache key

For example:

def update_product(product_id, changes):
    database.update_product(product_id, changes)

    cache.delete(f"product:{product_id}")

The next reader misses the cache and loads the latest committed value.

Deleting the cached value is often simpler than attempting to update every cached representation of the changed database row.

A product update may affect several keys:

product:42
category:phones:page:1
search:iphone
featured-products
user:123:recommendations

This is why cache invalidation becomes difficult as cached data becomes more derived.

Cache-Aside Trade-Offs

Advantage Trade-Off
Only requested data is cached First request after a miss is slower
Database remains source of truth Application owns cache logic
Cache failure can often fall back to database Fallback can overload database
Simple mental model Concurrent reads and writes can create races

For a focused explanation, see What Is Cache-Aside?.

Write-Back Caching

Write-back caching, also called write-behind caching, acknowledges writes after placing them in the cache or an intermediate durable layer and persists them to the database later.

Write-Back Caching
Write-Back Caching

This can remove database latency from the immediate write path.

Write-Back Flow

Suppose a service receives a large number of counter updates:

+1
+1
+1
+1
+1
 │
 ↓
Cache: +5
 │
 ↓
later flush
 │
 ↓
Database: +5

Instead of performing five immediate database writes, the system can combine or batch them.

This can be useful for write-heavy workloads where temporary delay before persistence is acceptable.

Durability and Failure Risk

Write-back moves the cache or write buffer closer to the role of primary write storage.

That changes the failure model dramatically.

Client WRITE
     ↓
Cache accepts write
     ↓
Client receives SUCCESS
     ↓
Cache node fails
     ↓
Database never received write

If losing that write is unacceptable, the intermediate system must provide appropriate durability, replication, recovery, and acknowledgement guarantees.

Write-back should therefore not be treated as merely "write-through but faster."

It creates a period where acknowledged state may exist outside the primary database.

Typical use cases can include:

  • high-frequency counters;
  • aggregations;
  • telemetry;
  • buffered state updates;
  • workloads where writes can safely be combined.

It is usually a poor default for correctness-critical data unless the buffering layer is explicitly designed as durable infrastructure.

Cache Consistency

Once a database and cache both contain representations of the same logical data, the system has a consistency problem.

Cache Consistency
Cache Consistency

The obvious case is simple staleness:

Database: price = $120
Cache:    price = $100

But concurrency can produce subtler failures even when the application correctly deletes cache entries after database writes.

The Stale Refill Race

Consider two application servers operating concurrently.

Web Server A               Web Server B

GET cache
   ↓
MISS
   ↓
READ database
   ↓
gets version 10
                              UPDATE database
                                  ↓
                              version 11
                                  ↓
                              DELETE cache

SET cache = version 10
   ↓
stale value restored

Server A began reading before the update. Server B then committed a newer database value and invalidated the cache. Server A finally completes its older request and repopulates the cache with stale data.

The result is surprising:

Database = version 11
Cache    = version 10

This is why cache consistency cannot always be reduced to "delete the key after every update."

TTL Is Not a Complete Consistency Strategy

A time to live limits how long a cached entry can survive.

SET product:42
TTL = 300 seconds
      ↓
5 minutes later
      ↓
key expires

TTL is useful as a safety boundary, but it does not prevent stale reads during those five minutes.

If correctness requires a value to change immediately after an update, waiting for expiration is insufficient.

TTL should therefore be chosen according to the application's acceptable staleness, origin capacity, update frequency, and invalidation design.

Versioned Cache Writes

One way to defend against stale overwrites is to associate data with a monotonically increasing version or another ordering mechanism.

Database:
value = X
version = 42


Cache:
value = X
version = 42

A delayed operation trying to write version 41 should not replace version 42.

Cache currently: version 42

Delayed SET:
version 41
    ↓
reject / ignore

The exact implementation depends on the cache, database, and consistency requirements. Other systems use locks, generation keys, event-driven invalidation, carefully ordered writes, or application-specific reconciliation.

Cache Invalidation Strategies

Invalidation determines when cached data is no longer safe to serve.

Common approaches include:

Strategy Behavior Typical Fit
TTL expiration Key expires automatically Bounded staleness is acceptable
Explicit delete Application removes key after change Cache-aside
Cache update Application replaces cached value Simple direct representations
Event-driven invalidation Data-change events invalidate affected keys Distributed services and derived caches
Versioned keys New data uses a new key generation Immutable or version-aware data

The correct strategy depends on the cost of stale data.

A cached avatar URL can tolerate different consistency behavior from a permissions record or product inventory count.

Detailed invalidation patterns are covered in Cache Invalidation Strategies for Production Systems.

Cache Stampedes and Hot Keys

A popular key expiring can cause many application instances to miss simultaneously.

Popular key expires
       ↓
 ┌─────┼─────┬─────┐
 ↓     ↓     ↓     ↓
MISS  MISS  MISS  MISS
 │     │     │     │
 └─────┴──┬──┴─────┘
          ↓
       Database

If 10,000 requests arrive immediately after expiration, the database may receive thousands of identical queries even though only one result is needed.

This is a cache stampede.

Typical defenses include request coalescing, distributed locking, stale-while-revalidate behavior, proactive refresh, and TTL jitter.

TTL jitter avoids many keys expiring at exactly the same time:

Instead of:

TTL = 300
TTL = 300
TTL = 300


Use:

TTL = 287
TTL = 316
TTL = 301

A hot key creates a related problem when a small number of keys receive a disproportionate amount of traffic. Even a well-partitioned cache cluster can become bottlenecked if one node owns an extremely popular key.

Both failure modes are covered in Preventing Cache Stampedes and Hot Keys.

Choosing a Caching Strategy

The best caching strategy depends on whether the workload is read-heavy, write-heavy, consistency-sensitive, and tolerant of temporary data loss.

Requirement Common Starting Point
Read-heavy application with ordinary database-backed objects Cache-aside
Data should already be cached after writes Write-through
Extremely high write rate with acceptable delayed persistence Write-back
Multiple application instances need the same cache Distributed cache
Stale values create correctness problems Explicit consistency and invalidation design
Data changes rarely and tolerates staleness Cache-aside with TTL may be sufficient

A caching strategy does not need to be global.

One application can deliberately use different strategies for different data:

Product descriptions
        ↓
Cache-aside + long TTL


Product price
        ↓
Cache-aside + explicit invalidation


Session metadata
        ↓
Distributed cache


Analytics counters
        ↓
Buffered / write-back style updates


Inventory reservation
        ↓
Database correctness path
with carefully controlled caching

The consistency requirement of the data should determine the strategy, not a preference for one cache pattern.

Production Design Example

Consider an e-commerce platform receiving 60,000 product reads per second during normal traffic and 150,000 reads per second during promotions.

The product database handles writes comfortably but becomes expensive under the read workload.

Normal traffic:

Product reads: 60k/s
Product writes: 500/s


Promotion:

Product reads: 150k/s
Product writes: 1.2k/s

A distributed cache is placed between the application and database.

Clients
   ↓
Load Balancer
   ↓
Application Fleet
   │
   ├────────→ Distributed Cache
   │                │
   │                └── HIT → response
   │
   └── MISS ───────→ Database
                         │
                         ↓
                    cache refill

Product descriptions, images, and other relatively stable metadata use cache-aside with a 30-minute TTL.

Price data also uses cache-aside, but updates explicitly invalidate the product key after the database transaction commits:

UPDATE database price
        ↓
COMMIT
        ↓
DELETE product cache
        ↓
next read repopulates cache

Inventory is treated differently.

The cache can expose inventory information for display purposes, but it does not decide whether an order can actually reserve stock:

Product page:
cache → "12 available"


Checkout:
database transaction
      ↓
conditional inventory update
      ↓
reservation succeeds or fails

This avoids turning potentially stale cache state into the authority for a correctness-sensitive operation.

The initial production metrics are:

Cache hit rate:        96%
Cache GET p95:          2 ms
Database read p95:     28 ms
API p95:               95 ms
Database reads:       2.4k/s

During a promotion, one popular product receives 25,000 requests per second.

The key expires and hundreds of application instances simultaneously query the database:

product:HOT expires
        ↓
hundreds of simultaneous misses
        ↓
hundreds of identical DB queries

Request coalescing is introduced so one worker refreshes the value while other requests wait briefly or receive an acceptable stale value.

TTL jitter is also added to related product keys so a large catalog segment does not expire at once.

A second problem appears during price updates. Monitoring occasionally observes:

Database price: $129
Cache price:    $119

Tracing reveals the stale refill race: an older database read completes after another request updates the product and invalidates the key.

The design is changed to prevent an older version from overwriting a newer cached representation.

The final production design therefore combines several techniques:

  • distributed caching for shared application state;
  • cache-aside for demand-driven population;
  • explicit invalidation for mutable product data;
  • TTL as a secondary safety mechanism;
  • request coalescing for cache stampedes;
  • TTL jitter for synchronized expiration;
  • version-aware writes for race protection;
  • database transactions for inventory correctness;
  • monitoring for cache and database behavior.

This illustrates an important principle: production caching is rarely one pattern used everywhere. It is usually a combination of strategies selected according to the semantics of each data type.

Common Caching Mistakes

  • Caching everything. Data that is cheap to compute or rarely reused may not justify cache complexity.
  • Treating the cache as automatically consistent with the database. Two copies of mutable state require synchronization rules.
  • Using TTL as the only invalidation strategy for correctness-sensitive data.
  • Updating the cache before the database transaction commits. A failed transaction can leave a value that never became authoritative.
  • Ignoring stale refill races. A delayed reader can repopulate an invalidated key with older data.
  • Using write-back without understanding durability. Acknowledged writes can be lost if the buffering layer fails.
  • Allowing unlimited database fallback during a cache outage. The fallback can overload the origin.
  • Using identical TTLs for huge numbers of keys. Synchronized expiration can create traffic spikes.
  • Ignoring hot keys. One key can overload a single cache partition despite plenty of total cluster capacity.
  • Ignoring negative caching. Repeated requests for nonexistent data can still overload the database.
  • Building enormous cache keys or values. Serialization, network transfer, and memory overhead still matter.
  • Monitoring only hit rate. A high hit rate does not prove that the cache is healthy or useful.

Production Checklist

  • Define why each data type is being cached.
  • Choose cache-aside, write-through, or write-back according to write semantics.
  • Define the database or another durable system as the source of truth unless the architecture intentionally says otherwise.
  • Define acceptable staleness for every cached data type.
  • Set TTLs intentionally.
  • Add explicit invalidation where expiration alone is insufficient.
  • Protect against stale refill races where concurrent updates matter.
  • Use bounded key and value sizes.
  • Design cache keys with predictable namespaces.
  • Plan key versioning when schemas or representations change.
  • Use TTL jitter for large groups of expiring keys.
  • Protect popular keys against cache stampedes.
  • Identify and monitor hot keys.
  • Define cache behavior when a node or the entire cluster fails.
  • Capacity-test the database fallback path.
  • Set cache operation timeouts.
  • Avoid excessive retries when the cache is unavailable.
  • Monitor hit rate, miss rate, latency, evictions, memory, errors, and connection usage.
  • Monitor database load alongside cache metrics.
  • Verify write-back durability before acknowledging critical writes.

Frequently Asked Questions

Caching patterns can look similar in architecture diagrams while having very different read, write, and failure semantics. The most important difference is who manages the cache and when the database becomes authoritative.

What Is the Difference Between Cache-Aside and Write-Through?

With cache-aside, the application populates the cache when a read misses. Writes commonly update the database and invalidate the cached value.

With write-through, cache population is part of the write path, so successfully written data is already available in the cache.

Is Write-Back Caching Safe?

It can be, but only if the system is designed around its durability requirements. If a write is acknowledged before reaching the database, the intermediate cache or buffer contains state that must not be casually lost.

Critical workloads therefore need appropriate persistence, replication, recovery, and failure semantics in the write-back layer.

Should the Cache Be the Source of Truth?

Usually not for ordinary cache-aside architectures. The database remains authoritative, and cached values can be discarded and rebuilt.

Some architectures deliberately use durable in-memory systems as authoritative storage, but that is a storage architecture decision rather than ordinary caching.

How Long Should Cache TTL Be?

There is no universal TTL. It depends on update frequency, acceptable staleness, database capacity, value popularity, and invalidation behavior.

Static metadata might tolerate hours. Frequently changing data might use seconds or explicit invalidation. Correctness-sensitive data may require stronger mechanisms than TTL.

When Is a Distributed Cache Necessary?

A distributed cache becomes useful when multiple application instances need access to shared cached data, when the dataset is larger than one process can reasonably hold, or when cache availability and partitioning need to be managed independently from application instances.

Small applications can often begin with no cache or a local in-process cache and introduce distributed caching only when measurements justify the additional infrastructure.

Conclusion

Caching strategies differ primarily in how cached data enters the cache and how changes reach persistent storage. Cache-aside populates data on demand, write-through couples cache population to writes, and write-back delays persistence to optimize the write path. Distributed caching makes cached state available across application instances.

The harder problem is consistency. Expiration, invalidation, concurrent reads and writes, cache stampedes, hot keys, node failures, and database fallback behavior all determine whether a cache remains an optimization or becomes a source of production incidents.

Key takeaway: select caching behavior per data type. Define the source of truth, acceptable staleness, read path, write path, invalidation mechanism, and failure behavior before optimizing for hit rate.

Comments (0)