Back to Blog

When a Cache Miss Becomes a Database Incident

How cache stampedes overload databases, and how single-flight loading, Redis locks, stale responses, TTL jitter, and observability prevent them.

Illustration for When a Cache Miss Becomes a Database Incident

The GET /api/trending-products endpoint is working well.

Redis serves the popular response in a few milliseconds. The database is healthy. Traffic is normal.

Then a five-minute cache entry expires.

Within a second, a large number of requests arrive for the same key. They all see a cache miss. They all run the same expensive database query. The connection pool fills, latency rises, and client retries add even more load.

Nothing was wrong with the cache while it existed. The problem began when the cache disappeared.

This is a cache stampede. It is a concurrency problem that can turn one expired hot key into a database incident.

A cache miss should not mean that every request is allowed to rebuild the same value.

What a cache stampede is

Assume this cache key:

Key: trending_products
TTL: 5 minutes

While the key exists, the request path is simple:

1,000 requests
      │
      ▼
    Redis
      │
      ▼
Cached response

At the moment it expires, the same traffic can become this:

Request A ─┐
Request B ─┤
Request C ─┼── cache miss ──► identical database queries
Request D ─┤
Request E ─┘

The database now receives the same expensive work many times. If a query takes 200 ms and hundreds of requests start it together, the impact is much larger than one slow query.

Database work increases
          ↓
Connection pool fills
          ↓
Requests wait for connections
          ↓
API latency rises
          ↓
Timeouts and retries create more traffic

That feedback loop is what makes a stampede dangerous.

The code that creates the problem

This cache-aside pattern looks correct in isolation:

import json

async def get_trending_products() -> list[dict]:
    cached = await redis.get("trending_products")

    if cached is not None:
        return json.loads(cached)

    products = await load_trending_products_from_database()

    await redis.set(
        "trending_products",
        json.dumps(products),
        ex=300,
    )

    return products

The race appears when many requests execute it at nearly the same time:

Request A → Redis GET → miss
Request B → Redis GET → miss
Request C → Redis GET → miss
              ↓
      All three call the database

The cache protects the database only while the value is present. When it expires, every request assumes that it is responsible for rebuilding it.

It is also useful to distinguish related problems:

Problem What happens Typical protection
Cache stampede Many requests rebuild one expired value Single-flight, distributed lock, stale-while-revalidate
Thundering herd Many requests wake up or arrive together and overload a shared resource Coalescing, queueing, backpressure
Cache penetration Repeated requests ask for data that does not exist Short negative-cache entries, validation, rate limits

The fix: one owner for one rebuild

For a hot cache key, only one request should do the expensive regeneration work. The other requests should reuse the result, wait briefly, receive an acceptable stale response, or fail in a controlled way.

This pattern is commonly called single-flight, request coalescing, or cache-stampede protection.

1,000 requests
      │
      ▼
Try to become refresh owner
      │
 ┌────┴────┐
 ▼         ▼
Owner    Waiters
 │         │
 ▼         │
Database   │
 │         │
 ▼         │
Populate cache
 └────┬────┘
      ▼
All requests receive the cached value

Why one local lock is not enough

Inside one application process, an in-memory lock or an in-flight promise can prevent duplicate work. It is not enough for a deployed service with multiple instances:

Load balancer
      │
  ┌───┼───┐
  ▼   ▼   ▼
API-1 API-2 API-3
  │     │    │
lock A lock B lock C

Each instance has its own memory and its own lock. Three instances can still run three identical database queries. For multiple instances, use a shared coordination mechanism such as Redis.

A Redis distributed-lock implementation

The following example uses a Redis lock so that only one instance refreshes a missing value. It uses a unique token and only releases a lock when the stored token still belongs to the caller.

import asyncio
import json
import random
import uuid
from typing import Any

from redis.asyncio import Redis

redis = Redis(host="localhost", port=6379, decode_responses=True)

CACHE_TTL_SECONDS = 300
# This must exceed normal refresh time; renew the lease if refreshes can run longer.
LOCK_TTL_SECONDS = 30
MAX_WAIT_SECONDS = 3

RELEASE_LOCK = """
if redis.call('get', KEYS[1]) == ARGV[1] then
    return redis.call('del', KEYS[1])
end
return 0
"""


class CacheRefreshInProgress(Exception):
    """The cache is being rebuilt by another request."""


async def load_trending_products_from_database() -> list[dict[str, Any]]:
    # Replace with a real, bounded database query.
    await asyncio.sleep(0.2)
    return [{"id": 1, "name": "Product A"}]


async def release_lock(lock_key: str, token: str) -> None:
    await redis.eval(RELEASE_LOCK, 1, lock_key, token)


async def wait_for_cache(cache_key: str) -> list[dict[str, Any]]:
    """Wait briefly for the refresh owner instead of querying the database."""
    deadline = asyncio.get_running_loop().time() + MAX_WAIT_SECONDS

    while asyncio.get_running_loop().time() < deadline:
        cached = await redis.get(cache_key)
        if cached is not None:
            return json.loads(cached)

        # Jitter stops all waiters from polling Redis at the same instant.
        await asyncio.sleep(0.04 + random.uniform(0, 0.04))

    raise CacheRefreshInProgress(cache_key)


async def get_trending_products() -> list[dict[str, Any]]:
    cache_key = "trending_products"
    lock_key = f"lock:{cache_key}"

    cached = await redis.get(cache_key)
    if cached is not None:
        return json.loads(cached)

    token = str(uuid.uuid4())
    acquired = await redis.set(lock_key, token, nx=True, ex=LOCK_TTL_SECONDS)

    if not acquired:
        return await wait_for_cache(cache_key)

    try:
        # Another request may have filled the cache just before lock acquisition.
        cached = await redis.get(cache_key)
        if cached is not None:
            return json.loads(cached)

        products = await load_trending_products_from_database()
        await redis.set(cache_key, json.dumps(products), ex=CACHE_TTL_SECONDS)
        return products
    finally:
        await release_lock(lock_key, token)

At the API boundary, map a wait timeout to a controlled retry response. Do not query the database again just because the wait ended:

from fastapi import HTTPException

async def trending_products_route():
    try:
        return await get_trending_products()
    except CacheRefreshInProgress as error:
        raise HTTPException(
            status_code=503,
            detail="Cache refresh in progress; retry shortly.",
            headers={"Retry-After": "1"},
        ) from error

A stale response is often a better choice when the data allows it.

Why this implementation works

The lock is acquired with Redis SET using two important options:

NX → create the lock only if it does not already exist
EX → expire the lock so a crashed process cannot hold it forever

The request that wins the lock becomes the refresh owner. Other requests poll only the cache for a short, jittered period. They do not fall back directly to the database, because that would recreate the stampede.

The second cache read after acquiring the lock is also deliberate. Another process may have filled the value in the tiny window between the first miss and lock acquisition.

Why the lock needs a unique token

Never release a distributed lock with an unconditional DEL.

Consider this sequence:

Request A acquires lock with token A
                ↓
The database call takes longer than the 30-second lock TTL
                ↓
Redis expires lock A
                ↓
Request B acquires a new lock with token B
                ↓
Request A finishes and blindly deletes the lock

Request A could accidentally delete Request B’s lock. The Lua script in the example deletes the lock only when its stored token still matches the request’s token.

The lock TTL must be longer than the normal refresh duration, with a safety margin. If refreshes can legitimately exceed it, consider lease renewal or redesign the refresh work. Do not simply set an extremely long TTL and forget about failure recovery.

Decide what waiting requests should receive

There is no single correct answer after a request loses the lock. Choose the response based on what the data means to users.

Situation Good response
Slightly old data is acceptable Serve a stale value and refresh in the background
Data must be current Wait for a short bounded time, then return a controlled retryable error
Refresh can run ahead of demand Warm or refresh the key in a background worker
Data is cheap and local A short wait may be enough; do not add distributed locking unnecessarily

The important point is to make the fallback explicit. A silent “wait timeout → query the database anyway” can create a second stampede.

Other useful protections

Serve stale data while refreshing

For homepages, product rankings, public content, and reference data, a slightly old result is often better than a slow or failed request.

Fresh TTL: 5 minutes
Stale window: 2 minutes

0–5 minutes → serve fresh value
5–7 minutes → serve stale value and trigger one background refresh
7+ minutes → wait briefly for refresh or use a controlled fallback

This is stale-while-revalidate. It removes refresh latency from the user path, but it is not suitable for data that must be immediately consistent.

Add TTL jitter

If many keys are created at the same time with the same TTL, they can expire together. Spread their expiration times:

base_ttl_seconds = 300
jitter_seconds = random.randint(0, 60)

await redis.set(cache_key, value, ex=base_ttl_seconds + jitter_seconds)

Jitter reduces synchronized expiration across many keys. It does not replace single-flight protection for one very hot key.

Warm predictable keys

For data such as homepage recommendations, popular products, configuration, or exchange rates, refresh the key before it expires:

Scheduler → background worker → database → Redis

Cache warming is helpful when traffic and refresh schedules are predictable. It adds background work, so only warm keys that are valuable enough to justify it.

Cache negative results briefly

Cache penetration is different from a stampede, but it can hurt the same database. If an ID does not exist, a short negative cache prevents repeated reads:

Key: user:99999999
Value: NOT_FOUND
TTL: 30 seconds

Use a short TTL because the record might be created later. Pair negative caching with input validation and rate limits for public endpoints.

Protect the database even when the cache fails

Redis is useful, but it should not be the only thing keeping the database healthy. Keep these guardrails as well:

  • bounded connection pools;
  • database query timeouts;
  • request rate limits or backpressure;
  • circuit breakers or a controlled degraded response;
  • bounded retry policies;
  • a deliberate plan for Redis unavailability.

If Redis is down, do not automatically let unlimited requests fall through to the database. Decide whether the endpoint can serve stale data, reject excess traffic, or safely allow a limited fallback.

What to measure in production

Do not add cache-stampede protection and assume it works. Track what happens around hot keys and refreshes:

cache_hit_rate
cache_miss_rate
cache_refresh_duration_ms
cache_refresh_count
lock_acquisition_success
lock_contention_count
cache_wait_duration_ms
stale_response_count
database_fallback_count

A 99% cache-hit rate can still hide a problem. A sudden rise in lock contention or refresh duration can mean the key is too expensive to rebuild, too short-lived, too hot, or invalidated too often.

A useful structured event could look like this:

{
  "cache_key": "trending_products",
  "event": "cache_refresh",
  "lock_acquired": true,
  "refresh_duration_ms": 1840,
  "cache_ttl_seconds": 300
}

A practical decision guide

One application process
  → in-flight task / single-flight can be enough

Multiple application instances
  → Redis-backed distributed lock or another shared coordination mechanism

Hot endpoint where stale data is acceptable
  → stale-while-revalidate + one refresh owner

Predictable, important data
  → cache warming

Many keys created together
  → TTL jitter

Repeated requests for missing data
  → short negative-cache entries

Final takeaway

A cache stampede is not mainly a cache problem. It is a concurrency problem.

The naive model is:

Cache miss → query database

The production model is:

Cache miss
      ↓
Is another request already rebuilding this value?
      ↓
Yes → reuse the result, wait briefly, or serve stale data
No  → become the single refresh owner

For a hot key, turn N identical expensive operations into one refresh and N consumers. That is the design decision that keeps an expired cache entry from becoming a database incident.