UUID-Based Sharding: Routing, Locality, and Hotspots

    28 October 2024Updated 10 September 2026
    6 min read
    Technical explainer
    Architecture
    uuid
    database
    scalability
    architecture

    Start with the queries

    Sharding places subsets of data on different nodes. A useful shard key lets common requests reach a small number of those nodes. Generating a unique identifier and deciding where its record lives are separate decisions.

    Consider an application storing documents for customers. Routing by document UUID spreads a customer's documents across shards. Routing by tenant UUID keeps that customer's data together, but a large or busy customer can overload its shard. Neither choice is automatically better: count the queries and transactions each requires.

    Routing keyUseful propertyMain trade-off
    Hash of document UUIDDirect lookup by document IDTenant-wide lists may query many shards
    Hash of tenant UUIDTenant records stay togetherOne large tenant can dominate a shard
    UUID v7 time rangeTime-window routingNew writes can concentrate in the newest range

    What UUID versions change

    UUID v4 contains 122 random bits after its fixed version and variant bits. UUID v7 includes a Unix millisecond timestamp and additional fields for uniqueness. Version 7 is standardized in RFC 9562, published in May 2024; it is no longer a draft or beta format. See RFC 9562.

    With range partitioning, increasing v7 values tend to reach the newest range. With hash partitioning, the hash spreads those values across the hash space, losing time-based shard routing. Keeping the original v7 as each shard's indexed key may still improve insertion locality within that shard. Clock behavior and generation order mean it is not a global transaction sequence.

    Do not choose a shard by the first few characters of a v7 UUID and expect random distribution: those characters belong to the timestamp.

    A reproducible routing example

    This Python example illustrates a fixed number of shards. It parses a UUID before hashing, so uppercase text, lowercase text, and text without hyphens route identically:

    python
    import hashlib
    import uuid
    
    
    def shard_for(identifier: str, shard_count: int) -> int:
        if isinstance(shard_count, bool) or not isinstance(shard_count, int):
            raise TypeError("shard_count must be an integer")
        if shard_count < 1:
            raise ValueError("shard_count must be positive")
        canonical_bytes = uuid.UUID(identifier).bytes
        digest = hashlib.sha256(canonical_bytes).digest()
        return int.from_bytes(digest, "big") % shard_count
    
    
    value = "550e8400-e29b-41d4-a716-446655440000"
    expected = shard_for(value, 16)
    assert shard_for(value.upper(), 16) == expected
    assert shard_for(value.replace("-", ""), 16) == expected
    print(expected)

    Parsing rejects malformed UUID input. This is routing code, not an authorization check or a production cluster manager. Other languages must use the same 16-byte representation, SHA-256, big-endian integer interpretation, and shard count. Python documents that byte representation in its UUID API.

    Why changing the shard count needs a migration

    The expression hash % shard_count changes assignments for many records when the count changes. If writers immediately use the new count while existing rows remain on the old nodes, lookups can miss records.

    A production design needs a versioned routing map and a migration procedure. One option assigns records to many logical buckets and maps those buckets to physical nodes. Rebalancing then moves selected buckets. Consistent hashing is another approach to limiting reassignment; it still requires data movement, failure handling, and coordinated routing updates.

    Prefer the database's supported routing and balancing mechanisms when available. For example, MongoDB hashed sharding supports hashed keys, but queries without the shard key may need to reach multiple shards. Do not substitute the illustrative Python hash for a database's own algorithm.

    Balanced keys do not guarantee balanced traffic

    A hash can spread distinct identifiers while requests remain concentrated on a few popular records. Check per-shard request rates, latency, storage growth, and the largest tenants. A thousand equally sized records are a different workload from one heavily updated record and 999 idle records.

    For the document application, decide how a tenant's listing, export, and deletion operations find its data. If a tenant spans shards, these operations need routing across all relevant partitions. If each tenant stays on one shard, define how an oversized tenant can move or split before it becomes urgent.

    Keep security separate from routing

    A request that reaches the correct shard must still pass tenant and object authorization. UUIDs can appear in logs, links, and shared documents; possession of an ID is not permission to read its record. See the multi-tenant UUID guide.

    Check the design with representative workloads

    Exercise direct lookups, tenant lists, recent-event queries, and cross-record transactions. Test a deliberately hot tenant and a node migration, including retries during movement. Measure these cases before adding shards; additional nodes also introduce coordination and operational work.

    Use the UUID v7 generator to inspect timestamp-ordered values, and the binary UUID guide to keep storage conversion consistent across clients.

    Generate Your Own UUIDs

    Ready to put this knowledge into practice? Try our UUID generators:

    Summary

    Choose a UUID shard key based on your queries. Compare hashing, tenant routing, and time ranges, with a Python example that normalizes UUIDs before routing.

    TLDR;

    UUIDs identify records; a routing strategy determines their shard. Random identifiers do not guarantee balanced traffic.

    Hash canonical UUID bytes for stable routing. UUID v7 can improve local index locality, but range routing can concentrate new writes. Plan for tenant queries, hot keys, and resharding before choosing a scheme.