Benchmarks¶
Numbers from the benchmarks/ harness. They shift from run to
run, so trust the ordering more than the exact values.
Setup¶
- AMD Ryzen 9 5950X (16C/32T) · 32 GiB RAM · Ubuntu 24.04 · Linux 7.0
- CPython 3.14.2 (GIL build), Django 6.1.1
- Redis 8 / Valkey 9 in local Docker, paired natively per adapter
(
redis-py→ redis,valkey-py/valkey-glide→ valkey,django (builtin)→ redis) - Every round trip 250 µs longer than on the local Docker network, as within a cloud availability zone
Each column from get to delete is a phase of 1,000 calls, or 100 calls of
10 keys for mget and mset. Throughput is in calls per second. A phase has
10 timed runs after a warmup pass. The cached value pickles to ~150 B, except
in the compressor benchmarks.
Sync direct¶
Cache calls such as cache.get(...), with no Django request and no asyncio.
The round trip takes most of each call, so the django-cachex adapters land
within 7% of each other. Django's built-in RedisCache builds a redis.Redis
client for every call, about 40 µs, and its incr makes two round trips
(EXISTS, then INCRBY).
| Adapter | get | get-miss | set | mget | mset | incr | delete | py-mem KiB |
|---|---|---|---|---|---|---|---|---|
| redis-py | 2,789 | 2,896 | 2,782 | 2,362 | 2,213 | 2,887 | 2,879 | 64 |
| redis-py+hiredis | 2,783 | 2,887 | 2,777 | 2,391 | 2,215 | 2,901 | 2,905 | 36 |
| valkey-py | 2,869 | 2,946 | 2,883 | 2,396 | 2,301 | 2,953 | 2,953 | 91 |
| valkey-py+libvalkey | 2,920 | 2,990 | 2,922 | 2,535 | 2,383 | 2,989 | 2,979 | 31 |
| valkey-glide | 2,733 | 2,837 | 2,743 | 2,337 | 2,297 | 2,840 | 2,834 | 29 |
| django (builtin) | 2,511 | 2,535 | 2,458 | 2,227 | 2,032 | 1,366 | 2,526 | 50 |
Serializers¶
The valkey-py+libvalkey adapter with each serializer.
| Serializer | get | get-miss | set | mget | mset | incr | delete |
|---|---|---|---|---|---|---|---|
| pickle | 2,906 | 2,976 | 2,924 | 2,499 | 2,383 | 2,984 | 2,980 |
| json | 2,883 | 2,967 | 2,861 | 2,395 | 2,215 | 2,963 | 2,962 |
| msgpack | 2,917 | 2,982 | 2,914 | 2,551 | 2,392 | 2,981 | 2,990 |
| orjson | 2,923 | 2,984 | 2,936 | 2,573 | 2,456 | 2,987 | 2,980 |
| ormsgpack | 2,921 | 2,956 | 2,897 | 2,549 | 2,368 | 2,966 | 2,957 |
Compressors (macro)¶
Cache calls through valkey-py+libvalkey with pickle, on a 14 KiB
queryset-shaped payload.
| Compressor | get | get-miss | set | mget | mset | incr | delete | srv-mem KiB |
|---|---|---|---|---|---|---|---|---|
| none | 2,372 | 2,976 | 2,711 | 918 | 1,480 | 2,988 | 2,982 | 1,269 |
| zlib | 2,495 | 2,983 | 2,195 | 1,105 | 680 | 2,971 | 2,981 | 169 |
| gzip | 2,471 | 2,988 | 2,099 | 1,027 | 608 | 2,980 | 2,978 | 169 |
| lzma | 2,384 | 2,985 | 682 | 920 | 88 | 2,988 | 2,982 | 167 |
| lz4 | 2,551 | 2,987 | 2,677 | 1,181 | 1,420 | 2,986 | 2,978 | 238 |
| zstd | 2,532 | 2,984 | 2,563 | 1,153 | 1,190 | 2,987 | 2,984 | 169 |
Compressors (micro)¶
compress and decompress alone on the same 14 KiB payload, with no
adapter and no network.
| Compressor | output ratio | compress (MB/s) | decompress (MB/s) |
|---|---|---|---|
| zlib | 11.9% | 155.4 | 1,074.4 |
| gzip | 11.8% | 128.4 | 766.2 |
| lzma | 11.4% | 12.2 | 419.3 |
| lz4 | 16.6% | 2,511.6 | 6,163.5 |
| zstd | 11.4% | 676.4 | 1,695.5 |
Django request cycle¶
The sync direct workload, with every cache call inside Client().get(url)
(URL resolve → CommonMiddleware → view → request_finished). The gap to
sync direct is Django's per-request overhead, about 115 µs here.
| Adapter | get | get-miss | set | mget | mset | incr | delete |
|---|---|---|---|---|---|---|---|
| redis-py | 2,105 | 2,168 | 2,090 | 1,854 | 1,738 | 2,010 | 2,081 |
| redis-py+hiredis | 2,150 | 2,215 | 2,142 | 1,900 | 1,806 | 2,058 | 2,097 |
| valkey-py | 2,184 | 2,235 | 2,191 | 1,902 | 1,803 | 2,072 | 2,246 |
| valkey-py+libvalkey | 2,201 | 2,257 | 2,205 | 1,963 | 1,862 | 2,084 | 2,236 |
| valkey-glide | 2,079 | 2,126 | 2,054 | 1,823 | 1,791 | 1,974 | 2,106 |
| django (builtin) | 1,955 | 1,987 | 1,934 | 1,778 | 1,647 | 1,126 | 1,875 |
Async serial¶
One awaited call at a time, such as await cache.aget(...). The gap to sync
direct is asyncio loop overhead. Django's built-in RedisCache has no native
async path, so it also pays for sync_to_async.
| Adapter | get | get-miss | set | mget | mset | incr | delete |
|---|---|---|---|---|---|---|---|
| redis-py | 2,709 | 2,805 | 2,666 | 2,316 | 1,924 | 2,793 | 2,787 |
| redis-py+hiredis | 2,684 | 2,774 | 2,662 | 2,335 | 1,889 | 2,776 | 2,790 |
| valkey-py | 2,678 | 2,768 | 2,675 | 2,318 | 1,997 | 2,756 | 2,746 |
| valkey-py+libvalkey | 2,670 | 2,762 | 2,671 | 2,318 | 2,006 | 2,766 | 2,759 |
| valkey-glide | 2,615 | 2,699 | 2,619 | 2,247 | 2,202 | 2,695 | 2,680 |
| django (builtin) | 2,005 | 2,031 | 1,999 | 1,826 | 1,689 | 1,205 | 2,021 |
Async concurrent (50 in flight)¶
asyncio.gather of 50 calls at a time. redis-py and valkey-py open a
connection per call in flight, and valkey-glide sends them all over one.
Django's built-in RedisCache runs every call through sync_to_async on one
thread, so it serves 50 in flight about as fast as one at a time. Connection
counts, which include the harness's own, stay flat between phases on every
adapter (Δ = 0).
| Adapter | get | get-miss | set | mget | mset | incr | delete | conns peak |
|---|---|---|---|---|---|---|---|---|
| redis-py | 16,411 | 18,362 | 15,447 | 8,937 | 5,466 | 18,488 | 18,481 | 51 |
| redis-py+hiredis | 16,885 | 18,484 | 15,476 | 8,678 | 5,383 | 18,569 | 18,543 | 51 |
| valkey-py | 20,514 | 22,735 | 20,088 | 10,133 | 6,427 | 22,510 | 22,838 | 51 |
| valkey-py+libvalkey | 20,541 | 22,882 | 20,318 | 10,108 | 6,430 | 22,563 | 23,058 | 51 |
| valkey-glide | 43,804 | 47,758 | 33,133 | 13,296 | 10,043 | 50,684 | 49,981 | 2 |
| django (builtin) | 2,244 | 2,281 | 2,220 | 2,000 | 1,863 | 1,292 | 2,188 | 2 |
ASGI full-stack¶
granian (4 workers) and httpx (100 concurrent clients, 20 s) against a
view that makes six async cache calls per request. The harness samples server
RSS and connected_clients every 5 s, and the connection counts include about
5 of its own. The req/s column is noisy, so read it in rough buckets: 500 to
600 for the django-cachex adapters, 350 for Django's built-in RedisCache.
Under ASGI, Django gives every request its own instance of a cache backend.
The built-in RedisCache builds a connection pool for each instance, and its
close() leaves the pool open, so its connections climb until the garbage
collector frees the pools.
| Adapter | req/s | avg ms | p99 ms | RSS peak (MiB) | conns peak | conns settled |
|---|---|---|---|---|---|---|
| redis-py | 487 | 204.4 | 1,522.8 | 364 | 105 | 105 |
| redis-py+hiredis | 597 | 167.1 | 1,270.4 | 368 | 105 | 105 |
| valkey-py | 613 | 162.6 | 1,317.0 | 369 | 104 | 104 |
| valkey-py+libvalkey | 563 | 177.0 | 1,383.5 | 368 | 104 | 104 |
| valkey-glide | 500 | 199.2 | 1,451.1 | 425 | 8 | 8 |
| django (builtin) | 355 | 279.6 | 1,769.4 | 459 | 1,371 | 180 |
ORM cache vs django-cachalot¶
The ORM cache derives from django-cachalot 2.9.1.
The contenders are the database alone (none), cachalot, the ORM cache
(cachex), and both over a
TrackingCache (+tracking), which
keeps local copies of what each process read.
- The machine above, with CPython 3.14.2 (free-threaded build, with the GIL that libvalkey turns back on), Django 6.1.1, psycopg 3.3.6
- PostgreSQL 18 with
fsync=offand Valkey 9 in local Docker, both libraries on thevalkey-py+libvalkeybackend - The same 250 µs added to every round trip, except in the clock skew runs
Correctness¶
Twelve cases where a cached result can differ from the database, each checked against an uncached read. They are the known differences between the two libraries, not a sample of typical queries. Cachalot got all 12 wrong, the ORM cache none:
- A read while an autocommit
UPDATEruns, and one by a process that anon_commit()hook hands the new row to. Cachalot invalidates before an autocommit statement runs, and inatomic()only after every hook has run. - A subquery or
Now()nested in afilter()expression, andorder_by(Random()). Cachalot looks for subqueries andNow()only as a filter's direct operands, and refusesorder_by("?")but notRandom(). - Two JSON filter values of the same length that agree in their first 35
characters. Cachalot's cache key holds each parameter's
str(), which psycopg 3 shortens to those characters and the length. DB_CASCADEdeletes, rawTRUNCATE, rawREFRESH MATERIALIZED VIEW, and raw writes to a mixed-case or unmanaged table. Cachalot doesn't invalidate every table these change.- Reads in a
REPEATABLE READtransaction from its old snapshot, which cachalot stores in the shared cache at the commit.
CACHALOT_FINAL_SQL_CHECK = True and listing the unmanaged table in
CACHALOT_ADDITIONAL_TABLES fix 2 of the 12.
Races¶
Four reader processes read a counter row while a writer updates it 500 times,
at random times 10 ms apart on average. A read is stale if it returns a
version older than one committed before the read began. Stale reads per
1,000, where the ORM cache column covers it with and without a
TrackingCache:
| Scenario | cachalot | cachalot+tracking | ORM cache |
|---|---|---|---|
| Autocommit | 132 | 504 | 0 |
| Autocommit, commits 2 ms slower | 755 | 602 | 0 |
atomic() |
8.0 | 512 | 0 |
atomic(), 1 ms of later on_commit() hooks |
117 | 567 | 0 |
- Under autocommit, cachalot invalidates before the statement runs, so a read before the commit stores the old row as newer than the write. The cache serves it until the next write: cachalot hid 146 in 1,000 writes that way, and with slower commits all of them. Writes in quick succession, as in a loop of updates, make it worse, since a raced read can store its old row after the next write's invalidation too.
- In
atomic(), cachalot invalidates only after allon_commit()hooks have run, so a process that a hook hands the new row to, say through a task queue, can get the old one. It also invalidates in two steps a round trip apart, and until the second it can serve a row read before the commit. cachalot+trackingserves hits from the process's memory. The thread that evicts a copy after a write needs the GIL, and readers that never pause make it wait up to 5 ms. At a read every 1 ms, it served 40 stale reads per 1,000.- The ORM cache sends a table's queries to the database while a write to it runs. That costs it 6 to 7 points of hit ratio against cachalot, and 20 to 22 with slower commits.
Clock skew¶
Cachalot stamps results and invalidations with the time.time() of the
process that makes them. If a reader's clock runs ahead by Δ, a result it
stores right after a write outlives every write in the next Δ. With writes at
random times and λ = Δ / mean time between writes, that hides λ / (1 + λ) of
the writes. Runs with Δ from 2 to 60 ms came within 34 per 1,000 of that.
Hidden writes per 1,000, by how far the reader's clock runs ahead and how
often the table is written:
| Skew | Write every 1 s | 1 min | 15 min | 1 h |
|---|---|---|---|---|
| 10 ms | 9.9 | 0.17 | 0.011 | 0.003 |
| 100 ms | 91 | 1.7 | 0.11 | 0.028 |
| 1 s | 500 | 16 | 1.1 | 0.28 |
| 10 s | 909 | 143 | 11 | 2.8 |
A hidden write stays hidden until the table's next write after the skew has passed. A writer's clock ahead costs only hits, keeping e^(−λ) of them, which the runs matched within 1.3 points: 1 s against a write every minute loses under 2%. The ORM cache keeps a generation per table on the cache server and reads no clock. With either clock 40 ms ahead, it hid no write.
Speed¶
Median latency in µs over 2,000 calls. For none, the hit and the miss are
the query alone.
| Contender | Hit | Miss | Autocommit write | atomic() write |
|---|---|---|---|---|
| none | 566 | 612 | 563 | 1,244 |
| cachalot | 500 | 1,453 | 990 | 2,072 |
| cachex | 541 | 1,489 | 1,354 | 2,057 |
| cachalot+tracking | 136 | 1,508 | 1,010 | 2,128 |
| cachex+tracking | 531 | 1,516 | 1,377 | 2,082 |
- A hit takes a round trip, as the query does, so it saves only the database's work. A hit on one row takes about 90% of the query's time, on 100 rows about half, and on 1,000 rows about a quarter.
- The ORM cache's hits take about 40 µs longer than cachalot's, for the Lua script that checks leases and generations. Its autocommit writes make two round trips to Valkey, a lease before the statement and its release after, where cachalot makes one.
cachalot+trackingserves hits from memory in about a quarter of the time, which is also what makes its reads stale. The ORM cache over aTrackingCachestill asks the server whether its copy is current.- In a mixed workload of 100-row reads by one of 20 queries, the caches serve
about 1.5 times the database's throughput at 1% writes, and
cachalot+trackingtwice. At 10% writes they serve about a fifth less, since every write makes the cached queries on its table miss, and a miss takes three round trips. - A
migratewith nothing to apply emptied cachalot's cache of 100 queries and left the ORM cache's intact.
Network delay¶
The same runs with no added delay and with 1 ms. Without delay, a round trip takes about 60 µs to Valkey and 80 µs to PostgreSQL. A hit saves the database's work but not the round trip, so with delay it saves less: at 1 ms, cachalot's one-row hit took 95% of the query's time, against 82% without delay. A race lasts a few round trips, so with delay more reads fall in one. Cachalot's results, and the ORM cache's hit ratio in the last row:
| Result | No delay | 250 µs | 1 ms |
|---|---|---|---|
| Throughput at 1% writes, × the database's | 1.9 | 1.5 | 1.2 |
| Throughput at 10% writes, × the database's | 0.94 | 0.79 | 0.62 |
| Stale reads per 1,000, autocommit | 104 | 132 | 247 |
Stale reads per 1,000, atomic() |
0.19 | 8.0 | 88 |
| Hit ratio at a read every 1 ms | 94% | 90% | 83% |
| Hit ratio at a read every 1 ms, ORM cache | 92% | 83% | 60% |
The ORM cache's throughput stayed within 4% of cachalot's, and it served no stale read at any delay. Its hit ratio falls faster, since its lease sends a table's reads to the database for as long as a write runs, and writes take longer with delay.
Reproducing¶
The harness starts its own Redis, Valkey and PostgreSQL containers, so a Docker daemon is the only host requirement:
# Everything, with 250 µs added to every round trip
BENCH_NET_DELAY_US=250 uv run pytest benchmarks/ -c benchmarks/pytest.ini
# A single slice
uv run pytest benchmarks/test_throughput.py::test_adapters_sync \
-c benchmarks/pytest.ini
# The ORM cache comparison alone
BENCH_NET_DELAY_US=250 uv run pytest benchmarks/test_orm.py \
-c benchmarks/pytest.ini
benchmarks/README.md lists the slices, the knobs (N_OPS, K_RUNS,
WARMUP_KEYS, MGET_BATCH) and the full methodology. BENCH_NET_DELAY_US
delays everything the server containers send, using tc netem from a helper
container, so it needs no host privileges.