← writing

From 3 GB to 764 MB: How We Fit 8 Million Credit Scores in Redis

Achieving a 75% reduction in memory footprint, an engineering story of optimizing our Redis caching layer through serialization benchmarks and jemalloc deep-dives

M

MELVIN LIU

October 4, 2026 · 13 min read

redisbackendperformanceoptimizations
All parameter names, variable names, field labels, and enum values used in this post are illustrative examples created for conceptual clarity. They do not reflect the actual naming conventions, schema design, or internal logic used in Kredivo's production codebase. All benchmark figures, memory measurements, and CPU throughput numbers were derived from real production-aligned datasets.

At Kredivo, every time a user opens the mobile app, the first thing that renders on the home screen is their credit score. Right at launch, before the user has done anything else. Not during checkout, not mid-application. With over 11 million active users in Indonesia, that single dashboard read generates enormous traffic on its own.

Kredivo dashboard showing credit score
Dashboard Screen Image from https://kredivo.id/

Before we built a caching layer, the scoring API hit the primary relational database directly on every request. Credit scores aren't real-time values; they're computed periodically through scheduled risk engine cycles and asynchronous underwriting pipelines, remaining stable over defined intervals. We were reading the same static rows millions of times per day, burning database connection pool slots and IOPS on queries that returned identical data to what had been returned a minute earlier.

The fix was obvious: add a Redis cache. What was not obvious was how much engineering work it would take to make that cache actually fit within our infrastructure constraints.

Note

The 8,000,000 record figure used throughout this post is a benchmark dataset we use to architect and stress-test high-throughput caching systems. It does not represent Kredivo's actual Daily Active Users or any internal business metrics, which remain strictly confidential.

The Constraint That Made This Hard

Our dedicated production Redis instance had a hard memory ceiling of 3.0 GB. No renegotiating with infrastructure, no easy vertical scaling path. Just 3.0 GB.

The first approach anyone would reach for is storing one key per user:

SET user_score:12345678 '{"metric_score": 742, "risk_grade": "GRADE_A_PLUS", ...}'

Our credit score JSON payload runs about 243 bytes. Eight million of those comes out to roughly 1.81 GB of raw data, which sounds comfortable. The problem is that Redis does not store raw bytes. It wraps every entry in C data structures and routes all allocations through jemalloc. Once you account for the full per-key overhead, 8 million standalone keys consumes 3.05 GB. That's already over the limit before any production traffic arrives.

This forced us to think carefully about two separate problems: how we serialize each payload, and how Redis stores that payload in memory.

8 million records compressed from 3GB to 764MB
The goal: fit 8 million scoring records into a 3.0 GB Redis instance with room to spare.

Phase 1: What You Store

Our credit score profile carries a mix of numeric metrics, risk tier enums, Unix timestamps, user category identifiers, and arrays of improvement factor codes. Here's what a typical record looks like in raw JSON (field names are generic mock representations to protect proprietary schema details, but the structure and byte counts reflect real production data):

{
  "metric_score": 742,
  "risk_grade": "GRADE_A_PLUS",
  "prior_grade": "GRADE_A",
  "assessment_epoch": 1775210000,
  "refresh_epoch": 1759500000,
  "valid_until_epoch": 1791000000,
  "profile_category": "TIER_GOLD",
  "action_factors": ["frequency_pattern_attr", "consistency_attr"],
  "primary_indicator": "activity_burst_attr"
}

243 bytes. The question was whether we could do better, and by how much.

We benchmarked five strategies on our production runtime across 8 million records: raw JSON, Zlib at Level 6, Zlib at Level 9, MessagePack with string keys, and MessagePack with integer keys and full enum mapping.

Why Integer Keys Change Everything

Before getting to the numbers, it helps to understand why integer-mapped keys produce such a dramatic reduction. The answer is in the MessagePack format specification.

MessagePack defines a positive fixint format: any integer from 0 to 127 encodes in a single byte. The high-order bit serves as the type marker; the remaining 7 bits carry the value. No length descriptor, no separate type header.

Contrast that with storing "metric_score" as a string key: 1 byte for the string header plus 12 bytes for the characters, totaling 13 bytes for a field name that could have been 1 byte. Across 9 schema fields, string keys alone consume 54 to 75 bytes of redundant metadata per record.

Map those 9 fields to integer keys 1 through 9, and map all enum values to small integers ("TIER_GOLD" becomes 11, "GRADE_A_PLUS" becomes 2), and the same record packs down to:

{1: 742, 2: 2, 3: 3, 4: 1775210000, 5: 1759500000, 6: 1791000000, 7: 11, 8: [20, 22], 9: 21}

That's 35 bytes.

FieldKeyEncoding
metric_score116-bit unsigned integer
risk_grade2Enum integer (A++=1, A+=2, A=3, B=4 ...)
prior_grade3Same grade scale
assessment_epoch432-bit Unix timestamp
refresh_epoch532-bit Unix timestamp
valid_until_epoch632-bit Unix timestamp
profile_category7Enum integer (Silver=10, Gold=11, Platinum=12 ...)
action_factors8List of enum integers
primary_indicator9Factor enum integer

Benchmark Results

StrategyAvg PayloadTotal (8M)vs JSONSerializeDeserializeRead Throughput
Raw JSON242.9 B1,853 MBbaseline25.15 µs8.28 µs~117,500 ops/s
Zlib L6174.9 B1,334 MB-28%37.85 µs16.20 µs~60,800 ops/s
Zlib L9174.9 B1,334 MB-28%41.94 µs14.68 µs~68,000 ops/s
MsgPack (strings)192.9 B1,472 MB-20.6%2.42 µs31.80 µs~31,400 ops/s
MsgPack (integers)35.0 B267 MB-85.6%2.49 µs9.88 µs~101,200 ops/s
Serialization strategy comparison chart
Payload size and read throughput across five serialization strategies.

The Decompression Trap

One number in that table deserves extra attention. MessagePack with string keys deserializes at 31.80 µs per record. Integer-mapped MessagePack does it in 9.88 µs. That's a 3.2x difference in read-path CPU cost. In a system where over 99% of operations are reads, serialization speed barely matters; deserialization speed is everything.

When the decoder encounters string keys, it has to allocate and instantiate a string object for every field name in every single payload. Under concurrent dashboard traffic, that object allocation thrashing saturates CPU cores and throttles read throughput to around 31,400 ops/s. That looks fine in isolation but becomes a bottleneck when millions of users are hitting the dashboard simultaneously.

Integer keys have no such cost. The decoder reads a numeric tag, maps it to an integer value, and moves on. Zero dynamic string allocation.

Zlib fails for a different reason. Even at maximum compression, DEFLATE only reduces our 243-byte JSON payloads by 28%. The algorithm's sliding window LZ77 matching gains almost nothing on payloads this small. To make it worse, every cache read requires two full parse passes: first inflate the bitstream, then parse the resulting JSON string. Pushing from Level 6 to Level 9 added a 10.8% CPU penalty with zero additional byte savings.

Integer-mapped MessagePack wins on every dimension that matters.

Phase 2: How Redis Stores It

After phase 1, the mental math looks fine: 8,000,000 × 35 bytes ≈ 280 MB, well within 3.0 GB.

That math is incomplete.

Redis wraps every key in C data structures and allocates all memory through jemalloc, which quantizes allocations into discrete size classes. The raw payload is only a fraction of what actually lands in RAM.

Per-key Redis memory overhead breakdown
A 35-byte payload expands to 176–192 bytes per key after Redis internal structs and jemalloc binning.

What a Standalone Key Actually Costs

For each SET user_score:<user_id> <payload> operation:

┌─────────────────────────────────────────────────────────┐
│ Global db→dict hash table bucket share       ~16 Bytes  │
│ dictEntry (key ptr, val ptr, next ptr)         32 Bytes  │
│ Key redisObject (type, encoding, lru, refcnt)  16 Bytes  │
│ Key SDS string ("user_score:12345678")         32 Bytes  │
│ Value redisObject                              16 Bytes  │
│ Value SDS (35 B payload → 48 B jemalloc bin)  48 Bytes  │
│ db→expires dictEntry + TTL timestamp           32 Bytes  │
└─────────────────────────────────────────────────────────┘
  TOTAL: ~176–192 bytes  (payload is 35 B; overhead is 81%)

Three forces drive that overhead:

Double dictionary entries. Every key that carries a TTL requires an entry in the primary key dictionary and a second dictEntry in the expires dictionary for its timer. That's 64 bytes of pure bookkeeping per record, and you cannot avoid it with standalone keys.

redisObject wrappers. Both the key and the value are wrapped in 16-byte robj structs carrying type, encoding, LRU clock, and reference count. These are non-negotiable parts of Redis's internal object model.

jemalloc binning. jemalloc allocates in size classes: 8, 16, 24, 32, 40, 48, 56, 64 bytes, and so on. Our 35-byte payload plus a 3-byte sdshdr8 header and a null terminator totals 39 bytes. That rounds up to the 48-byte bin. The key string "user_score:12345678" is 20 bytes with its header, which rounds up to 32. You pay for bin headroom you cannot use.

The result: 8 million standalone keys, even with our 35-byte MessagePack payloads, consumes 1,403 MB.

Bucketed Hashes

Instead of 8 million top-level keys, we partition users into 10,000 buckets:

bucket_id = user_id % 10000
HSET score_bucket:{bucket_id} {user_id} {payload}

Each bucket holds roughly 800 user fields. The per-field cost under this layout:

┌─────────────────────────────────────────────────────────┐
│ Inner hash dict bucket array share           ~12 Bytes  │
│ Inner dictEntry (key ptr, val ptr, next ptr)  32 Bytes  │
│ Field SDS "12345678" → 16-byte jemalloc bin  16 Bytes  │
│ Value SDS (35 B → 48 B jemalloc bin)          48 Bytes  │
│ Per-record robj wrappers                       0 Bytes  │
│ Per-record expires dictEntry                   0 Bytes  │
└─────────────────────────────────────────────────────────┘
  TOTAL: ~96–108 bytes  (~80 bytes saved per record)

Three meaningful changes happen here. Fields inside a Redis Hash are direct SDS strings rather than redisObject structs, eliminating 32 bytes of robj wrappers per record. The TTL is set on the 10,000 bucket keys rather than 8 million individual entries, wiping out the expires dictionary footprint entirely. That's roughly 256 MB of timer bookkeeping gone. The field name also shrinks from "user_score:12345678" (which bins to 32 bytes) down to plain "12345678" (which bins to 16 bytes).

Across 8 million records, those 80 bytes per entry add up to 640 MB saved.

The TTL Trade-off

This architecture comes with one constraint worth naming clearly. In Redis versions before 7.4, TTL is managed at the key level, not the hash field level. The HEXPIRE command that enables per-field expiration was introduced in Redis 7.4; in older environments, you cannot automatically expire a single user's entry inside a bucket.

For our domain, this trade-off is straightforward. Credit scores update periodically on predictable cycles. We set a 24-hour TTL on each bucket key and issue explicit HDEL calls whenever the scoring engine recalculates a specific user's record. The application handles individual invalidation; Redis handles bulk cleanup. The memory saved by eliminating 8 million per-record TTL entries was worth accepting this constraint by a wide margin.

Phase 3: The Full Picture

Combining both dimensions (serialization format and storage method) against our 3.0 GB ceiling:

FormatStorageTotal RAM% of 3 GBStatus
Raw JSONStandalone SET3,051.8 MB101.7%OOM crash on deploy
MsgPack stringsStandalone SET2,807.6 MB91.4%Eviction storm risk
Raw JSONBucketed HSET2,412.8 MB78.5%No headroom for BGSAVE
Zlib L6Standalone SET2,563.5 MB83.4%High CPU latency
MsgPack integersStandalone SET1,403.8 MB45.7%Functional
MsgPack integersBucketed HSET764.8 MB24.9%Optimal
Redis memory utilization across all format and storage combinations
Only the integer-mapped MessagePack combined with bucketed hashes leaves meaningful headroom.

The gap between 45.7% and 24.9% is not cosmetic. When Redis triggers a background snapshot (BGSAVE) or rewrites the AOF file, the Linux kernel's copy-on-write semantics can temporarily spike the effective memory footprint by 50–80% during heavy write periods. A Redis instance running at 45.7% utilization has almost no room for that spike before evictions start firing. Running at 24.9% gives us stable headroom for both routine background saves and future scale. At that utilization, 764 MB for 8 million users means our current cluster can absorb 25 million users without requiring vertical infrastructure changes.

Phase 4: The Final Architecture

The encoding pipeline for every write operation:

Score object (strings, enums, lists)
  → Map field keys to integers 1–9
  → Map enum values to fixint integers
  → MessagePack binary pack → 35-byte blob
  → bucket_id = user_id % 10,000
  → HSET score_bucket:{bucket_id} {user_id} {blob}

And the read path:

Dashboard request for user_id
  → bucket_id = user_id % 10,000
  → HGET score_bucket:{bucket_id} user_id → 35-byte blob
  → MessagePack unpack
  → Reverse integer-to-field mapping
  → Hydrated score object returned (~0.8 ms)

The schema mapping is a static bidirectional lookup table initialized at service startup. Encoding and decoding each record is pure integer arithmetic: no string hashing, no dynamic allocation. The entire transformation adds negligible CPU overhead per operation.

Results

  • 764.8 MB for 8 million scoring records (24.9% of our 3.0 GB limit)
  • >450,000 read ops/second sustained throughput
  • ~0.8 ms p99 latency for cache hits, down from 15–45 ms for database reads
  • 99.4% of dashboard read traffic diverted away from the primary database
  • 75% reduction in memory footprint compared to the naive JSON approach

What to Take From This

Serialization is a data architecture decision. JSON is readable and convenient, but for high-volume internal caching it carries enormous overhead. Converting string keys and enumerated values to small integers before passing them to MessagePack cut our payload size by 85.6% with no additional algorithmic complexity. The whole mapping is just a lookup table.

Raw byte count is not Redis memory. The actual RAM consumed by a key is 4–5x the payload size once you account for dictEntry structs, redisObject wrappers, SDS string headers, and jemalloc bin quantization. Calculate what Redis will actually allocate, not just the data you're storing.

Bucketed hashes consolidate overhead at scale. Grouping millions of small records into a few thousand hash keys eliminates per-key robj wrappers and the entire expires dictionary footprint for those records. For any dataset where records cluster naturally (user profiles, session data, feature flags), hash bucketing is one of the most effective memory optimizations available without changing the underlying data.

Operating below 30% memory utilization is load-bearing infrastructure. Redis instances above 75% utilization become fragile: BGSAVE forks spike memory, replication backlogs consume unallocated RAM, and eviction storms can cascade. Headroom is not wasted capacity; it's the margin that keeps the system stable when background operations compete for resources.

Discussion

Sign in to leave a comment.

or continue with

Loading comments...