Lesson 5 � Intermediate

Database Scaling: Sharding, Replication

Database scaling sabse mushkil challenge hai system design mein. Sharding aur replication dono alag problems solve karti hain. Chalo samajhte hain.

? 25 min✓ Intermediate✓ Caching basics

Database scaling kyun zaroori hai?

WHAT

Ek single database server limited hai � disk space, CPU, connections. Jab data bahut badh jaaye ya traffic bahut zyada ho, toh single database handle nahi kar paata. Database scaling ka matlab hai data ko distribute karna ya load share karna.

WHEN

10GB se zyada data ho, 10K+ queries per second ho, ya 99.9% uptime chahiye. Single database mein bottleneck hota hai � slow queries, timeouts, crashes.

WHERE

MySQL, PostgreSQL � traditional databases jo vertical scaling limit pe aa jaati hain. Sharding/replication se horizontally scale karte hain.

Database Replication

Replication ka matlab hai data ki copies rakhna multiple servers pe. Ek primary (master) server hai jo writes handle karta hai, aur multiple replica (slave) servers hain jo reads handle karte hain.

# Primary-Replica Architecture

Primary DB (Master) ✓ Handles ALL writes
 ? (async/sync replication)
Replica 1 (Slave) ✓ Handles reads
Replica 2 (Slave) ✓ Handles reads
Replica 3 (Slave) ✓ Handles reads

# Flow:
Write: Client ✓ Primary DB ✓ Replicated to replicas
Read: Client ✓ Load Balancer ✓ Any Replica

# Pros:
✓ Read scalability � 10 replicas = 10x read capacity
✓ High availability � primary down, promote replica
✓ Data redundancy � 3 copies = crash-proof

# Cons:
✓ Write bottleneck � still single primary
✓ Replication lag � replicas might have stale data
✓ Complexity � failover logic chahiye

# Use case: Read-heavy apps (95% read, 5% write)
# Real: MySQL Primary-Replica, PostgreSQL Streaming Replication
Mental Model: Replication jaise photocopy. Original book (Primary) mein kuch likha toh photocopies (Replicas) mein bhi aa jaayega. Reading ke liye photocopies use karo � original safe rahega. But original mein kuch naya likhna ho toh sirf original pe likho.

Database Sharding

Sharding ka matlab hai data ko horizontally divide karna multiple databases mein. Har database ka apna subset of data hota hai.

# Sharding Example: User Data

Without Sharding:
 1 Database ? 100 million users ✓ SLOW!

With Sharding (by User ID):
 Shard 1: Users 1-25M ✓ DB Server 1
 Shard 2: Users 25M-50M ✓ DB Server 2
 Shard 3: Users 50M-75M ✓ DB Server 3
 Shard 4: Users 75M-100M ✓ DB Server 4

# Sharding Key: User ID (decides which shard stores which data)

# Pros:
✓ Write scalability � writes distribute ho jaate hain
✓ Storage scalability � data distribute ho jaata hai
✓ No single point of failure � 1 shard down, baaki kaam karte rahein

# Cons:
✓ Complex queries � cross-shard joins mushkil
✓ Hotspot risk � agar ek shard pe zyada data aaye
✓ Resharding � naya shard add karna mushkil
✓ Distributed transactions � ACID maintain karna hard

# Use case: Massive data (billions of rows)
# Real: Instagram (user data sharded), LinkedIn

Sharding Strategies

# 1. Range-Based Sharding
✓ Data ranges ke basis pe divide
✓ User 1-1M ✓ Shard A, 1M-2M ✓ Shard B
✓ Simple to implement
✓ Hotspot risk (new users sab ek shard pe)

# 2. Hash-Based Sharding
✓ Hash function lagao shard key pe
✓ shard = hash(user_id) % num_shards
✓ Even distribution
✓ Resharding mushkil (hash change = data reshuffle)

# 3. Directory-Based Sharding
✓ Ek lookup table jo batata hai kaunsa data kis shard pe
✓ Flexible � data move karna easy
✓ Lookup table itself bottleneck ban sakta hai

# 4. Geo-Based Sharding
✓ Geography ke basis pe divide
✓ India users ✓ India DB, US users ✓ US DB
✓ Low latency � nearby server
✓ Cross-region queries mushkil

# Best Practice: Start with hash-based, move to geo-based at scale

Consistency Models

# 1. Strong Consistency
✓ Har write ke baad sab replicas updated
✓ Lagta hai: Zyada time (synchronous replication)
✓ Use: Banking, inventory systems

# 2. Eventual Consistency
✓ Eventually sab replicas updated honge
✓ Lagta hai: Kam time (asynchronous replication)
✓ Use: Social media, analytics

# 3. Read-Your-Writes Consistency
✓ Tum apna latest write hamesha dekh sakte ho
✓ Use: User profile, settings

# CAP Theorem (next lesson mein detail):
- Consistency + Availability ✓ No Partition Tolerance
- Consistency + Partition Tolerance ✓ No Availability
- Availability + Partition Tolerance ✓ No Consistency

# Real world: Most systems use eventual consistency
# Users ko 1-2 second delay acceptable hai

Exercise

Question: Sharding mein data ko kaunse key ke basis pe divide karte hain? (2 words)

Question: Primary-Replica architecture mein writes kaun handle karta hai? (1-2 words)

Common mistakes

Lesson complete?

Database Scaling samajh aa gayi✓ Ab SQL vs NoSQL seekhte hain � kaunsa database kab use karna hai.