Blob Storage Explained: Architecture, APIs, Cost & CDN Strategies
Vidhya Sagar ThakurSeptember 14, 202618 min read0 views
Introduction
Blob storage (Binary Large Object storage) is purpose-built for storing unstructured binary data — images, videos, PDFs, backups, logs, ML model weights, and anything that isn't rows in a relational table. Unlike file systems that organize data in hierarchical directories, blob storage uses a flat namespace of key-value pairs: you give it a key (like users/123/avatar.png) and a blob (the raw bytes), and it stores them durably.
Services like AWS S3, Azure Blob Storage, and Google Cloud Storage have made blob storage the default for virtually all unstructured data in modern systems.
GIF via GIPHY
Core Architecture
text
┌────────────────────────────────────────────────────────┐
│ Blob Storage Service │
│ │
│ ┌──────────┐ ┌───────────────┐ ┌──────────────┐ │
│ │ API │ │ Metadata │ │ Storage │ │
│ │ Gateway │───▶│ Service │───▶│ Backend │ │
│ │ │ │ │ │ │ │
│ │ Auth, │ │ Key → object │ │ Actual bytes │ │
│ │ routing │ │ mapping, ACLs │ │ across disks │ │
│ │ │ │ size, type │ │ │ │
│ └──────────┘ └───────────────┘ └──────────────┘ │
│ │
│ Metadata DB: lightweight (key, bucket, size, etag, │
│ created, storage_class, acl) │
│ │
│ Storage backend: raw bytes on disk, replicated across │
│ availability zones │
└────────────────────────────────────────────────────────┘
GIF via GIPHY
Data Model
text
┌──────────────────────────────────────────────┐
│ Bucket: "my-app-uploads" │
│ │
│ Key Size Class │
│ ─────────────────── ────── ────────── │
│ users/1/avatar.png 48 KB Standard │
│ users/1/resume.pdf 2.1 MB Standard │
│ videos/promo.mp4 850 MB Standard │
│ backups/2024-01.tar 12 GB Glacier │
│ logs/2024/01/01.gz 400 MB Infrequent │
│ │
│ Flat namespace — "folders" are just key │
│ prefixes. users/1/ is not a real directory. │
└──────────────────────────────────────────────┘
GIF via GIPHY
API Operations
Python
import boto3
from botocore.config import Config
s3 = boto3.client("s3", config=Config(
retries={"max_attempts": 3, "mode": "adaptive"}
))
# Upload a blob
def upload_blob(bucket, key, file_path, content_type):
s3.upload_file(
file_path, bucket, key,
ExtraArgs={
"ContentType": content_type,
"ServerSideEncryption": "AES256",
"CacheControl": "max-age=31536000", # immutable assets
}
)
# Download
def download_blob(bucket, key, dest_path):
s3.download_file(bucket, key, dest_path)
# Generate pre-signed URL (time-limited access without exposing credentials)
def get_presigned_url(bucket, key, expires_in=3600):
return s3.generate_presigned_url(
"get_object",
Params={"Bucket": bucket, "Key": key},
ExpiresIn=expires_in,
)
# Returns: https://bucket.s3.amazonaws.com/key?X-Amz-Signature=...&Expires=...
# Multipart upload for large files
def upload_large_file(bucket, key, file_path):
"""Files > 100MB should use multipart upload."""
config = boto3.s3.transfer.TransferConfig(
multipart_threshold=100 * 1024 * 1024, # 100 MB
multipart_chunksize=100 * 1024 * 1024, # 100 MB per part
max_concurrency=10,
)
s3.upload_file(file_path, bucket, key, Config=config)
# Uploads 10 chunks in parallel → faster upload for large objects
GIF via GIPHY
Storage Classes and Cost Optimization
text
Access Storage Retrieval Min Storage
Class Frequency Cost Cost Duration
──────────────────────────────────────────────────────────────────
Standard Frequent $$$ Free None
Infrequent Access Monthly $$ $ 30 days
Glacier Quarterly $ $$ 90 days
Deep Archive Yearly ¢ $$$ 180 days
Lifecycle policy example:
Day 0-30: Standard (active use)
Day 30-90: Infrequent (still accessible, cheaper)
Day 90-365: Glacier (archived, retrieval takes minutes)
Day 365+: Deep Archive (long-term, retrieval takes hours)
JSON
{
"Rules": [{
"ID": "lifecycle-policy",
"Status": "Enabled",
"Transitions": [
{ "Days": 30, "StorageClass": "STANDARD_IA" },
{ "Days": 90, "StorageClass": "GLACIER" },
{ "Days": 365, "StorageClass": "DEEP_ARCHIVE" }
],
"Expiration": { "Days": 2555 }
}]
}
GIF via GIPHY
Serving Blobs at Scale: CDN Integration
text
Direct from blob storage:
Client → S3 (single region) → Response
Latency: 50-200ms depending on region
With CDN:
Client → CloudFront edge (nearest PoP) → Cache hit → Response (5ms)
→ Cache miss → S3 → Response (60ms)
┌──────┐ ┌─────────────┐ ┌──────────────┐
│Client│────▶│ CDN Edge │────▶│ Blob Storage │
│ │ │ (cache hit) │ │ (origin) │
│ │◀────│ 5ms │ │ │
└──────┘ └─────────────┘ └──────────────┘
For user-uploaded content:
Upload: Client → API → Blob Storage (direct upload via pre-signed URL)
Serve: Client → CDN → Blob Storage (CDN caches on first request)
Purge: When blob is updated, invalidate CDN cache
GIF via GIPHY
Access Control Patterns
text
1. Pre-signed URLs (most common for user content)
Server generates a time-limited, signed URL.
Client uses URL directly — no credentials needed.
URL expires after N seconds.
2. Bucket policies (public static assets)
"All objects in /public/* are readable by anyone."
Used for static websites, public assets.
3. IAM roles (server-to-server)
Application server has an IAM role granting access.
No credentials in code — role attached to EC2/ECS/Lambda.
4. Signed cookies (streaming media)
Multiple files under a path, all authorized at once.
Used for video streaming: sign access to /videos/course-1/*
GIF via GIPHY
Durability and Replication
text
AWS S3 durability: 99.999999999% (11 nines)
→ You'd lose 1 object per 10 million years storing 10M objects.
How:
- Data replicated across ≥ 3 availability zones
- Each AZ has independent power, cooling, networking
- Checksums verified on read and write
- Bit rot detected and auto-repaired
Cross-region replication:
Primary bucket (us-east-1) ──replication──▶ Replica bucket (eu-west-1)
Use cases:
- Disaster recovery
- Lower latency for global users
- Compliance (data in specific regions)
GIF via GIPHY
Performance Patterns
text
Throughput optimization:
1. Randomize key prefixes
Bad: logs/2024/01/01/file001.gz (hotspot on "logs/2024" partition)
Good: a3f2/logs/2024/01/01/file001.gz (hash prefix distributes load)
(Modern S3 handles this automatically since 2018, but other
blob stores may still benefit from prefix randomization.)
2. Multipart upload for large files
Upload parts in parallel → saturate network bandwidth.
3. Range reads for partial content
GET /video.mp4 Range: bytes=0-1048575
Only download the first 1MB. Used for video streaming, PDF viewers.
4. Batch operations
Delete 10,000 objects → don't make 10,000 API calls.
Use batch delete API: single request, 1000 keys at a time.
GIF via GIPHY
Key Takeaways
- Blob storage is for unstructured binary data — images, videos, backups, logs; it uses a flat key-value namespace, not a hierarchical file system
- Use pre-signed URLs for secure direct upload/download — the server generates a time-limited signed URL; the client talks directly to blob storage without proxying through your server
- Lifecycle policies automate cost optimization — transition infrequently accessed blobs to cheaper storage classes automatically; archive old data to Glacier/Deep Archive
- Put a CDN in front of blob storage — cache frequently accessed blobs at edge locations to reduce latency from 200ms to 5ms and offload origin traffic
- Multipart upload for large files — upload chunks in parallel for faster uploads; resume interrupted uploads without restarting
- 11 nines of durability comes from multi-AZ replication — data is replicated across 3+ availability zones with checksums to detect and repair bit rot
- Design key schemas to avoid hot partitions — distribute load by using varied prefixes; avoid sequential keys that concentrate writes on a single partition
GIF via GIPHYWhat did you think?