PrepZone Logo
PrepZone

S3 Object Storage

Design blob storage with durability, versioning, lifecycle policies and multipart uploads.

Why this matters

  • S3-compatible object stores underpin StreamHub video uploads, user avatars, static assets, and ML training datasets.
  • Understanding multipart upload, versioning, and lifecycle policies is essential for designing any blob-heavy system.
  • Object storage optimises for throughput and durability, not low-latency random access — choose the right tier.

S3 core concepts

  • Bucket — top-level container with regional placement and access policies.
  • Object key — unique identifier within a bucket (videos/2026/04/clip_8f2a.mp4).
  • Metadata — user-defined key-value pairs and system metadata (content-type, etag).
  • Versioning — retain previous object versions on overwrite; protect against accidental deletion.
  • Lifecycle rules — auto-transition to cheaper tiers or expire objects after N days.

Request flow

S3 object storage architecture

PUTCLIENT
Client upload
COMPUTE
EKS APIpresigned URL
STORAGE
Amazon S3Std → IA → Glac…
NETWORK
CloudFrontdelivery
CLIENT
Downloaders
Presigned uploads → lifecycle tiers → CloudFront delivery → Glacier archive.
Java
PUT /streamhub-media/videos/2026/04/clip_8f2a.mp4
Host: s3.us-east-1.amazonaws.com
Content-Type: video/mp4
Content-Length: 524288000
x-amz-storage-class: STANDARD

<binary body>
Java
GET /streamhub-media/videos/2026/04/clip_8f2a.mp4
Host: s3.us-east-1.amazonaws.com

→ 200 OK
  ETag: "d41d8cd98f00b204e9800998ecf8427e"
  Content-Type: video/mp4

Multipart upload for large objects

Objects over 100 MB should use multipart upload: split into parts, upload in parallel, then assemble.

Java
POST /streamhub-media/videos/large_clip.mp4?uploads
→ <UploadId>upload_xyz</UploadId>

PUT /streamhub-media/videos/large_clip.mp4?partNumber=1&uploadId=upload_xyz
PUT /streamhub-media/videos/large_clip.mp4?partNumber=2&uploadId=upload_xyz
...

POST /streamhub-media/videos/large_clip.mp4?uploadId=upload_xyz
<CompleteMultipartUpload>
  <Part><PartNumber>1</PartNumber><ETag>"abc"</ETag></Part>
  <Part><PartNumber>2</PartNumber><ETag>"def"</ETag></Part>
</CompleteMultipartUpload>

Durability and consistency

AspectStorage classUse case
STANDARD11 nines durability, ms latencyHot data: active videos, thumbnails
STANDARD_IALower cost, retrieval feeInfrequently accessed archives
GLACIERMinutes-to-hours retrievalCompliance backups, old logs
INTELLIGENT_TIERINGAuto-moves by access patternUnpredictable access frequency
  • STANDARD

    Storage class11 nines durability, ms latency
    Use caseHot data: active videos, thumbnails
  • STANDARD_IA

    Storage classLower cost, retrieval fee
    Use caseInfrequently accessed archives
  • GLACIER

    Storage classMinutes-to-hours retrieval
    Use caseCompliance backups, old logs
  • INTELLIGENT_TIERING

    Storage classAuto-moves by access pattern
    Use caseUnpredictable access frequency

S3 provides read-after-write consistency for new objects; overwrites are eventually consistent in older APIs — use version IDs for critical data.

Lifecycle policy example

Java
{
  "Rules": [
    {
      "ID": "archive-old-uploads",
      "Filter": { "Prefix": "uploads/raw/" },
      "Status": "Enabled",
      "Transitions": [
        { "Days": 30, "StorageClass": "STANDARD_IA" },
        { "Days": 90, "StorageClass": "GLACIER" }
      ],
      "Expiration": { "Days": 365 }
    }
  ]
}

Presigned URLs for secure access

Never expose bucket credentials to clients. Generate time-limited signed URLs server-side.

Java
import boto3
from datetime import timedelta

s3 = boto3.client("s3")
url = s3.generate_presigned_url(
    "get_object",
    Params={"Bucket": "streamhub-media", "Key": "videos/clip_8f2a.mp4"},
    ExpiresIn=3600  # 1 hour
)

Scale estimates

MetricEstimate
Total objects10B+
Average object size5 MB (heavy tail of large videos)
Total storage50 PB
PUT rate peak100K objects/s
GET rate peak1M objects/s (CDN absorbs 90%)
Durability target99.999999999% (11 nines)
  • Total objects

    Estimate10B+
  • Average object size

    Estimate5 MB (heavy tail of large videos)
  • Total storage

    Estimate50 PB
  • PUT rate peak

    Estimate100K objects/s
  • GET rate peak

    Estimate1M objects/s (CDN absorbs 90%)
  • Durability target

    Estimate99.999999999% (11 nines)

Front S3 with CloudFront CDN for read-heavy workloads; S3 handles writes and origin fetches.

Event-driven processing

S3 event notifications trigger downstream pipelines on object creation.

Java
{
  "Records": [{
    "eventName": "ObjectCreated:Put",
    "s3": {
      "bucket": { "name": "streamhub-media" },
      "object": { "key": "uploads/raw/clip_8f2a.mp4", "size": 524288000 }
    }
  }]
}

Lambda or Kafka consumers pick up the event to transcode, generate thumbnails, or index metadata — the pattern behind StreamHub's upload-to-play pipeline.

Quick recall

Everything you need if you only revisit this box.

  • Object storage = flat key namespace, opaque blobs, optimised for throughput not random I/O.
  • Multipart upload for large files: parallel parts, then CompleteMultipartUpload.
  • Storage classes trade cost vs retrieval latency; lifecycle rules automate transitions.
  • Presigned URLs grant time-limited access without exposing credentials.
  • S3 events trigger async processing (transcode, index) on object creation.

Test yourself

Answer these before moving on — recall is what makes it stick.