Why this matters
- S3-compatible object stores underpin StreamHub video uploads, user avatars, static assets, and ML training datasets.
- Understanding multipart upload, versioning, and lifecycle policies is essential for designing any blob-heavy system.
- Object storage optimises for throughput and durability, not low-latency random access — choose the right tier.
S3 core concepts
- Bucket — top-level container with regional placement and access policies.
- Object key — unique identifier within a bucket (
videos/2026/04/clip_8f2a.mp4). - Metadata — user-defined key-value pairs and system metadata (content-type, etag).
- Versioning — retain previous object versions on overwrite; protect against accidental deletion.
- Lifecycle rules — auto-transition to cheaper tiers or expire objects after N days.
Request flow
S3 object storage architecture
PUT /streamhub-media/videos/2026/04/clip_8f2a.mp4
Host: s3.us-east-1.amazonaws.com
Content-Type: video/mp4
Content-Length: 524288000
x-amz-storage-class: STANDARD
<binary body>
GET /streamhub-media/videos/2026/04/clip_8f2a.mp4
Host: s3.us-east-1.amazonaws.com
→ 200 OK
ETag: "d41d8cd98f00b204e9800998ecf8427e"
Content-Type: video/mp4
Multipart upload for large objects
Objects over 100 MB should use multipart upload: split into parts, upload in parallel, then assemble.
POST /streamhub-media/videos/large_clip.mp4?uploads
→ <UploadId>upload_xyz</UploadId>
PUT /streamhub-media/videos/large_clip.mp4?partNumber=1&uploadId=upload_xyz
PUT /streamhub-media/videos/large_clip.mp4?partNumber=2&uploadId=upload_xyz
...
POST /streamhub-media/videos/large_clip.mp4?uploadId=upload_xyz
<CompleteMultipartUpload>
<Part><PartNumber>1</PartNumber><ETag>"abc"</ETag></Part>
<Part><PartNumber>2</PartNumber><ETag>"def"</ETag></Part>
</CompleteMultipartUpload>
Durability and consistency
| Aspect | Storage class | Use case |
|---|---|---|
| STANDARD | 11 nines durability, ms latency | Hot data: active videos, thumbnails |
| STANDARD_IA | Lower cost, retrieval fee | Infrequently accessed archives |
| GLACIER | Minutes-to-hours retrieval | Compliance backups, old logs |
| INTELLIGENT_TIERING | Auto-moves by access pattern | Unpredictable access frequency |
STANDARD
Storage class11 nines durability, ms latencyUse caseHot data: active videos, thumbnailsSTANDARD_IA
Storage classLower cost, retrieval feeUse caseInfrequently accessed archivesGLACIER
Storage classMinutes-to-hours retrievalUse caseCompliance backups, old logsINTELLIGENT_TIERING
Storage classAuto-moves by access patternUse caseUnpredictable access frequency
S3 provides read-after-write consistency for new objects; overwrites are eventually consistent in older APIs — use version IDs for critical data.
Lifecycle policy example
{
"Rules": [
{
"ID": "archive-old-uploads",
"Filter": { "Prefix": "uploads/raw/" },
"Status": "Enabled",
"Transitions": [
{ "Days": 30, "StorageClass": "STANDARD_IA" },
{ "Days": 90, "StorageClass": "GLACIER" }
],
"Expiration": { "Days": 365 }
}
]
}
Presigned URLs for secure access
Never expose bucket credentials to clients. Generate time-limited signed URLs server-side.
import boto3
from datetime import timedelta
s3 = boto3.client("s3")
url = s3.generate_presigned_url(
"get_object",
Params={"Bucket": "streamhub-media", "Key": "videos/clip_8f2a.mp4"},
ExpiresIn=3600 # 1 hour
)
Scale estimates
| Metric | Estimate |
|---|---|
| Total objects | 10B+ |
| Average object size | 5 MB (heavy tail of large videos) |
| Total storage | 50 PB |
| PUT rate peak | 100K objects/s |
| GET rate peak | 1M objects/s (CDN absorbs 90%) |
| Durability target | 99.999999999% (11 nines) |
Total objects
Estimate10B+Average object size
Estimate5 MB (heavy tail of large videos)Total storage
Estimate50 PBPUT rate peak
Estimate100K objects/sGET rate peak
Estimate1M objects/s (CDN absorbs 90%)Durability target
Estimate99.999999999% (11 nines)
Front S3 with CloudFront CDN for read-heavy workloads; S3 handles writes and origin fetches.
Event-driven processing
S3 event notifications trigger downstream pipelines on object creation.
{
"Records": [{
"eventName": "ObjectCreated:Put",
"s3": {
"bucket": { "name": "streamhub-media" },
"object": { "key": "uploads/raw/clip_8f2a.mp4", "size": 524288000 }
}
}]
}
Lambda or Kafka consumers pick up the event to transcode, generate thumbnails, or index metadata — the pattern behind StreamHub's upload-to-play pipeline.
Quick recall
Everything you need if you only revisit this box.
- Object storage = flat key namespace, opaque blobs, optimised for throughput not random I/O.
- Multipart upload for large files: parallel parts, then CompleteMultipartUpload.
- Storage classes trade cost vs retrieval latency; lifecycle rules automate transitions.
- Presigned URLs grant time-limited access without exposing credentials.
- S3 events trigger async processing (transcode, index) on object creation.
Test yourself
Answer these before moving on — recall is what makes it stick.