Why this matters
- StreamHub creators upload thumbnails, clips, and project assets — the same chunked-upload and sync patterns power Google Drive, Dropbox, and S3-backed drives.
- Resumable uploads prevent a 2 GB video failing at 99% due to a network blip.
- Block-level deduplication saves storage when users share files or sync the same content across devices.
Drive architecture layers
- Metadata service — file names, paths, ownership, sharing ACLs, version history pointers.
- Block/chunk store — content-addressed blobs (hash as key) in object storage.
- Sync engine — detects local changes, computes diffs, uploads only modified chunks.
- Notification service — pushes change events to connected clients via long-polling or WebSocket.
- Search indexer — full-text index over file names and extracted document content.
Metadata vs content split
Cloud file storage
CREATE TABLE files (
file_id UUID PRIMARY KEY,
owner_id UUID NOT NULL,
parent_id UUID, -- folder reference
name VARCHAR(255),
mime_type VARCHAR(128),
size_bytes BIGINT,
content_hash VARCHAR(64), -- SHA-256 of assembled content
version INT DEFAULT 1,
created_at TIMESTAMPTZ,
updated_at TIMESTAMPTZ
);
CREATE TABLE file_chunks (
file_id UUID,
chunk_index INT,
chunk_hash VARCHAR(64), -- content-addressed
size_bytes INT,
PRIMARY KEY (file_id, chunk_index)
);
Chunked resumable upload
Large files upload in fixed-size chunks (e.g. 8 MB). Each chunk is content-addressed and stored independently.
POST /v1/uploads/init
Content-Type: application/json
{ "name": "stream-recording.mp4", "size_bytes": 2147483648, "mime_type": "video/mp4" }
→ 201 { "upload_id": "upl_9k2m", "chunk_size": 8388608, "total_chunks": 256 }
PUT /v1/uploads/upl_9k2m/chunks/0
Content-Range: bytes 0-8388607/2147483648
Content-MD5: d41d8cd98f00b204e9800998ecf8427e
<binary chunk data>
POST /v1/uploads/upl_9k2m/complete
→ 200 { "file_id": "fil_a8x2", "content_hash": "sha256:abc123..." }
Sync and conflict resolution
Clients maintain a local manifest of (path, content_hash, modified_at). On sync:
| Aspect | Scenario | Resolution |
|---|---|---|
| Local newer, server unchanged | Upload diff chunks | Standard sync |
| Server newer, local unchanged | Download diff chunks | Pull update |
| Both modified | Conflict detected | Keep both as version N and N+1, or prompt user |
| Same content, different name | Hash match | Deduplicate — no re-upload |
Local newer, server unchanged
ScenarioUpload diff chunksResolutionStandard syncServer newer, local unchanged
ScenarioDownload diff chunksResolutionPull updateBoth modified
ScenarioConflict detectedResolutionKeep both as version N and N+1, or prompt userSame content, different name
ScenarioHash matchResolutionDeduplicate — no re-upload
Content-hash dedup means two users sharing the same 1 GB file stores it once in the block store.
Sharing and permissions
{
"file_id": "fil_a8x2",
"acl": [
{ "principal": "user:usr_42", "role": "owner" },
{ "principal": "user:usr_99", "role": "editor" },
{ "principal": "link:public", "role": "viewer", "expires_at": "2026-05-01T00:00:00Z" }
]
}
Permission checks happen at the metadata layer before generating a signed URL for chunk download from object storage.
Scale estimates
| Metric | Estimate |
|---|---|
| Registered users | 100M |
| Average files per user | 500 |
| Total metadata rows | 50B |
| Average file size | 2 MB (skewed — few huge files) |
| Storage after dedup | ~30% savings from shared content |
| Upload throughput peak | 50 GB/s globally |
Registered users
Estimate100MAverage files per user
Estimate500Total metadata rows
Estimate50BAverage file size
Estimate2 MB (skewed — few huge files)Storage after dedup
Estimate~30% savings from shared contentUpload throughput peak
Estimate50 GB/s globally
Metadata in a sharded SQL cluster; content in regional S3 buckets with cross-region replication for durability.
Change notification pipeline
When a file changes, publish an event so connected clients and collaborators update.
{
"event": "file.updated",
"file_id": "fil_a8x2",
"version": 3,
"changed_by": "usr_99",
"timestamp": "2026-04-01T15:00:00Z"
}
Fan out via WebSocket to subscribed clients. For offline devices, sync on next connection using the version cursor.
Quick recall
Everything you need if you only revisit this box.
- Separate metadata (SQL, queryable) from content (object store, content-addressed chunks).
- Resumable uploads: init → PUT chunks → complete; each chunk keyed by hash.
- Content-hash dedup saves storage when files are shared or synced across devices.
- Conflict resolution: compare versions; keep both or prompt on simultaneous edits.
- ACL checks at metadata layer; serve content via time-limited signed URLs.
Test yourself
Answer these before moving on — recall is what makes it stick.