Why this matters
- StreamHub uses ML for content moderation (flag inappropriate streams), recommendation ranking (suggest videos), and thumbnail quality scoring — each needs a production inference path.
- A model that scores 95% accuracy in Jupyter is worthless if inference takes 5 seconds or crashes under 100 concurrent requests.
- GPU autoscaling, model versioning, and shadow deployments are production ML engineering, not data science.
Inference serving components
- Model registry — versioned model artifacts (weights, config, preprocessing) in S3 or MLflow.
- Inference server — TorchServe, Triton, or BentoML serving predictions via REST/gRPC.
- Feature store — pre-computed features for low-latency inference (Feast, Tecton).
- Autoscaling — scale GPU pods on queue depth or request rate; scale to zero for batch models.
- A/B routing — split traffic between model versions; compare accuracy and latency.
- Monitoring — prediction latency, throughput, data drift, model accuracy degradation.
Serving architecture
ML inference on AWS
# FastAPI inference endpoint
from fastapi import FastAPI
import torch
app = FastAPI()
model = torch.jit.load("models/content-moderator-v3.pt")
model.eval()
@app.post("/v1/moderate")
async def moderate(request: ModerateRequest):
features = preprocess(request.thumbnail_url, request.title)
with torch.no_grad():
logits = model(features)
scores = torch.softmax(logits, dim=-1)
return {
"model_version": "content-moderator-v3",
"scores": {
"safe": float(scores[0]),
"nsfw": float(scores[1]),
"violence": float(scores[2])
},
"action": "block" if scores[1] > 0.85 else "allow"
}
Kubernetes deployment with GPU
apiVersion: apps/v1
kind: Deployment
metadata:
name: content-moderator
spec:
replicas: 3
template:
spec:
containers:
- name: inference
image: registry.streamhub.io/moderator:v3.2.0
resources:
limits:
nvidia.com/gpu: 1
memory: 8Gi
requests:
nvidia.com/gpu: 1
memory: 4Gi
cpu: 2
ports:
- containerPort: 8080
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 60
env:
- name: MODEL_PATH
value: /models/content-moderator-v3.pt
- name: BATCH_SIZE
value: "32"
- name: MAX_BATCH_WAIT_MS
value: "10"
Batching for throughput
Individual inference is GPU-inefficient. Dynamic batching groups concurrent requests.
class DynamicBatcher:
def __init__(self, model, max_batch=32, max_wait_ms=10):
self.model = model
self.max_batch = max_batch
self.max_wait_ms = max_wait_ms
self.queue = []
async def predict(self, features):
future = asyncio.Future()
self.queue.append((features, future))
if len(self.queue) >= self.max_batch:
await self._flush()
else:
asyncio.get_event_loop().call_later(
self.max_wait_ms / 1000, self._flush
)
return await future
async def _flush(self):
if not self.queue:
return
batch = self.queue[:self.max_batch]
self.queue = self.queue[self.max_batch:]
features = torch.stack([f for f, _ in batch])
with torch.no_grad():
results = self.model(features)
for (_, future), result in zip(batch, results):
future.set_result(result)
| Aspect | Serving framework | Best for |
|---|---|---|
| TorchServe | PyTorch models, dynamic batching | PyTorch-heavy shops |
| NVIDIA Triton | Multi-framework (TF, ONNX, PyTorch) | Mixed model types on shared GPU |
| BentoML | Python-native, easy packaging | Small teams, rapid iteration |
| SageMaker Endpoints | Fully managed AWS | AWS-native, less ops burden |
| vLLM | LLM-specific, PagedAttention | Large language model serving |
TorchServe
Serving frameworkPyTorch models, dynamic batchingBest forPyTorch-heavy shopsNVIDIA Triton
Serving frameworkMulti-framework (TF, ONNX, PyTorch)Best forMixed model types on shared GPUBentoML
Serving frameworkPython-native, easy packagingBest forSmall teams, rapid iterationSageMaker Endpoints
Serving frameworkFully managed AWSBest forAWS-native, less ops burdenvLLM
Serving frameworkLLM-specific, PagedAttentionBest forLarge language model serving
StreamHub uses Triton for moderation (ONNX) and vLLM for the creator support RAG generator.
A/B model testing
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: recommender-routing
spec:
hosts: [recommender]
http:
- route:
- destination:
host: recommender
subset: model-v2
weight: 90
- destination:
host: recommender
subset: model-v3
weight: 10
Track per-version metrics: click-through rate, watch time, inference latency. Promote v3 when CTR improves ≥ 2% with no latency regression.
Scale estimates
| Model | QPS peak | Latency p99 | GPU count |
|---|---|---|---|
| Content moderator | 2,000 | 45ms | 6× T4 |
| Recommendation ranker | 15,000 | 12ms | 4× T4 (CPU fallback) |
| Thumbnail scorer | 500 | 80ms | 2× T4 |
| Support RAG (LLM) | 50 | 2.5s | 2× A100 (vLLM) |
Content moderator
QPS peak2,000Latency p9945msGPU count6× T4Recommendation ranker
QPS peak15,000Latency p9912msGPU count4× T4 (CPU fallback)Thumbnail scorer
QPS peak500Latency p9980msGPU count2× T4Support RAG (LLM)
QPS peak50Latency p992.5sGPU count2× A100 (vLLM)
Recommendation ranker uses CPU-optimised ONNX for cost; GPU reserved for vision and LLM models.
Quick recall
Everything you need if you only revisit this box.
- Inference serving = model registry + GPU-autoscaled pods + batching + monitoring.
- Dynamic batching groups requests for GPU efficiency (max_batch=32, max_wait=10ms).
- Triton for multi-framework; vLLM for LLM serving; BentoML for rapid prototyping.
- A/B route traffic between model versions; promote on accuracy + latency metrics.
- Monitor data drift — input distribution changes degrade model accuracy silently.
Test yourself
Answer these before moving on — recall is what makes it stick.