PrepZone Logo
PrepZone

ML Inference Serving

Model registries, GPU autoscaling, A/B routing and latency budgets for production ML.

Why this matters

  • StreamHub uses ML for content moderation (flag inappropriate streams), recommendation ranking (suggest videos), and thumbnail quality scoring — each needs a production inference path.
  • A model that scores 95% accuracy in Jupyter is worthless if inference takes 5 seconds or crashes under 100 concurrent requests.
  • GPU autoscaling, model versioning, and shadow deployments are production ML engineering, not data science.

Inference serving components

  • Model registry — versioned model artifacts (weights, config, preprocessing) in S3 or MLflow.
  • Inference server — TorchServe, Triton, or BentoML serving predictions via REST/gRPC.
  • Feature store — pre-computed features for low-latency inference (Feast, Tecton).
  • Autoscaling — scale GPU pods on queue depth or request rate; scale to zero for batch models.
  • A/B routing — split traffic between model versions; compare accuracy and latency.
  • Monitoring — prediction latency, throughput, data drift, model accuracy degradation.

Serving architecture

ML inference on AWS

CLIENT
API request
DATABASE
Feature svcRedis cache
COMPUTE
SageMakerGPU endpoint
CLIENT
Responsefraud score
Feature cache → SageMaker endpoint → low-latency prediction API.
Java
# FastAPI inference endpoint
from fastapi import FastAPI
import torch

app = FastAPI()
model = torch.jit.load("models/content-moderator-v3.pt")
model.eval()

@app.post("/v1/moderate")
async def moderate(request: ModerateRequest):
    features = preprocess(request.thumbnail_url, request.title)
    with torch.no_grad():
        logits = model(features)
        scores = torch.softmax(logits, dim=-1)

    return {
        "model_version": "content-moderator-v3",
        "scores": {
            "safe": float(scores[0]),
            "nsfw": float(scores[1]),
            "violence": float(scores[2])
        },
        "action": "block" if scores[1] > 0.85 else "allow"
    }

Kubernetes deployment with GPU

Java
apiVersion: apps/v1
kind: Deployment
metadata:
  name: content-moderator
spec:
  replicas: 3
  template:
    spec:
      containers:
        - name: inference
          image: registry.streamhub.io/moderator:v3.2.0
          resources:
            limits:
              nvidia.com/gpu: 1
              memory: 8Gi
            requests:
              nvidia.com/gpu: 1
              memory: 4Gi
              cpu: 2
          ports:
            - containerPort: 8080
          livenessProbe:
            httpGet:
              path: /health
              port: 8080
            initialDelaySeconds: 60
          env:
            - name: MODEL_PATH
              value: /models/content-moderator-v3.pt
            - name: BATCH_SIZE
              value: "32"
            - name: MAX_BATCH_WAIT_MS
              value: "10"

Batching for throughput

Individual inference is GPU-inefficient. Dynamic batching groups concurrent requests.

Java
class DynamicBatcher:
    def __init__(self, model, max_batch=32, max_wait_ms=10):
        self.model = model
        self.max_batch = max_batch
        self.max_wait_ms = max_wait_ms
        self.queue = []

    async def predict(self, features):
        future = asyncio.Future()
        self.queue.append((features, future))

        if len(self.queue) >= self.max_batch:
            await self._flush()
        else:
            asyncio.get_event_loop().call_later(
                self.max_wait_ms / 1000, self._flush
            )
        return await future

    async def _flush(self):
        if not self.queue:
            return
        batch = self.queue[:self.max_batch]
        self.queue = self.queue[self.max_batch:]
        features = torch.stack([f for f, _ in batch])
        with torch.no_grad():
            results = self.model(features)
        for (_, future), result in zip(batch, results):
            future.set_result(result)
AspectServing frameworkBest for
TorchServePyTorch models, dynamic batchingPyTorch-heavy shops
NVIDIA TritonMulti-framework (TF, ONNX, PyTorch)Mixed model types on shared GPU
BentoMLPython-native, easy packagingSmall teams, rapid iteration
SageMaker EndpointsFully managed AWSAWS-native, less ops burden
vLLMLLM-specific, PagedAttentionLarge language model serving
  • TorchServe

    Serving frameworkPyTorch models, dynamic batching
    Best forPyTorch-heavy shops
  • NVIDIA Triton

    Serving frameworkMulti-framework (TF, ONNX, PyTorch)
    Best forMixed model types on shared GPU
  • BentoML

    Serving frameworkPython-native, easy packaging
    Best forSmall teams, rapid iteration
  • SageMaker Endpoints

    Serving frameworkFully managed AWS
    Best forAWS-native, less ops burden
  • vLLM

    Serving frameworkLLM-specific, PagedAttention
    Best forLarge language model serving

StreamHub uses Triton for moderation (ONNX) and vLLM for the creator support RAG generator.

A/B model testing

Java
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: recommender-routing
spec:
  hosts: [recommender]
  http:
    - route:
        - destination:
            host: recommender
            subset: model-v2
          weight: 90
        - destination:
            host: recommender
            subset: model-v3
          weight: 10

Track per-version metrics: click-through rate, watch time, inference latency. Promote v3 when CTR improves ≥ 2% with no latency regression.

Scale estimates

ModelQPS peakLatency p99GPU count
Content moderator2,00045ms6× T4
Recommendation ranker15,00012ms4× T4 (CPU fallback)
Thumbnail scorer50080ms2× T4
Support RAG (LLM)502.5s2× A100 (vLLM)
  • Content moderator

    QPS peak2,000
    Latency p9945ms
    GPU count6× T4
  • Recommendation ranker

    QPS peak15,000
    Latency p9912ms
    GPU count4× T4 (CPU fallback)
  • Thumbnail scorer

    QPS peak500
    Latency p9980ms
    GPU count2× T4
  • Support RAG (LLM)

    QPS peak50
    Latency p992.5s
    GPU count2× A100 (vLLM)

Recommendation ranker uses CPU-optimised ONNX for cost; GPU reserved for vision and LLM models.

Quick recall

Everything you need if you only revisit this box.

  • Inference serving = model registry + GPU-autoscaled pods + batching + monitoring.
  • Dynamic batching groups requests for GPU efficiency (max_batch=32, max_wait=10ms).
  • Triton for multi-framework; vLLM for LLM serving; BentoML for rapid prototyping.
  • A/B route traffic between model versions; promote on accuracy + latency metrics.
  • Monitor data drift — input distribution changes degrade model accuracy silently.

Test yourself

Answer these before moving on — recall is what makes it stick.