Why this matters
- "Works on my machine" disappears when the same Docker image runs in dev, staging, and production — only environment variables change.
- Kubernetes rolling updates deploy new BookStore versions with zero downtime by gradually replacing pods while readiness probes gate traffic.
- Graceful shutdown ensures in-flight checkout requests complete before the pod terminates, preventing lost orders during deploys.
StreamHub production architecture (AWS)
Deployment essentials
- Multi-stage Dockerfile — build with Maven in one stage, run a slim JRE image in the final stage.
- Liveness probe — restarts the pod if the JVM is wedged; hits
/actuator/health/liveness. - Readiness probe — removes the pod from load balancing until it can serve traffic.
- Resource limits —
requestsandlimitsfor CPU and memory prevent one pod from starving the node. - Graceful shutdown —
server.shutdown: graceful+terminationGracePeriodSecondslet requests drain.
Production Dockerfile
# Build stage
FROM eclipse-temurin:21-jdk-alpine AS build
WORKDIR /app
COPY mvnw pom.xml ./
COPY .mvn .mvn
RUN ./mvnw dependency:go-offline -B
COPY src ./src
RUN ./mvnw package -DskipTests -B
# Runtime stage
FROM eclipse-temurin:21-jre-alpine
RUN addgroup -S bookstore && adduser -S bookstore -G bookstore
WORKDIR /app
COPY --from=build /app/target/bookstore-api-*.jar app.jar
USER bookstore
EXPOSE 8080
ENTRYPOINT ["java", \
"-XX:+UseContainerSupport", \
"-XX:MaxRAMPercentage=75.0", \
"-Djava.security.egd=file:/dev/./urandom", \
"-jar", "app.jar"]
Graceful shutdown configuration
server:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 30s
@Bean
public TomcatConnectorCustomizer gracefulShutdownCustomizer() {
return connector -> connector.setProperty("connectionTimeout", "20000");
}
Kubernetes deployment manifest
apiVersion: apps/v1
kind: Deployment
metadata:
name: bookstore-api
labels:
app: bookstore-api
spec:
replicas: 3
selector:
matchLabels:
app: bookstore-api
template:
metadata:
labels:
app: bookstore-api
spec:
terminationGracePeriodSeconds: 45
containers:
- name: bookstore-api
image: registry.example.com/bookstore-api:2.4.1
ports:
- containerPort: 8080
env:
- name: SPRING_PROFILES_ACTIVE
value: prod
- name: SPRING_DATASOURCE_URL
valueFrom:
secretKeyRef:
name: bookstore-db
key: url
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 1000m
memory: 1Gi
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 60
periodSeconds: 10
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
initialDelaySeconds: 30
periodSeconds: 5
---
apiVersion: v1
kind: Service
metadata:
name: bookstore-api
spec:
selector:
app: bookstore-api
ports:
- port: 80
targetPort: 8080
type: ClusterIP
Build and deploy pipeline
docker build -t registry.example.com/bookstore-api:2.4.1 .
docker push registry.example.com/bookstore-api:2.4.1
kubectl apply -f k8s/deployment.yaml
kubectl rollout status deployment/bookstore-api
Quick recall
Everything you need if you only revisit this box.
- Multi-stage Dockerfiles keep production images small — JRE only, no Maven or source code.
- Enable
server.shutdown: gracefuland setterminationGracePeriodSecondsabove your longest request. - Liveness restarts wedged pods; readiness controls whether the pod receives traffic.
- Set CPU/memory requests and limits; use
-XX:MaxRAMPercentage=75.0so the JVM respects container memory. - Point probes at
/actuator/health/livenessand/actuator/health/readinesswith adequateinitialDelaySeconds.
Test yourself
Answer these before moving on — recall is what makes it stick.