DevOps Engineer Interview Questions
Core Overview
Prepare for DevOps engineer interviews covering Linux administration, networking, infrastructure, CI/CD, containers, Kubernetes, cloud platforms, infrastructure as code, security, observability, reliability, and incident response.
Ready to test your knowledge?
Launch a focused practice session to review questions without distraction.
What are Linux processes and signals, and how should an operator stop a running service safely?
Direct Answer
A process is a running program with its own identifier and resources. Signals request actions such as graceful termination, reload, interruption, or forced termination.
Detailed Explanation
A Linux process is an executing program identified by a process ID, or PID. A process has resources such as memory, environment variables, open file descriptors, credentials, and signal state.
Useful process-inspection tools include:
ps for listing process informationtop or similar tools for resource activity/proc/<pid> for process-specific kernel informationsystemctl status for services managed by systemdjournalctl for service logslsof for open files and socketsA signal is a mechanism used to notify a process that an event occurred or request a particular action.
Common signals include:
SIGTERM: Requests graceful termination and can be handled by the process.SIGINT: Commonly sent by Ctrl+C for an interactive interruption.SIGHUP: Often used by services to request configuration reload, though behavior is application-specific.SIGKILL: Immediately terminates the process and cannot be caught or ignored.An operator should normally use the service manager rather than killing a PID directly. For example, systemctl stop application.service allows the service manager to send the configured termination signal, wait for graceful shutdown, and track the resulting state.
A reasonable escalation path is:
1. Confirm the correct process or service.
2. Review its current state and logs.
3. Request graceful termination with the service manager or SIGTERM.
4. Wait for cleanup within the defined timeout.
5. Use SIGKILL only when the process cannot terminate safely through normal mechanisms.
Immediately using SIGKILL can interrupt active writes, skip shutdown hooks, leave temporary resources behind, and hide the underlying reason the service became unresponsive.
Code Example
# Inspect the service and recent logs.
systemctl status resumeloop-api.service
journalctl --unit resumeloop-api.service --since "15 minutes ago"
# Ask the service manager to stop it gracefully.
sudo systemctl stop resumeloop-api.service
# Inspect an unmanaged process.
ps -fp 2481
cat /proc/2481/status
ls -l /proc/2481/fd
# Request graceful termination.
kill -TERM 2481
# Force termination only after investigation
# and an appropriate grace period.
kill -KILL 2481Common Interview Pitfalls
- Sending SIGKILL before attempting a graceful service shutdown.
- Killing a process without confirming its PID and ownership.
- Assuming every application interprets SIGHUP as a configuration reload.
- Restarting a failing service repeatedly without reviewing logs or dependencies.
- Managing a systemd service by killing its child process directly.
- Ignoring active requests and cleanup requirements during termination.
How do Linux file permissions, ownership, and file descriptors affect application operation?
Direct Answer
Ownership and permission bits control filesystem access, while file descriptors are per-process references to open files, sockets, pipes, and other I/O resources.
Detailed Explanation
Linux permissions are evaluated using the file owner, group, mode bits, access-control mechanisms, and the credentials of the requesting process.
Traditional mode bits define permissions for:
The common permissions are:
r: Readw: Writex: Execute for files or traverse for directoriesDirectory permissions have distinct meanings:
A file descriptor is a small process-local integer referring to an open file description. Descriptors can represent regular files, directories, pipes, terminals, sockets, and devices.
The conventional descriptors are:
0: Standard input1: Standard output2: Standard errorApplications can fail even when the executable starts correctly if they cannot read configuration, write logs, bind sockets, access certificates, or traverse parent directories.
Process execution may also exhaust its descriptor limit because of leaked files or sockets. Symptoms can include Too many open files, failed network connections, or inability to open logs.
Troubleshooting should verify the effective service user, ownership, full path permissions, mount mode, descriptor usage, and configured resource limits. Broadly applying chmod 777 weakens security and usually conceals the actual ownership or deployment problem.
Code Example
# Inspect ownership and permissions.
ls -ld /opt/resumeloop
ls -l /opt/resumeloop/config
namei -l /opt/resumeloop/config/app.env
# Inspect the service identity.
systemctl show resumeloop-api.service --property=User --property=Group
# Inspect open descriptors.
ls /proc/2481/fd | wc -l
lsof -p 2481
# Inspect process limits.
cat /proc/2481/limits
# Apply narrow ownership and permissions.
sudo chown root:resumeloop /opt/resumeloop/config/app.env
sudo chmod 640 /opt/resumeloop/config/app.envCommon Interview Pitfalls
- Using chmod 777 instead of correcting ownership and required access.
- Checking only the target file while ignoring permissions on parent directories.
- Assuming an application runs as the interactive shell user.
- Ignoring leaked file descriptors during connection failures.
- Making private keys or environment files readable by every user.
- Changing permissions without considering container or service-manager identities.
How do TCP, UDP, ports, and sockets relate when troubleshooting application connectivity?
Direct Answer
IP identifies hosts, ports identify transport endpoints, sockets combine addressing and protocol state, TCP provides reliable streams, and UDP sends independent datagrams.
Detailed Explanation
Networked applications communicate through protocol layers.
An IP address identifies an interface or network endpoint. A transport-layer port identifies an application endpoint on that address.
A socket is an operating-system object used for network communication. A listening server socket is commonly identified by protocol, local address, and local port. An established TCP connection includes both local and remote address and port information.
TCP provides an ordered byte stream with connection establishment, acknowledgements, retransmission, flow control, and congestion control. Applications such as HTTPS commonly use TCP, although newer application protocols may use other transports.
UDP sends independent datagrams without TCP’s connection and delivery guarantees. Applications must tolerate loss, duplication, and reordering or provide required reliability themselves.
When a service is unreachable, distinguish among several conditions:
Tools such as ss, ip, ping where permitted, traceroute, nc, and curl provide evidence at different layers.
A successful TCP connection proves only that the transport endpoint accepted the connection. It does not prove that TLS validation, HTTP routing, authentication, or application dependencies are healthy.
Code Example
# Show listening TCP and UDP sockets.
ss -lntup
# Confirm the service binding.
ss -lntp '( sport = :3000 )'
# Inspect local addresses and routes.
ip address show
ip route show
# Test transport connectivity.
nc -vz api.example.com 443
# Test the application and TLS path.
curl --verbose --connect-timeout 5 https://api.example.com/health
# Test the application directly on the host.
curl --verbose http://127.0.0.1:3000/healthCommon Interview Pitfalls
- Assuming an open TCP port proves that the application is healthy.
- Binding a service only to loopback when remote traffic must reach it.
- Testing only with ping and concluding that every TCP service is unavailable.
- Confusing a connection refusal with a connection timeout.
- Ignoring outbound network rules for dependency connections.
- Troubleshooting HTTP before confirming DNS, route, and transport connectivity.
How should a DevOps engineer troubleshoot DNS and TLS failures along an HTTPS request path?
Direct Answer
Verify DNS records and resolver behavior, network reachability, server identity, certificate validity, trust chains, hostname matching, and protocol negotiation in sequence.
Detailed Explanation
An HTTPS request depends on several systems operating correctly.
A simplified path includes:
1. The client resolves the hostname through DNS.
2. The client selects an address from the result.
3. Routing and network policy allow the connection to the destination port.
4. A TCP or other supported transport connection is established.
5. The client and server negotiate TLS.
6. The server presents its certificate chain.
7. The client validates the trust chain, validity period, hostname, and other certificate requirements.
8. The client sends the HTTP request.
9. A proxy, load balancer, ingress, or application routes the request.
DNS troubleshooting should check:
TLS troubleshooting should check:
Tools such as dig, getent, curl, and openssl s_client help isolate these stages.
Temporarily disabling certificate verification may confirm that trust validation is involved, but it is not an acceptable production fix. The underlying chain, hostname, or trust-store problem must be corrected.
Code Example
# Resolve through the configured system resolver.
getent ahosts api.example.com
# Inspect DNS records directly.
dig api.example.com A
dig api.example.com AAAA
dig api.example.com CNAME
# Inspect the HTTPS request path.
curl --verbose --resolve api.example.com:443:203.0.113.10 https://api.example.com/health
# Inspect the certificate chain and SNI behavior.
openssl s_client -connect api.example.com:443 -servername api.example.com -showcerts
# Show key certificate fields.
openssl s_client -connect api.example.com:443 -servername api.example.com </dev/null 2>/dev/null |
openssl x509 -noout -subject -issuer -dates -ext subjectAltNameCommon Interview Pitfalls
- Assuming every hostname-resolution failure is an authoritative DNS problem.
- Testing an HTTPS endpoint by IP without sending the expected hostname.
- Disabling TLS verification permanently instead of correcting the trust problem.
- Checking only certificate expiration while ignoring hostname and chain validation.
- Forgetting that DNS caches can preserve an old result until expiration.
- Ignoring differences between IPv4 and IPv6 connection paths.
- Troubleshooting the application before confirming TLS termination and routing.
How do subnets, route tables, security groups, and network ACLs control traffic in a cloud virtual network?
Direct Answer
Subnets divide address space, route tables choose next hops, security groups filter resource traffic statefully, and network ACLs filter subnet traffic using ordered stateless rules.
Detailed Explanation
A cloud virtual network provides isolated address space and routing for deployed resources.
A subnet represents a range of IP addresses within the virtual network. Subnets are commonly distributed across availability zones and classified according to their routing and exposure requirements.
A subnet is not public merely because it has a public-sounding name. Its effective behavior depends on routing, address assignment, gateway configuration, and security policy.
A route table determines where traffic matching a destination prefix is sent. Routes may target:
A security group is associated with network interfaces or resources and controls allowed inbound and outbound traffic. AWS security groups are stateful: return traffic for an allowed connection is automatically permitted.
A network ACL applies at the subnet boundary. AWS network ACLs use ordered allow and deny rules and are stateless, meaning return traffic must also be permitted explicitly.
Troubleshooting connectivity requires evaluating the complete forward and return path:
1. Source address and route
2. Source egress policy
3. Intermediate gateways or load balancers
4. Destination subnet route
5. Network ACL rules in both directions
6. Destination security-group rules
7. Operating-system firewall
8. Listening application socket
Opening inbound traffic alone may not solve the issue if the return route or ephemeral-port policy is missing.
Code Example
# Example conceptual routing:
# Public application subnet
10.0.10.0/24 -> local
0.0.0.0/0 -> internet-gateway
# Private worker subnet
10.0.20.0/24 -> local
0.0.0.0/0 -> nat-gateway
# Database security policy:
# Inbound TCP 5432 only from application
# and worker security groups.
# Troubleshooting from a Linux workload:
ip address show
ip route show
ss -lntp
curl --connect-timeout 5 https://dependency.example.com
traceroute dependency.example.comCommon Interview Pitfalls
- Assuming a subnet is public solely because of its name.
- Opening a security group to every source instead of the required workload.
- Ignoring return-path routing and stateless network ACL rules.
- Confusing a resource security group with a subnet-level network ACL.
- Allowing direct public access to a database that only application workloads require.
- Checking cloud firewall rules without confirming that the process is listening.
- Adding a default route without understanding which gateway receives the traffic.
How would you systematically troubleshoot an intermittent production service failure across Linux, networking, and cloud infrastructure?
Direct Answer
Define the symptom and timeline, follow one failing request across layers, compare healthy and unhealthy instances, gather evidence, test one hypothesis at a time, and preserve recovery options.
Detailed Explanation
Intermittent infrastructure failures require a structured investigation because restarting components too early can destroy the evidence needed to find the cause.
A practical method includes the following stages.
Establish impact and timeline
Confirm the symptom
Distinguish among DNS failure, connection timeout, connection refusal, TLS failure, HTTP error, slow response, process crash, resource exhaustion, and dependency failure.
Follow the request path
Trace one failing request through:
1. DNS resolution
2. Client or CDN
3. Load balancer or proxy
4. Network route and security policy
5. Host and operating-system firewall
6. Listening process
7. Application logs
8. Database, cache, queue, or external dependency
Inspect Linux resources
Review:
Compare healthy and unhealthy instances
Differences in image version, configuration, secret version, route, subnet, certificate, limits, or workload can narrow the search quickly.
Form and test hypotheses
Change one variable at a time where possible. Use packet capture, system-call tracing, application traces, or temporary diagnostics only when they answer a defined question and can be used safely in production.
Mitigate while investigating
Remove an unhealthy instance from traffic, reduce load, disable a failing optional feature, roll back a release, or increase capacity without presenting the mitigation as the confirmed root cause.
Complete the investigation
Record the triggering condition, contributing factors, why detection or safeguards failed, and corrective actions. Corrective work should address the systemic weakness rather than only restarting the affected process.
Code Example
# Process and service evidence.
systemctl status resumeloop-api.service
journalctl --unit resumeloop-api.service --since "2026-08-05 13:30:00"
# Kernel and resource evidence.
dmesg --time-format iso | tail -n 100
free -m
df -h
df -i
cat /proc/2481/limits
# Socket and connection evidence.
ss -s
ss -lntp
ss -tan state time-wait
ss -tan state established
# Process-level evidence.
ps -o pid,ppid,state,%cpu,%mem,etime,cmd -p 2481
lsof -p 2481 | wc -l
# Network-path evidence.
dig api.example.com
curl --verbose https://api.example.com/health
# Use system-call tracing only for a defined,
# controlled diagnostic question.
sudo strace -f -p 2481 -e trace=network,fileCommon Interview Pitfalls
- Restarting the affected instance before preserving logs and runtime evidence.
- Changing several network and application settings simultaneously.
- Assuming correlation with a deployment proves the deployment caused the failure.
- Investigating only average metrics that hide instance-specific failures.
- Treating a temporary mitigation as the confirmed root cause.
- Running invasive tracing or packet capture without scope and data-handling controls.
- Ignoring kernel logs during memory, disk, or network resource failures.
- Troubleshooting each layer independently without following one request end to end.
What is the difference between continuous integration, continuous delivery, and continuous deployment?
Direct Answer
Continuous integration verifies frequent code changes, continuous delivery keeps verified releases deployable, and continuous deployment automatically releases every qualifying change.
Detailed Explanation
Continuous integration, or CI, is the practice of merging small changes frequently and validating each change through automated checks.
A CI workflow commonly performs:
The goal is to detect integration problems close to the change that introduced them.
Continuous delivery extends CI by ensuring that a verified build is always in a deployable state. Production deployment may still require a business decision, approval, maintenance window, or manual trigger.
Continuous deployment automatically releases every qualifying change to production after required verification and deployment controls pass.
The distinction is therefore largely about the final production-release decision:
A pipeline that only builds code but cannot reproduce a production artifact is not a complete delivery system. Similarly, automatically copying files to a server without adequate testing is automation, but it is not necessarily safe continuous deployment.
Teams should choose the production-release model according to product risk, recovery capability, regulatory requirements, test confidence, and operational maturity.
Code Example
name: Application pipeline
on:
pull_request:
push:
branches:
- main
jobs:
verify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@<pinned-commit>
- name: Install dependencies
run: npm ci
- name: Verify
run: |
npm run lint
npx tsc --noEmit
npm test -- --runInBand
npm run build
deploy-production:
if: github.ref == 'refs/heads/main'
needs: verify
environment: production
runs-on: ubuntu-latest
steps:
- name: Deploy verified release
run: ./scripts/deploy.shCommon Interview Pitfalls
- Using continuous delivery and continuous deployment as interchangeable terms.
- Calling a workflow continuous integration when developers integrate infrequently.
- Deploying automatically without adequate testing or recovery controls.
- Rebuilding the release differently after continuous-integration verification.
- Treating pipeline automation as a substitute for deployment observability.
- Creating one large change that is difficult to test, review, and roll back.
Why should a delivery pipeline build an immutable artifact once and promote it between environments?
Direct Answer
Building once ensures that staging and production use the same verified bytes, while immutable versions provide traceability, reproducibility, rollback, and auditability.
Detailed Explanation
A build artifact is the deployable output created by the pipeline. It may be a container image, package, binary, archive, serverless bundle, or static-site output.
An artifact should normally be:
Rebuilding separately for each environment creates a risk that production receives different code or dependencies from the artifact tested in staging. Even when the source commit is identical, dependency resolution, build timestamps, base images, or build-tool behavior can produce different output.
Environment-specific configuration should be injected separately where possible. Database addresses, credentials, feature configuration, and environment identity should not require rebuilding the application.
Container deployments should prefer immutable digests or unique content-addressed versions over mutable tags such as latest. A mutable tag may point to different content later, making investigation and rollback unreliable.
Artifact retention should support the organization’s rollback, audit, and incident-investigation requirements. Access to publish or overwrite artifacts should be tightly controlled.
Code Example
# Build one immutable image.
docker build --tag registry.example.com/resumeloop-api:${GITHUB_SHA} .
# Push the content-addressed artifact.
docker push registry.example.com/resumeloop-api:${GITHUB_SHA}
# Capture its immutable digest.
IMAGE_DIGEST="$(
docker inspect --format='{{index .RepoDigests 0}}' registry.example.com/resumeloop-api:${GITHUB_SHA}
)"
echo "Deploying ${IMAGE_DIGEST}"
# Use the same digest in staging and production.
./deploy.sh staging "${IMAGE_DIGEST}"
./deploy.sh production "${IMAGE_DIGEST}"Common Interview Pitfalls
- Rebuilding the application independently for staging and production.
- Deploying mutable latest tags without recording the image digest.
- Embedding production secrets into the build artifact.
- Allowing published artifacts to be overwritten under the same version.
- Retaining no known-good release for rollback.
- Recording the source commit without recording the actual deployed artifact.
How should pipeline stages, required status checks, branch protection, and deployment environments be used together?
Direct Answer
Stages produce progressively stronger evidence, required checks block unsafe merges, branch protection controls source changes, and environments gate sensitive deployments and secrets.
Detailed Explanation
A delivery pipeline should produce evidence in stages while avoiding unnecessary work after an early failure.
Typical stages include:
1. Source validation: Formatting, linting, policy checks, and dependency lockfile verification.
2. Build: Compilation, type checking, packaging, and artifact generation.
3. Testing: Unit, integration, contract, end-to-end, and migration tests.
4. Security: Dependency, code, secret, image, and infrastructure scanning.
5. Release: Artifact publication, provenance, signing, and release metadata.
6. Deployment: Environment-specific rollout and configuration.
7. Verification: Smoke tests, health checks, telemetry review, and exposure decisions.
A quality gate is a condition that must pass before the workflow proceeds. Examples include a successful production build, no prohibited critical vulnerability, or an approved migration review.
Required status checks connect pipeline evidence to branch governance. A protected branch can require selected checks, pull-request review, resolved discussions, and an up-to-date branch before merging.
Production environments can add a separate control layer by:
Not every gate should be manual. Excessive approval steps can become routine clicks that add delay without meaningful risk reduction. Manual approval is most useful when it represents a real business, security, compliance, or production-risk decision.
Pipeline jobs should use least privilege. Pull requests from untrusted forks should not receive production credentials or unrestricted write permissions.
Code Example
name: Verified production release
on:
push:
branches:
- main
permissions:
contents: read
jobs:
build-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@<pinned-commit>
- run: npm ci
- run: npm run lint
- run: npx tsc --noEmit
- run: npm test -- --runInBand
- run: npm run build
deploy:
needs: build-and-test
environment: production
permissions:
contents: read
id-token: write
runs-on: ubuntu-latest
steps:
- name: Authenticate using workload identity
run: ./scripts/authenticate.sh
- name: Deploy
run: ./scripts/deploy-production.shCommon Interview Pitfalls
- Allowing protected branches to merge while required verification is failing.
- Giving every pipeline job production-level credentials.
- Exposing production secrets to untrusted pull-request workflows.
- Adding manual approvals that provide no meaningful risk decision.
- Running expensive end-to-end tests before quick validation checks.
- Allowing the person initiating a sensitive deployment to be its only approver.
- Failing to pin or control third-party pipeline dependencies.
How do rolling, blue-green, and canary deployment strategies differ?
Direct Answer
Rolling replaces instances gradually, blue-green switches between complete environments, and canary exposes a new version to a small audience before increasing traffic.
Detailed Explanation
Deployment strategies control how a new version replaces the current production version and how much user impact a defect may create.
Rolling deployment
Blue-green deployment
Canary deployment
These strategies are not complete rollback plans by themselves. A previous application version may fail after a destructive schema migration, incompatible event change, or irreversible external side effect.
The strategy should match the application’s state model, infrastructure cost, traffic-routing capability, compatibility requirements, and acceptable blast radius.
Code Example
apiVersion: apps/v1
kind: Deployment
metadata:
name: resumeloop-api
spec:
replicas: 6
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app: resumeloop-api
template:
metadata:
labels:
app: resumeloop-api
version: v2
spec:
containers:
- name: api
image: registry.example.com/api@sha256:<digest>
readinessProbe:
httpGet:
path: /health/ready
port: 3000Common Interview Pitfalls
- Using rolling deployment when old and new versions cannot coexist.
- Calling a deployment canary without measuring the exposed cohort separately.
- Assuming blue-green automatically rolls back database changes.
- Switching all traffic before the new environment passes readiness checks.
- Keeping no capacity for healthy traffic during a rolling update.
- Increasing canary exposure without predefined success and stop criteria.
- Selecting a deployment strategy without considering stateful dependencies.
How should database migrations be designed for zero-downtime releases and safe application rollback?
Direct Answer
Use backward-compatible expand-and-contract changes, separate long data migrations, verify constraints later, and keep old and new application versions compatible during rollout.
Detailed Explanation
Application rollback is only safe when the previous application version can operate against the current database schema and data.
A destructive migration deployed simultaneously with new code can prevent rollback even when the application artifact itself can be restored quickly.
A safer approach is expand and contract.
Expand phase
Migration phase
Transition phase
Contract phase
Large index creation, table rewrites, constraint validation, and data backfills should be evaluated for their locking and resource impact.
Rollback should normally mean redeploying a known-good compatible application version. Automatically attempting a destructive reverse migration during an incident may create additional data loss or downtime.
Code Example
-- Release 1: expand.
ALTER TABLE job_applications
ADD COLUMN archived_at TIMESTAMPTZ;
-- Deploy code that understands both the old status
-- representation and the new archived_at column.
-- Backfill in bounded, resumable batches.
UPDATE job_applications
SET archived_at = updated_at
WHERE id IN (
SELECT id
FROM job_applications
WHERE status = 'archived'
AND archived_at IS NULL
ORDER BY id
LIMIT 1000
);
-- Release 2: switch reads and writes.
-- Release 3: validate and contract only after
-- all application versions use the new model.Common Interview Pitfalls
- Dropping a column in the same release that stops writing to it.
- Adding a required field before old application versions can supply it.
- Running a large backfill in one long production transaction.
- Assuming application rollback also reverses data mutations safely.
- Creating an index or constraint without evaluating production locking.
- Removing old schema while background workers still depend on it.
- Writing to old and new fields without defining conflict behavior.
How would you design a secure, reliable, auditable release pipeline for a production platform?
Direct Answer
Protect source changes, isolate builds, create one signed artifact, use short-lived credentials, enforce environment gates, deploy progressively, verify health, and retain rollback evidence.
Detailed Explanation
A production release pipeline is part of the organization’s software supply chain and should be treated as a privileged system.
A secure and reliable design includes the following controls.
Source governance
Build isolation
Artifact integrity
Credentials and environments
Release safety
Recovery
Auditability
Record:
Pipeline success should not be defined merely as a deployment command returning zero. The release is successful only after the intended version is healthy and serving expected traffic.
Code Example
name: Production release
on:
workflow_dispatch:
concurrency:
group: production-release
cancel-in-progress: false
permissions:
contents: read
jobs:
verify-and-package:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@<pinned-commit>
- run: npm ci
- run: npm run lint
- run: npx tsc --noEmit
- run: npm test -- --runInBand
- run: npm run build
- name: Build immutable artifact
run: ./scripts/build-image.sh "${GITHUB_SHA}"
- name: Scan and sign artifact
run: ./scripts/verify-artifact.sh "${GITHUB_SHA}"
deploy-canary:
needs: verify-and-package
environment: production
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- name: Obtain temporary deployment identity
run: ./scripts/cloud-auth-oidc.sh
- name: Deploy verified digest to canary
run: ./scripts/deploy-canary.sh
- name: Verify canary
run: ./scripts/check-canary.sh
- name: Expand production exposure
run: ./scripts/promote-release.shCommon Interview Pitfalls
- Giving untrusted pull-request workflows production credentials.
- Using permanent administrator credentials for routine deployments.
- Rebuilding the artifact after its verification or approval.
- Allowing pipeline actions to float to uncontrolled mutable versions.
- Declaring success before readiness and production smoke checks complete.
- Running overlapping production deployments without coordination.
- Deploying an incompatible database or event schema without rollback planning.
- Recording the commit but not the actual deployed artifact digest.
- Giving one pipeline identity unrestricted access to every environment.
What is the difference between a container image and a running container, and what isolation does a container runtime provide?
Direct Answer
An image is an immutable packaged filesystem and metadata, while a container is a running process created from that image with isolated namespaces and controlled resources.
Detailed Explanation
An image is a packaged, read-only template containing an application, runtime libraries, filesystem content, and metadata describing how the application should start.
A container is a running instance created from an image. The runtime adds a writable layer, configures isolation, applies resource controls, and starts the container process.
Images are commonly stored in a registry and identified by a repository plus tag or immutable digest. Tags are convenient labels, but they can be changed to reference different content. Digests identify exact image content and are more reliable for production traceability.
Linux containers use operating-system mechanisms such as:
Containers share the host kernel. They do not provide the same isolation boundary as a separate virtual machine with its own kernel.
The process running inside the container is still a normal host process from the kernel’s perspective. If that process runs as root and receives broad capabilities, host mounts, or privileged mode, the security boundary is substantially weakened.
Production images should normally run as a non-root user, include only required runtime files, avoid embedded secrets, use a read-only filesystem where practical, and drop capabilities that the workload does not need.
Code Example
# Pull an image by an immutable digest.
docker pull registry.example.com/resumeloop-api@sha256:<digest>
# Inspect image metadata and layers.
docker image inspect registry.example.com/resumeloop-api@sha256:<digest>
docker image history registry.example.com/resumeloop-api@sha256:<digest>
# Run with reduced privileges and bounded resources.
docker run --rm --read-only --user 10001:10001 --cap-drop ALL --memory 512m --cpus 1 --publish 3000:3000 registry.example.com/resumeloop-api@sha256:<digest>Common Interview Pitfalls
- Using the terms image and container as though they describe the same object.
- Deploying mutable tags without recording the resolved image digest.
- Running application containers as root without a demonstrated requirement.
- Treating containers as equivalent to virtual machines with separate kernels.
- Giving a container privileged mode or broad host mounts unnecessarily.
- Embedding credentials or private keys into image layers.
- Assuming a deleted file is removed from every earlier image layer.
How should Dockerfile layers, build caching, and multi-stage builds be used to create efficient production images?
Direct Answer
Order stable build steps before frequently changing files, minimize build context and layers, and copy only required runtime artifacts into a small final stage.
Detailed Explanation
A Dockerfile defines how an image is assembled. Many instructions create image layers, and build systems can reuse cached results when the instruction and its relevant inputs remain unchanged.
Efficient Dockerfile design generally includes:
.dockerignore to exclude repositories, credentials, local builds, logs, and other unnecessary context.A multi-stage build contains multiple FROM instructions. One stage can contain compilers, development dependencies, and test tooling, while a later runtime stage receives only the built application and required production dependencies.
This provides several benefits:
Layer order influences cache efficiency. If the full source tree is copied before dependency installation, changing one source file may invalidate the dependency layer and force an unnecessary reinstall.
Build secrets must be passed through a secure build-secret mechanism rather than written through ARG, ENV, or copied files that can remain in image metadata or layers.
Code Example
# syntax=docker/dockerfile:1
FROM node:22-alpine AS dependencies
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
FROM dependencies AS build
WORKDIR /app
COPY . .
RUN npm run build
RUN npm prune --omit=dev
FROM node:22-alpine AS runtime
WORKDIR /app
ENV NODE_ENV=production
RUN addgroup --system --gid 10001 appgroup && adduser --system --uid 10001 --ingroup appgroup appuser
COPY --from=build --chown=appuser:appgroup /app/package.json /app/package-lock.json ./
COPY --from=build --chown=appuser:appgroup /app/node_modules ./node_modules
COPY --from=build --chown=appuser:appgroup /app/.next ./.next
COPY --from=build --chown=appuser:appgroup /app/public ./public
USER appuser
EXPOSE 3000
CMD ["npm", "start"]Common Interview Pitfalls
- Copying the complete source tree before installing unchanged dependencies.
- Including source-control history, local secrets, and build output in the build context.
- Installing compilers and development tools in the final runtime image.
- Using build arguments to pass secrets that can remain in image metadata.
- Running the final image as root without a specific requirement.
- Using a mutable base-image tag without an update and verification policy.
- Deleting sensitive files in a later layer after copying them into an earlier layer.
- Installing packages without cleaning temporary package-manager data in the same layer.
How do Kubernetes Pods, Deployments, ReplicaSets, and Services work together?
Direct Answer
Pods run colocated containers, Deployments manage desired replica and rollout state through ReplicaSets, and Services provide stable networking to selected Pods.
Detailed Explanation
A Pod is Kubernetes’ smallest deployable workload unit. It contains one or more tightly coupled containers that share networking and can share volumes.
Containers in the same Pod:
A Pod should usually represent one application instance rather than an entire distributed system.
A ReplicaSet maintains a specified number of matching Pods. Operators rarely create ReplicaSets directly for ordinary stateless applications.
A Deployment manages ReplicaSets and provides declarative updates, scaling, rollout status, and revision history. When the Pod template changes, the Deployment normally creates a new ReplicaSet and gradually replaces the previous Pods according to its update strategy.
Pods are replaceable. Their names and IP addresses can change after rescheduling or rollout.
A Service provides stable network identity and traffic routing to a selected group of Pods. It commonly selects Pods through labels and forwards traffic to ready endpoints.
Common Service types include:
ClusterIP: Internal cluster accessNodePort: Exposes a port on cluster nodesLoadBalancer: Requests external load-balancer integration where supportedExternalName: Provides DNS alias behaviorA Service selector must match the labels on the target Pods. If the selector is wrong or Pods are not ready, the Service may have no usable endpoints even though the Service object exists.
Code Example
apiVersion: apps/v1
kind: Deployment
metadata:
name: resumeloop-api
spec:
replicas: 3
selector:
matchLabels:
app: resumeloop-api
template:
metadata:
labels:
app: resumeloop-api
spec:
containers:
- name: api
image: registry.example.com/api@sha256:<digest>
ports:
- name: http
containerPort: 3000
---
apiVersion: v1
kind: Service
metadata:
name: resumeloop-api
spec:
selector:
app: resumeloop-api
ports:
- name: http
port: 80
targetPort: http
type: ClusterIPCommon Interview Pitfalls
- Deploying standalone Pods for workloads that require automated recovery and rollout.
- Treating Pod names or Pod IP addresses as stable service identities.
- Creating a Service selector that does not match the workload labels.
- Placing unrelated application services inside one Pod.
- Assuming a Deployment sends traffic directly without a Service or other routing layer.
- Changing Deployment selector labels after creation without understanding immutability rules.
- Assuming a running Pod is automatically ready to receive traffic.
How should Kubernetes ConfigMaps, Secrets, health probes, and resource requests and limits be configured?
Direct Answer
Use ConfigMaps for non-sensitive configuration, Secrets for confidential values, distinct probes for startup, readiness, and liveness, and measured requests and limits.
Detailed Explanation
ConfigMaps store non-confidential configuration separately from container images. They can be consumed through environment variables, command arguments, or mounted files.
Secrets are intended for confidential data such as credentials, tokens, or keys. A Kubernetes Secret is not automatically protected merely because it uses the Secret object type. Cluster encryption at rest, RBAC, secret-distribution controls, log redaction, and external secret management may still be required.
Configuration consumed as environment variables is normally fixed for the lifetime of the container. Mounted configuration may update eventually, but the application must support reloading and the rollout behavior should be explicit.
Kubernetes health probes have different purposes:
Liveness should not fail merely because an optional downstream dependency is temporarily unavailable. Otherwise Kubernetes may restart healthy application processes during a dependency incident.
Resource management includes:
CPU limits can cause throttling. Exceeding a memory limit can result in container termination. Requests that are too low can cause poor scheduling and contention, while requests that are too high waste capacity or leave Pods pending.
Resource values should be based on observed workload behavior, performance testing, and expected bursts rather than copied blindly from another service.
Code Example
apiVersion: apps/v1
kind: Deployment
metadata:
name: resumeloop-api
spec:
replicas: 3
selector:
matchLabels:
app: resumeloop-api
template:
metadata:
labels:
app: resumeloop-api
spec:
containers:
- name: api
image: registry.example.com/api@sha256:<digest>
envFrom:
- configMapRef:
name: resumeloop-api-config
- secretRef:
name: resumeloop-api-secrets
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
cpu: "1"
memory: 512Mi
startupProbe:
httpGet:
path: /health/startup
port: 3000
periodSeconds: 2
failureThreshold: 30
readinessProbe:
httpGet:
path: /health/ready
port: 3000
periodSeconds: 5
failureThreshold: 3
livenessProbe:
httpGet:
path: /health/live
port: 3000
periodSeconds: 10
failureThreshold: 3Common Interview Pitfalls
- Storing confidential credentials in a ConfigMap.
- Assuming Kubernetes Secrets are automatically encrypted and inaccessible to cluster users.
- Using the same dependency-heavy endpoint for startup, readiness, and liveness.
- Making liveness fail when an optional external service is temporarily unavailable.
- Setting memory limits below normal peak usage and causing repeated termination.
- Omitting requests and allowing the scheduler to overcommit a node unknowingly.
- Setting requests far above measured usage and leaving workloads unschedulable.
- Expecting environment-variable configuration to update inside an existing process.
How do Kubernetes scheduling, autoscaling, disruption controls, and rollout settings support availability?
Direct Answer
The scheduler places Pods from resource and placement rules, autoscalers adjust capacity, disruption budgets limit voluntary loss, and rollout settings preserve healthy replicas.
Detailed Explanation
Kubernetes scheduling selects a suitable node for each unscheduled Pod.
Scheduling decisions can consider:
A workload can remain Pending when no node satisfies all requirements or has enough allocatable resources.
The Horizontal Pod Autoscaler, or HPA, adjusts the replica count of a scalable workload based on observed metrics. CPU and memory utilization depend on resource requests, while custom or external metrics can represent queue depth, request rate, or another workload signal.
Autoscaling does not create instant capacity. New Pods need scheduling, image download, startup, and readiness time. Scaling policy should account for this delay.
A PodDisruptionBudget, or PDB, limits how many matching Pods may be unavailable during voluntary disruptions such as node draining. It does not prevent every failure and does not guarantee application availability by itself.
Deployment rollout settings include:
maxUnavailable: How many desired replicas may be unavailable during rolloutmaxSurge: How many additional replicas may temporarily existAvailability also requires spreading replicas across appropriate nodes or zones. Running three replicas on one node does not protect the workload from that node’s failure.
Anti-affinity and topology-spread rules should be balanced carefully. Rules that are too strict can prevent scheduling when the cluster lacks sufficient topology or capacity.
Code Example
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: resumeloop-api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: resumeloop-api
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: resumeloop-api
spec:
minAvailable: 2
selector:
matchLabels:
app: resumeloop-api
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: resumeloop-api
spec:
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
template:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: resumeloop-apiCommon Interview Pitfalls
- Configuring CPU-based autoscaling without defining meaningful CPU requests.
- Assuming the Horizontal Pod Autoscaler adds node capacity automatically.
- Using a PodDisruptionBudget as protection against every involuntary failure.
- Running several replicas on the same node without considering failure domains.
- Creating required anti-affinity rules that make the workload impossible to schedule.
- Setting maxUnavailable too high for a small critical workload.
- Scaling on a delayed metric without considering queue age or startup time.
- Expecting autoscaling to compensate for an inefficient or failing application instantly.
How would you design and troubleshoot a secure, reliable production workload running on Kubernetes?
Direct Answer
Use immutable images, least privilege, measured resources, health and disruption controls, safe rollouts, observability, and a layered troubleshooting process from workload to cluster.
Detailed Explanation
A production Kubernetes workload requires controls across image supply, identity, scheduling, networking, runtime security, availability, configuration, and observability.
Image and runtime security
Identity and access
Availability and lifecycle
SIGTERM and complete graceful shutdown within the termination period.Resources and scaling
Configuration and storage
Networking
Troubleshooting process
1. Determine impact, timeline, namespace, workload, and release version.
2. Inspect Deployment rollout status, events, and desired versus available replicas.
3. Inspect Pods, scheduling status, restarts, termination reasons, probes, and resource use.
4. Review current and previous container logs.
5. Check Service selectors and endpoint membership.
6. Test DNS and network paths from an appropriate debugging container.
7. Review node capacity, pressure, kubelet state, and cluster events.
8. Compare healthy and unhealthy Pods, nodes, zones, and image versions.
9. Mitigate through rollback, traffic removal, or scaling while preserving evidence.
Common states provide clues:
Pending: Scheduling, resource, volume, or placement issueImagePullBackOff: Registry, image name, credentials, or network issueCrashLoopBackOff: Repeated process failure after startOOMKilled: Memory limit or node-memory pressureA restart may restore service temporarily, but the investigation should identify the reason the workload became unhealthy.
Code Example
# Deployment and rollout evidence.
kubectl --namespace production rollout status deployment/resumeloop-api
kubectl --namespace production describe deployment resumeloop-api
# Pod state, placement, events, and restarts.
kubectl --namespace production get pods --selector app=resumeloop-api --output wide
kubectl --namespace production describe pod <pod-name>
# Current and previous container logs.
kubectl --namespace production logs <pod-name> --container api
kubectl --namespace production logs <pod-name> --container api --previous
# Confirm Service selectors and endpoints.
kubectl --namespace production get service resumeloop-api --output yaml
kubectl --namespace production get endpointslices --selector kubernetes.io/service-name=resumeloop-api
# Inspect resource use where metrics are available.
kubectl --namespace production top pod --selector app=resumeloop-api
# Add an ephemeral debugging container when permitted.
kubectl --namespace production debug --interactive <pod-name> --image=busybox:1.36Common Interview Pitfalls
- Restarting Pods before capturing events, previous logs, and termination reasons.
- Troubleshooting only the container while ignoring Service selectors and endpoints.
- Giving workloads cluster-admin access for operational convenience.
- Using liveness checks that restart Pods during temporary dependency failures.
- Setting no resource requests and then relying on predictable scheduling.
- Running all replicas in one failure domain.
- Using mutable image tags that make unhealthy releases difficult to identify.
- Treating CrashLoopBackOff as a root cause rather than a retry condition.
- Assuming NetworkPolicy rules are enforced without confirming network-plugin support.
- Using privileged debugging containers without authorization and data-handling controls.
What are cloud Regions and Availability Zones, and how does the shared-responsibility model affect system design?
Direct Answer
Regions are separate geographic areas, Availability Zones are isolated locations within a Region, and customers remain responsible for securing and operating their workloads.
Detailed Explanation
A Region is a separate geographic area in which a cloud provider operates infrastructure and services.
An Availability Zone, or AZ, is an isolated infrastructure location within a Region. A Region normally contains multiple Availability Zones connected through provider-managed networking.
Availability Zones are designed to reduce the chance that one infrastructure failure affects every zone in the Region. They can help protect against failures involving power, cooling, networking, or other localized infrastructure.
Deploying several replicas does not create zone resilience if all replicas run in one Availability Zone. A highly available regional design may distribute:
A multi-zone architecture does not protect against every Region-wide incident, application defect, configuration error, credential compromise, or destructive data operation.
The shared-responsibility model separates provider and customer responsibilities.
The cloud provider is generally responsible for security of the cloud, including physical facilities, foundational hardware, and managed infrastructure layers.
The customer remains responsible for security in the cloud, including areas such as:
The exact boundary changes according to the service model. A managed database transfers more infrastructure operations to the provider than a database installed directly on a virtual machine, but the customer still owns data access, schema, credentials, backup configuration, and workload design.
Code Example
# Conceptual regional architecture
Region: us-east-1
Availability Zone A:
- Application replica
- Private subnet
- Database standby
Availability Zone B:
- Application replica
- Private subnet
- Database primary
Availability Zone C:
- Application replica
- Private subnet
Regional services:
- Load balancer
- Object storage
- Monitoring
- Managed DNS
Customer responsibilities:
- IAM policies
- Network rules
- Encryption settings
- Backup policy
- Application security
- Capacity and failover testingCommon Interview Pitfalls
- Running several replicas in one Availability Zone and calling the design highly available.
- Assuming a multi-zone deployment protects against every regional or application failure.
- Treating managed cloud services as though the provider owns all security configuration.
- Failing to reserve enough remaining capacity for an Availability Zone failure.
- Deploying across zones without testing failover and dependency behavior.
- Ignoring data backup because the database already has redundant infrastructure.
What is Infrastructure as Code, and what does declarative infrastructure configuration mean?
Direct Answer
Infrastructure as Code represents infrastructure in versioned files, while declarative configuration describes the desired end state instead of every procedural step.
Detailed Explanation
Infrastructure as Code, or IaC, is the practice of managing infrastructure through machine-readable configuration rather than relying primarily on manual console changes.
IaC can define resources such as:
A declarative configuration describes the desired outcome. The IaC tool compares that desired configuration with known infrastructure state and determines which operations are required.
For example, configuration may declare that three application instances, one load balancer, and two private subnets should exist. The tool determines whether it must create, update, replace, or remove resources to reach that state.
Benefits include:
IaC does not automatically make infrastructure safe. A reviewed configuration can still contain an overly broad IAM policy, public database access, destructive lifecycle behavior, or an incorrect network route.
Infrastructure changes should normally pass formatting, validation, security-policy checks, planning, review, controlled application, and post-change verification.
Manual emergency changes may sometimes be necessary. They should be recorded and reconciled into code afterward so that the declared configuration and real infrastructure do not remain inconsistent.
Code Example
terraform {
required_version = ">= 1.8.0"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
}
resource "aws_s3_bucket" "application_artifacts" {
bucket = "resumeloop-production-artifacts"
tags = {
Environment = "production"
ManagedBy = "terraform"
}
}
resource "aws_s3_bucket_versioning" "artifacts" {
bucket = aws_s3_bucket.application_artifacts.id
versioning_configuration {
status = "Enabled"
}
}Common Interview Pitfalls
- Treating Infrastructure as Code as a generated backup of manual console work.
- Applying infrastructure changes without reviewing the execution plan.
- Assuming declarative configuration prevents destructive changes automatically.
- Making emergency manual changes without reconciling them into source control.
- Copying complete environment configurations instead of creating reusable boundaries.
- Storing credentials directly inside infrastructure configuration files.
- Using infrastructure code without validation or policy checks in continuous integration.
How do Terraform state, plans, modules, locking, and drift work together?
Direct Answer
State maps configuration to real resources, plans preview proposed operations, modules package reusable infrastructure, locking prevents concurrent writes, and drift records external changes.
Detailed Explanation
Terraform needs state to map configured resource addresses to real infrastructure objects and to retain metadata required for planning changes.
Terraform state can contain:
State must therefore be protected as operationally sensitive data.
For team environments, state should normally use a remote backend that supports appropriate access control, encryption, versioning, and state locking.
State locking prevents two Terraform operations from modifying the same state concurrently. Without locking, simultaneous applies can overwrite state or create conflicting infrastructure changes.
Terraform plan refreshes relevant information, compares configuration and state with infrastructure, and previews actions such as:
A reviewed saved plan can be applied so that the approved operations, rather than a newly calculated plan, are executed. However, infrastructure changes between planning and application may still affect the result.
A module is a reusable collection of Terraform resources with defined inputs and outputs. Good modules represent coherent infrastructure capabilities rather than wrapping every individual provider resource without adding policy or abstraction value.
Drift occurs when real infrastructure changes outside the expected Terraform workflow. A future plan may detect the difference and propose either restoring the declared configuration or accepting the external change through an intentional configuration update or import.
Operators should not edit state files manually. Supported state commands, import blocks, moved blocks, or backend recovery procedures should be used carefully, with backups and exclusive access.
Code Example
terraform {
backend "s3" {
bucket = "resumeloop-terraform-state"
key = "production/network/terraform.tfstate"
region = "us-east-1"
encrypt = true
dynamodb_table = "terraform-state-locks"
}
}
module "application_network" {
source = "../../modules/application-network"
environment = "production"
vpc_cidr = "10.20.0.0/16"
availability_zones = [
"us-east-1a",
"us-east-1b",
"us-east-1c"
]
}
# Pipeline sequence:
# terraform fmt -check
# terraform init
# terraform validate
# terraform plan -out=tfplan
# review and approve tfplan
# terraform apply tfplanCommon Interview Pitfalls
- Committing Terraform state files to a public or broadly accessible repository.
- Running concurrent applies against one state without locking.
- Applying a newly generated plan after reviewing a different saved plan.
- Editing Terraform state JSON manually.
- Using one state file for unrelated infrastructure with different ownership.
- Ignoring drift until an unrelated deployment attempts destructive reconciliation.
- Creating modules that expose every provider argument without adding a meaningful boundary.
- Moving resource addresses without preserving the state association.
How should IAM, least privilege, roles, temporary credentials, and workload identity be designed?
Direct Answer
Grant only required actions on required resources, prefer assumable roles and temporary credentials, constrain trust policies, and give each workload a distinct identity.
Detailed Explanation
Identity and Access Management controls which principals may perform actions on which resources and under what conditions.
An authorization decision may combine:
Least privilege means granting only the permissions required for the task. Policies should restrict:
An IAM role is an identity that can be assumed by an authorized principal. Roles normally issue temporary credentials rather than permanent access keys.
Temporary credentials reduce the lifetime of exposed credentials and can carry session context. They should be preferred for:
A role has two important policy dimensions:
Both must be restricted. A narrowly permissioned role with an overly broad trust policy may still be assumable by an unintended principal.
Each workload should receive its own identity rather than sharing one broad application credential. For Kubernetes, cloud identity can be associated with a specific service account so that only Pods using that account receive the permissions.
Policies should be reviewed using access logs, last-accessed information, policy simulator, and automated analysis. Temporary incident permissions should expire or be removed after the incident.
Code Example
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ReadOnlyRequiredSecret",
"Effect": "Allow",
"Action": [
"secretsmanager:GetSecretValue"
],
"Resource": [
"arn:aws:secretsmanager:us-east-1:123456789012:secret:production/resumeloop/database-*"
]
},
{
"Sid": "DecryptThroughSecretsManager",
"Effect": "Allow",
"Action": [
"kms:Decrypt"
],
"Resource": [
"arn:aws:kms:us-east-1:123456789012:key/example-key-id"
],
"Condition": {
"StringEquals": {
"kms:ViaService": "secretsmanager.us-east-1.amazonaws.com"
}
}
}
]
}Common Interview Pitfalls
- Granting wildcard actions and resources for deployment convenience.
- Using one long-lived access key across several workloads.
- Restricting permission policies while leaving the role trust policy overly broad.
- Giving every Kubernetes Pod the node or cluster-wide cloud permissions.
- Keeping temporary emergency permissions after the incident ends.
- Assuming a role name alone proves which workload assumed it.
- Granting read access to every secret when one named secret is required.
- Creating policies without reviewing actual access patterns later.
How should secrets, encryption keys, envelope encryption, access, and rotation be managed?
Direct Answer
Store secrets in managed systems, encrypt them with controlled keys, grant narrow access, audit retrieval, rotate credentials safely, and keep plaintext out of code and logs.
Detailed Explanation
A secret is confidential information whose disclosure could grant access or create material risk. Examples include:
Secrets should not be stored in source code, container images, Terraform variables committed to a repository, shared documents, or ordinary application logs.
A managed secrets system can provide:
Envelope encryption protects data using a data-encryption key, while a higher-level key-encryption key protects that data key. This allows the key-management service to control access to master key material without directly processing or storing every large plaintext payload.
Access to a secret may require both:
Encryption must be considered both at rest and in transit. Encrypting a stored secret does not make it safe to send over an unprotected connection or print in logs.
Secret rotation replaces the credential in both the secrets store and the target system that accepts the credential. A safe rotation design may support an overlap period so that active application instances can transition without downtime.
Encryption-key rotation and secret-value rotation solve different problems. Rotating a KMS key changes the cryptographic key material or version used to protect data. Rotating a database password changes the credential itself.
Applications should retrieve secrets through workload identity and should cache them only according to defined security, availability, and rotation requirements. A secrets-service outage policy must be explicit: an application may continue temporarily with an already loaded secret, but it should not retain plaintext indefinitely or bypass access controls.
Code Example
async function loadDatabaseCredentials(): Promise<{
username: string;
password: string;
}> {
const response = await secretsManager.getSecretValue({
secretId:
'production/resumeloop/database'
});
if (!response.secretString) {
throw new Error(
'Database secret has no string value'
);
}
const parsed: unknown =
JSON.parse(response.secretString);
if (
typeof parsed !== 'object' ||
parsed === null ||
!('username' in parsed) ||
!('password' in parsed) ||
typeof parsed.username !== 'string' ||
typeof parsed.password !== 'string'
) {
throw new Error(
'Database secret has an invalid structure'
);
}
return {
username: parsed.username,
password: parsed.password
};
}
// Never log the returned credentials.Common Interview Pitfalls
- Committing secrets to source control and deleting them only in a later commit.
- Embedding secrets in container-image layers or build arguments.
- Granting every workload permission to retrieve every production secret.
- Rotating a stored secret without updating the target database or service.
- Logging complete configuration objects containing secret values.
- Confusing encryption-key rotation with credential rotation.
- Allowing applications to use permanent credentials when workload identity is available.
- Changing a secret without considering already running application instances.
How would you design secure, repeatable cloud infrastructure across development, staging, and production environments?
Direct Answer
Separate trust boundaries, reuse versioned modules, isolate state and credentials, enforce policy in pipelines, promote reviewed plans, and verify security and resilience continuously.
Detailed Explanation
Multi-environment cloud infrastructure should balance consistency with strong isolation and environment-specific risk controls.
Account and trust separation
Production should normally have a distinct trust boundary from development. Depending on the provider, this may use separate accounts, subscriptions, or projects.
Separation reduces the risk that:
Infrastructure modules
Reusable, versioned modules can define approved patterns for:
Modules should provide secure defaults while allowing intentional, reviewed differences. Production and staging should not use unrelated copies of infrastructure code that drift independently.
State isolation
Each environment and major ownership boundary should have isolated Terraform state, locking, encryption, backups, and narrow access permissions. Development users should not automatically have read access to production state.
Identity
Pipeline controls
A safe workflow includes:
1. Format and validate configuration.
2. Run static security and policy checks.
3. Generate a plan using read access to the target environment.
4. Review destructive, security-sensitive, and high-cost changes.
5. Apply a saved approved plan using a separate deployment role.
6. Run post-deployment checks.
7. Record the source version, module versions, plan, approver, and outcome.
Security controls
Resilience
Production architecture should reflect explicit recovery objectives. Multi-zone deployment, backups, cross-Region copies, and disaster-recovery environments should be selected according to business requirements rather than copied into every environment automatically.
Testing and promotion
Infrastructure modules should be tested in disposable or lower-risk environments. Module versions should be promoted intentionally rather than allowing production to consume an unreviewed moving source.
Environment parity means using the same architectural patterns and release processes where useful. It does not require development to have the same scale, cost, data, or access as production.
Code Example
infrastructure/
├── modules/
│ ├── application-network/
│ ├── kubernetes-platform/
│ ├── managed-database/
│ └── workload-identity/
└── environments/
├── development/
│ ├── backend.tf
│ ├── main.tf
│ └── variables.tf
├── staging/
│ ├── backend.tf
│ ├── main.tf
│ └── variables.tf
└── production/
├── backend.tf
├── main.tf
└── variables.tf
# Example pipeline sequence:
terraform fmt -check
terraform init
terraform validate
terraform plan -out=tfplan
security-policy-check tfplan
manual-production-approval tfplan
terraform apply tfplan
post-deployment-verificationCommon Interview Pitfalls
- Keeping development and production in one unrestricted administrative trust boundary.
- Using one Terraform state file for every environment and platform component.
- Copying modules into each environment and allowing them to diverge.
- Giving the planning job the same broad write permissions as the apply job.
- Allowing production to consume an unpinned module branch.
- Assuming environment parity requires production-scale development infrastructure.
- Sharing production secrets or customer data with lower environments.
- Applying infrastructure plans without preserving the reviewed plan and audit record.
- Using multi-zone infrastructure without testing failure and recovery behavior.
- Treating drift detection as a substitute for preventing unauthorized manual changes.
What is observability, and how do logs, metrics, and distributed traces complement each other when troubleshooting production systems?
Direct Answer
Metrics reveal system trends, logs record discrete events, and traces follow requests across services; together they help engineers explain production behavior.
Detailed Explanation
Observability is the ability to understand a system’s internal behavior from the signals it produces.
Monitoring usually asks predefined questions such as whether CPU usage, request latency, or error rate exceeded a threshold. Observability supports deeper investigation when engineers do not know the exact question in advance.
Metrics
Metrics are numerical measurements collected over time.
Examples include:
Metrics are efficient for dashboards, alerting, capacity analysis, and long-term trends because they aggregate many events into compact time-series data.
However, a metric showing that error rate increased does not necessarily explain which requests failed or why.
Logs
Logs record discrete events and associated context.
Useful structured log fields may include:
Structured logs are generally easier to search and aggregate than free-form text.
Logs should avoid secrets and unnecessary sensitive information.
Distributed traces
A trace represents the path of a request through a distributed system.
A trace contains spans representing individual operations such as:
Traces help identify which dependency or operation contributes to latency or failure.
Correlation
The signals become significantly more useful when they can be correlated.
For example:
1. A latency dashboard shows P99 latency increasing.
2. A trace identifies a slow database span.
3. Correlated logs show connection-pool exhaustion.
4. Infrastructure metrics show database connection saturation.
This provides much stronger diagnostic evidence than any one signal alone.
OpenTelemetry defines telemetry signals including traces, metrics, and logs and provides common instrumentation concepts for collecting and exporting them.
Telemetry design
Instrumentation should answer operational questions while controlling cardinality, storage, and cost.
For example, a metric label containing a unique user ID can create an extremely high number of time series and should generally be avoided.
The goal is not to collect every possible event indefinitely. The goal is to collect enough high-quality telemetry to understand service health, diagnose incidents, and improve reliability.
Code Example
from opentelemetry import trace
tracer = trace.get_tracer(
"job-service"
)
def get_job(job_id: str):
with tracer.start_as_current_span(
"get_job"
) as span:
span.set_attribute(
"job.id",
job_id
)
return repository.find(job_id)Common Interview Pitfalls
- Treating monitoring and observability as exactly the same concept.
- Collecting logs without correlation identifiers across distributed services.
- Using metrics with extremely high-cardinality labels such as user IDs.
- Logging secrets or unnecessary sensitive data.
- Collecting traces without preserving useful service and deployment context.
- Assuming one telemetry signal can answer every troubleshooting question.
- Keeping unlimited telemetry without considering storage and operational cost.
- Instrumenting infrastructure while ignoring user-facing application behavior.
What are SLIs, SLOs, SLAs, and error budgets, and how do they help teams make reliability decisions?
Direct Answer
An SLI measures service behavior, an SLO defines its target, an SLA is a service commitment, and an error budget quantifies allowable unreliability.
Detailed Explanation
Reliability engineering works best when teams define measurable expectations rather than trying to make every system perfectly available.
Service Level Indicator — SLI
An SLI is a quantitative measurement of service behavior.
Examples include:
An SLI should represent something meaningful to users or downstream consumers.
A server being reachable is not necessarily a useful availability indicator if requests return incorrect data.
Service Level Objective — SLO
An SLO specifies the desired level of an SLI over a defined measurement window.
Example:
99.9% of eligible API requests succeed during a rolling 30-day window.
SLOs should be difficult enough to protect users but realistic enough to leave room for maintenance, incidents, and product development.
Google SRE guidance recommends defining SLOs from the user perspective rather than relying only on internal component health.
Service Level Agreement — SLA
An SLA is a business or contractual commitment regarding service performance.
It may specify consequences when the service fails to meet defined commitments.
An internal SLO is usually designed to be stricter than the corresponding external SLA so the team has operational margin before contractual commitments are violated.
Error budget
An error budget is the amount of unreliability permitted by an SLO.
For a 99.9% success SLO:
error budget = 100% - 99.9% = 0.1%
The error budget can be consumed by failures, excessive latency, or whichever condition defines the SLI.
Google SRE describes error budgets as a mechanism for balancing reliability with the pace of change. When significant budget remains, teams can accept normal deployment risk. When the budget is exhausted, reliability work may take priority over risky feature releases.
Why not target 100%?
Attempting perfect reliability can create excessive cost and prevent useful changes while still being impossible to guarantee across dependencies and networks.
The appropriate target depends on user expectations, architecture, downstream dependencies, cost, and consequences of failure.
Multiple SLOs
A service may require several objectives, such as:
Meeting one does not compensate automatically for violating another.
SLOs should therefore reflect the most important dimensions of the actual user experience.
Code Example
from dataclasses import dataclass
@dataclass
class ReliabilityWindow:
eligible_requests: int
successful_requests: int
def success_ratio(
window: ReliabilityWindow,
) -> float:
if window.eligible_requests == 0:
return 1.0
return (
window.successful_requests
/ window.eligible_requests
)
def remaining_error_budget(
actual_success_ratio: float,
target_success_ratio: float,
) -> float:
allowed_failure = (
1.0 - target_success_ratio
)
actual_failure = (
1.0 - actual_success_ratio
)
return allowed_failure - actual_failureCommon Interview Pitfalls
- Defining an SLO using infrastructure metrics that users do not experience.
- Treating an SLO and an SLA as interchangeable concepts.
- Setting every service objective to 100 percent without evaluating cost.
- Defining availability while ignoring latency and correctness.
- Creating an error budget but never using it in release decisions.
- Changing SLO definitions whenever the service performs poorly.
- Counting requests that should be excluded from the SLI denominator.
- Using a single SLO for services with several important user-facing behaviors.
How should a DevOps engineer design production alerting so that on-call engineers receive actionable signals without excessive alert fatigue?
Direct Answer
Alert on meaningful user impact or imminent risk, define ownership and response actions, deduplicate noise, and continuously tune alerts using incident evidence.
Detailed Explanation
The purpose of production alerting is to cause an appropriate human or automated response when action is necessary.
An alert that does not require any action is usually better represented as a dashboard, log entry, or diagnostic signal.
Google SRE distinguishes monitoring outputs based on urgency: conditions requiring immediate human action should page, less urgent work should become tracked work, and diagnostic information can remain available for later analysis.
AWS similarly recommends that alerts have defined response processes, clear ownership, and escalation paths.
Alert from symptoms when possible
A symptom describes user-visible failure, such as:
A cause describes an internal condition such as high CPU usage.
High CPU may be completely harmless when the service is healthy, while moderate CPU can still accompany a dependency outage.
Page primarily on meaningful symptoms or conditions strongly predictive of imminent failure.
Define actionability
Every paging alert should answer:
Alert severity
Not every threshold violation requires waking an engineer.
Possible destinations include:
Severity should reflect urgency and impact rather than the name of the metric.
Avoid static threshold mistakes
A fixed CPU threshold such as 80% may generate noise during healthy traffic growth.
Better signals can include:
Error-budget burn alerts
Burn-rate alerting measures how quickly the service is consuming its error budget.
A fast burn can detect severe outages quickly, while slower windows can identify persistent degradation without paging on brief noise.
Deduplication and grouping
A single dependency failure can trigger alerts in dozens of downstream services.
Group related alerts and identify the highest-value signal so responders are not flooded with duplicate pages.
Alert fatigue
Repeated false or low-value alerts cause engineers to stop trusting the system.
Track metrics such as:
Review alerts after incidents and remove or redesign signals that did not contribute to detection or recovery.
Code Example
groups:
- name: api-reliability
rules:
- alert: HighRequestFailureRate
expr: |
(
sum(rate(http_requests_total{
status=~"5.."
}[5m]))
/
sum(rate(http_requests_total[5m]))
) > 0.05
for: 10m
labels:
severity: page
annotations:
summary: >
API error rate above 5%
runbook: >
https://internal/runbooks/api-errorsCommon Interview Pitfalls
- Paging on every infrastructure threshold regardless of user impact.
- Creating alerts with no owner or documented response procedure.
- Using the same urgency level for every operational event.
- Allowing one dependency failure to produce dozens of duplicate pages.
- Keeping noisy alerts because they might become useful someday.
- Alerting on instantaneous spikes without considering persistence.
- Using email as the primary mechanism for urgent operational events.
- Measuring alert quantity without measuring whether alerts are actionable.
How should a DevOps team manage a major production incident from detection through recovery and post-incident learning?
Direct Answer
Establish incident command, assess impact, mitigate quickly, communicate clearly, preserve evidence, verify recovery, and convert postmortem findings into owned improvements.
Detailed Explanation
Incident management should optimize for restoring safe service while maintaining clear coordination and preserving enough evidence to understand what happened afterward.
AWS recommends documented incident-management processes, tested response plans, simulations, and mechanisms for learning from incidents.
1. Detect and acknowledge
The incident may be detected through:
The on-call engineer should acknowledge the incident and quickly determine whether it requires escalation.
2. Determine severity and impact
Assess:
Severity should be based primarily on impact rather than technical complexity.
3. Assign roles
For significant incidents, separate responsibilities can include:
The incident commander coordinates rather than personally debugging every subsystem.
4. Mitigate before perfect diagnosis
When a safe mitigation exists, restoring service may be more important than immediately identifying the deepest root cause.
Possible actions include:
Avoid making multiple uncontrolled changes simultaneously because this makes recovery and diagnosis harder.
5. Communicate
Maintain a shared timeline and provide updates appropriate to stakeholders.
Useful updates state:
Avoid unsupported speculation.
6. Verify recovery
Do not close an incident because one dashboard turned green.
Confirm:
7. Preserve evidence
Capture:
8. Conduct a post-incident review
The review should identify why the system allowed the incident and how detection, mitigation, and recovery can improve.
Useful sections include:
Google SRE postmortem practices emphasize learning from incidents and connecting findings to reliability improvements rather than treating a postmortem as an exercise in individual blame.
9. Track actions to completion
A postmortem without implemented actions has limited value.
Each action should have:
Repeated incidents should trigger examination of whether previous corrective actions were ineffective or never completed.
Code Example
from dataclasses import dataclass
from enum import Enum
class Severity(str, Enum):
SEV1 = "sev1"
SEV2 = "sev2"
SEV3 = "sev3"
@dataclass
class Incident:
severity: Severity
user_impact: str
commander: str
mitigation: str | None = None
resolved: bool = False
def can_close_incident(
service_healthy: bool,
backlog_recovered: bool,
data_integrity_verified: bool,
) -> bool:
return all([
service_healthy,
backlog_recovered,
data_integrity_verified,
])Common Interview Pitfalls
- Having every engineer make changes independently during a major incident.
- Focusing on perfect root-cause diagnosis before applying a safe mitigation.
- Declaring recovery from one infrastructure metric without checking user impact.
- Failing to maintain a shared incident timeline.
- Publishing speculative incident updates as established facts.
- Writing postmortems that focus primarily on individual blame.
- Creating corrective actions without owners or deadlines.
- Closing incidents while significant backlogs or integrity issues remain.
How should DevOps engineers test resilience, plan capacity, and design recovery so systems continue operating through component and regional failures?
Direct Answer
Design for failure domains, maintain capacity headroom, automate safe failover, define recovery objectives, and test failure and disaster scenarios regularly.
Detailed Explanation
Reliability requires more than monitoring failures after they happen. Systems should be designed and tested so expected failures can occur without causing unacceptable user impact.
AWS reliability guidance recommends monitoring workload components, failing over to healthy resources, automating recovery where appropriate, testing resilience, performing post-incident analysis, and conducting game days.
Failure domains
Identify which failures can occur independently.
Examples include:
Redundancy within one failure domain may not protect against failure of that domain.
Capacity headroom
A service should normally retain enough capacity to survive expected traffic variation and infrastructure failures.
For example, if an application runs across three zones but cannot handle normal traffic after losing one zone, the architecture is not truly tolerant of one-zone failure.
Capacity planning should include:
Graceful degradation
Not every feature must remain fully functional during severe failure.
Examples include:
Prioritize critical user journeys.
Recovery objectives
Disaster-recovery planning commonly defines:
The architecture and backup strategy should be designed around these objectives.
Backups are not recovery tests
A successful backup job does not prove restoration works.
Test:
Chaos and failure testing
Resilience testing deliberately introduces controlled failures to verify assumptions.
Possible scenarios include:
Experiments should define:
Do not begin with uncontrolled production destruction.
Game days
Game days exercise technical systems and human response together.
They can reveal weaknesses in:
Reliability improves when failure assumptions are repeatedly tested rather than merely documented.
Code Example
from dataclasses import dataclass
@dataclass
class RecoveryRequirements:
recovery_time_minutes: int
recovery_point_minutes: int
survives_zone_failure: bool
restore_tested: bool
def recovery_ready(
requirements: RecoveryRequirements,
) -> bool:
return all([
requirements.recovery_time_minutes <= 30,
requirements.recovery_point_minutes <= 5,
requirements.survives_zone_failure,
requirements.restore_tested,
])Common Interview Pitfalls
- Calling a system highly available because it has several replicas in one failure domain.
- Running production at full capacity with no headroom for instance or zone failures.
- Assuming successful backups prove that disaster recovery works.
- Designing failover without testing dependency limits in the failover region.
- Requiring every noncritical feature to remain fully available during severe failures.
- Running chaos experiments without defining blast radius and abort conditions.
- Writing disaster recovery plans that are never exercised.
- Ignoring additional capacity required during deployments and failovers.
How would you design an enterprise observability, SRE, and incident-response operating model for hundreds of services owned by many engineering teams?
Direct Answer
Standardize telemetry and reliability primitives, assign service ownership, define tiered SLOs, automate safe operations, and preserve team accountability for production outcomes.
Detailed Explanation
An enterprise reliability model should make healthy operational practices repeatable across teams without turning one central SRE or platform group into the operator of every application.
1. Establish service ownership
Every production service should have metadata describing:
An unowned service creates operational risk because nobody is clearly accountable during an incident.
2. Define service tiers
Not every system requires the same reliability controls.
For example:
Higher tiers may require stricter SLOs, redundancy, disaster recovery, on-call coverage, and change controls.
3. Standardize telemetry
Provide shared instrumentation conventions using common telemetry standards.
Standard attributes might include:
OpenTelemetry supports common collection and export of traces, metrics, and logs, which can reduce application-specific instrumentation fragmentation.
4. Provide a paved observability path
Shared tooling can provide:
Teams should not need to assemble an entirely new observability stack for every service.
5. Define user-centered SLOs
Each critical service should define indicators around meaningful user behavior rather than infrastructure health alone.
Examples include:
Google SRE guidance emphasizes user-centered SLOs and error budgets as a shared mechanism for reliability decisions.
6. Connect error budgets to change policy
Define what happens when reliability degrades.
Possible policies include:
The policy should be defined before an outage rather than negotiated during one.
7. Standardize alerting
Require paging alerts to have:
AWS recommends having defined processes for alerts and reducing alert overload so responders can act consistently.
8. Build incident-management capability
Provide reusable tooling and procedures for:
Critical incidents should not depend on teams inventing coordination procedures while production is failing.
9. Create learning loops
Significant incidents should produce post-incident analysis and tracked corrective work.
Aggregate findings can expose organization-wide patterns such as:
This allows the platform team to solve recurring classes of failure once rather than requiring every product team to rediscover them.
10. Control observability cost
Telemetry can become a significant platform expense.
Manage:
Cost controls should preserve diagnostically valuable data rather than imposing arbitrary global limits.
11. Automate safe operations
AWS recommends safely automating operational responses where possible while using guardrails such as rate controls, thresholds, and approvals.
Examples include:
Automation should stop or escalate when assumptions are violated.
12. Measure reliability program health
Useful organization-level measures include:
The objective is not merely more monitoring. It is a system in which teams can detect, understand, mitigate, learn from, and ultimately prevent production failures.
Code Example
from dataclasses import dataclass
from enum import Enum
class ServiceTier(str, Enum):
CRITICAL = "tier-0"
USER_FACING = "tier-1"
STANDARD = "tier-2"
INTERNAL = "tier-3"
@dataclass
class ServiceReliabilityMetadata:
service_name: str
owner: str
tier: ServiceTier
slo_defined: bool
on_call_defined: bool
runbook_defined: bool
dashboard_defined: bool
recovery_tested: bool
def production_ready(
service: ServiceReliabilityMetadata,
) -> bool:
base_requirements = all([
bool(service.owner),
service.dashboard_defined,
service.runbook_defined,
])
if not base_requirements:
return False
if service.tier in {
ServiceTier.CRITICAL,
ServiceTier.USER_FACING,
}:
return all([
service.slo_defined,
service.on_call_defined,
service.recovery_tested,
])
return TrueCommon Interview Pitfalls
- Making a central SRE team operationally responsible for every product service.
- Applying identical availability requirements to critical and low-impact workloads.
- Allowing teams to invent incompatible telemetry naming conventions.
- Creating dashboards without defining who owns the underlying services.
- Measuring infrastructure health while omitting user-centered SLOs.
- Defining error budgets without connecting them to release policy.
- Centralizing every incident decision and creating an operational bottleneck.
- Collecting unlimited telemetry without cardinality and retention controls.
- Automating remediation without rate limits, abort conditions, or escalation paths.
- Tracking incident counts without tracking repeated failures and corrective-action completion.
Want to tailer your resume for DevOps Engineer roles?
Import your resume, scan it for critical DevOps Engineer keywords, and compare it against ATS standards instantly.
