24 Interview Questions for Hiring a DevOps Engineer

By Lee Flanagan

27th Jul. 2026  |  Last Updated: 27th Jul. 2026

Why DevOps Hiring Is Broken (And How to Fix It)

DevOps as a job title has become almost meaningless. It stretches from the sysadmin who wrote some Terraform scripts three years ago to the platform engineer who owns a multi-region Kubernetes substrate and mentors five engineers. You’re being asked to calibrate across a three-fold salary range on essentially shared keywords. The result? Hiring managers interview CV keyword-checkers instead of builders, and recruiters spend cycles screening for “five years of Terraform” instead of “someone who can own production reliability and coach the team through chaos.”

Here’s what we see go wrong. First: tool-list matching on CVs instead of probing what the candidate actually built with those tools. A DevOps engineer who spent eighteen months owning a single AWS account and migrating it to multi-region is worth far more than someone who’s dabbled in Kubernetes, Terraform, and Datadog at three different companies. Second: hiring for “keeps the lights on” when your team actually needs someone to rebuild your CI/CD pipeline from scratch, or vice versa. You’ll burn out a platform engineer doing on-call escalations, and you’ll waste a shift operator’s time asking them to design a self-service deployment model. Third, and most damaging: skipping the incident-response signal. A fifteen-minute conversation about how someone commanded a production incident reveals more about their thinking, judgment, and communication than two hours of architecture questions. It’s the single highest-leverage signal for this role.

This page covers 24 interview questions across 6 dimensions: CI/CD & Pipeline Engineering, Incident Command & Reliability, Infrastructure-as-Code & Automation, Cloud Cost & Security Posture, Platform-as-Product Mindset and Cross-team Collaboration & Developer Experience. We’ve split them into sections: how you spot their track record, how you probe their thinking under pressure, and how you calibrate whether they’ll thrive in your on-call culture.


Core Competencies for a Strong DevOps Engineer Hire

The best DevOps hires combine technical depth with ownership mentality. You’re looking for: Technical Skill (deep cloud or Kubernetes fluency, not breadth across everything), Ownership (they own outcomes, not tickets), Attention to Detail (observability and runbooks prevent fires), Strategic Thinking (they influence product teams on reliability trade-offs), Intuition (debugging production systems is pattern recognition under pressure), and Service Orientation (they see infrastructure as a platform, not a tax on product).


CI/CD & Pipeline Engineering

What good looks like

A strong DevOps engineer treats the pipeline as a product, not a script — they understand build determinism, test signal quality, and deployment safety as engineering problems. They make deliberate trade-offs between speed and safety, and they can show how their pipeline changes reduced lead time, failure rate, or rollback pain in measurable terms.

Behavioural questions

  1. Tell me about a CI/CD pipeline you inherited that was unreliable.
    • What were the specific failure modes — flaky tests, slow builds, broken caching, something else?
    • How did you decide what to fix first, and what did you consciously leave alone?
    • What was the change that had the biggest impact, and why did it work?
    • Where did the metrics land six months later compared to where you started?
  2. Walk me through a deployment you designed for a service with significant blast radius.
    • What did the rollout strategy look like — canary, blue-green, percentage-based, something else?
    • What signals did you wire in to decide whether to continue or roll back automatically?
    • Who else did you need to bring in, and what did you negotiate with them?
    • When did you last actually exercise the rollback path, and what happened?
  3. Describe a time you said no to a pipeline change someone wanted.
    • What were they asking for, and what was the risk you saw that they didn’t?
    • How did you frame the pushback so it landed?
    • What did you offer as an alternative path?
    • How did that decision look in hindsight?

Situational scenario

Your team’s main deployment pipeline takes 45 minutes end-to-end. Engineers are bypassing it for “small fixes” by SSH-ing onto boxes and patching directly. The lead engineer thinks the pipeline is too slow; the security team thinks the bypass is the bigger problem.

Walk me through how you’d diagnose and unblock this.

Then introduce a constraint: “You discover the 45 minutes is mostly an integration test suite that's caught two production-breaking bugs in the last quarter. Cutting it speeds the pipeline but reintroduces real risk. What do you do?”

What to listen for

whether they separate the social problem (bypassing) from the technical problem (slow pipeline), whether they reach for parallelisation and test pyramid thinking before deletion, and whether they own the security conversation rather than punting it.

Incident Command & Reliability

What good looks like

A strong DevOps engineer can hold a sev-1 together under pressure — they separate stabilisation from root cause, communicate clearly to non-technical stakeholders mid-incident, and run post-mortems that change the system rather than blame the person. They treat reliability as a product investment, not a heroics culture.

Behavioural questions

  1. Tell me about the worst production incident you’ve been on the ground for.
    • What were you actually doing in the first fifteen minutes — be specific about the commands or dashboards?
    • At what point did you bring in others, and who?
    • How did the team land on the root cause, and how confident were you in that diagnosis?
    • What concrete change to the system came out of the post-mortem, and is it still in place?
  2. Describe a recurring alert or incident pattern you eliminated.
    • What was the symptom people kept seeing, and what was the underlying cause?
    • Why hadn’t it been fixed already by the time you got to it?
    • What was your fix, and what did it cost to put in?
    • How did you verify the pattern was actually gone rather than just quiet?
  3. Tell me about a time you pushed back on a post-mortem that blamed an individual.
    • What was the narrative the room was settling on?
    • What did you say, and how did people react?
    • Where did the post-mortem end up landing?
    • What changed about how that team ran incidents afterwards?

Situational scenario

You’re on-call. At 02:40 your primary database alert fires — replication lag at 90 seconds and climbing. Customer-facing latency is starting to spike. Your senior engineer is on holiday. You have a junior engineer also paged in who is asking what they should do.

Take me through the next ten minutes.

Then introduce a constraint: “Twenty minutes in, the lag stabilises but doesn't recover, and your VP of Engineering joins the call asking whether to send a public status update. What's your answer, and what's it based on?”

What to listen for

whether they prioritise stabilisation over diagnosis early, whether they give the junior a concrete task rather than leaving them watching, how they handle the comms decision under uncertainty, and whether they track time-to-mitigate as a distinct thing from time-to-resolve.

Infrastructure-as-Code & Automation

What good looks like

A strong DevOps engineer writes infrastructure code the same way good software engineers write application code — modular, tested, reviewed, and versioned. They understand state management, drift, and blast radius, and they actively replace manual toil with durable automation rather than running scripts repeatedly.

Behavioural questions

  1. Tell me about an IaC migration you led or did substantial work on.
    • What was the starting state — click-ops, shell scripts, an older tool?
    • How did you scope the migration so you weren’t trying to boil the ocean?
    • What did you do about the resources that existed in production but weren’t in code yet?
    • What broke along the way, and how did you recover?
  2. Describe a piece of manual operational work you automated away.
    • What was the task, and how often was it being done before you touched it?
    • What did your automation actually do — show me the seams?
    • How did you make sure the automation was safe to run unattended?
    • How much time did it actually save, and how do you know?
  3. Tell me about a time your IaC change caused an outage or near-miss.
    • What did you change, and what was the unintended consequence?
    • How did you catch it — was it the pipeline, a review, or production?
    • What did you put in place afterwards to stop that class of error?
    • Has that guardrail been tested in anger since?

Situational scenario

Your team manages a Terraform monorepo with 30+ modules across three cloud accounts. A junior engineer opens a PR that refactors a shared networking module. The plan output looks clean in their dev account, but you notice the module is also imported by the production VPC stack.

Walk me through what you do before approving or rejecting that PR.

Then introduce a constraint: “They tell you they need this merged today because it's blocking another team's launch. The other team's lead is already messaging you. What do you do?”

What to listen for

whether they reach for plan output in the production workspace before anything else, whether they think about state, dependency graphs and blast radius rather than just code quality, and how they handle the time-pressure social dynamic without rubber-stamping.

Cloud Cost & Security Posture

What good looks like

A strong DevOps engineer treats cloud spend and security posture as ongoing engineering disciplines, not annual clean-up projects. They can identify where money or risk is leaking, prioritise fixes by impact rather than ease, and influence other teams to change behaviour without owning the bill or the threat model entirely themselves.

Behavioural questions

  1. Tell me about a cloud cost reduction you drove.
    • How did you find the spend in the first place — what tooling or queries?
    • What were the top two or three line items, and what was driving them?
    • What did you change, and who else did you need on board?
    • What was the saving in real money, and did it hold up six months later?
  2. Describe a security finding you raised that other people didn’t want to act on.
    • What was the finding, and how severe was it in practice?
    • Why was the team reluctant — effort, cost, ownership, something else?
    • How did you make the risk concrete enough that it got prioritised?
    • What ended up happening?
  3. Tell me about a time you had to balance a security control against developer friction.
    • What was the control being proposed, and what was the friction it created?
    • How did you work out where the right line was?
    • What did you ship in the end, and what got cut?
    • How do you know it was the right call?

Situational scenario

A finance review flags that your team’s AWS bill grew 40% last quarter while user numbers grew about 8%. You’ve got a week before you need to present findings and a plan to your director. You don’t currently have cost allocation tags on most resources.

Walk me through your first 48 hours on this.

Then introduce a constraint: “Two days in, you find that 60% of the increase is a single service team running an experimental ML workload with no budget approval. They're a peer team, not yours. How do you handle that?”

What to listen for

whether they start with data — cost explorer, usage reports — before reaching for tags, whether they distinguish one-time spikes from structural growth, and whether they raise the cross-team issue through their director or try to fix it themselves.

Platform-as-Product Mindset

What good looks like

A strong DevOps engineer thinks of internal platform work as a product with users, adoption metrics, and a roadmap — not a ticket queue. They invest in self-service, paved roads, and good documentation, and they can show how a platform decision they made changed the behaviour of the engineers using it.

Behavioural questions

  1. Tell me about a piece of internal platform tooling you built that got real adoption.
    • Who were the users, and how did you work out what they actually needed?
    • What was the first version like, and what was deliberately left out?
    • How did you measure whether people were actually using it?
    • What did the second version look like, and what drove those changes?
  2. Describe a platform feature you built that didn’t get used.
    • What did you build, and what did you expect would happen?
    • When did you realise adoption wasn’t there?
    • What did you learn was actually getting in the way?
    • What did you do about it — iterate, deprecate, something else?
  3. Tell me about a time engineers were going around your platform rather than through it.
    • What were they doing instead, and what did that tell you?
    • How did you find out — metrics, conversations, an incident?
    • What did you change, and was it the platform or the conversation?
    • Where did usage land afterwards?

Situational scenario

You’ve built a deployment platform that abstracts away Kubernetes for product teams. Adoption is good for new services, but one senior team refuses to migrate their legacy services, saying the platform doesn’t support their needs. Your director wants the migration done this quarter.

Walk me through how you handle this.

Then introduce a constraint: “When you dig in, you discover their actual blockers are reasonable — they need a feature your platform genuinely doesn't have, and building it would push out two other roadmap items. What do you tell your director?”

What to listen for

whether they go talk to the team before defending the platform, whether they distinguish "won't" from "can't" use cases, and whether they manage upwards honestly rather than promising the migration date will hold.

Cross-team Collaboration & Developer Experience

What good looks like

A strong DevOps engineer operates as a force multiplier across product teams rather than a gatekeeper. They build relationships before they need them, can hold a firm line on risk without becoming the team of no, and they measure their own success partly through the productivity and confidence of the engineers around them.

Behavioural questions

  1. Tell me about a time a product team asked you to do something risky to hit a deadline.
    • What were they asking for, and what was the deadline pressure?
    • What was the specific risk you saw?
    • How did the conversation go — what did you say, what did they say?
    • Where did it land, and would you handle it the same way again?
  2. Describe a piece of developer experience work you did that wasn’t on anyone’s roadmap.
    • What pain were you trying to fix, and how did you spot it?
    • How did you justify spending time on it?
    • What did you build or change?
    • What did the engineers using it say afterwards?
  3. Tell me about a working relationship with a product team that was difficult.
    • What was the friction about — process, priorities, personalities, something else?
    • What did you try first that didn’t work?
    • What eventually shifted things?
    • Where’s that relationship now?

Situational scenario

A product team has been silently maintaining their own parallel deployment scripts because they think your team’s standard pipeline is too restrictive. You only find out when one of their hand-rolled deploys causes an incident. Their tech lead is defensive and frames it as a platform failure, not a process failure.

Walk me through how you handle the conversation and what you do next.

Then introduce a constraint: “Your own manager wants to escalate this to their director and frame it as the product team going rogue. You disagree. What do you do?”

What to listen for

whether they go into the conversation curious rather than righteous, whether they own the platform's share of the failure honestly, and whether they handle the disagreement with their manager directly rather than going along with it.

Red Flags in DevOps Engineer Interviews

Watch for candidates who list every tool on their CV but can only describe one deeply—they’re generalizers, not specialists, and generalists flame out on-call. Be wary of anyone who describes incidents in blameful terms (“the frontend team deployed a bad build”) rather than learning-oriented ones (“we didn’t have automated rollback for that scenario”). If they treat security and cost as someone else’s job—”that’s the security team’s problem” or “I just run what I’m told”—they won’t own platform outcomes. Red flag, too: they’ve never written a runbook or automation they’re proud of. If they can’t walk you through a concrete piece of infrastructure they built and maintained, they’re probably a ticket-closer, not a builder. Finally, if they describe their on-call experience only as fire-fighting—paged, fixed the thing, moved on—rather than “I built X to prevent that happening again,” they haven’t developed the ownership mindset that keeps systems running.


How to Structure a DevOps Engineer Interview Loop

A solid interview loop for DevOps engineers looks like this: a recruiter screen that probes their on-call experience and biggest infrastructure project; a hiring manager conversation focused on ownership, communication, and one significant incident; a practical task (30–60 minutes) where they write a small piece of Terraform or a deployment script alongside your team, not leetcode, so you see how they think about infrastructure; a system-design conversation centred on reliability—how would they design a multi-region system, or rebuild a failing deployment pipeline—rather than generic architecture; an incident-response scenario where you describe a degraded system and they talk through their triage and communication; and a quick behavioural round on how they’ve handled pressure and pushed back on risky changes.

Most organisations botch this by pattern-matching on tools in the recruiter screen, skipping the incident-response signal altogether, and testing generic coding ability instead of infrastructure-as-code thinking. The result: you hire someone technically competent but without the judgment to run your platform. Standardise your questions across your hiring managers.


Book an Interview Intelligence demo

Frequently asked questions

What's the difference between a DevOps engineer and a Site Reliability Engineer?

DevOps is primarily about building the tooling, pipelines, and infrastructure that allow product teams to deploy safely and often. SRE is broader—they own reliability across the whole system, including application architecture, database design, and observability. In practice, they overlap significantly. A good screening question: "How much of your time goes to 'building systems' versus 'keeping systems running'?" You want DevOps engineers who spend most of their time building.

Should I ask tool-specific questions?

Avoid it. A question like "How do you set up a VPC?" tests memory, not judgment. Instead, ask them to walk you through a real infrastructure decision they made. The tool becomes incidental—you'll learn what they know by asking them to teach you a concrete example.

How do I tell if they've actually been on-call?

Ask specific questions. "What was the worst page you've ever gotten?" If they can't answer with a real story, they haven't been paged. A candidate with no incident-response experience will tell you—don't hire someone into a role that requires it and assume they'll learn. On-call is a skill you develop over years.

What's a deal-breaker?

If they can't walk you through one concrete production system they built and owned, they're not a builder. If they describe problems as "they broke it" rather than "we need to prevent this," they don't own outcomes. If they've never experienced on-call, don't expect them to handle it calmly.

What should I ask about their current role if they're happy there?

Don't. Instead, ask what their next level looks like. "What's the biggest infrastructure challenge your current company faces that you're not working on?" This reveals whether they're looking to grow or just looking for a pay rise.