hiring revops

RevOps interview questions that predict performance

Most RevOps interview questions are answerable from a template. These are not, plus what a strong answer sounds like and what each one is actually measuring.

Rareix · · updated · for employers

RevOps interview questions that predict performance

Most published RevOps interview questions are answerable from a template. “Tell me about a difficult stakeholder” has a correct-shaped answer that any competent candidate can produce, and it tells you nothing.

But the questions are the smaller half of the problem. The evidence is fairly clear that structured interviews are the best-validated selection method in common use, at an operational validity of .42 against job performance on the corrected 2022 estimates, ahead of job knowledge tests at .40 and work sample tests at .33, while unstructured interviews sit at .19. That is a very large difference between two literatures, and almost none of it comes from which questions get asked.

It comes from the scoring key. So this post carries one.

The short answers

  • Structure is what predicts, not the question list. Same wording for every candidate, a key written in advance, independent scores, and a fixed bar rather than a ranking. Most processes do the first and none of the rest.
  • Write what a 1, a 3 and a 5 sound like before you meet anyone. In observable behaviour, not adjectives. This is the artefact that does the work and it is the one nobody writes.
  • Four or five questions with real follow-ups beats twelve. Depth separates candidates; breadth produces answers everyone can give.
  • The follow-up is the measurement. The first answer is often rehearsed. “You check that and it is fine, now what” is not.
  • Score against the bar, not the field. A ranking tells you who was best of three. The bar tells you whether to hire any of them.
  • Ask what they ruled out and what would have changed their mind. Neither is answerable from fluency, which is the thing interviews over-reward.
  • Cut anything with a correct-shaped answer. Greatest weakness, difficult stakeholder, brainteasers.
  • Expect the bar to drift down over a long search. Keep the anchors visible and re-read two early scorecards before each debrief.

Write the scoring key first

A scoring key is one page per question. For each, you write what a weak, adequate and strong answer actually sounds like, in behaviour you could observe rather than in adjectives you could apply afterwards.

The test of an anchor is whether two people who both heard the same answer would land on the same score. “Thoughtful” fails that test. “Asked what the discrepancy was measured against before proposing a cause” passes it.

Three anchors are enough. Five is false precision, and the middle three collapse into each other under pressure anyway.

ScoreWhat it meansHow you know
1Would not do this part of the jobNamed behaviour is absent, and prompting does not produce it
3Would do this part of the job adequatelyNamed behaviour appears, with a prompt
5Would raise the standard of this part of the jobNamed behaviour appears unprompted, with a reason attached

Two rules that make the key survive contact with a real panel.

Written before you meet anyone. A key written afterwards describes the candidates you liked, which is a ranking with extra steps.

Anchored to your scorecard, not to a generic competency model. The outcomes you wrote when you scoped the role are what the interview is measuring. If a question does not map to one of them, it is there out of habit.

The five questions

Each one below has the question, what it measures, the anchors, and the follow-up that does the separating. Ask them in this order: diagnostic first, because it is the one a candidate is least likely to have rehearsed and it sets the register for everything after.

1. Diagnostic reasoning

Closed-won revenue is reporting about 20 per cent below what finance sees. You have access to the instance. What do you check first, and why that first?

Measures: whether they investigate before concluding, and the order they work in. This is the highest-signal question in the set for an operations role, because the ordering is the job.

What a 1 sounds like. Names a cause immediately and starts solving it. Reaches for a tool before establishing what the numbers are being compared against. Cannot say what would prove them wrong.

What a 3 sounds like. Asks at least one clarifying question. Works through plausible causes in a sensible order. Gets to currency, date boundaries or record types without prompting, but treats the first plausible answer as the answer.

What a 5 sounds like. Establishes what is being compared to what before anything else: which system, which date range, which definition of closed. Checks the boring causes first, out loud, and says why they are checking them in that order. Names what they are ruling out as they rule it out. Volunteers what would change their mind.

Follow-up that separates people: “You check that and it is fine. Now what?” The candidates who have actually done this have a second and third hypothesis ready. The ones who have read about it do not. A second follow-up worth having ready: “How would you know this had happened again, without anyone telling you?”

Context worth having. This scenario is not exotic. Gartner’s survey work reports fewer than half of sales leaders and sellers holding high confidence in forecast accuracy, and a similar share believing their data quality is high. That is a release with no disclosed sample size, so read it as an indication of how common the complaint is rather than a measurement. Almost every candidate has met a version of this. What differs is what they did about it.

2. Saying no

A VP asks for six new required fields at close stage. What do you do?

Measures: whether they understand that an operations function which ships every request becomes the bottleneck it was hired to remove.

What a 1 sounds like. Implements it. Or refuses flatly without offering an alternative, which is the same failure wearing a different hat.

What a 3 sounds like. Pushes back on the volume, negotiates down to two or three fields, understands that required fields at close produce fiction entered at speed.

What a 5 sounds like. Asks what question the fields are meant to answer, then proposes the cheapest way to answer it, which is frequently not a field. Says no in writing, with a reason, and offers an alternative in the same message. Mentions who maintains the field afterwards and what happens when a rep does not know the answer.

Follow-up: “The VP says they need it anyway, and they outrank you.” A strong answer escalates with a written trade-off rather than either capitulating or digging in. Look for whether they name the cost to someone else, usually the reps and sometimes the forecast, rather than framing it as their own preference.

3. Being wrong in public

Tell me about a number you reported that turned out to be wrong.

Measures: the core professional skill in operations, which is saying what the data does not support.

What a 1 sounds like. Cannot think of one. That is the answer, and at senior level it is close to disqualifying: everyone who has done this job at any depth has published a wrong number. The other 1 is an example where the error was someone else’s fault throughout.

What a 3 sounds like. A specific example, honestly told, with what the error was and roughly what it cost.

What a 5 sounds like. All of that, plus what they changed so it could not recur, without being prompted for it. The strongest answers describe the control they added, not the apology they made, and are specific about how they found out, because the detection mechanism is usually the interesting part.

Follow-up: “How did you find out?” If the answer is that someone else noticed, ask what they built afterwards so that they would notice first.

4. Design under conflict

Sales want lead routing by territory and named account. Marketing want it by score. Design it.

Measures: whether they resolve conflicts or pick a side. This is the question that most directly tests the part of the job that is not technical.

What a 1 sounds like. Produces an elegant design that assumes the conflict away. Or asks you which team wins, which is asking you to do the job.

What a 3 sounds like. Establishes what each side is optimising for, proposes a workable compromise, describes how it would be built.

What a 5 sounds like. Starts with what each side is actually optimising for and says so out loud. Designs for the failure case, meaning what happens when a lead matches both rules, or neither. States the trade-off they accepted and names who would have to agree it. Asks who maintains it after they leave, and whether the rule can be explained to a rep in one sentence.

Follow-up: “Six months later it is not being followed. What happened?” Strong candidates go to enforcement and comprehensibility rather than to the technical design, because that is where routing rules actually die.

5. Explaining to someone who does not care

Explain what you just worked through to a VP of Sales who does not care how the CRM works.

Measures: the most common failure in operations, which is not technical. Ask it immediately after question 1 or 4, about that specific answer, so it cannot be prepared.

What a 1 sounds like. Narrates the method chronologically. Uses tooling vocabulary without noticing. Cannot say what the business should do differently.

What a 3 sounds like. Leads with the answer, drops most of the jargon, states a recommendation.

What a 5 sounds like. Leads with the answer and the decision it implies. Drops the tooling vocabulary without dumbing the content down. Is explicit about confidence: what they know, what they are inferring, what they would need to be sure. Says what they would do next and what it would cost.

Follow-up: “The VP disagrees and says the number is wrong.” What you are watching for is whether they defend the number, revisit the method, or ask what the VP is seeing. The third is the right instinct and the rarest.

Three more, for specific situations

For a GTM engineering role: Tell me about something you built that broke. How did you find out? Instrumentation is what separates builders from configurers, and the answer “a rep told me” is a 1 for a role whose output is automation. The wider context is real: the martech landscape mapped 15,384 products in 2025, up 9% on the year, and a builder who cannot monitor their own work is adding to that surface rather than controlling it. The role comparison covers where this hire fits.

For a first ops hire: It is week one. You have no admin access yet and no documentation. What do you do? You are testing for whether they generate their own mandate or wait to be given one, which is the single most predictive thing about a first ops hire. Strong answers start with people and definitions rather than systems. See what a first ops hire should look like.

For a manager-level hire: What have you stopped doing, and how did you decide? Operations accumulates obligations: reports nobody reads, fields nobody uses, integrations nobody owns. A candidate who has never removed anything has never run the function, only serviced it. FoundHQ’s analysis of Yelp’s Salesforce team, where 3,000 users are supported by a core team containing five administrators against a commonly quoted practice of seven to ten per 1,000 users, is an example of what deliberate subtraction looks like at scale.

A completed scorecard, filled in

The abstract version of a scoring key is easy to agree with and hard to act on, so here is one as it looks after an actual conversation. This is a real shape, with the candidate and company invented.

Role: RevOps Manager. Scorecard outcome under test: a single agreed definition of a qualified opportunity, in use, by week 8.

QuestionScoreEvidence
1. Diagnostic4Established which two systems were being compared before proposing anything. Checked date boundaries and record types out loud. Did not volunteer what would change their mind until asked
2. Saying no5Asked what question the six fields answered. Proposed a report off existing data instead. Named the cost to reps unprompted
3. Being wrong3Specific example, honestly told. Needed a prompt to get to what they changed afterwards
4. Design under conflict4Named both objectives before designing. Handled the both-rules-match case. Did not address who maintains it
5. Explaining2Narrated the method chronologically. Two uses of “basically”. Did not state a recommendation

Recommendation: hire, with the communication gap named in the first ninety days.

Three things about that sheet are worth copying.

The evidence column is not optional. A score with no evidence beside it is an impression that has been given a number, and it is unusable in a debrief because nobody can check it. The rule that works: if you cannot quote or paraphrase what they actually said, you have not scored it.

A low score does not veto. A 2 on communication against a scorecard whose outcomes are all about definitions and agreement is a serious finding. The same 2 for a role whose first task is a systems migration is a note. The scorecard decides which, and it decided before the interview, which is the point.

Disagreement lands on a row, not on a candidate. When the second interviewer scores question 5 a 4, the conversation is about one exchange that both people heard, which is a resolvable disagreement. Without the sheet it is a conversation about whether they were impressive, which is not.

Questions to cut

Not because they are unpleasant, but because they have a correct-shaped answer that carries no information about performance.

  • Greatest weakness. Measures preparation.
  • Difficult stakeholder. Measures narrative skill. If you want the underlying thing, ask question 2 or the follow-up to question 4.
  • Where do you see yourself in five years. Measures nothing, and reliably annoys senior candidates.
  • Brainteasers and puzzles. Measure whether someone has seen that puzzle.
  • Anything about your business they have no information to answer. “How would you fix our attribution” invites a guess and then rewards confidence in it.
  • Why do you want to work here. Ask it if you like, but do not score it. The answer is a courtesy.

One category worth handling deliberately rather than cutting: questions about level and ambition, for a candidate who looks over-qualified for the seat. The research synthesis on perceived overqualification associates it with lower satisfaction and with turnover, and identifies autonomy and empowerment as substantial moderators. So the useful question is not “will you be bored”, which nobody answers honestly, but “what would you want to be able to decide without asking anyone?” That answer tells you whether the seat as designed will hold them.

Running the panel

The questions do less than the process around them.

Two rounds, not four. One structured conversation on the scorecard, one panel built from what the candidate actually did in an assessment. Four unstructured interviews is the same weak measurement taken four times, and the agreement between them feels like corroboration when much of it is a shared first impression propagating. Every additional round also costs you the candidates who have options.

Same wording, every time. A question rephrased is a different question, and scores from differently-worded questions are not comparable. Read them out if necessary.

Independent scores, submitted before discussion. Then reveal at once. This is the single cheapest intervention in the whole process and the one most often skipped because it feels bureaucratic in a room of three people who like each other.

Explore disagreement rather than averaging it. A 2 and a 5 on the same question means two people heard different things. Finding out which is the entire value of the debrief.

Score against the bar. If nobody clears it, nobody clears it. A weak field otherwise produces a hire and a strong field produces a rejection, which is exactly backwards.

Scoring drift, and how to catch it

Keys drift, and they drift in one direction. After eight weeks of a search that is not going well, a 3 starts to feel like a 4.

That matters because searches for these roles are not short: median time to first fill runs to 71 days for senior roles and 75 for technical ones on Ashby’s benchmark data, across 93,000 jobs. Two months is long enough for a bar to move without anyone deciding to move it.

Three cheap controls. Keep the anchors physically visible during scoring rather than in a document nobody opens. Re-read two early scorecards before each new debrief, so the comparison is against the standard rather than against last week. And write down, at the start, what you will do if nobody clears the bar, because that decision is much harder to make honestly in week ten than in week one.

What this cannot measure

Worth stating, because a question list that only advertises its strengths is a sales document.

An interview measures how someone describes and reasons about work, in a conversation, under observation. It does not measure whether they will still be doing it well in year two, how they behave when nobody is watching, or whether they will get on with a team they have not met. At .42 the best-validated method in common use is a long way from deterministic, and any process claiming to eliminate hiring risk is selling something.

The evidence on conventional hiring practice is unflattering for exactly this reason. Peter Cappelli’s survey of the field is largely a catalogue of processes that measure what is easy to measure. The 2023 research on selection design points the constructive way out: job-specific measures outrank general construct measures, so the highest-value stage is usually the one closest to the actual work.

Which is the argument for watching the work rather than only discussing it. If you want to see what that produces, the Rareix assessment is published in full, including what it deliberately does not test.

Questions

What people ask about this.

Do interview questions actually predict job performance?
Structured ones do, better than anything else commonly used. On the corrected 2022 meta-analytic estimates, structured interviews reach an operational validity of .42 against job performance, ahead of job knowledge tests at .40 and work samples at .33. Unstructured interviews sit at .19. The question list is not what produces that gap. The scoring key is.
What makes an interview structured?
Four things, and most processes do only the first. The same questions in the same wording for every candidate. A scoring key written before you meet anyone, with behavioural anchors describing what each score looks like. Scores recorded independently before any discussion. And scoring against a published standard rather than against the other candidates. Skipping the key is what turns a structured process back into an unstructured one with a list.
How many questions should an ops interview cover?
Four or five, with follow-ups, beats twelve. Depth is what separates candidates: a wide sweep of shallow questions produces answers everyone can give. Budget about twelve minutes per question including the follow-up, which is what turns a rehearsed answer into an unrehearsed one, and hold the fifth slot for whatever the assessment or the CV actually raised.
How do I write a scoring key?
For each question, write what a 1, a 3 and a 5 sound like in behavioural terms, using observable actions rather than adjectives. Not thoughtful, but asked what the discrepancy was measured against before proposing a cause. Three anchors are enough; five is false precision. Write them before you meet anyone, because a key written afterwards describes the candidates you liked.
What is the biggest interviewing mistake in ops hiring?
Rewarding fluency. A confident, well-structured answer sounds like competence and is not evidence of it, particularly now that every candidate has rehearsed against a model. The correction is mechanical: ask what they ruled out and why, and ask what would have changed their mind. Neither is answerable from fluency alone.
Should candidates get the questions in advance?
Send the themes, not the wording. Telling a candidate the conversation will cover a diagnostic problem, a time they were wrong and a design trade-off removes anxiety without removing the measurement, because these questions are not answerable by preparation alone. It also levels the field between candidates who have been interviewing recently and those who have been doing the job.
Do these questions work for GTM engineering roles too?
The diagnostic and communication questions transfer directly. For GTM engineering, add one about something they built that broke and how they found out, because instrumentation is what separates builders from configurers. For a first ops hire, add the one about what they would do in week one with no access and no documentation.
How do we stop the panel from anchoring on the first opinion?
Everyone writes and submits their scores before any discussion, and the scores become visible at the same moment. A panel that debriefs before scoring produces one opinion held by three people, and the most senior voice in the room sets the anchor inside about ninety seconds. Then explore disagreement rather than averaging it: a 2 and a 5 on the same question means two people heard different things, and finding out which is the entire value of the meeting.
Should we ask candidates to do a take-home?
A short live exercise beats a multi-day take-home. The take-home measures the output, which is the part a model now produces competently, and it filters for candidates with spare evenings rather than for capability. Sixty to ninety minutes is the ceiling for anything unpaid, and the 2023 research on selection design supports the underlying instinct: job-specific measures outrank general construct measures.
What questions should we cut?
Anything with a correct-shaped answer available from preparation. Greatest weakness, difficult stakeholder, where do you see yourself, and any question whose honest answer is a guess about your business that the candidate has no information to make. Also cut brainteasers, which measure whether someone has seen that brainteaser.
How do we interview for a role we do not understand ourselves?
Recognise it as the situation you are in, which is common: operations roles are frequently assessed by people who have never done them. Two responses work. Bring in one person who has done the job, from anywhere, including outside the company. And weight the evidence toward something you can evaluate without expertise, which is what watching someone work gives you that a technical conversation does not.
Do scoring keys drift over a long search?
Yes, and they drift downward as a search runs on. The mechanism is well understood: after eight weeks of a hard search, a 3 begins to feel like a 4. Catch it by keeping the anchors visible during scoring, by re-reading two early scorecards before each new debrief, and by treating the bar as fixed. A search that lowers its standard has not found a candidate, it has changed the question.

Tell us the role. We will tell you honestly whether we can fill it.

Nothing owed until someone starts.

Book a call