TL;DR: Most reference checks for IT staffing engagements stop at the happy-path question “would you hire them again?” and produce almost no signal. Per Forrester‘s 2026 Vendor Trust Outlook, 78% of staffing references are warm contacts who default to positive answers, producing roughly 15% predictive value over no reference at all. Deep-diligence references that cover operational fit, mid-engagement adjustments, pricing reality, exit experience, and renewal intent produce 4-5x more useful signal.
References are the diligence stage almost every buyer schedules and almost no buyer runs well. The default 15-minute “would you hire them again?” call produces a predictable warm answer and leaves the buyer no better informed than before the call.
For the broader vendor-evaluation context, see our checklist on 15 questions to ask IT staffing providers.
Methodologies and benchmarks in this article are sourced from Forrester‘s 2026 Vendor Trust Outlook, Staffing Industry Analysts (SIA) 2025 IT Staffing Report, Gartner‘s 2026 Reference Diligence Index, Robert Half‘s 2026 Hiring Quality Report, and the McKinsey 2025 Talent Diligence study. Reference benchmarks assume 30-45 minute structured calls rather than email-based reference forms.
Why Happy-Path Reference Checks Almost Never Work
The default reference call follows a predictable shape: the provider hands the buyer two or three named contacts, the buyer schedules a 15-minute call, the buyer opens with “how was the engagement?” and the contact responds with positive sentiment. The call ends with “would you hire them again?” answered yes, and the buyer marks the diligence stage complete.
Per Forrester, this script produces roughly 15% predictive value over no reference at all. The reason is structural. Providers hand-pick references from satisfied clients (selection bias). The buyer’s questions invite generic praise (anchor effect). The contact has a relationship with the provider and will not volunteer negative signal without an explicit prompt (politeness gradient). The format produces warm reassurance, not diligence.
The fix is not to skip references. SIA’s 2025 report shows that buyers who run structured references see 31% fewer bad engagements than buyers who skip the step. The fix is to change the script. Deep-diligence questions cover the dimensions where signal lives and the dimensions where provider-selected references cannot easily generate warm answers without revealing useful information.
The Five Categories That Produce Real Signal
Forrester’s 2026 outlook identifies five reference categories that consistently produce useful signal, ranked by predictive value. Happy-path questions cover none of them in depth; deep-diligence questions cover them all.
Category one: operational fit. How did the engagement actually integrate with the buyer’s team? Did the sprint plan have to change after the engineer joined? Was on-call coverage adjusted? Were timezone overlap windows hit or missed? Operational fit questions reveal the real friction of the engagement, not just the headline outcome.
Category two: mid-engagement adjustments. Almost every engagement has at least one mid-flight surprise: a scope shift, a personnel substitution, a pricing line-item discussion, a missed deadline. How the provider handled the surprise is more predictive than the absence of surprises. References will rarely volunteer this; targeted questions surface it.
Category three: pricing reality. Did the quoted all-inclusive rate hold across the engagement? Did line items appear (equipment, replacement fees, EOR setup, conversion fees)? Were renewal escalators discussed transparently? Pricing-reality questions are where the politeness gradient breaks down because clients have specific numbers to refer to.
Category four: exit experience. How did the last 30 days of the engagement run? Was knowledge transfer structured? Were access credentials cleanly transferred? Did the engineer ramp down or check out early? The exit experience is the single strongest indicator of whether the provider treats the engagement as a contract or a relationship.
Category five: renewal intent. Is the client renewing or not? If yes, on the same terms or adjusted terms? If no, what specifically would they change? Renewal intent is harder to dodge than “would you hire again” because the contract is either renewing or not. Soft answers (“we are evaluating”) are themselves signals worth interpreting.

The Deep-Diligence Question Bank
A 30-45 minute structured reference call covers the five categories above with specific question patterns. The questions are designed to produce concrete examples rather than abstract praise.
Skill and Role Quality
Replace “Are they good?” with these patterns:
- “Where did this engineer struggle in the first 90 days and how did they recover?”
- “Tell me about a moment you remember when their work surprised you in either direction.”
- “Was there a sub-skill (system design, code review, debugging) where they were noticeably stronger or weaker than the rest of your team?”
- “How did they compare on the AI tooling axis specifically? Cursor or Claude Code fluency?”
Communication and Async Fit
Replace “Were they easy to work with?” with:
- “Show me one written async update from them that you remember well. What was good or bad about it?”
- “How did they handle a moment when a decision had to be made without you in the room?”
- “Did time zone overlap go as expected or did you need to adjust?”
Operational Fit
Replace “Did they meet deadlines?” with:
- “How did your sprint plan change after they joined? Faster, slower, or different shape?”
- “Did the engineer integrate with your existing on-call rotation? How smoothly?”
- “Was there a specific operational rhythm (standups, retros, planning) that took longer than usual to settle?”
Mid-Engagement Adjustments
Replace “Anything come up?” with:
- “What scope shift surprised you in the first six months and how did the provider adjust?”
- “Did the provider ever propose a personnel change or substitution? How did that conversation go?”
- “Was there a single moment where you considered ending the engagement early? What changed your mind?”
Pricing Reality
Replace “Was the price fair?” with:
- “Did the all-inclusive monthly rate hold across the engagement or did line items appear?”
- “How did the renewal escalator conversation go? Was the new rate within 5% of the original or further?”
- “Did the provider invoice match the contract every month, or did you have to reconcile mismatches?”
- “If you converted an engineer to FTE, what was the conversion fee and how was it negotiated?”
Exit Experience
Replace “Did the engagement end well?” with:
- “Walk me through the last 30 days of the engagement. What did knowledge transfer look like?”
- “How quickly did the provider cleanly remove the engineer’s access after the engagement closed?”
- “Was there any post-engagement support window or follow-up call?”
Renewal Intent
Replace “Would you hire them again?” with:
- “Are you renewing the engagement? If yes, same terms or adjusted? If no, what specifically would you change?”
- “If a peer asked you for a recommendation on this provider, what would you say without naming names?”
- “What would have to change for you to expand the engagement to 5 engineers from 1?”
How to Interpret Evasive vs Direct Answers
Even structured questions produce noisy signal if the interpretation framework is loose. Per the McKinsey 2025 Talent Diligence study, the most useful signals are not the words in the answer but the patterns around the answer. Seven patterns to watch:

A long pause before answering a specific question (about pricing, about exit) is almost always a signal worth probing. Direct readers are thinking through the answer. Evasive readers are weighing what to omit. The follow-up question is “what are you weighing right now?” which often unlocks the real answer.
“They were great” with no example is the single most common warm-reference pattern. The follow-up is “great in what specific way?” If the contact cannot produce a concrete example in 10 seconds, the warm answer is rehearsed cover, not lived experience.
Naming a peer to validate is a positive signal in most cases. References who name a teammate or a backup contact are offering corroboration, which is a hallmark of honest reference experience. Evasive references redirect to a different topic instead.
The strongest direct signal is “I would not renew” stated cleanly. It is rare because contacts that strongly negative usually decline the reference call entirely. When it appears, weight it heavily. The strongest evasive signal is “we are evaluating” or “we are in the middle of a decision” when the contract end date has passed. The math is one-sided: renewing references say so directly. Non-renewing references soften the answer.
Common Reference Biases and How to Correct For Them
Three biases run through almost every reference call. Naming them in advance and structuring the call to correct for them is the difference between a useful diligence and a marketing exercise.
Selection bias. Providers hand-pick references. Per Gartner, 87% of provided references rate as 4 or 5 stars when surveyed. The bias is structural and cannot be eliminated. The correction is to ask for an additional reference type the provider does not pre-select: a back-channel contact (a peer at another company who has used the same provider), or a recent off-board (an engineer the provider placed who has since rolled off, where the buyer asks the engineer’s perspective rather than the client’s).
Anchor effect. The first answer to “how was the engagement?” sets the tone for the rest of the call. Warm openers produce warm follow-ups. The correction is to open with a specific operational question rather than a general “how was it?” Starting with “tell me about a moment you remember well” or “walk me through the first 90 days” produces concrete narrative immediately and resists anchoring to abstract praise.
Politeness gradient. References who have a relationship with the provider are reluctant to volunteer negative signal. The correction is to explicitly invite it. The phrase “I am not looking for marketing pitch, I’m trying to make a real decision; what would I want to know that’s harder for you to say?” gives the contact permission to share the difficult parts. Per Robert Half, this phrasing increases substantive disclosure rates by 42%.
When to Skip Reference Checks Entirely
Reference checks are not the right diligence stage for every engagement. Three scenarios where skipping is justified:
Short trial engagements. If the engagement starts with a 30 or 60-day trial and the contract terms allow $0-cost early termination, the trial itself is a more reliable diligence signal than any reference call. The trial period replaces references for short engagement profiles.
Roster-replacement substitutions. When the provider is substituting an engineer mid-engagement (a replacement under guarantee, for example), references on the replacement engineer are usually not feasible or useful. The provider’s track record on the original engagement is the relevant signal.
Highly specialized roles where references are scarce. For very specific roles (a niche MLOps specialty, an unusual hardware integration), provider-side references on the same role may not exist. In these cases, the live technical assessment carries more weight, and reference time is better spent on the provider’s process quality than on the individual engineer.
SIA’s 2025 report finds that 22% of high-quality engagements proceed without a formal reference call because one of these three scenarios applies. The right question is not “do we run references?” but “do references produce more signal than the alternative diligence stage we would skip to make room?”
Reference Format: Call, Form, or Hybrid
Forrester benchmarks four reference formats by signal strength.
Email reference forms. Lowest signal. The contact fills out a form on their own time with no follow-up loop. Useful only for verifying employment dates and basic facts.
Templated phone calls. Slightly stronger but still low signal. The interviewer reads a fixed script, gets short answers, and produces a checklist outcome. The format does not allow for the follow-up questions that produce real signal.
Structured 30-45 minute calls. The recommended format. The interviewer uses the deep-diligence question bank as a starting point but adapts based on the contact’s answers. Follow-up questions are where the strongest signals appear.
Multi-reference hybrid. The strongest format. Two structured calls (one provider-supplied, one back-channel) plus one email form to a former engineer. The triangulation across three perspectives surfaces inconsistencies that any single call would miss.
Multi-reference hybrid takes 90-120 minutes of interviewer time per engagement. For engagements above 12 months in expected duration, the ROI is strong. For shorter engagements, the structured single call is usually sufficient.
Putting References in Context
Reference checks are one diligence stage among several. They produce 15-20% of total decision signal for most engagements per Gartner. Live technical assessment, the provider’s funnel report, contract terms, and trial periods carry the rest. A common mistake is treating references as the deciding stage; a stronger pattern is treating references as the dimension that fills in operational, pricing, and exit gaps that the technical assessment cannot reach.
For more on the other diligence stages, see our companion guides on how to evaluate IT staffing companies and 15 questions to ask IT staffing providers.
Start an Engagement With Transparent Reference Access
Second Talent publishes named client references on request for every RFP we participate in, including back-channel contacts where the client has authorized them. Our standard reference packet includes two structured-call contacts and one written reference covering operational fit, pricing reality, and renewal intent. We do not pre-screen references for warm answers; we provide a representative sample of recent engagements, including ones where pricing or scope was renegotiated mid-flight.
Common starting points:
- How to evaluate IT staffing companies
- 15 questions to ask IT staffing providers
- Developer rate card by country
- AI Staffing & Augmentation Services overview
Matching in 24 hours. $0 upfront. Pay only when you make a hire.

