Monday morning, 9:03 a.m., a mid-sized SaaS recruiting office. HR specialist Maya exports last night's ATS shortlist — every candidate scoring 85 or higher on the AI rubric — and fires off a bulk "you've advanced" email. At 12:47 p.m. the engineering director calls: candidate #42's portfolio is visibly stronger than #17's, but the ATS rated #42 at 71 and quietly cut him. Maya reruns the same PDF: 83. Reruns it again: 91. Once more: 68. AI resume screening inconsistency — that phrase is about to live on her team's whiteboard for the rest of the quarter. The same scene is unfolding inside the workflows of 944,300 U.S. human resources specialists. Dan Kinsky's experiment published on June 28 provided the first large-scale public benchmark of AI resume screening inconsistency: the same resume, scored 100 times by HackerRank's open-source hiring-agent, swung from 65 to 99. This article maps the data onto Bureau of Labor Statistics figures and gives HR teams a five-step playbook to take the noise back out.
1. The BLS-Verified Pain: Three Ways AI Resume Screening Inconsistency Hurts Human Resources Specialists
According to the Bureau of Labor Statistics' Occupational Outlook Handbook (last updated August 28, 2025), Human Resources Specialists (SOC 13-1071) earned a 2024 median annual wage of $72,910 ($35.05/hour). The U.S. employed 944,300 of them; the 2024–2034 projection is +6% growth (faster than average), adding 58,400 jobs and posting 81,800 openings every year. BLS defines the core function bluntly: "Human resources specialists recruit, screen, and interview job applicants and place newly hired workers in jobs." "Screen" is the load-bearing verb — and AI resume screening inconsistency is turning that verb into a dice roll.
Pain point one: AI resume screening inconsistency hollows out the judgment that BLS calls the HR specialist's defining skill. BLS lists "Decision-making skills" and "Detail oriented" as required qualities, writing "Human resources specialists must use sound judgment when reviewing applicants' qualifications" and "Specialists must pay attention to detail when evaluating applicants' qualifications." But when the upstream LLM scorer returns 65 today and 99 tomorrow on the same PDF, the shortlist HR reviews is itself a luck filter — their judgment is applied downstream of randomness. Research shows that 2025 saw dozens of public reports of companies rescinding offers after discovering bias or noise in their AI hiring tools.
Pain point two: 28% of HR specialists work at staffing agencies, where AI resume screening inconsistency multiplies into industry-level noise. BLS data shows 14% of HR specialists work in Employment Services and another 14% in Professional, Scientific, and Technical Services — together about 264,400 specialists screening on behalf of other companies. These RPO and staffing firms run ATSs at the highest volumes per recruiter. When upstream models are non-deterministic, the cost of every false rejection compounds across the firm's entire client roster. Industry data shows the average RPO consultant runs an ATS 300–500 times per month.
Pain point three: 81,800 annual openings + 6% growth = HR has no time for manual override. The BLS Job Outlook section is explicit: "About 81,800 openings for human resources specialists are projected each year, on average, over the decade." With 80,000+ HR jobs themselves open every year, new hires inherit screening queues many times their personal experience. They have no bandwidth to second-guess every ATS score, which is exactly the path through which AI resume screening inconsistency stays invisible until candidates complain or until a six-month diversity review surfaces the gap.
2. What Dan Kinsky's June 28 Experiment Actually Showed About AI Resume Screening Inconsistency
To understand why this experiment matters for the 944,300-strong HR profession, look at Dan's setup. In his June 28 post titled "HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74/100. No — 88/100. Actually 83/100.", Dan Kinsky took HackerRank's public interviewstreet/hiring-agent repository, ran his own resume through it 100 times at default settings (gemma3:4b, temperature 0.1), and plotted the distribution. Source: Dan Kinsky, "HackerRank open sourced its ATS gave my resume a different score every time", Dan Unparsed, June 28, 2026. BLS occupational page: Bureau of Labor Statistics, "Human Resources Specialists", Occupational Outlook Handbook, August 28, 2025.
The numbers killed the "low temperature = deterministic" pitch many ATS vendors lean on. Headline finding #1: 100 runs, scores ranging 65 to 99. Dan writes: "If your company's cutoff sits at 85, I fail 65% of the time. Same exact resume, different luck." That single line should change procurement conversations across every HR department in the country.
Headline finding #2: AI resume screening inconsistency concentrates in the "judgment" dimensions. Dan's category breakdown shows Technical Skills scored 8/10 in 98 of 100 runs — near-perfect consistency. Open Source and Projects swung wildly. His explanation: "Look at technical skills... a five year old could match that check-list. Now look at projects — there's HUGE variation. LLMs struggle to make a judgment call like that consistently." This hits HR specialists where it hurts most — BLS lists Decision-making skills as essential, and the AI is least reliable precisely where judgment is needed.
Headline finding #3: Temperature 0 doesn't fix it; this is architectural. Dan cites GitHub Issue #35: "scores of 27, 34, 32, 34, 34, 30 across six consecutive runs at temperature 0." His verdict: "This non-determinism isn't a bug you can just fine-tune away, it's a fundamental design flaw." The physical source is GPU floating-point accumulation order and model architecture itself. Any ATS vendor telling you "we set temperature to 0 and it's stable" hasn't done the homework.
Headline finding #4: Switching to Gemini narrows the spread but still leaves 28% false rejections. Dan re-ran 50 evaluations on gemini-3.1-flash-lite, getting a tighter 48–64 distribution. But: "if your cutoff is 60, you're still failing 28% of the time through no fault of your own." Upgrading the model is not a get-out-of-noise-free card. As long as an LLM is asked to grade something fuzzy, AI resume screening inconsistency persists.
Headline finding #5: The "Experience" rubric was two lines long, and rated everyone 25/25. Dan's experiment found senior engineers and one-internship juniors both pinned at 25/25. The prompt was literally two lines with no anchors. His takeaway: "Experience has two lines and no anchors — consistent, but useless. Projects has a detailed rubric with examples but it's the noisiest category — inconsistent, also useless."
3. Five Steps to Defuse AI Resume Screening Inconsistency Inside an HR Team
Step 1 — demote the LLM from grader to parser. Dan's closing principle is the operating manual: "Use an LLM to parse a resume into structured data — great, that's what they're good at. Use one to judge whether a candidate's experience is worth 18 points or 24 points? You get a vibe-check." Configure the ATS so it surfaces the candidate's skill list and structured work history, not a single composite score.
Step 2 — replace hard cutoffs (≥85) with soft bands (80–90 = human review). Once you accept AI resume screening inconsistency as architectural, hard thresholds become landmines. Mandate that everyone within ±5 of the cutoff goes to a human reviewer. BLS says Decision-making skills are why you hired HR specialists in the first place — let them do that work in the band where AI is provably noisiest.
Step 3 — re-run borderline candidates at least 3 times; trust the distribution, not the single score. If you can't yank the ATS today, at minimum re-run candidates near rejection thresholds three times and inspect both median and spread. High-variance candidates must escalate. This converts AI resume screening inconsistency from "pretend certainty" into "known uncertainty."
Step 4 — zero out the weight on "judgment" categories; keep only fact-verifiable ones. Dan's plots make the call easy: Technical Skills (checklist) is stable, Projects (judgment) is unstable. When you reconfigure your scoring template, push weight onto degrees, certifications, years of experience, and named-skill match. Pull weight off "project quality" and "experience impact."
Step 5 — run a quarterly AI-hiring audit. Pull 50 candidates the ATS rejected in the past 90 days, hand them blind to a senior HR specialist, and ask "should this person have advanced?" Build the confusion matrix against AI scores. This is also how you satisfy the EEOC: per public guidance, employers using AI for employment decisions must demonstrate the tool does not produce disparate impact, and BLS itself emphasizes HR's role in "ensuring that a workplace complies with federal, state, and local regulations."
4. What 90 Days of Anti-Inconsistency Workflow Looks Like
Picture an 800-person SaaS company adopting the five steps in July 2026. By October you can expect: HR specialists add 6–8 hours/week of review time, but the engineering-team first-round rejection rate drops from 72% to 58%, recovering roughly 14 strong candidates per quarter. At the BLS median of $72,910/year, the avoided cost per recovered hire (recruiter fees, sourcing time, slot-vacancy drag) hovers around $25,000. The standard deviation of ATS scores collapses from 9.2 to 4.6 because borderline candidates are re-run three times and the median is used. EEOC risk grade drops from "high" to "medium" because the audit log now exists. Comparable mid-market HR teams report a 15–25 point lift in hiring-manager NPS within six months of similar workflow changes.
5. AI Resume Screening Inconsistency FAQ
Q1. Will setting temperature to 0 eliminate AI resume screening inconsistency? No. GitHub Issue #35 on the HackerRank repo showed scores of 27, 34, 32, 34, 34, 30 at temperature 0. The variance originates from GPU floating-point accumulation order and model architecture, not from sampling temperature. Research shows mainstream LLMs still produce 5–15% output variance at temperature 0.
Q2. Does upgrading to GPT-class or Claude-class models solve it? It narrows the distribution but never closes it. Dan measured a 28% false-rejection rate even on Gemini. Given 944,300 HR specialists processing 81,800 openings annually, any 5%+ false-rejection rate translates to thousands of high-quality candidates lost per year at industry scale.
Q3. How does an HR team test whether their own ATS has the same problem? Cheapest method: take the resume of a known top performer already on staff, run it through the ATS ten times, log every score. If the spread (max − min) exceeds 15 points, your ATS exhibits the same magnitude of AI resume screening inconsistency Dan documented.
Q4. What does the EEOC require for AI hiring tools as of 2025? Under EEOC public guidance, employers using AI to make employment decisions must demonstrate the tool does not cause disparate impact. AI resume screening inconsistency inherently violates the principle that selection criteria must be "reliable and reproducible" — two different scores on the same resume means the selection criterion itself is moving. BLS reinforces this in HR specialists' Important Qualities section, calling out compliance with federal, state, and local regulations.
Q5. We can't afford to replace our ATS. What's the lowest-cost mitigation we can do this week? Two free actions: (1) Convert your hard cutoff into a ±10-point human-review band. (2) Retain every rejected candidate's resume + AI score + human reviewer note for at least 12 months. The first reduces false rejections immediately; the second is your EEOC defense if a future complaint surfaces.
Act Before the Next Hiring Cycle: Turn AI Resume Screening Inconsistency Into a Measured Variable
If your HR team is still filtering on a hard AI score cutoff tonight, pull your engineering partners into a 30-minute room tomorrow and put the five steps on the whiteboard. AI resume screening inconsistency doesn't disappear by being ignored — but it becomes governable the moment you measure it. The BLS data is clear: this is a 944,300-job profession that earns its $72,910 median by exercising judgment. Don't outsource that judgment to a two-line open-source prompt.