AI Now Helps Write Performance Reviews. Somebody Still Has to Own the Verdict.
AI now sits inside the performance review before a manager ever opens the document. Companies are folding AI into how reviews get drafted, how ratings get calibrated, and how promotion recommendations get justified. A recent SHL survey of more than a thousand US workers found only 27 percent fully trust their employer to use AI responsibly in decisions like these, and 59 percent believe AI is making workplace bias worse, not better. That gap between what companies are adopting and what employees actually trust is not a communications problem. It is a leadership problem, and it is showing up in review cycles running right now.
Is AI Already Writing Performance Reviews?
In a growing number of companies, yes, at least in part. AI tools now draft review language from a manager’s notes, summarize a year of Slack messages and project tickets into a “performance narrative,” and flag outliers for calibration meetings. Some of this genuinely helps. A manager who used to spend a weekend staring at a blank text box now gets a starting draft. The risk is not that AI writes a first draft. The risk is when nobody edits it with real judgment, and the draft becomes the verdict.
Research on AI-assisted performance management points both ways at once. Some studies show measurable drops in inconsistency when AI standardizes how ratings get applied across a team. Other research is blunter about the danger: AI systems trained on years of past reviews can encode the same biases that were already sitting in those reviews, favoring familiar communication styles, constant visibility, and work patterns that photograph well in a dashboard. The tool does not invent bias from nothing. It launders whatever bias was already there, and hands it back with the confidence of a data point.
Where the Bias Actually Comes From
Here is the part worth sitting with. Most AI performance tools are only as good as what they can see, and what they can see is activity, not contribution. Messages sent. Tickets closed. Meetings attended and spoken in. That is a proxy for visibility, not for value.
The Spotlight Machine’s whole case rests on this exact gap. The work that keeps an organization running is usually the quietest work in the building: the person who prevents the fire nobody sees, the employee who fixes a problem before it ever reaches a dashboard, the veteran who does not narrate their process because they are too busy actually doing it well. An AI system summarizing “signals of impact” from a digital footprint will miss every one of them, and it will miss them with total confidence, because nothing in its training data taught it to look for the work that leaves no trace.
Layer a promotion or a rating on top of that blind spot, and the unfairness compounds. The loudest, most visible contributor gets read as high-performing. The quiet, essential one gets read as average, not because the work was average, but because the tool was never built to see it. AI did not create this bias. It just gave an old, human blind spot a spreadsheet and called it objective.
What Fair Leadership Actually Requires
Never let AI be the final signature. Use it to draft, summarize, or flag for a conversation, never to close the loop alone. If a rating or a promotion recommendation comes out of an AI process without a human who can explain the reasoning in plain language, that is not a review. It is an unexplained decision wearing a review’s clothes.
Ask what the tool cannot see before you trust what it did see. Before a calibration meeting, name the work in your building that would never show up in a chat log or a ticket count. If your AI-assisted process has no way to surface that person, the process has a hole, not the employee.
Separate visibility from value on purpose. A person who talks about their work constantly and a person who quietly does excellent work are not the same, and a tool trained on activity data will consistently prefer the first one. A manager’s job is to correct for that, every cycle, not to let the dashboard do the correcting.
Say out loud where AI is involved. Employees trust the process more, not less, when they know exactly where a machine touched their review and where a human made the actual call. Silence about the tool’s role is what erodes the 27 percent trust number further. Naming it plainly is what starts to rebuild it.
Audit who keeps landing in the middle of the curve. If the honest answer, cycle after cycle, involves people who do not self-promote, do not narrate their AI use in meetings, or simply do their job without performing it for an audience, that is a pattern an AI tool will not flag for you. Only a leader who is actually looking will catch it.
The Tension Worth Naming
Aziz Aghayev has spent enough rooms with leaders to know how this sounds, coming from someone who teaches AI for a living: use the tool to save the manager time, and refuse to let it decide who gets seen. That is not a contradiction. It is the whole point. AI is supposed to serve a human outcome, in this case a fair and accurate read on a person’s contribution. The moment a company lets an activity dashboard stand in for that judgment, the relationship has flipped. The technology is now deciding who gets recognized, instead of helping a manager recognize the right people faster and more consistently.
That is a leadership decision hiding behind a software feature. Nobody voted to let “generates a lot of visible AI-summarized activity” become a stand-in for “does great work.” It happened by default, because visible activity was the easiest thing for a tool to measure, and nobody in the room stopped to ask what the tool was missing.
Before the Next Review Cycle Lands
The uncomfortable truth is that AI-assisted reviews are not a future risk to plan for. They are already running inside performance cycles this year, in the United States, the UK, and across Europe, often without employees fully understanding where the machine’s judgment ends and a human’s begins. Leaders who assume the tool is neutral are not protecting anyone from bias. They are just letting an old blind spot run at scale, with better formatting.
The fix is not refusing the tool. It is training the humans who sit above it to know exactly what it can see, what it cannot, and when to override it in favor of the quiet, essential work a dashboard will never learn to notice. That is a skill, and like any real skill, it can be taught deliberately instead of assumed.
If your leadership team is heading into review season without a clear, shared answer for where AI ends and human judgment begins, see how The Spotlight Machine’s training builds that standard before the next calibration meeting decides it for you by default.