Why most virtual assistant reviews go badly
The usual pattern is familiar. You hire someone, the first few weeks feel busy and vaguely positive, and then somewhere around month three you get an uneasy sense that you are not getting what you expected. You cannot point at a number, because there is no number. So the review turns into a conversation about impressions, the assistant leaves it unsure what to change, and the same uneasy feeling comes back a month later. Eventually one of you ends the arrangement and both of you conclude that offshore hiring did not work, when what actually failed was the absence of a scorecard.
The underlying problem is clarity, and it is not unique to remote teams. Gallup's 2025 measurement of U.S. employee engagement put engagement at 31%, down from a peak of 36% in 2020, and found that clarity about what is expected had fallen nine percentage points over the same period. Their reading of the open-ended responses was blunt: a large share of employees cannot say what excellent performance in their own role would actually look like. If that is the norm for people sitting in the same building, it is the default for someone you have never met in person, working from another continent, whose only view of your expectations is what you happened to say on a call.
A scorecard fixes that at the cheapest possible moment, which is before the work starts. It converts a vague hope into four to six numbers with targets attached, gives the assistant something to self-correct against without waiting for you, and gives you a basis for saying either "this is working, let us expand scope" or "this specific measure is short, here is the plan" instead of a conversation neither of you enjoys.
What to measure, and what to leave alone
The single most common mistake is measuring activity instead of outcomes. Hours logged, screenshots captured, and keystrokes counted all tell you that a chair was occupied. None of them tell you whether the work was correct, whether it was on time, or whether a problem was raised before it became expensive. They also carry a real cost: a remote working relationship runs on trust, and surveillance is a withdrawal from that account every single day.
A better test for any candidate metric is whether you could produce the number today, from a tool you already pay for, without asking anyone to do extra admin. Your project tool knows what was due and what was late. Your help desk knows response and resolution times. Your CRM knows how fast a lead was contacted. Your accounting software knows when the close landed. If a metric needs a new spreadsheet that someone has to maintain by hand, it will be abandoned by the second month.
These five apply to almost every role. The generator above adds the role-specific ones on top.
| Metric | How to measure it | The trap |
|---|---|---|
| On-time completion | Assigned tasks finished by the agreed due date, from your project tool | Counting tasks that were never given a due date in the first place |
| Rework rate | Completed work you had to send back, as a share of everything completed | Fixing it yourself quietly, which keeps the number clean and the problem alive |
| Response time | Median reply time to you during their working hours, not yours | Measuring across your night hours and calling their overlap gap a delay |
| Escalation quality | Share of blockers raised with a proposed option attached | Rewarding silence, so problems only surface once they are expensive |
| Documentation coverage | Recurring tasks they own that have a current, followable SOP | Counting documents that exist but have not been updated in six months |
Keep the total between four and six. Ten metrics is not a scorecard, it is a monitoring system, and nobody improves against ten numbers at once. If a measure has not changed a decision in two quarters, drop it.
Measure response time inside their hours, not yours
This one detail causes more unfair reviews than any other. If you sit in New York and your assistant works South African hours, a message you send at 6pm your time arrives at midnight theirs. Measured naively, every evening message looks like a twelve-hour delay and the assistant's response time looks terrible while their actual behaviour is fine. Measure the clock only while their shift is running, and agree separately and explicitly what happens to anything that arrives outside it. The time zone overlap calculator will show you how many working hours you actually share, which is the honest denominator for this metric.
Where the work is genuinely time-sensitive, anchor the target to outside evidence rather than instinct. SuperOffice's benchmark study of 1,000 companies found an average customer service response time of 12 hours and 10 minutes, that 62% of companies never replied at all, and that only 20% resolved the question fully on the first reply. Those numbers are worth knowing before you set a support target, in both directions: a four-hour first response inside a covered shift is already better than most of the market, and a target of five minutes on email is a target nobody will hit sustainably.
How often to review, and what belongs in each review
Frequency matters more than formality. Gallup found that 61% of employees are engaged where feedback and recognition from a manager arrive at least once a week, compared with 38% where weekly feedback arrives without regular recognition, and that only about one in four employees strongly agree they receive genuinely valuable feedback from the people they work with. The lesson for a small team is not that you need an HR process. It is that fifteen honest minutes every week beats an hour every quarter, and that the weekly check-in should include something specific that went well, not only the list of what slipped.
| Cadence | What it covers | Why it exists |
|---|---|---|
| Weekly, 15 minutes | Response times, blockers, anything that slipped, one thing that went well | Catches drift while it is still one week of drift and not one quarter of it |
| Monthly, 30 minutes | Output metrics: completion rates, resolution rates, accuracy, volume | A month is the shortest window where output numbers stop being noise |
| At 30 days | Ramp check: is the role clear, are the tools working, what is still confusing | The last point where a bad start can be fixed cheaply |
| At 90 days | Full scorecard, scope confirmation, and a pay or hours conversation if earned | The point where you decide whether this is a long-term hire |
| Quarterly after that | Scored review against the previous quarter, plus growth and next scope | Keeps a good hire growing instead of quietly stalling |
Send the scorecard to the assistant before the review, never during it. Someone who sees a target for the first time in the meeting where they are being scored against it has been set up to fail, and they know it. Sending it ahead also changes what the meeting is for: it stops being a verdict you deliver and becomes a conversation about two or three specific gaps that both of you already knew were coming.
Setting the 30-day bar differently from the 90-day bar
The generator uses two different sets of targets on purpose. At 30 days someone is still learning your systems, your customers, and the unwritten rules of how you like things done. The right thing to reward at that stage is reliability and honest flagging: work arriving when promised, questions asked rather than guesses made, confusion surfaced early. An assistant who is a little slower but tells you when something is unclear will cost you far less over a year than one who is quick and quietly wrong.
By 90 days the bar moves to ownership. The question is no longer whether they can do the tasks but whether the outcome happens without you thinking about it. That is why the 90-day targets are written as outcomes rather than activities: the close lands in five working days, the queue is clear at shift end, overdue invoices get chased without a reminder from you, conflicts are caught before they reach your calendar. If you are still the person noticing the problem at day 90, the score is not a four regardless of how much activity you can see.
Handled well, that ninety-day mark is also the moment to talk about money and scope. A strong hire who has hit their targets and hears nothing about pay or growth starts looking around, and replacing them costs you the whole ramp again. Do not let that conversation drift because it is uncomfortable.
When the scorecard comes back short
A missed target is information, not a verdict. Before concluding anything about the person, check the three explanations in order, because the first two are far more common than most founders expect. Was the expectation ever actually clear, in writing, with a number attached? Is something structural in the way, such as missing access, a broken tool, a dependency on someone who never replies, or an overlap window too small to do the job? Only when both of those are clean is the third explanation, fit, worth considering.
If it is genuinely fit, be direct and quick. Name the measure, name the gap, write a plan with a date no more than 30 days out, and say plainly what happens if it is missed again. That is a fairer thing to do than another vague quarter, and it is the only version of the conversation that gives the person a real chance to fix it. On a managed plan with Cherry Assistant you are not doing this alone: if a match is not working, we help you replace it, and you pay nothing if you do not hire.
Use this with the rest of the hiring toolkit
A scorecard is the last step in a chain, and it works best when the earlier steps line up with it. Scope the role first with the job description generator, screen against that scope with the interview questions generator, and put the terms in writing using the contract template generator. Once someone starts, the onboarding checklist generator sequences the first 90 days and the SOP generator turns each process they learn into something the next hire can follow. The targets on this page then measure exactly the scope those tools defined, which is the whole point: a scorecard that measures something the job description never mentioned is just a trap.
To size the decision itself, the cost calculator compares a dedicated assistant with an in-house hire, the ROI calculator puts a number on the hours you reclaim, and the salary comparison shows what the same role costs across twelve markets. Browse the roles we source, the industries we support, and common use cases to see how other teams scoped the work before they hired.
Sources
The engagement and feedback figures cited above come from Gallup's published workplace research, and the customer service response benchmarks come from SuperOffice's study of 1,000 companies. The targets in the generator are our own, drawn from the roles we place and review every week, and they are meant to be edited to fit your business.
- Gallup: U.S. Employee Engagement Declines From 2020 Peak (engagement at 31% in 2025, and clarity of expectations down nine points since 2020)
- Gallup: Organizations Can Redefine Feedback by Including Recognition (61% engaged where feedback and recognition arrive at least weekly, against 38% for weekly feedback alone)
- SuperOffice Customer Service Benchmark Report (1,000 companies: average response time 12 hours 10 minutes, 62% never reply, 20% resolve on the first reply)
- U.S. Bureau of Labor Statistics: American Time Use Survey news release