Published: August 13, 2026
Updated: August 13, 2026
Every founder who is about to hire their first assistant now asks a question that did not exist a few years ago: "Do I actually need a person for this, or can AI just do it?"
It is a fair question and almost nobody answers it honestly. The AI vendors tell you the person is obsolete. The staffing agencies tell you the software is a toy. Both are selling, and both are answering at the wrong level. "Assistant" is not a task. It is a bundle of thirty or forty different tasks that happen to have been given to one person for historical reasons, and the bundle contains work with wildly different properties. Some of it is now genuinely better done by software. Some of it will not be done by software for a long time, and a few pieces of it should never be, for reasons that have nothing to do with capability.
So the useful question is not "AI or a person?" It is "which of these thirty tasks is which?" This guide gives you a test you can run on any individual task in about a minute, sets out what the research actually measures (which is narrower than the headlines suggest), and describes the split that works in practice for a small team.
It is written for employers. If you have not yet worked out how many hours of work you are actually holding, the virtual assistant hours calculator is a better first stop than this article, because the answer to "AI or human?" changes depending on whether you have four hours a week of work or thirty.
The short version
- Specification is the dividing line, not difficulty. AI handles hard tasks you can fully describe far better than it handles easy tasks you cannot. The test is not "is this complex?" but "could I write down every step?"
- Verification cost decides whether automation actually saves you anything. If checking the output takes as long as producing it, you have moved the work rather than removed it.
- Accountability does not transfer to software. A tribunal has already ruled that a company owns what its chatbot told a customer. Delegating a task to AI does not delegate the liability.
- The measured gains are real but concentrated. The best evidence on AI in a support setting found a 14% average productivity gain, with most of it going to novices and close to none to experienced staff.
- The winning pattern is not either/or. It is a person who runs the loop and uses AI inside it. That person needs different qualities than the assistant you would have hired before, and you should hire for those qualities deliberately.
What the research actually measures
Before the framework, it is worth being precise about the evidence, because the gap between what gets measured and what gets claimed is where most bad hiring decisions come from.
Autonomous task length is rising fast, on a narrow slice of work. The research group METR tracks what it calls the time horizon of frontier models: the length of task, measured by how long a skilled human needs, that a model can complete on its own with 50% reliability. That horizon has been doubling roughly every seven months for several years, which is a genuinely steep curve and worth taking seriously.
The caveats METR itself publishes matter more than the curve for a founder making a hiring decision. The tasks are drawn largely from software engineering, machine learning and cybersecurity. They are, in METR's words, "self-contained and well-specified" with "clear success criteria that can be automatically evaluated." METR states plainly that "most jobs are not composed of well-specified, algorithmic tasks" and instead "tend to require interacting with other people and involve success metrics that cannot be algorithmically scored." It also notes that the measurements reflect what someone with low or no prior context could do, comparable to a new hire or a freelance contractor, rather than an experienced person who knows your business.
Read that carefully and it is close to a definition of the work an assistant actually does. Chasing a supplier who has stopped replying, deciding which of two annoyed customers to call first, noticing that a calendar request from a particular investor should be accepted even though the calendar says no: none of these are self-contained, well-specified, or automatically scoreable.
Simple real-world assistant tasks are harder for AI than professional exams. The GAIA benchmark was built to test exactly this. It poses questions that a competent human nearly always gets right but that require stringing together browsing, reading a file, and following several steps. When it was published, humans scored 92% and the best AI system with plugins scored 15%. Scores have risen a great deal since, and any specific number is out of date quickly, but the design point holds: models that pass professional exams stumble on conceptually simple multi-step errands. The failure is rarely a lack of intelligence. It is a step skipped, a file not opened, a tool not used.
The productivity gains are real, and unevenly distributed. The strongest study on AI in an assistant-shaped job is Generative AI at Work by Brynjolfsson, Li and Raymond, which followed 5,179 customer support agents through a staggered rollout of an AI conversational assistant. Access raised issues resolved per hour by 14% on average. That average conceals the finding that matters: a 34% gain for novice and lower-skilled workers, and minimal impact on experienced and highly skilled ones. The tool was spreading the practices of the best agents to the newest ones. It was not replacing the best agents, and it had little to give them.
If you are hiring one assistant, that result should change what you look for. AI compresses the gap between a weak assistant and an average one. It does very little to the gap between an average one and a genuinely good one, which means the value of hiring well went up, not down.
Adoption is lower than the discourse implies. The Census Bureau's Business Trends and Outlook Survey, which asks a large probability sample of firms whether they used AI in producing goods or services, put overall use between 17% and 20% across panels running to early May 2026, with a national average of 19.8%. Use rises with headcount: about 37% at firms with 250 or more employees, 32% at 100 to 249, and under 20% at firms with fewer than 20 people, a figure that had not moved much over the period. Information sat at 39.7% and finance and insurance at 33.9%, while retail trade was near 14%.
Take that as calibration rather than as permission to ignore the technology. If you run a small firm and have not automated much, you are not behind your peers. You may still be leaving something on the table.
The four-part test
Run this on one task at a time. It takes about a minute per task and it is more reliable than any general opinion about what AI can do, because it asks about your task rather than about the technology.
1. Can you specify it completely?
Could you write down every step, every input, and every rule for the edge cases, such that a stranger following the document would produce what you want? Not "could you explain it to a smart person," which invites them to fill gaps with judgment. Could you write it down.
If yes, it is a strong automation candidate regardless of how technical or tedious it looks. If you find yourself writing "and then use your judgment" or "it depends on the client," you have located the exact point where a human is doing something you have not noticed you were paying for.
This is why the intuition of "give AI the hard stuff, give people the easy stuff" gets the answer backwards so often. Reformatting a 300 row spreadsheet against a complicated rule set is hard and fully specifiable. Deciding whether a customer's complaint is the kind that needs a refund or the kind that needs a phone call is easy for a person and not specifiable at all.
2. What does it cost to check the output?
Automation only pays if verification is cheaper than production. Three cases:
- Self-evident errors. A draft email that reads wrong, a summary that missed the point, code that fails to run. You see the problem instantly. Automate freely.
- Expensive to check. A list of thirty prospect email addresses, a set of figures pulled from statements, a competitor pricing sheet. Every item looks equally plausible whether right or wrong, and confirming them takes about as long as gathering them did. Automate only with a sampling routine and a person who owns the result.
- Errors surface later, at the customer. Anything that goes out under your name without a person reading it. This is the category that costs real money, and the arithmetic is not about the average case. A process that is 97% right sounds excellent until the 3% is a wrong price quoted to a customer in writing.
Most people underprice this test badly. If you are checking the AI's work carefully, you have not saved the hour. You have converted an hour of doing into an hour of reviewing, which is often more tiring and is the thing you were trying to hand off in the first place.
3. Who is accountable when it is wrong?
This is settled enough to cite. In Moffatt v. Air Canada, the British Columbia Civil Resolution Tribunal heard a claim from a passenger who had asked Air Canada's website chatbot about bereavement fares, been given incorrect information about how to claim the discount, relied on it, and been refused the refund. Air Canada argued, in effect, that the chatbot was responsible for its own statements. The tribunal rejected this and found the airline liable for negligent misrepresentation, holding that Air Canada was responsible for all the information on its website whether it came from a static page or a chatbot, and that it had failed to take reasonable care that the information was accurate. The passenger was awarded $650.88.
The sum is small and the principle is not. Handing a task to software does not hand over the consequences. If a task carries the ability to commit you to something, quote a price, promise a delivery date, state a policy, agree a term, then a named person should own the output even if software drafted it. That is not an argument against using AI for the drafting. It is an argument that "the bot said it" is not a defence you have.
4. Does it require chasing a human?
A large fraction of what an assistant does is applying friction to other people. The supplier who has not sent the invoice. The candidate who went quiet after the second interview. The client who needs to be told the thing they do not want to hear, in the right tone, at the right time.
This work is judged on outcomes with no clean specification. It requires reading a relationship, choosing the moment, escalating carefully, and knowing when persistence turns into damage. It is also the work founders most want off their plate, because it is emotionally taxing rather than intellectually hard. It is close to the last thing you should try to automate, and it is the single best argument for hiring a person.
How this sorts a real assistant's work
Applying the four tests to the tasks that actually make up an assistant role gives a fairly consistent split.
| Task | Best owner | Why |
|---|---|---|
| Transcribing and summarizing calls | AI | Fully specifiable, errors are visible, no external commitment |
| First-draft emails and documents | AI drafts, human sends | Cheap to produce, but sending commits you |
| Reformatting and cleaning data | AI with spot checks | Specifiable, but verification is expensive at volume |
| Research briefs on public information | AI drafts, human verifies | Plausible wrong answers are the standard failure |
| Routine scheduling within clear rules | AI | Specifiable once the rules are written down |
| Scheduling that involves priority calls | Human | "Is this person worth moving the board meeting for?" is not specifiable |
| Inbox sorting and triage | AI sorts, human decides | Categories are learnable, consequences of a miss are not symmetric |
| Tier one customer questions | AI with a human path out | Volume suits automation, but the escalation route has to be real |
| Complaints and refund decisions | Human | Accountability, tone, and commitment all present |
| Chasing suppliers and late payers | Human | Persistence and relationship judgment, no specification |
| Recruiting coordination and candidate care | Human | Your employer brand is being formed in these messages |
| Bookkeeping entry and reconciliation prep | AI drafts, human reviews | Specifiable, but errors surface late and cost money |
| Social posting to a set calendar | AI drafts, human approves | Published under your name |
| Managing a project across several people | Human | Almost entirely interacting with humans, no scoreable success metric |
| Building and maintaining the automations above | Human | This is now part of the job, and it is the highest leverage part |
The pattern in that middle column is the actual finding. The most common right answer is not "AI" or "human." It is "AI does the production step, a person owns the judgment step and the send button." Very few tasks are pure.
The pattern that works: a person who runs the loop
The teams getting real leverage are not choosing between an assistant and a set of tools. They hire one capable person and expect that person to use AI heavily inside their own work. The shape looks like this:
- The person owns an outcome, not a task list. "Inbox at zero by 10am with anything needing me flagged" rather than "read emails."
- They automate their own repetitive work. The assistant who notices they format the same report every Monday and builds a way to stop doing it by hand is producing compounding value. They are also the person best placed to notice, because they are the one doing it.
- They stay accountable for output quality. Nothing goes out unread because software wrote it. The name on the work is theirs.
- They escalate the genuinely ambiguous cases to you. This is the part that never automates and the reason the role exists.
This is roughly what the Brynjolfsson study observed from the other direction. The tool made newer agents behave more like experienced ones. It did not make the experienced ones unnecessary, and it gave them relatively little. If AI raises the floor of assistant performance, then the value of a hire sits in the ceiling: judgment, ownership, and the willingness to tell you something you do not want to hear. Those are the qualities to interview for.
Which means the practical change is in the hiring bar, not the hiring decision. Typing speed, formatting ability, and raw task throughput matter less than they did. Judgment, written communication, initiative, and comfort with tools matter more. If you are writing a role description, our job description generator and interview questions tool are built around competency rather than task lists, which is the right emphasis for this.
Where AI genuinely wins, without qualification
To be even-handed, there are categories where a person is now the wrong answer and holding on is just cost:
- Volume transcription and summarization. Paying an hourly person to write up call notes is difficult to justify.
- First drafts of anything. The blank page problem is solved. A person editing a decent draft is faster than the same person starting cold.
- Translation and tone adjustment. Good enough for internal use and for a first pass on external copy.
- Structured extraction from documents. Pulling fields out of invoices, contracts and forms, with sampling.
- Round-the-clock acknowledgement. A customer at 2am getting an instant, accurate, honest "we have this and a person will reply by 9am" beats silence, provided the promise is kept.
- Anything you were never going to hire for anyway. The work that simply was not getting done. This is the largest and least discussed category of genuine gain.
Note what unites them. Every one is either specifiable, self-checking, or purely internal. None commits you to anything in front of a customer without a person reading it first.
Doing the cost math honestly
The comparison people make is a software subscription against a monthly salary, which flatters the software by leaving out its real costs. A fair comparison includes:
- Setup and maintenance. Automations are not free once built. They break when a tool changes, and someone has to notice and fix them.
- Your review time. If you are checking output, price your own hour into the total. This is usually the largest hidden cost and the one that makes founders feel that automation did not help.
- The failure tail. Not the average error rate but what the worst plausible error costs. One mispriced quote or one badly handled complaint can exceed a month of salary.
- What does not get done. The chasing, the relationship work, the noticing. If nobody owns it, it does not happen, and the cost shows up later as churn or a supplier problem.
Run the same discipline on the hiring side and compare like for like. The cost calculator gives a fully loaded monthly figure for a dedicated assistant, and the ROI calculator takes the hours you would hand over and what your own time is worth. For the market rates behind those numbers, how much a virtual assistant costs breaks the pricing down by country, role and hiring model.
For most small teams the honest answer is that the two budgets are not competing. Software handles the volume floor, a person handles the judgment ceiling, and the combined cost is lower than either doing the whole job alone would be.
What to do this week
- List the tasks, not the role. Write down every recurring thing you want off your plate. Aim for twenty or more items, specific enough to argue about.
- Run the four tests on each. Specifiable? Cheap to verify? Does it commit you to anything? Does it require chasing a person? Mark each one automate, delegate, or split.
- Automate the top three automate items first. They are usually reporting, transcription and first drafts. Do this before you hire, because it changes what you hire for.
- Total the hours left in the delegate and split columns. That is your actual role. The hours calculator converts it into a weekly commitment and an engagement shape.
- Hire for the judgment tasks, not the volume tasks. The volume is going to shrink. The judgment is not.
If that list comes to ten or more hours a week of work that failed the specification test, you have a genuine role and no amount of tooling will cover it. That is the point to talk to us. You can request candidates and see profiles matched to the judgment work you have identified, or book a meeting and walk through your task list with a recruiter who can tell you which parts of it we would not staff. If you are still working out which country to hire from, the Philippines versus South Africa comparison covers the trade-offs, and tasks to delegate to a virtual assistant is a useful prompt if your list came up short.
The founders who get this right are not the ones who picked the correct side of an argument. They are the ones who stopped treating "assistant" as a single indivisible thing, sorted the work honestly, and then hired a good person for the half that was never going to sort itself.