Org Design

The HR Handbook Your AI Agents Never Got

AI agents are joining 2026 growth teams as headcount. The HR handbook they need: job descriptions, scorecards, escalation paths, reviews, and firing criteria.

On this page

Your newest team member started Monday. It has no LinkedIn profile, skips every standup, and will cheerfully work through the holidays. It also told a prospect your pricing was 40% lower than it actually is, twice, before anyone noticed.

Welcome to the 2026 growth team, where the roster is mixed and the handbook hasn't caught up. In Salesforce's April–May 2026 survey of 100 UK and Ireland CMOs, 77% said AI agents will require a fundamental redesign of the marketing department. Meanwhile McKinsey's State of AI research found 23% of organizations already scaling agentic AI in at least one function, with another 39% experimenting. The deployment question is settled. The management question is wide open.

Here's the uncomfortable part: Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. Read that list again. Those are not model failures. Those are management failures. Nobody wrote the job description, nobody set the KPIs, and nobody defined what firing looks like.

So let's write the handbook.

Write the job description before the system prompt

Every failed agent deployment we've audited shares a birth defect: the agent was scoped like a feature ("add AI to support") instead of hired like a role ("triage tier-1 billing questions, escalate everything involving refunds over $200"). A feature has requirements. A role has a charter.

The charter is one page, and it forces the arguments you'd rather skip:

  • Role title and single job. One agent, one job. "Marketing agent" is not a job. "Draft first-pass ad variants from the winning-creative brief" is.
  • Inputs and outputs. What it reads, what it produces, in what format, delivered where.
  • Tools and permissions. Every API, every credential, scoped to least privilege.
  • Non-goals. The section everyone skips and the one that matters most. What this role explicitly does not do: touch pricing, contact customers directly, publish without review.
  • Hiring manager. A named human accountable for this agent's output. Shared ownership is no ownership.

If you're fuzzy on where the line sits between a bot that answers and an agent that acts, that distinction is the whole hiring decision; we broke it down in chatbots vs AI agents. You give job descriptions to things that do work, and FAQ scripts to things that answer questions.

people sitting on chair in front of table while holding pens during daytime
Photo by Dylan Gillis on Unsplash

Onboarding week is a probation period

You wouldn't give a new hire the production database and the company card on day one. Agents get exactly that, constantly, and it's why the incident channel fills up by Thursday.

Run onboarding as graduated autonomy:

  1. Shadow mode (week 1–2). The agent produces outputs; humans do the actual work. You're building an evidence base, comparing agent drafts to human decisions on identical inputs.
  2. Supervised mode. The agent acts, but every action requires human approval. Touch rate is 100% by design.
  3. Bounded autonomy. The agent acts alone inside explicit thresholds (dollar limits, confidence floors, approved channels) and escalates everything else.

Onboarding also means context. Brand voice docs, the current pricing sheet, the escalation directory, the compliance no-fly list. An agent hallucinating your pricing is usually an agent that was never handed the pricing. The Salesforce study found 69% of marketing leaders say hallucination risk makes agent deployment harder in marketing than in back-office functions, precisely because the output is customer-facing. Probation ends when the agent hits pre-agreed exit criteria on the scorecard below, with the criteria agreed before day one.

This is also where most teams discover their real blocker is plumbing: fragmented data, no audit logging, credentials nobody can rotate. Our agentic AI automation work usually starts there, in the unglamorous access layer, before any agent gets a badge.

The agent scorecard

Performance reviews need numbers, and the wrong numbers are worse than none. "Messages handled" is the agent equivalent of measuring a salesperson by calls dialed. Score outcomes.

Quarterly agent scorecard (illustrative thresholds)
MetricWhat it measuresGreenYellowRed (PIP)
Task success rate% of tasks completed to spec with zero human rework≥ 90%75–90%< 75%
Escalation precision% of escalations a human reviewer agrees were necessary≥ 80%60–80%< 60%
Missed escalationsHigh-stakes actions taken where the charter required handoff01 per quarter> 1 per quarter
Error / brand-safety rateFactual or policy errors per 100 shipped outputs< 11–3> 3
Cost per completed taskFully loaded model + tooling cost vs. human baseline≤ 30% of human cost30–70%> 70%
Human touch rate% of outputs edited before shippingFalling month over monthFlatRising
Thresholds are starting points; calibrate per role and risk level. Cost baselines should use fully loaded human cost for the same task.

Two notes on reading it. First, cost per completed task is the metric finance will actually respect, and it's the one that decides renewals; pressure-test yours with our AI ROI calculator before the budget review, because token costs at scale surprise almost everyone. Second, human touch rate is directional by design. A mature agent's touch rate falls. A flat line after month three means the agent stopped improving or the humans stopped trusting it, and both are review topics.

Escalation paths, or teaching the agent to raise its hand

The best junior hires know what they don't know. Agents have to be told, numerically.

An escalation policy defines the triggers that force a handoff: confidence below a floor, dollar value above a ceiling, detected negative sentiment, legal or medical topics, anything the agent has never seen before. Then it defines where the handoff goes, to which named human, in which channel, with what context attached. An escalation that arrives as a bare transcript is a resignation letter in disguise.

Grade escalation in both directions. Under-escalation is the scary one, the agent issuing a refund it had no authority to issue. Over-escalation is the sneaky one: an agent that escalates 60% of tasks is a very expensive routing rule, and its cost-per-task line will eventually say so. Support teams have run this exact playbook on human agents for decades, which is why support automation is where the discipline is most mature; we covered how the data side feeds it in how audience data powers B2B customer support automation.

The performance review runs weekly, then monthly

Human reviews are annual because humans change slowly. Agents change every time you touch a prompt, swap a model version, or let the context drift. So the review cadence is weekly during probation, monthly at steady state, and immediately after any model or prompt change.

The mechanics look pleasingly familiar to anyone who has run a real review:

  • The evidence file. A frozen "golden set" of representative tasks, rerun after every change. If task success on the golden set drops, the change rolls back. No exceptions for exciting new models.
  • The work samples. The owner brings five random transcripts, unfiltered, to a 30-minute review. Reading raw transcripts is the single highest-signal management act available; summaries hide the weird stuff.
  • The outcome. Promotion means expanded permissions or higher thresholds. Demotion means tightened ones. Either way it's documented, because next quarter someone will ask why this agent can approve $500 refunds.

If this sounds like a heavier version of the QA you already run on lifecycle campaigns, it is, and deliberately so. Agents are the next chapter of the same discipline we mapped in what marketing automation means in the AI era: the workflows got smarter, so the governance has to match.

40%+of agentic AI projects Gartner expects to be canceled by end of 2027, mostly for cost, value, and risk-control failures

Firing an agent is a process you write before hiring it

Here's where the Gartner number stops being trivia and starts being a mirror. Projects rarely get canceled cleanly; they get abandoned expensively, after months of quiet underperformance nobody was accountable for measuring. The fix is to define offboarding criteria at hire time, in the charter, when nobody is emotionally invested:

  • Two consecutive red-zone reviews on the scorecard.
  • Any compliance or brand-safety incident without a credible remediation.
  • Cost per completed task exceeding the human baseline for a full quarter.
  • The workflow itself being retired (agents die with their jobs, and that's fine).

Offboarding itself is boring, which is the point: revoke credentials, archive logs for audit, run a post-mortem, update the charter template with what you learned. Sometimes the "termination" is really a reorg, splitting one overloaded agent into two narrower ones. Demand for the work keeps growing either way; Digital Commerce 360 reports Salesforce data projecting AI agent use to grow 67% by 2027. The teams that can fire agents confidently are the teams that can hire them aggressively.

Someone has to manage the roster

All of this implies a role that barely existed in 2024: the agent manager. In practice it's an ops-minded senior marketer who owns the charter library, chairs the reviews, holds the credential inventory, and answers one question on demand: which agents can do what, and who approved it?

Span of control is real here. One competent owner can genuinely manage six to ten narrow agents; past that, the transcripts stop getting read, and unread transcripts are how you end up explaining a rogue discount to the CFO. When headcount planning season arrives, agents belong in the plan, with a line for their management overhead, exactly like contractors.

If you're not sure whether your team could pass its own audit, our free AI readiness scorecard is a blunt first look at where the gaps are, from data access to governance.

The strange truth of the mixed roster is that none of this is new management theory. Job descriptions, probation, scorecards, escalation, documented exits: it's the ordinary machinery of running a team, applied to workers who happen to be software. The companies in Gartner's doomed 40% skipped the machinery. The 77% of CMOs redesigning their departments are, mostly, building it. Print the handbook. The agents will never read it. Every human who manages one should.

Sources

Frequently asked questions

Do AI agents really need job descriptions?
Yes, and skipping them is the most common failure mode. A one-page role charter defines the agent's single job, its inputs and outputs, the tools and permissions it holds, its explicit non-goals, and the named human accountable for its output. Gartner attributes most agentic AI project cancellations to unclear business value and weak risk controls, both of which trace back to undefined scope. The charter forces those decisions before launch, when they are cheap to make.
What KPIs should an AI agent be measured on?
Score outcomes rather than activity. The core scorecard covers task success rate (completed to spec with zero rework), escalation precision (were handoffs actually necessary), missed escalations, error and brand-safety rate per 100 outputs, cost per completed task versus the human baseline, and human touch rate over time. Volume metrics like messages handled reward busyness. A useful benchmark: a mature agent's touch rate should fall month over month while its cost per task stays well under the human equivalent.
Who should manage AI agents on a marketing team?
A named agent manager, usually an ops-minded senior marketer, who owns the charter library, chairs performance reviews, maintains the credential inventory, and can answer which agents can do what and who approved it. Span of control matters: one owner can realistically manage six to ten narrow agents before transcript review quality collapses. Shared or absent ownership is the fastest route into the 40% of projects Gartner expects to be canceled.
When should you retire or fire an AI agent?
When it hits offboarding criteria you wrote at hire time: two consecutive red-zone scorecard reviews, a compliance or brand-safety incident without credible remediation, cost per completed task exceeding the human baseline for a full quarter, or the underlying workflow being retired. Offboarding means revoking credentials, archiving logs for audit, and running a post-mortem. Sometimes the right move is a reorg, splitting one overloaded agent into two narrower roles.

Free tools for this topic

FREE TOOLAI Brand Visibility MonitorDoes ChatGPT recommend you — or your competitor?CALCULATORAI & Automation ROI CalculatorPut a payback date on every automation idea.FREE TOOLAI Readiness ScorecardTwelve questions. Your automation roadmap, scored.

Keep reading

ComparisonsChatbots vs AI Agents: Answering vs DoingRead →GlossaryWhat Is Marketing Automation in the AI Era?Read →DataHow Audience Data Powers B2B Customer Support AutomationRead →
CATALIST NEWSLETTER

Monthly dose of growth marketing.

Get marketing tips, narratives, guides, and playbooks delivered to your inbox.