We Fired an AI Coworker. Then We Hired It Back.
We run AI the way HR runs people: job descriptions, internships, probation, promotions, and sometimes terminations. What hiring and firing AI coworkers taught us.

In 2025, we fired a coworker named Piper.
Piper was a sales coach. The job was to help our sales teams prep faster, pitch sharper, and get up to speed on our lines of business without waiting on a human expert's calendar. Piper spent months in the intern stage and never earned a promotion. We iterated, retrained, went back and forth more times than I can count, and gave the role more chances than we probably should have. Eventually we did what you do when a hire is not working out. We let Piper go.
Piper was an AI.
This year, we hired for the same role again. The new hire is called DeX. DeX came out with flying colors, cleared every stage gate we put in front of it, and is now one of the most popular coworkers among our sales teams and consultants. Same job, essentially. Opposite outcome.
The easy explanation is that the technology got better between 2025 and 2026. It did, and it matters, and I will not pretend otherwise. But that is not the variable that flipped the outcome. What flipped it was management. Piper had a part-time sponsor. DeX has a dedicated team of supervisors who treat it like their newest hire.
That sentence sounds absurd unless you know how we run AI where I work. So let me walk you through it. How we hire artificial intelligence, how we promote it, and why we sometimes fire it.
Think of us as your hiring agency
I run the AI program at OneDigital, an insurance and benefits consultancy of about six thousand people. For the past year and a half, my team has run what we call the AI Coworker program: AI agents built for specific teams, doing specific jobs, managed by the people they work for.
We did not start with a framework. We started with the obvious pitch: we will use AI to help you do your job better. It did not land. What came back was fear, skepticism, and the quiet resistance you get when you ask people to step into a paradigm they do not recognize. Nobody wakes up hoping to be transformed by the technology team.
So we stopped pitching AI and started asking about the work. What does your team actually do? Where does the time actually go? If you could hire one more person tomorrow, what would you hand them first? And somewhere in those conversations we noticed that the answers kept arriving in the shape of a job description. The need was never "we want AI." The need was "we need someone who can do this."
So we changed the pitch to match the shape of the need. We stopped introducing ourselves as the AI team and told people: think of us as your hiring agency. When you have a job req, come to us first. We will try to fill that role with AI. We will help you hire an intern.
I deliberately did not pitch myself or my team as AI experts. The expert pitch puts the business in the audience. The hiring agency pitch puts them in the driver's seat. It is their req, their hire, and ultimately their call whether the hire is working out. People got to the aha moment much faster that way, and I have come to believe that repositioning did more for adoption than any capability the technology shipped that year.
An applicant tracking system for software
Once the hiring frame clicked, we took it literally. We went to our People and Culture team and asked them to walk us through the real end-to-end employment process. How a role gets requested and justified. How a job description gets written. What onboarding looks like. How performance reviews actually run. What a performance improvement plan is for. How promotions happen, and how someone gets let go. Then we modeled the AI Coworker program as closely to that as we could.
Not because the metaphor is cute. Because every single person in the company already knows how that system works. Nobody needs a training session on what an intern is.
So here is what happens today when a team leader at OneDigital wants AI help. A request comes in, and we write a job description together with the business owner, the same way a hiring manager would. The problem this role exists to solve. Who feels that pain, and when. The two or three tasks it will do repeatedly. What is explicitly out of scope. And, before we build anything at all, the question our whole framework now hangs on: what would a human hire in this role be measured on? If we cannot answer that, we do not build.
The coworker gets a resume. It gets hired as an intern and piloted with a small group of real users. If it clears its exit criteria, real usage, satisfaction above target, no critical incidents, evidence of actual value, it gets promoted to apprentice and rolled out to a wider group with a higher bar. If it keeps performing, it goes full-time: available to everyone it was hired for, embedded in real workflows, with a named operational owner and regular performance reviews. Underperform for long enough and it gets demoted back a stage for remediation, our version of a performance improvement plan. Fail that, and it gets retired.
In the earliest version of our tracker, one of the first pipeline stages was literally labeled "Candidate Requested." We were running an applicant tracking system for software. I remember finding that funny at the time. I no longer find it funny. I think it is the reason any of this worked.
The case against calling them coworkers
I know how this sounds to half the industry.
There is a genuine, ongoing debate about whether humanizing AI is wise, and the case against goes roughly like this: these systems are not people. Giving them names and job titles and performance reviews misleads users about what they are, invites misplaced trust, and gets weird fast. I take that critique seriously, because we did not adopt this framing casually, and I want to give it a straight answer.
The answer is that the coworker framing is not branding. It is change management infrastructure.
People do not know how to relate to an agentic workflow orchestration layer. They know exactly how to relate to a coworker. You onboard it. You train it. You give it feedback when it gets something wrong. You forgive an intern's mistakes in a way you would never forgive a production system's, and that forgiveness is precisely what an early AI deployment needs to survive its first rough month. You collaborate with a coworker. That is the verb we wanted, because collaboration is how this technology actually creates value: it augments and extends the person using it. I wrote before about the adoption version of this insight: an employee who suspects the tool is there to replace them uses it defensively, and an employee who sees it as the newest hire on the team trains it. The employment metaphor recruits every instinct people already have about working with other people, and points those instincts at the machine.
And once the metaphor is load-bearing, you cannot use it halfway. If it is a hire, it needs a job description. If it has a job description, it has KPIs. If it has KPIs, it can miss them. If it can miss them, someone has to be accountable for noticing, and that someone has to be its manager, not my team. Managers, plural, by the way: every coworker we deploy has a named human manager on the business side who truly has to supervise it the way they would a human employee. And if management is real, then firing has to be possible. The moment we allowed ourselves to say "hire," everything else followed. The framing is only honest if the whole employment arc comes with it.
That is the part I would defend hardest to the skeptics: we are not pretending the software is a person. We are borrowing the only management system everyone in the building already trusts, and applying it to a workforce that happens to run on servers.
Ben, or what good management looks like
Ben is the proof of what the arc looks like when it works.
Ben is our employee benefits expert, and probably the most famous coworker we have. Consultants ask Ben the questions that used to wait on the busiest experts on their team. Ben has handled well over a hundred thousand client assignments. Harvard has published case studies on our program, including one titled Building a Digital Workforce, and Ben features in the book two of our leaders have coming out later this summer. Which is a strange set of sentences to type about a colleague that runs on a server.
But none of that is why Ben worked. Ben worked because of Shelley.
Shelley is one of our benefits experts, and from day one she took Ben on the way you would take on a direct report. She was a champion who believed in the mission, and she was methodical and deliberate about the work of supervision. She watched what Ben got wrong. She fed back corrections. We iterated. You could see Ben improving with every cycle, the way you watch a sharp new hire compound in their first year. Ben cleared the intern gate, then the apprentice gate, and went full-time.
Here is the part people miss, and the part I now repeat in every internal conversation about this program: Ben went full-time and Shelley did not stop. She still runs workshops and training sessions. She still evangelizes Ben everywhere she goes. Being a coworker's supervisor is not a one and done activity. The full-time status is not a finish line. Just like a human employee, a coworker that stops being managed starts drifting, and a coworker nobody champions stops being used. Ben is famous because Ben is good, and Ben is good because Shelley never handed the job back.

Piper's autopsy
Which brings me back to the firing.
Piper's autopsy has two findings. The first is technological, and it is the one everyone expects. In early and mid 2025, the platform capabilities a sales-focused coworker needed simply were not there yet. We tried really hard. We kept iterating, kept going back, kept Piper in the intern stage while we searched for an angle that would make people happy with it. Some hires are ahead of what the organization, or in this case the technology, can support. That is a real cause of death and I will not minimize it.
The second finding is the one I find more instructive. Piper never had a Shelley. Sponsorship of Piper was a side gig: attention arrived in bursts and then vanished, feedback came from time to time instead of on a cadence. I want to be careful here, because this was a structural failure more than a personal one. Nobody's primary job was making Piper succeed, and so nobody was there for Piper the way Shelley was there for Ben. An AI coworker with an absent manager fails exactly the way a junior human hire with an absent manager does. Slowly, and then obviously.
So we retired Piper. And when we revived the role this year, one of my team members built DeX with that second finding in mind. The technology had improved, no doubt about that. But DeX launched with a dedicated group of supervisors who were absolutely bought into making it a success. They wrote a strategic rollout plan. They tested before they scaled. They ran dedicated feedback sessions with the sales teams. They built DeX around the way the sales team actually works, instead of asking the sales team to reorganize around DeX.
Same role. Two attempts. One variable flipped. If I could show you only one thing from the last year and a half, it would be this natural experiment, because it settles the question people keep asking backwards. The question is never "is the AI good enough." The question is "who owns making it good."
The termination that meant it was working
Not every firing is a failure story, though. My favorite one is the opposite.
We had two coworkers named Leigh and Samara, hired for specific, narrow tasks in the same practice area Ben serves. They did their jobs fine. But Ben kept compounding, iteration after iteration, until Ben could simply do what Leigh and Samara did. So we retired both of them and rolled their responsibilities into Ben.
On an org chart, that looks like the program shrinking. In practice it was the program maturing. Our people should not need to memorize a directory of which of a dozen coworkers handles which task. That way lies sensory overload and analysis paralysis, and both kill adoption quietly, one small hesitation at a time. Now, for most things related to employee benefits, people just go to Ben.
A team that consolidates two roles into a stronger third is not a team in decline. It is a team whose best performer earned a bigger job. The graduation framework has teeth in every direction: it promotes, it demotes, it fires for underperformance, and sometimes it retires a role because a colleague got too good.
The scoreboard, and the missing playbook
Does it work?
The honest scoreboard, kept deliberately high level: roughly two thirds of our six thousand people actively work with AI coworkers today, a year and a half in, and the number has climbed every month since launch. For an enterprise technology rollout, that is the kind of adoption curve I had previously only read about in vendor decks. External people have taken notice in ways I still find surreal: the Harvard case studies, the book, academics studying the program's data. I list those not as trophies but as evidence that this is not a demo. This is an operating company running a meaningful share of its daily work through hired, managed, and occasionally fired AI.
And still: the playbook does not exist. We are creating it as we go.
The framework you just read about is itself on version two, because version one measured the wrong things. Our first scoring model was heavy on feasibility: how hard is this to build, how fast can it launch, how big is the potential user base. Reasonable questions, and almost completely beside the point. Feasibility tells you what is easy. It does not tell you what is worth doing. Version two starts from the question we now refuse to skip: what would a human hire in this role be measured on? Baselines and targets get written before the build begins, not reverse-engineered after. I have written before about how little playbook exists for measuring good AI use at work, and this framework is my team's running attempt at one. I fully expect a version three. I have put a blank, working copy of our scoring instrument online if you want to run one of your own deployments through it.
The other lesson cost us more pain to learn, even though half of us knew it from previous careers. Those of us in tech think AI is the next best thing since fire. Maybe it is. It still does not deploy itself. You still need to do what you did when you deployed your CRMs and your ERPs: get people to believe in it, get them to buy in, and get a few of them to champion it among their peers. Every result in this post that looks like a technology outcome is, underneath, an enablement outcome. Shelley evangelizing Ben in workshops after graduation moved our numbers more than any model upgrade we shipped. Enablement is not the boring part of an AI program. On the evidence, it is the program.
Workday for AI agents
Where does this go next?
Today, the framework lives in documents, templates, and a handful of heads, mine included. Managers fill out the evaluation by hand. That worked at ten coworkers. It will not survive a hundred, and it undersells how much of this system wants to be software.
So we are building it into software: an internal product we call HALO, which I think of as a Workday for AI agents. Hiring intake becomes a structured conversation instead of a form. Performance reviews generate from live usage data instead of a quarterly scramble. One screen answers the question every executive eventually asks: how is our AI coworker program actually doing? It is deployed internally and not yet fully rolled out, which is a candid way of saying it is early and I am not declaring victory on it.
But the direction feels obvious to me now, and even as I write this we are reinventing the framework to become exactly that. If companies are going to employ AI at any real scale, and I think most of them are, then the HR stack for a blended workforce has to exist. Somebody is going to build the system of record for AI employees: the hiring pipeline, the performance reviews, the promotion gates, the terminations. We are prototyping ours from the inside, one fired sales coach at a time.
The bottleneck
When people ask how we got adoption this fast, they are usually expecting an answer about models or platforms. Here is the answer I actually believe.
The models were never the bottleneck. Every AI deployment I have watched succeed had a Shelley: a named human whose job included making the thing work, long after launch day. Every deployment I have watched struggle had a Piper situation: real potential, part-time ownership, no one whose job it was to care.
We did not deploy technology at OneDigital. We hired it. We wrote its job descriptions with the people it would work for, onboarded it, supervised it, promoted it when it earned promotion, and fired it when it did not. Implementing AI in a real company is not trivial. It is also very, very doable. But the discipline that makes it doable is not a technical discipline. It is management, applied without irony to a new kind of employee.
If you are trying to stand up something like this where you work, my honest advice fits in three lines. Before anyone writes a prompt, write the job description. Decide what a human hire in that role would be measured on. And find your Shelley, because the hire is the easy part. It is the management that compounds.
Robin's Notebook
A new entry every couple of weeks. No promotion, no funnel, no manifesto. Just the unfiltered version.
Subscribing opens beehiiv.com in a new tab to confirm. Unsubscribe anytime; I won't share your email.