What I've Learned Running an AI Program at a 6,000-Person Company

The two loudest stories about enterprise AI, that it deploys itself or that it replaces everyone, are both wrong. The real work is the boring middle. Notes from inside a 6,000-person AI program.

enterprise AIAI adoptionoperator noteschange managementAI workforce
A document styled like an employee onboarding file, held by a coral paperclip on a light wooden desk beside a coffee cup and a closed laptop in soft morning light. Magazine-editorial illustration.

A few weekends ago I was reviewing a late draft of a chapter in the book Vinay is writing. Vinay is my boss at OneDigital, where I run the AI program. He'd been writing about the ninety minutes of hands-on training every new employee gets before they're allowed anywhere near an AI Coworker, and he asked me for the line that explains what actually has to happen in that room for the training to work. I gave him this:

AI adoption happens when you help people do today's work better and with more confidence, rather than when you ask them to believe in the future.

I gave him that line because it's the most honest summary I have of what we've actually learned in eighteen months running this program. It's also not the sentence I would have predicted in 2024, when I was reading the same vendor decks and analyst reports everyone in this work reads. Their thesis was that AI deploys itself. Buy the licenses. Get out of the way. Watch the productivity gains show up on their own.

That is not what happened. AI does not deploy itself. It also does not take everything. The work happens in the boring middle, and almost nobody writing about enterprise AI right now has actually stood in it.

This is what eighteen months in the boring middle looks like.

Why both stories about AI are wrong

The two stories you hear most about AI in the enterprise are mirror images of each other, and both of them are wrong.

The first story comes from the vendor and analyst class. It says: AI is a horizontal productivity boost. Pick a model. Deploy it. Save 30% on cognitive labor. This is the story sold in keynotes. Ethan Mollick, in his Substack writing and in Co-Intelligence, has spent the last few years documenting what he calls the jaggedness of AI capability: it is superhuman on some tasks and useless on others, and the variance across people inside a workforce is wider than the variance across the model versions on offer. Whether your enterprise lands in the high-leverage or no-leverage cluster is not a function of which model you bought.

The second story comes from the doom and labor-displacement commentary. It says: AI takes everything. The agents replace the workers, the workers replace themselves, the firm reduces headcount until there is no firm. I have not seen that play out in any practice I work with at OneDigital. Not one knowledge-work function has been fully replaced. Plenty of routine sub-tasks inside knowledge work have moved to AI. The function around them has not.

The two stories share a quiet assumption: that the AI is the variable. Either the right AI saves the business, or the wrong AI sinks it. The truth is closer to what Andrej Karpathy was getting at with Software 3.0: when the AI is the foundation, the programming layer becomes the prompt and the workflow. The bottleneck moves from "do we have the right software" to "can we manage the new kind of teammate this software became."

That bottleneck is a people-and-process problem. It is the same bottleneck that determined whether enterprise CRM rollouts in 2014 succeeded or failed. It is the same bottleneck that determined whether ERP rollouts in the 1990s succeeded or failed. The bet of our AI program at OneDigital is that we have to treat it that way. As an HR transformation, not a technology rollout.

The methodology, in plain language

I run an AI Coworker program. We call them Coworkers on purpose, not agents, not bots, not assistants. "Agent" describes what the software does. "Coworker" describes what the person sitting across from it is supposed to feel: this has a job, it can be trained, it can be wrong, and it isn't here to take my seat. That's not a branding choice. An employee who believes the thing in front of them exists to replace them uses it defensively, or not at all. An employee who believes it's the newest hire on the team actually trains it, corrects it, delegates to it, the way they would a junior colleague. None of what follows, the job description, the knowledge base, the named human supervisor, the probation cycle, works without that first move landing. Coworkers are conversational systems anyone in the company can talk to through our internal tools, practice-specific or cross-practice, and every one goes through the same three-phase lifecycle: Internship, Apprenticeship, Full-time.

Internship is where a small group of users stress-tests a brand-new Coworker against real work, surfaces the failure modes, and feeds the corrections back to a human supervisor. I tell every cohort the same thing on day one: the thing in front of you is going to make mistakes; your job is to find them, not to be impressed by it. If people show up expecting a finished product, the program fails, because the bar is set wrong. If they show up expecting an intern, the bar gets met, and the Coworker gets better because of it.

Here's what that actually looks like, not the version that fits on a slide. The Coworker most people at OneDigital would recognize handles benefits questions; everyone just calls it Ben. The supervisor who owns Ben and I ran the same loop, daily, for eight months. She'd surface a case where Ben got something wrong. I'd trace it to a gap in the knowledge base or an instruction that needed rewriting. We'd fix it. The next day's testing would expose the next weakness. Eight months. Not eight weeks, not a two-week beta. That is what Internship-grade iteration actually costs when you're doing it instead of demoing it.

Apprenticeship is the validation phase. The Coworker is opened to a wider but still curated audience for two or three weeks. The supervisor tightens the prompts, expands the knowledge base, retrains the system on the edges that came up in Internship.

Full-time is when the Coworker goes available to the whole company. Conversations are retained for context. Users share threads with each other the way they'd forward a good email.

The structurally important detail, the one that does most of the actual lifting, is the supervisor. Every Coworker has a human supervisor, and that supervisor sits in the business unit, not in the tech organization. The supervisor is whoever is already best at the relevant job. That choice is expensive. Pulling a top performer off active work to encode their craft into an AI is an opportunity cost the program owns directly, on purpose. It is also the only way the Coworker becomes worth using. The Coworker is, at the end of Internship, a high-fidelity transmission of the supervisor's expertise. Mediocre supervisor, mediocre Coworker. The model is interchangeable. The supervisor is not.

Three folders fanned left to right, growing thicker, with a coral tab on the last, suggesting a progression from intern to full-time

There's a second layer under this that took me longer to see, because it isn't visible from inside any single Coworker's rollout. Every new employee gets that same ninety minutes of hands-on training before they're let near a Coworker at all. For a long time I assumed that if adoption still lagged after that, the problem was motivation: people who'd sat through the training but weren't convinced, or hadn't built the habit yet. When we finally pulled the numbers across every practice this year instead of going by feel, the actual gap was earlier and duller than that. The pattern was always the same: people show up, find nothing built for their actual job, and leave. Not because they don't want to use AI. Because there's nothing there for them yet. Everyone gets the same training. Not everyone gets something built for their actual work. Coverage, not motivation, is the real gap.

What gets paved over in keynote decks

A meaningful share of Coworkers don't survive Internship. They get cut. They didn't perform. We may revisit them in six months when the underlying models improve. Right now they're on the cutting room floor. The internal joke in the AI Programs team is that I have fired enough Coworkers by now that when AI becomes sentient I will be first in its firing line. The joke is sticky because it is operationally true.

A non-trivial share of Coworkers that survived Internship are currently on Performance Improvement Plans. The Supervisor and I review their KPIs together; we agree on what better looks like; we set a timeline; we either retire the Coworker or we graduate it. None of this is hypothetical. None of it shows up in the brochure.

The cheapest adoption lever we've found has nothing to do with the AI, and nothing to do with the Coworker's supervisor either. It's the person's own manager. When a team's manager actually uses the thing, and brings the team together for twenty minutes every couple of weeks to work a real, hard case live instead of forwarding a link, that team's usage can climb from around a third to about three-quarters inside a quarter. When the manager stays out of it, the team stalls, no matter how good the Coworker is. And I didn't just eyeball that across the company as a whole, where strong teams and strong managers travel together and you can't tell which one is doing the work. I checked it inside each practice separately. The manager gap holds within every one of them. The Coworker didn't change. The manager did.

I am about a year into knowing this, and I still have to remind myself of it weekly. The instinct is to chase a new model release. The discipline is to get one more manager to run one more session.

The next problem I do not have an answer to

For all the methodology, there's one problem here I got wrong for a while, and a second one I still haven't solved.

The one I got wrong: I used to think the goal was moving people from transactional use of a Coworker (quick question, quick answer) to collaborative use (a real back-and-forth that produces something neither of us would have produced alone). I built training and messaging around that distinction for months before I trusted the data enough to admit the distinction itself was the mistake, not the direction people were failing to move in. I've written the longer version of how that fell apart and what replaced it elsewhere. The short version for here is that the real question was never how long the conversation ran. It was whether the person's time actually came back to them.

The one I still haven't solved: most people, even once they know a Coworker can handle the harder, more collaborative ask, don't reach for it. The ninety minutes gets them in the door. The tool can do more than they're asking it to do. The habit of asking for more doesn't form on its own, and I don't yet have a lever for it the way I have a lever for coverage or for supervisor cadence. I'd rather say that plainly than pretend the metric fix solved it.

Why I trust this is real

The honest test, when you're inside a program and can't tell whether you're looking at a real pattern or a flattering one, isn't whether outside experts noticed. It's whether the thing survives someone actively trying to break it.

That's the test I apply to my own numbers first. Every dollar-value claim the program makes runs through a model I built and rebuilt, and when I tightened the methodology this year, the program's own headline value came out meaningfully lower, because the earlier estimate leaned on research that didn't hold up once someone actually read it. I published the correction anyway, with a section that does nothing but list what we got wrong. Same discipline as the Coworker lifecycle, aimed at my own numbers instead of the AI. A claim that can't survive a hostile reviewer isn't a number. It's a slide.

The second test is closer to home. Vinay wrote the book about this program's methodology, and more than once, when I've pushed back on how he wanted to frame our own adoption metric, he's told me the critique was fair and asked me to go build a better one instead of defending his. In a lot of companies, the person who wrote the book about the program is the last one who wants to hear the methodology has a hole in it. Vinay isn't.

The outside validation is real too, and it matters, just not as much as the two tests above. Harvard published a case study on what we did this spring; I wrote about that when it landed. Vinay's book about the methodology comes out this summer, with the same set of mechanics described here at much greater depth. A joint Harvard, MIT, and UCLA research paper studying the same topic is in motion for later this year.

None of that is why I trust the work. It's confirmation that the work was worth studying in the first place. What I actually trust is that the number survived me trying to break it, and that my boss would rather I broke his framing than protect it.

What this means if you are buying AI

If you're at an enterprise buying AI right now, my honest read is that picking the right model isn't what determines whether you succeed. The frontier models are converging fast enough that whichever one you standardize on this quarter, GPT, Claude, Gemini, whatever's ahead by the time you read this, matters less than whether you run a real change-management program around it.

Write the job description for the Coworker you actually want to hire. Find the top performer in the relevant role inside your company and ask them to become its supervisor. Build the Internship as a real quality-assurance cycle, not as a beta test. Run the biweekly facilitations. Cut the Coworkers that aren't earning their seats. Resist the temptation to skip the supervisor cadence because users want access faster.

None of this is futuristic. None of it is hard to explain. It is, on the inside, deeply unsexy. It is also the reason a 6,000-person traditional consulting firm has reached 64% workforce-wide adoption eighteen months in, while most enterprise AI programs stall out with only a small fraction of the workforce ever really using the thing.

None of this is permanent. The models will keep improving. The supervisor model will shift as the tools get better at the parts supervisors currently do by hand. The word "Coworker" itself might not be the noun we're using for this in five years.

But the discipline underneath it doesn't need the model to be right, and it doesn't need the noun to survive. It needs someone in the business unit willing to own a mediocre first draft, a training program people actually sit through, and a habit of showing up to the cadence instead of chasing the next release. That part isn't going anywhere, because it was never really about AI. It's the same thing that decided which CRM rollouts worked in 2014. We just keep forgetting it every time the technology changes costume.

Robin's Notebook

A new entry every couple of weeks. No promotion, no funnel, no manifesto. Just the unfiltered version.

Subscribing opens beehiiv.com in a new tab to confirm. Unsubscribe anytime; I won't share your email.