Episode 001 · 10 min
After the demo
An AI tool turns three policy documents into a briefing in twenty-six seconds. The next afternoon, the same tool meets the real version of that job. What the demonstration left out is the episode.
This episode has not been narrated yet. The full script is below — it reads the way it is meant to sound.
What this episode says
- A demonstration is an edited truth: the task was chosen, the inputs prepared, and the run you are watching is the one that worked.
- AI capability is jagged. The same tool can be excellent at a hard-looking task and wrong about one that looks almost identical.
- Fluency hides the boundary. A weak answer arrives with good grammar and a confident tone, so finished reads as correct.
- Four questions for any demo: what was prepared beforehand, what could the system not see, what has to happen after, and who owns the result.
Transcript
Every word, in full. Nothing is said in the audio that is not here.
Okay, picture this with me.
It's about ten past ten on a Tuesday morning. Someone's in a leadership meeting, sharing their screen. They drop three policy documents into an AI tool and ask it to turn them into a one-page briefing for staff.
And it comes back in about twenty-six seconds. Headline. Clean summary. Four action points. Even the tone's right.
And everyone in that room sits up a bit. Somebody laughs. Somebody says, "that takes us half a day." And someone else asks whether the whole thing could be automated by next month.
Five minutes later the demo's done, and everyone's a bit more optimistic than they were when they walked in.
Now come back with me to the next afternoon. Because this is the part I actually want to talk about.
A manager sits down to do the real version of that task — the actual change they've got to tell the team about.
And here's what they're working with. One of the policies is out of date. There's an exception that only exists in an email from six months ago. Two of the numbers don't match each other. The team was promised something last year that isn't written down anywhere. Oh, and one of the source documents has information in it that the wider team is absolutely not meant to see.
The AI writes the briefing anyway. Confidently. It's clear, it's well structured, and it's quietly wrong in about six different ways.
Hello, I'm Bit. This is In Real Life, from The Human Bit — where we look past what AI can do in a demo, and talk about what happens when it meets actual work.
That scene I just described is made up, by the way. It's a composite. But I don't think there's anything in it you haven't seen a version of.
So let's start with the demo itself, because I want to be fair to it.
That demo wasn't fake. The tool genuinely read those documents. It genuinely wrote that briefing. Twenty-six seconds was twenty-six seconds.
But here's the thing about a good demo — it's an edited truth. It's built to make one thing easy to see. Somebody picked the task. Somebody prepared the inputs. The goal was clean. And the run you're watching is the one that worked.
What you don't see is everything around it.
Somebody decided which documents could be trusted. Somebody took out the stuff that didn't belong. Somebody knew the old policy shouldn't go in. Somebody worked out what a good answer would even look like. Somebody checked nothing private was about to end up in a staff email. And somebody understood that organisation well enough to spot the sentence that's technically true and completely impossible in practice.
None of that shows up on screen. It doesn't arrive in a dramatic burst of text. It looks like reading things. Asking someone. Checking. Comparing. Remembering. Deciding.
It's slower than the demo. It's also usually the reason the demo worked.
And this is why I get a bit protective about this, honestly — because when you try the same task with your own messy material and get a worse result, what do you conclude? Most people conclude they're bad at it. That they're not technical enough. That everyone else has figured out something they haven't.
You weren't shown the whole method. You were shown the visible half.
Now, that would matter less if there wasn't so much pressure attached to it. But there is.
Microsoft did some research in Australia last year — their Work Trend Index — and sixty-eight per cent of workers said they were afraid of falling behind if they didn't adapt fast. Sixty-eight per cent.
In the same research, only twenty-eight per cent said their organisation was actually clear on its AI strategy.
Now — Microsoft sells AI, so I'd take their research with that in mind. I'm not going to pretend a provider's survey is neutral. But that gap? That one rings true. People are being told this is urgent well before anyone's explained where it actually fits in their job.
Pew found something similar in the States. Around a third of workers said they felt overwhelmed about AI at work. More than half said they were worried about it. But when you ask what they're actually using — most of them said little or none. Sixteen per cent said any real part of their work was being done with AI.
Sixteen per cent. And yet if you spend any time online you'd think everyone had rebuilt their entire job around agents by about March.
The truth is much messier. Some people are doing genuinely clever things with this. Plenty are poking at the edges. A lot are waiting for a use that makes sense. And a big chunk of the confidence you're seeing out there is confidence about a demo, not evidence of something that actually works week after week.
There's one more reason this all feels confusing, and I think it's the most useful thing in this episode.
AI is uneven. Properly uneven. The same tool can be brilliant at something that looks hard, and then fall over on something that looks almost identical.
There's a study — seven hundred and fifty-eight consultants, so a decent size — and they called this the jagged frontier. Lovely phrase. On the tasks that sat inside what the model was good at, people using AI did more work, faster, and better.
But on a harder task, deliberately placed outside that frontier? The people using AI were nineteen per cent *less* likely to get the right answer than the people who weren't using it at all.
Now, that's not "AI is bad". That's not the lesson. The lesson is that you can't see the edge. Fluency hides it.
Think about it — when a spreadsheet formula breaks, it usually looks broken. When an AI answer is weak, it turns up with perfect grammar, a sensible structure, and the calm confident tone of something that's checked itself.
The better it's written, the easier it is to mistake finished for correct.
Right. Let me make this concrete, because I don't want to leave it abstract.
Say you're a manager and you need to tell your team everyone's expected in the office one extra day a week.
From a distance that's just writing an email, isn't it. Give the AI the decision, ask for a professional tone, read the draft. That's the demo version.
Here's the real version. The business has a reason but the evidence for it is a bit thin. Two people have caring responsibilities. Someone joined on a different flexibility agreement entirely. Another team's being treated differently because they're customer-facing. Nobody's agreed how exceptions get handled. And the last big change was announced badly, so trust isn't exactly full.
Now — AI can help you here. Genuinely. It can pull the approved facts together. It can check your decision against the actual policy. It can find the contradictions. It can tell you which questions you haven't answered. It can give you the email, the manager briefing and the FAQ page. It can flag the sentence that sounds defensive.
That's real help. That's hours.
But it can't settle the disagreement underneath the decision. It can't work out which exceptions are fair. It doesn't know what your team was promised. It can't approve anybody's flexibility arrangement. And it isn't in the room when someone says this has made their life harder.
So you're not adding value by typing every sentence yourself. You're adding value by making the decisions those sentences represent.
And I think that's a more honest split than "AI does the work, human checks it". Because "check the output" is advice that means nothing. AI can prepare a serious chunk of the work. You still have to establish what's actually true, what's actually fair, and who's answerable when it goes wrong.
So. Next time you're watching a demo — and this is the thing I'd genuinely suggest you take away — there are four questions worth asking.
One. What happened before the prompt? Which documents, which examples, which corrections, how many attempts that didn't make the video?
Two. What couldn't the system see? What history, what relationships, what unwritten rules, what consequences?
Three. What has to happen after? What facts need checking, what permissions, what privacy, what tone?
And four — this is the one people skip — who owns what happens next?
Sometimes the answer to all that is "great, let's use the tool." Sometimes it's "our information isn't good enough yet." Sometimes it's "we need a human approval step here." And sometimes, honestly, it's "this one shouldn't be handed over at all."
All four of those are fine answers.
So — next time a demo makes an hour of work vanish, let yourself be impressed. I mean it. That's genuinely impressive and I don't want to be the person who sniffs at it.
Just stay curious about the part they cut. Where did the information come from? What had already been decided? What didn't it know? What will someone have to check? And who answers for it when real life doesn't match the tidy example?
Those questions don't take anything away from what AI can do. They're how it becomes useful.
The demo shows you what the model can produce. The real work is deciding what deserves to leave the screen.
That's In Real Life. Your work is human. Your AI guide should be too. See you next time.