Twenty Years of 1:1s Taught Me What AI Can't Do. And What It Should.

I spent a year researching AI for 1:1 preparation. The science convinced me one feature should never be built. Here's why - and what I built instead.

I've been leading design teams for twenty years, and in that time I've probably sat through a few thousand one-to-ones - on both sides of the table.

I've run good ones that changed the direction of someone's career. I've run bad ones that were status updates wearing a nicer name. And a handful of times, I've walked out of one and only understood months later - usually reading a resignation email - what the person had actually been trying to tell me.

Those are the ones that stay with you. Early in my career, I watched a business I'd built start to unravel because a key person left, then another, and I hadn't seen either coming. The signals were there. I just wasn't prepared enough, or asking well enough, to catch them.

So when AI tools started promising to solve exactly that - *we'll read your notes and tell you who's disengaging, who's burning out, who's about to leave* - I wanted it to be true more than most people. I've spent the past year building an AI tool for 1:1 preparation, and I went into the research fully expecting to build that feature.

I came out convinced it should never be built. Not by me, not by anyone. Here's why.

What the research actually says

I spent months going through the evidence - the computational linguistics studies, the performance-appraisal research, the regulatory guidance, and the post-mortems of companies that tried this before. Three things kept coming up, and each one changed how I think about AI at work.

**First: the science of "detecting" states from text was built on people describing themselves.** The credible studies all analyse words written by the person actually experiencing burnout or stress - their own survey answers, their own words, in volume. A manager's two-line prep note about someone else is a completely different thing, and there is no published evidence that machines can reliably read a person's inner state through it. Even the most famous marker in the field - pronoun use predicting depression - turns out to be far weaker than the headlines suggested.

**Second, and this is the one that humbled me as a leader: my notes about my team say more about me than about them.** One of the most cited studies in performance research (Scullen, Mount & Goff, *Journal of Applied Psychology*, 2000) decomposed thousands of workplace ratings and found that over half the variation came from the *rater* - our moods, our standards, our blind spots - while the person being rated accounted for roughly a fifth. Twenty years in, I believe that completely. I can tell you which of my old feedback notes were written on a good day and which were written after a rough steering meeting. Any AI reading those notes wouldn't have been analysing my team. It would have been analysing me.

**Third: the maths punishes exactly the people you're trying to protect.** Genuine disengagement at any moment is rare, and rare events break prediction. The best academic burnout detector published to date is wrong roughly three times out of four when it raises a flag. Now imagine that flag landing on a manager's screen before a meeting: *something feels off about this person.* The employee is fine - but the manager walks in primed with suspicion, and the employee spends thirty minutes wondering why the temperature changed. A missed signal keeps things as they were. A false one damages a relationship that was healthy.

Sero error and validation states - UI showing how Sero surfaces factual gaps and missing information in preparation data rather than inferred psychological flags
Where Sero does surface warnings: gaps in preparation data, missing follow-ups, overdue actions - facts, not inferences about how someone is feeling.

The industry has already run this experiment at scale, publicly. [Microsoft](https://www.theverge.com/2020/11/25/21722440/microsoft-productivity-score-feature-employee-monitoring-privacy) had to strip individual employee data out of its Productivity Score within days of launch after the surveillance backlash in 2020. [IBM's famous claim](https://techcrunch.com/2018/06/05/ibm-can-predict-with-95-accuracy-when-employees-are-about-to-quit/) of predicting resignations with "95% accuracy" was never backed by published figures. And regulators - the ICO in the UK, the EU with the AI Act - have drawn increasingly firm lines around inferring emotional states at work. In a regulated enterprise environment, this isn't an edge case. It's a compliance question with a very short answer.

Sero Universe - a constellation graph of all team members, sessions, meeting types, and engine concepts, visualising the full network of 1:1 relationships and data Sero tracks
The Sero Universe: a live graph of every person, session, and concept the engine holds - none of it involves guessing at anyone's inner state.

What I decided to build instead

Here's the thing though - the underlying problem is still real, and it's still costing organisations their best people. Gallup attributes around 70% of team engagement variance to the manager. The 1:1 is the highest-leverage tool that manager has. And the honest research finding (Steven Rogelberg has spent a career on this) is that 1:1s mostly fail for an unglamorous reason: preparation. We walk in cold. We run through status. We leave with nothing agreed. I've done it myself, in busy quarters, more times than I'd like to admit.

That's where AI genuinely helps - not as a judge of people, but as preparation for the human conversation.

Sero home screen - 'Welcome to Sero. Prep for your next 1:1 in a few minutes.' with a simple three-step flow: tell Sero who you're meeting, answer a few short questions, get a briefing.
Sero's home screen: the whole experience is built around one idea - walk in prepared, not cold.
Sero Guide screen - a structured AI-generated briefing with topic suggestions, conversation starters, and carry-forward items from previous 1:1 sessions
The Sero Guide: an AI-prepared briefing based on past sessions and agreed actions - turning manager prep time from thirty minutes to five.

There *is* a category of signal a system can hold with complete integrity, because none of it involves guessing at anyone's inner life: the action we both agreed on has now rolled over three meetings running. The fortnightly 1:1 has quietly become monthly. The thing I promised to unblock is still blocked. Those are facts about the work, visible to both people, and they're the fairest possible prompts for a real conversation.

Sero Library view showing past 1:1 preparation sessions listed by date, team member name, and focus area - all unreviewed, ready to carry forward
The Library: every past session, what was agreed, what's unreviewed - facts about work, not inferences about people.
Sero user management screen - showing team members, their roles, and session history without any sentiment scores, risk ratings, or inferred psychological states
Sero's people view: names, roles, history. No scores, no risk ratings, no inferred states. The only data is what verifiably happened.

So the tool I've been building, Sero, is designed around a rule I've made non-negotiable: **no inferred psychological states, ever.** No scores on people, no trend lines, no dashboards of human beings. It tracks what verifiably happened, remembers what was promised, and turns a manager's rough notes into sharper questions and a clearer plan. The judgement - the actual leadership - stays with the human. Where twenty years of experience tells me it belongs.

Sero Tasks planner - a kanban board with columns for Ideas, To do, Doing, and Done, showing work-in-progress cards like 'Design cleanups', 'Run QA fixes', and 'Live-data cleanup'
Sero's planner tracks what was agreed and what's still open - the kind of accountability a 1:1 should surface, not a surveillance score.
Sero design system - component library and visual language documentation showing the design principles and UI building blocks used across the Sero product
The Sero design system: building consistent, trustworthy UI is part of the product promise - a tool that earns trust should look like it deserves it.

The bit I keep coming back to

The people I failed to keep, all those years ago, didn't need an algorithm to detect their disengagement. They needed me to walk into our 1:1s prepared, ask a better question, and follow through on what I promised.

AI can help with all three of those. It can't do the fourth thing - care - and we should stop buying, and stop building, software that pretends it can.

That's not a limitation of the technology. That's the design.

Key sources

  • [Scullen, Mount & Goff, Journal of Applied Psychology (2000)](https://doi.org/10.1037/0021-9010.85.5.956)
  • [Edwards & Holtzman, Journal of Research in Personality (2017)](https://doi.org/10.1016/j.jrp.2016.06.022)
  • [Kurpicz-Briki et al., Frontiers in Big Data (2022)](https://doi.org/10.3389/fdata.2022.786054)
  • [Gallup, State of the American Manager](https://www.gallup.com/services/182138/state-american-manager.aspx)
  • [Rogelberg, Glad We Met (OUP, 2024)](https://global.oup.com/academic/product/glad-we-met-9780197641873)
  • [ICO guidance on monitoring workers (2023)](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/employment/employment-practices-and-data-protection-monitoring-workers/)
  • [EU AI Act](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689)