The Museum of Meaningless Metrics Is Opening a Design Wing

Design is about to adopt AI activity metrics the way engineering adopted tokens spent. Here is why the next design metric should not exist, and what design leaders should put on the scoreboard instead.

There's a cartoon doing the rounds: a museum guide showing visitors past four glass cases. Lines of code. Story points. Pull requests. And the newest acquisition, lit like a Fabergé egg: tokens spent. "Our newest exhibit," he says. Everyone in the room laughs, and then everyone in the room goes back to the office and installs the fifth plinth.

I've spent this year watching engineering leaders live the consequences. Uber reportedly burned through its entire 2026 AI budget in four months after ranking engineers on internal usage leaderboards. Appian's CEO compared token-counting to the Soviet Union grading chandeliers by weight - a policy that produced chandeliers heavy enough to pull down the ceilings they hung from. Targets hit, building condemned. Salesforce has responded with a metric of its own invention - agentic actions per employee - which is simply tokens spent with a more flattering name.

Design leaders are watching all this and, mostly, feeling smug about it. We shouldn't. Because design is about to make exactly the same mistake with different nouns, and the data says it's already started.

Design doesn't need a new metric for the AI era. The moment designers ship to production, design inherits the product's outcome metrics - conversion, retention, revenue, resolution time. Any design-specific activity metric, whether screens generated or "AI fluency" scored, is a new exhibit in the museum.

The design wing is being built by default, not decision

Here's the finding that should stop every design leader mid-scroll. In this year's AI in Design survey by Designer Fund and Foundation Capital - over 900 designers across startups, enterprises and agencies - 73% of designers said they feel rising expectations around output, quality and speed. And on the other side of the table? Only 28% of leaders said their companies have made any formal update to evaluation, compensation or hiring. Just 8% have changed performance metrics.

Sit with that gap for a second. Nearly three-quarters of designers are being pushed for more, faster, against measurement systems that nobody has updated. That vacuum doesn't stay empty. It fills with whatever is easiest to count: screens generated this sprint, prototypes shipped this quarter, percentage of UI that's AI-made, a "fluency" score from the enablement team. Meta is already grading every employee on "AI-driven impact". The design version of that review is coming to a calibration call near you.

Nobody commissions the museum. It accumulates.

Same shape, different nouns

Every exhibit in the engineering wing got there the same way. It was measurable, it correlated loosely with effort, and it was available before anyone had done the harder work of connecting activity to value. Lines of code in the seventies. Story points when agile arrived. Pull request counts when GitHub made them free to query. Tokens when the AI invoice made them impossible to ignore.

Design's candidates have precisely the same shape:

  • Screens generated. Volume of the single most gameable artefact in the stack. Nielsen Norman Group's research on AI prototyping carries the perfect title for this: Good from Afar, But Far from Good. AI output looks finished while lacking the structure, logic and usability thinking underneath. Counting it is counting varnish.
  • Prototypes per sprint. Rewards the part AI just made nearly free, ignores the part that got harder: choosing. The hard part now is no longer producing options - it's picking the right one and proving it works.
  • Percentage of UI that's AI-generated. The design cousin of "percentage of code written by AI", which engineering commentators were calling a vanity metric by January. It measures tool adoption. It says nothing about whether the product got better.
  • Designer AI fluency scores. The most dangerous of the four, because it sounds like capability building. But it still measures the designer's relationship to the tool, not the work's relationship to the business. Measure activity instead of outcomes and people will, quite rationally, perform the activity.

To be fair to design leaders - and my rule is the sceptic always gets a fair hearing - the same survey shows only 5% are placing less emphasis on execution quality. Leaders keep saying taste and judgement matter more now that "good enough" is free. They know counting screens is daft. What they haven't done is operationalise the alternative. That's the 8%.

The excuse AI just removed

Now the part most of the current debate misses entirely, and the reason design's answer is different from engineering's.

Design has spent decades with a ready-made excuse for unmeasurable impact: our work never reached production directly. It went through a translation layer. Engineering rebuilt it, shipped it, and owned the telemetry, so when the outcome data came back, it belonged to someone else. Design was left holding proxies - satisfaction scores, heuristic audits, the occasional triumphant usability video. Convenient, in a way. You can't be held to a number you were never connected to.

That excuse is gone. Roughly half of designers have now pushed AI-generated code to production. Companies like Vercel, Ramp, Linear and Anthropic have formalised the design engineer role. When a designer ships through a real pull request - reviewed by engineering, merged on the same standard as anyone else's code, live in the experiment - something quietly enormous happens to measurement. Design inherits the product's metrics. Conversion. Retention. Revenue per session. Time to resolution. Change failure rate. These were already the numbers the business ran on. They're now the numbers design runs on too.

I've felt this personally building Sero. As a solo founder who designs the product and ships it, there is no translation layer left to hide behind. Nobody asks me how many screens I generated last week, least of all me. The only questions with any weight are whether managers run better one-to-ones after using it and whether they come back. Everything else is decoration. It's a smaller stage than an enterprise, but the physics are identical.

The moment a designer ships through a reviewed pull request, design inherits the product's metrics. That's not a burden - it's the argument design has always needed.

The two-column test

The practical move for a design leader isn't inventing the next metric. It's running every metric you currently report through one question: is this activity, or is this an outcome someone was already accountable for before AI arrived?

Two honest caveats, because the evidence demands them. First, the tooling to connect design activity to shipped outcomes is immature - engineering is finding the same thing with tokens, where the consensus is that combined metric sets beat any single number and the connective tissue doesn't fully exist yet. You will be building some plumbing. Second, outcome metrics only work at the level of teams and initiatives. The moment you point them at individuals, you've built a leaderboard. Leaderboards produce tokenmaxxing. Use them at the wrong level and you've simply rebranded the museum.

Neither caveat rescues the activity metrics. A hard-to-measure truth still beats an easy-to-measure fiction.

Where to start this week

Take your current design dashboard - the one that goes in the QBR deck - and sort every number into the two columns above. Most leaders I work with find the left column embarrassingly full. Then pick one initiative shipping this quarter and agree with your product and engineering counterparts, in writing, which product metric it will be judged on. Not a design metric. Theirs. When it moves, design moved it, on a scoreboard nobody can dismiss as marking its own homework. That's one initiative. That's enough to start.

The museum will keep acquiring. It always does. Your job is to make sure design's contribution is on the scoreboard, not behind the rope.

In short

The next design metric shouldn't exist. When design ships to production, it shares the product's outcome metrics - and every design-specific activity metric is a museum piece waiting for a plinth.

Practical checklist

  • Audit your current design metrics: label each one activity or outcome.
  • Refuse volume metrics for AI-assisted work - screens, prototypes, generation counts.
  • Attach every design initiative to one existing product KPI before work starts.
  • Measure at team and initiative level, never individual leaderboards.
  • Treat "AI fluency" as a training input, not a performance output.
  • Revisit quarterly: the 8% who updated their metrics will set the norms for everyone else.

When to use this approach (and when not to)

  • Use it when designers on your team already ship to production, or will within two quarters.
  • Use it when leadership is asking "how do we measure AI adoption in design?" - redirect the question before a dashboard answers it for you.
  • Use it when design's budget case depends on demonstrating impact in business units.
  • Do not use outcome metrics as a stick during the transition - teams need runway while the plumbing gets built.
  • Do not measure exploration work (research, early discovery) on shipped outcomes; measure it on decision quality and speed.
  • Do not assume a vendor-defined metric is neutral. Whoever defines the unit controls the benchmark.

If this is the conversation you're stuck in

If your leadership team is currently deciding how to measure AI-era design work, this is the cheapest possible moment to get it right - metrics are far easier to install than to remove. I run focused sessions with design and product leadership teams on exactly this translation. The museum is lovely to visit. You don't want to work there.

Frequently asked questions

What's wrong with measuring how much AI designers use?

Usage is an adoption signal, not a performance measure. It tells you the tool is switched on. The moment it becomes a target, people optimise for the interaction rather than the result - the same dynamic that produced token leaderboards and burned budgets in engineering.

Should design teams measure anything new in the AI era?

Mostly no. The defensible move is inheriting existing product outcome metrics - conversion, retention, task success, change failure rate on design-shipped work. The one genuinely new number worth watching is cost per shipped, validated improvement, and even that is an efficiency lens on an old outcome.

Doesn't this only apply to designers who code?

It applies fastest to them, but the principle is general: tie design work to an outcome the business already tracks. Designers shipping through reviewed pull requests simply remove the last structural excuse, because the telemetry connects directly to their work.

How do I push back when leadership asks for an AI adoption dashboard?

Give them one - as a temporary adoption view with a stated expiry, clearly separated from performance evaluation. Then present the outcome scoreboard alongside it. Leaders rarely insist on the vanity version once a credible alternative is on the table.

What should design leaders do about "AI fluency" assessments?

Use them for development conversations and training investment, never for performance ratings. Fluency predicts capability; it does not demonstrate delivered value. Grading it as performance is how the museum gets its next exhibit.