← Index
A point of view · Design leadership

The Grammar Shift

For twenty years, design shipped screens. When an agent is the one assembling the interface, that stops working — the screen isn't drawn until runtime, by something that doesn't read your guidelines. So the deliverable changes underneath us: we stop shipping screens and start shipping the grammar the machine is held to. This is the transformation I lead a design org through — and how I'd run the first ninety days of it.

For
Design leaders & the execs who sponsor them
Claim
Screens → specs → a machine-checkable system
Proof
A shipped agent, not a slidethis page has one docked to the right
Read
~7 minutesor ~2 with the agent
Agent-led walkthrough Short on time? Let the agent take you through the thesis and the 30 / 60 / 90 — five chapters, each an idea you can operate rather than a wall of text. ~2 min · start →

The one idea, stated plainly

Most design orgs are meeting the agentic era by making their existing process faster — AI in the prototyping loop, a copilot in the file, prettier decks sooner. That's optimising the old job. The old job is changing shape.

A design system used to be a library of parts, held together because designers read the guidance. When an agent assembles the UI, the same conventions have to become semantics a machine can be checked against. The library becomes a grammar. The shift this whole thesis turns on

Say that out loud and three things follow, and they're the actual transformation — not tooling, but what the craft is. Design stops producing the final artifact and starts producing the decision the artifact must satisfy. The design system stops being documentation and starts being enforcement. And quality assurance stops being a human clicking through happy paths and starts being an adversary that never behaves the same way twice.

The org that wins isn't the one that adopts AI fastest. It's the one that rewires its practice around specs and evals before its component library quietly stops meaning anything.

Three shifts, from → to

This is the whole transformation on one screen. Each of these is a change in what a designer is accountable for — not a new tool they learn on Friday, a new definition of the job.

The deliverable

Screens → specs

The source artifact becomes the written decision — behaviour, states, semantics, the a11y contract — and the interface is generated against it and answerable to it. Not a prompt that produced a plausible screen. A decision the code has to satisfy, and a test can prove it did.

The design system

Library → grammar

Conventions move off the guidance page and onto the primitives themselves: what a card means, what it may contain, when it's the wrong choice. The agent selects and arranges within that grammar — it never invents a component — and a bad composition fails a check instead of reaching a user.

Quality

Click-through → red-team

You can't recruit your way to the edges of a non-deterministic system. Synthetic personas and AI-simulated testing pressure agent trust, transparency and override behaviour before launch — adversarial evaluation for a product that never runs the same way twice.

Notice what's underneath all three: the spec is the one thing a designer writes, an agent builds from, and an eval checks against. One source of truth that humans, machines and tests all answer to. That single artifact is the spine of the whole operating model — and the reason this is a design leadership problem and not an engineering one. Someone has to decide what the grammar means, and that someone is design.

Why this is design's mandate to claim, not engineering's to inherit

Engineering can express every rule in this thesis — schemas, type systems, CI checks. What it can't do is decide what a surface should mean to the person accountable for it, or which compositions are legible versus merely valid. A machine-checkable design system is still a set of human judgements about meaning; it's just enforced instead of suggested.

If design doesn't own the grammar, it gets written by default — by whoever ships the component, in whatever shape was expedient — and design goes back to redlining the output after the fact. Owning the spec layer is how design moves upstream of the screen for good. That's the strategic prize, and it's why I'd lead an org toward it deliberately rather than let it happen to them.

Where a UX team is going

Before the plan, the destination — because most orgs are climbing the wrong ladder. Bolting a copilot onto the old process moves you one rung and stops. An AI-frontier team is a different rung entirely: the design decision is the source artifact, the system enforces it instead of documenting it, and evals do the QA a human never could against a non-deterministic product.

MATURITY ↑ ▼ YOU ARE HERE Screen factory hand-drawn screens, redlined after the fact AI-assisted copilots speed up the old process Spec-driven the decision is the source artifact THE GOAL AI-frontier team machine-checkable system · evals as QA
Maturity ladder — the rung most “AI adoption” stops on, and the one this plan is for.

The unit that scales: spec → build → red-team

The whole transformation is one loop, installed once and then multiplied. A designer writes the spec — the decision, states, semantics, the a11y contract. An agent builds against it and is answerable to it. A red-team of synthetic personas tries to refute it before a user does, and what it finds rewrites the spec. Same artifact, three jobs. Every phase below is just this loop reaching one more surface.

ONE ARTIFACT three jobs SPEC the decision BUILD agent generates RED-TEAM eval refutes
The loop — one source of truth a human writes, an agent builds from, and an eval checks against.

How I'd lead it — the 30 / 60 / 90

A thesis a team can't act on is a keynote. Here's the version that survives contact with a real org: narrow, evidence-first, and built to produce one undeniable proof before it asks anyone to change how they work. Fear is the real blocker, and you don't argue people out of it — you show them one loop that made their work better and let them reach for the next. The arc is deliberate: prove, then systemise, then get out of the way.

TEAM LEVERAGE → today · screen factory frontier team PROVE GRAMMAR SELF-SERVE DAY 30 DAY 60 DAY 90 one loop, in the open grammar + red-team runs without me
The 90-day arc — each milestone raises how autonomously the team runs, not how much I personally ship.
DAYS 0–30

Prove the loop in the open

Land one undeniable win before asking anyone to change how they work. A working loop beats a manifesto every time.

  • Read the org, not the roadmap: find where the design system already lies to the code, and pick the single surface closest to being agent-assembled.
  • Recruit the two designers who are curious, not threatened — build the pilot with them so it spreads by envy, not mandate.
  • Run one real component through the full spec → build → red-team loop where the team can watch it happen.
  • Bank the win in their language: faster handoff, fewer redlines, an eval that caught what review didn't.

By day 30 One shipped surface running the full loop — and a team that watched it work.

DAYS 31–60

Turn the library into a grammar

Move conventions off the guidance page and onto the primitives — one at a time — so the system enforces what it used to merely suggest.

  • Attach machine-checkable semantics to the ten primitives that surface already uses. Don't boil the whole system.
  • Wire the catalogue both ways to the design file (Figma MCP) so it can never drift from what ships.
  • Prove a bad composition fails a check instead of reaching a user, and make that failure visible in review.
  • Start a shared spec library so the second and third components cost a fraction of the first.

By day 60 A machine-checkable grammar holding on one real surface, wired both ways to the design file.

DAYS 61–90

Make the practice self-serve

Make it spread without me in the room: a written standard, designers who can teach it, and a red-team that runs itself.

  • Stand up synthetic-persona red-teaming against the pilot surface; let it find what recruiting couldn't.
  • Write the trust-and-transparency bar as a reviewable standard — the tie-breakers design holds when autonomy and safety pull apart.
  • Hand the loop to the two pilot designers to teach the next cohort; I move to unblocking, not doing.
  • Name the metric that proves it stuck, and put it on a dashboard leadership already reads.

By day 90 A written standard, a self-running red-team, and two designers who can teach the loop without me.

The measure of this role isn't what I build — it's how quickly the team stops needing me to. Leverage, not a hero project. The numbers below are from running exactly this loop on shipped work, not projections.

~25%Faster design-to-dev handoff, via two-way MCP workflows
~30%Fewer live-research cycles, via synthetic-persona testing
1Spec that a human writes, an agent builds from, a test checks

Why I can lead this and not just describe it

Most people making this argument have the talk and none of the receipts. The rarer half — the one I'd bring — is that I've already built each piece of it, in production, before writing the thesis about it.

On Autonomous Security at Microsoft I ship the agentic experience layer, and I built the composable UI proof-of-concept spec-first — the composition rules and catalogue semantics written as a spec, the implementation generated against it, a Figma MCP server wired both ways, and Playwright driving the agent through composed layouts to check they hold when a machine assembles them. On Frictionary I make the case that the spec is the source artifact the schema, components, a11y semantics and tests all derive from. And the agent docked to the right of this page is one I engineered end to end — grounded in my real case studies, guarded so it never invents a fact, rate-limited before it went live. It argues my work in my voice. I also build agent skills for the design work itself: a Six Thinking Hats critique agent in Copilot Studio, built solo, that runs a flow or a spec through de Bono's six lenses and returns a prioritised critique ending in the single change worth making next — so I've argued the opposite case before I commit. That convergence — shipped agentic UX, an explicit thesis on machine-legible design, eval-based red-teaming, and agent skills that do real design work — is the thing a typical “AI designer” can't show all of.

I don't want to be the leader who read about the shift. I want to be the one who already ran the loop, and can hand the team the grammar. The job I'm actually applying for

Want the long version, including the parts that didn't work? Email me, or ask the agent on the right anything on this page.

Agent

Live modelPowered by GPT-4o Grounded in Sayena's case studies, not a canned bot. Billed per question — so make them good ones.