Design system, foundations to governance
Closing the gap between what shipped and what we drew
Components disagreed with each other on size, spacing and typography, and having two dashboards meant common components and platform specific ones had quietly diverged. Nothing was broken enough to escalate, which is why it had persisted.

The system in use. One promotion card assembled entirely from shared components, surrounded by the primitives and pieces it is built from.
Context
We run two dashboards on a shared Material UI base. Over time each grew its own conventions, so a component that existed in both places behaved differently in each, and the Figma library described neither accurately. Typography was not aligned to the system, and the typeface was a paid licence with proportional figures, so numbers in tables never lined up. Spacing, radius and shadow had no tokens, so cards carried a blurry over softened look that came from nobody having defined elevation.

Before. Hellix against Figtree on the same card. Three reasons drove the change: a paid licence replaced by an open one, figures that now align in columns, and text greys that finally separate secondary from disabled.
Tension
There was no single visible failure to point at. Every individual inconsistency was small enough to ignore, and the cost only appeared in aggregate: engineers rebuilding components that already existed, designers picking values by eye, and me spending three to four hours a week checking whether frontend output matched intent. The other difficulty was that this had to happen under active development on two products, with frontend engineers who report into other teams and could only take work as sprint side items.
Insight
The gap was not a documentation problem, it was an absence of source of truth in either direction. Production could not be trusted and neither could the library. So I stopped treating this as a design task and started treating it as an audit: pull production into Figma, compare it against what we claimed, and close the difference item by item. Once the foundations were tokens, component correctness became something a machine could verify rather than something I checked by hand every week.
Nothing was broken enough to escalate. That is exactly why it had survived for so long.
Decisions, and what each one cost
Impact
The QA figure is the one I would defend. Reviewing frontend tasks used to mean checking every component by hand. Now that components are shared rather than custom built per screen, review is about UX and flow, which is roughly 130 hours a year returned to actual design work.
More: the system underneath, agent readability, screens, leadership, step by step and what comes next
The system underneath
Four layers. A value has one definition and one owner, and components never reference a primitive directly.
- Cool grey ramp
- Red, orange, blue, green
- Type scale, Figtree
- Primary, success, danger
- Warning, information
- Main, light, dark, strong
- Rest at 0%
- Hover at 4%
- Pressed at 8%
- Spacing tokens
- Radius tokens
- Elevation, redefined
The interaction layer and the structural layer are the two that did not exist before. Together they are the reason a component can now declare its own states instead of an engineer deciding them per screen.
Built to be read by agents
The system is built to be read by coding agents, not only by people. That is a deliberate design goal rather than a side effect.
- Start by asking what it needsI begin by asking the model how it wants to read a component and which cases it will check. Sometimes its expectations match our structure. Where they do not, the gap tells me our definitions are ambiguous.
- Teach it the parts we ownWhere our conventions differ from what it assumes, I document the difference explicitly rather than renaming things to suit the model.
- Fetch state, not screenshotsAn MCP connection pulls component states and specifications directly into the codebase, so generated code references the real system instead of approximating a picture of it.
Foundations

How the layers connect. A primitive step becomes a semantic role, and the role drives a component state. Nothing skips a layer.

Greys before and after. The original ramp mixed cool and neutral tones with no rule, so two adjacent surfaces could disagree.

Five semantic roles, four tones each, defined once so no component derives a state colour on its own.

One overlay of the role colour, 4% on hover and 8% when pressed, holding up on every surface it sits on. Previously each component improvised its own.

Spacing, radius and elevation. The old cards looked soft and blurred because elevation had never been defined, only approximated.

Icons, mostly hand drawn. I set the ruleset and removed every size except 16 and 24, because inconsistent icon sizing was a recurring source of misalignment.
Components, and the one I pulled

Button. Primary, secondary and ghost hierarchies across brand, danger and neutral, each with default, hover, pressed, focus, disabled and loading. It had the largest gap from best practice and became the proving ground for the whole colour system.

Segmented control. Returned from development planning when our VP disliked the treatment, so I produced alternatives and tested them with other leads before settling on this one.

A sample of the refactored set. Each component shipped as a paired design and development task and was verified in Storybook and a sandbox before release.

The focus ring, and why it is paused. Wrappers sit flush to component edges, so the ring is clipped whenever a component is nested inside a dialog or filter panel.
Leadership and ownership
- Getting a typeface change approvedA new typeface touches every screen, so it needed management level agreement rather than a design decision. I mapped the system, made the case on licence cost, numeric alignment and text hierarchy, and got sign off from nine leaders.
- Working through teams I do not manageFrontend engineers sit in other teams and could only take this as sprint side work, so every component needed to be specified well enough to be picked up without me in the room.
- Absorbing a late reversalThe segmented control came back from development planning because our VP disliked it. Rather than defend the original, I produced alternatives and tested them with other leads. The revised version is better and it cost a cycle.
- Pulling my own featureI shipped a keyboard focus ring that looked correct in isolation. In real layouts, wrappers sat flush to component edges so rings were clipped inside dialogs and filter modals. I paused it rather than leave a broken accessibility affordance in place.
How it works, step by step
- Pull production into the fileHTML to Figma to bring live screens in as artefacts, so the comparison was against what actually ships rather than what we remembered.
- First pass with OpenCodeAn automated analysis to surface systemic inconsistencies, then a manual review to catch what it missed.
- Ship the foundationsTypography, colour, spacing, radius and elevation pushed as V2 libraries, correct everywhere at once.
- Refactor components in pairsEach component became a design task and a development task, picked up by frontend engineers across teams as sprint side items.
- Verify before releaseEvery component checked in Storybook and in a sandbox environment before it went out, because a shared component failing is worse than a one off failing.
What comes next
Two threads open. The focus ring is an accessibility commitment waiting on a container fix, and the AI workflow is the part I am actively building out.
Looking back
The visible output is a token library and ten components. The two parts I would defend are the interaction layer, because it is why engineers stopped inventing colours, and the QA number, because three hours a week returned is the clearest evidence the system does work people used to do by hand. The focus ring is the honest counterweight: a good idea that failed on contact with real layouts, and pausing it was better than shipping accessibility theatre.



ME
DÖ
P
UY