Algolia · 2026 · Design lead, author and owner of the design rules, built with an engineer

AI UX review agent.

Too long, didn’t read.

I built a design reviewer, then gave it away.

Algolia's dashboard is where its customers configure search. Every change to it arrives as a front end pull request and nobody was checking those against the design rules.

I turned five design rules into an automated reviewer that comments on every front end pull request, backtested it before asking anyone to trust it, then merged it into the always on bot another team already ran.

One always on reviewer now covers every front end pull request and every rule survived the merge. Four of the nine findings that got an answer merged with nobody acting on them, which I publish as the failure metric rather than hide.

The principal move was giving away the tool, its name and my own metric, because none of them was the standard. The rules were.

Introduction.

Every product has a gap between the design that was approved and the code that ships. Inside Algolia's dashboard that gap is where an accessibility affordance quietly disappears and a component gets rebuilt by hand instead of taken from the design system. Nobody reviews it, so those regressions ship in silence.


I turned five design rules into an automated reviewer that comments on every front-end pull request, then gave the tool away to the team already running a bot so the rules would reach everything. I backtested it before asking anyone to trust it, rewrote my own headline number when it was challenged and binned my own metric when I saw it could not detect its own silence. One always-on reviewer now covers every front-end pull request. The rules inside it are mine to maintain.

My role.

Design lead
Author and owner of the design rules

Scope.

Five design rules
Reviewer prompt and voice
Read-only backtest
Three-week pilot
Merge into the always-on bot
Measurement and reporting

Influenced.

How design quality is checked at the moment of coding
Which reviewer the front end runs
What gets measured and how a finding stays traceable

Outcome.

One always-on reviewer on every front-end pull request
Every rule survived the merge
A published failure metric

Design can be right and still ship wrong. Nobody reviews that gap. A pull request, the moment an engineer proposes a change, is the last point where a mistake is cheap to fix. Between the design handover and the customer, an accessibility affordance disappears, an error branch gets built to log to the console and show the user nothing, a component that already exists in the design system gets rebuilt by hand. None of that is carelessness. It is what happens when the last checkpoint on design quality is a designer's attention. There are far more pull requests than designers.

The cost lands on the customer, who meets the unhandled error. It lands on the design system, which decays one hand-built component at a time. And it lands on the design team, who meet the regression months later and argue for a fix that would have been a one-line comment at the time.


The obvious answer was more guidance. Guidance only works on people who go looking. Documentation, a checklist, a pattern page about error states. The person about to ship an unhandled error branch is not looking. So I put the standard where the work happens, as rules an automated reviewer applies to every front-end pull request while the code is still cheap to change. Not a document engineers should read. A comment on their diff.

One design review on one pull request A design review summary scores five UX promises for one pull request, two passing and three with findings. Below it, on the changed line of code that adds a destructive menu item, the reviewer posts a blocker explaining that deleting a row has no confirmation or undo and pointing to the confirmation pattern the feature already uses. One review on one pull request Review bot left a comment Design review, five UX promises This pull request adds a setup modal and a table of saved items. The overall interaction design is solid and the destructive flow for the whole feature is gated behind a confirmation. Three issues need attention. Findings 3 Blocker · 0 Consider · 0 Ask Promise 1 Feedback pass Save, add and remove all show pending states. Promise 2 Safety finding Deleting a row commits permanently with no confirmation or undo. Promise 3 Recovery finding A failed fetch is swallowed and looks like an empty library. Promise 4 Accessibility finding An icon button has an empty label, so screen readers get nothing. Promise 5 Consistency pass Design system components are used throughout. …/features/items/ItemsTable.tsx 245 + })} 246 + </MenuButton.Item> 247 + <MenuButton.Item 248 + variant="destructive" Review bot on the changed line Blocker · Promise 2. Deleting a row commits permanently with no confirmation or undo Choosing Delete from the row menu removes the record at once. There is no confirmation, undo toast or cancel path. Apply the confirmation pattern this feature already uses for removing the whole setup.
One review landing on one pull request. The summary scores the five promises, then the finding sits on the exact line that caused it. Redrawn from the pull request with neutral names.

I called this the floor, not the ceiling. I said so from the start. Catching mistakes at scale is not preventing them. The structural fix, a design system with fewer ways to get it wrong, is slower and worth more. I took the floor first because its findings show where the ceiling needs the investment.


I owned the pilot, chose its first intervention and wrote the rules it still runs on. The idea of a review gate came from more than one direction, so I do not claim I dreamt it up alone. What is mine is the call to make it the pilot's first move, the five rules, the backtest, the pilot groups, the measurement and, by explicit agreement at the merge, the rules' content going forward. An engineer I built it with tuned and landed the first version and later routed the merge.

The five rules are held deliberately to the judgment calls a linter cannot make.

Two further rules were designed and held back, because a reviewer nobody trusts yet should say less.


A reviewer that cries wolf gets muted. The accurate version was unaffordable. The obvious accuracy fix was a verification pass on every pull request. At roughly eight dollars a pull request it was ruled out on cost at this repository's volume, so accuracy had to come from scoping the rules better instead. The second tension was tone. A gate that reads like a compliance form gets ignored, so I gave the tool a name and a voice on purpose. Personality turned out to be the first thing a merge costs you.


1. Backtested it before asking anyone to trust it

The first conversation was about evidence, not an idea. Before proposing rollout I ran the reviewer read-only against eight already-merged dashboard pull requests, roughly seven thousand lines of diff. It produced ten findings. Two were genuinely useful. The one real bug was a setup wizard whose completion step failed silently on error, no message, no recovery, nothing, exactly the "only built the happy path" failure the exercise existed to catch. I published the backtest's limits alongside it. One manual pass rather than the live model, a small sample that never tested the destructive-action rule and full file context the live bot would not always have, so the real noise floor could run higher.

An automated design review comment on a pull request diff, with the rule that produced it named on the comment A compressed diff shows a list rendering straight from a fetch, with a loading state handled but no empty state. Beneath it sits a posted design review comment asking for an empty state. At the foot of the comment a provenance line names Rule 4, handle loading, empty and error states, together with the rule file it came from. AUTOMATED DESIGN REVIEW · FRONT END PULL REQUEST src/components/list/DataList.tsx const { items, isLoading } = useItems(query) if (isLoading) return <Spinner /> + return <List>{items.map(renderRow)}</List> Design review·Consider This list renders straight from the fetch. If the query returns no matches, the user sees an empty box with no explanation and nothing to do next. Add an empty state that says why the list is empty and what to try. Rule 4 · handle loading, empty and error states design-rules/unhappy-paths.md
Every finding carries the rule that fired and a link to the rule itself. Schematic, not a screenshot.

Outcome.

Rollout was argued from a results table rather than a slide, which is also what made the next challenge possible.

2. Rewrote the headline number when it was challenged, then the measure behind it

I had claimed a low false-positive rate. Read plainly against my own table, the figure was nearer 80% false positives. The number was challenged. I checked it against the table and the challenge was right. I had been counting in a way that flattered me. The assertive tier held, the single Blocker was correct, but precision across all ten findings was about 20%. So I rewrote the headline rather than defend it, then rewrote the measure too. A single pass or fail number was hiding the shape of the failure, so I split every outcome three ways. Wrong. Correct but not worth a comment. A fair question that still spends a reader's attention. With that split the cause fell out at once. The rules were firing on patterns a pull request merely touched rather than introduced. I scoped every rule to new code only and reported the 20% as a pre-mitigation baseline instead of quietly replacing it with a better number.

Backtest findings sorted three ways Ten squares, one per backtest finding. Two are useful, the one real bug and one other. Four are fair questions that still cost attention. Four were wrong or correct but not worth a comment. Precision was about 20 percent before rules were scoped to new code. About 20% precision across the ten backtest findings, before rules were scoped to new code Useful Useful A fair question A fair question A fair question A fair question Wrong, or correct but not worth a comment Wrong, or correct but not worth a comment Wrong, or correct but not worth a comment Wrong, or correct but not worth a comment Useful 2 Useful The one real bug and one Consider. A fair question 4 A fair question Asks that were not wrong but still cost a reader's attention. Wrong, or correct but not worth a comment 4 Wrong, or correct but not worth a comment The remaining Considers. Read-only backtest against eight already-merged pull requests, about seven thousand lines of diff.
The ten backtest findings sorted three ways. Two were useful, four were fair questions that still cost attention and four were wrong or not worth a comment. This is the baseline before rules were scoped to new code.

Outcome.

The reviewer got more accurate and so did I. The proposal went forward carrying its worst number on the front page.

3. Folded it into another team's bot rather than defend my own

When a second reviewer appeared on the same pull requests, one voice was the only workable answer. It was not going to be mine. Two bots on one thread get each other muted. Two copies of the same guidance drift apart the moment either is edited. The other bot already ran on every push, while mine ran on an opt-in label, only when someone remembered. So I argued against running both and folded mine into theirs. Every rule survived intact. I declined the easy version, which was to carry my opt-in trigger across, because the trigger was the problem. I gave up the tool, its name, its voice and my trigger. I kept the standard, with a split agreed at the merge. They own and run the bot. I own and maintain the design rules inside it.

Three versions of the review comment Three review comments side by side. The first is a formal design review scoring five UX promises. The second is the same review with a name, the design goblin, and a playful voice. The third is the team's AI code review bot after the merge, with no name or voice, where the finding still ends with the design rule it came from. Version 1 My reviewer Review bot left a comment Design review, five UX promises The overall interaction design is solid. Three issues need attention before this ships. Findings 3 Blocker · 0 Consider · 0 Ask Promise 1 Feedback pass Promise 2 Safety finding Promise 3 Recovery finding Promise 4 Accessibility finding Promise 5 Consistency pass This automated review is advisory. Its labels and comments never block merging. Five promises, a formal voice. Version 2 The design goblin Review bot left a comment The design goblin took a look The goblin sniffs through the diff and finds mostly empty space where a banner used to lurk. Nothing broken, nothing missing. The goblin is almost disappointed there's nothing to scold. 0 Blocker · 0 Consider · 0 Ask · the goblin found nothing to grumble about. What the goblin poked at 1 Feedback pass 2 Safety pass 3 Recovery pass 4 Accessibility pass 5 Consistency pass The goblin is advisory only. It grumbles, it never blocks your merge. Same rules, now with a name and a voice. Version 3 The team's bot, after the merge Review bot left a comment Code review (AI) Findings 0 blocker · 1 consider · 0 nit …/features/new/Provider.tsx 138 + const options = useMemo(() => { consider · A fetch error silently hides every option The provider returns an empty list for both loading and error, so the user sees no options, no message and cannot continue. Show a skeleton or an error while the list loads. Rule ux/error-and-recovery-states Name and voice gone. The rule tag carried across. Open full size
Three versions of the same review. Mine, then the design goblin with a name and a voice, then the team's bot after the merge. The name and voice went, the rule on every finding stayed. Redrawn from the pull requests with neutral names.

Outcome.

One always-on reviewer on every front-end pull request instead of two competing ones, an order of magnitude more coverage in a fraction of the time. I verified that against the live repository with my own counting script rather than taking it on report.


My own measurement went in the bin with the tool, because it could not detect its own silence. The pilot metric counted comments actioned off a label I controlled, so it could only measure what the reviewer had already decided to say. A real problem the rules never fired on leaves no label and no trace, so the metric was blind by construction to the failures that matter most and would have looked healthiest when it was most wrong. So every finding now carries a machine-readable tag naming the rule that produced it. A script I wrote counts from those instead. Inside a system I do not own, my contribution would have disappeared the moment measurement moved off my labels, so I got that requirement built in before the new reporting shipped and checked it in the code myself.

A pull request label beside a rule tag On the left, the reviewer adds a single design review pass label to the whole pull request. On the right, one finding ends with the name of the design rule that produced it. What a label records Review bot added design-review:pass and removed design-review One label for the whole pull request, set by what the reviewer chose to say. What a rule tag records consider · The radio group lost its accessible name The replaced element carried the label, the new one drops it. Rule ux/accessible-interactions One tag on every finding, so each one traces back to the rule behind it.
A label can only record what the reviewer chose to say about a whole pull request. A rule tag on every finding makes each one traceable. Redrawn from the pull requests.

Ten days after the merge I caught my own counting script understating the results. It capped the pull request list with no warning when the real population ran past the cap. Run deep enough to cover every pull request, findings went from 15 to 21 and findings that had drawn a human reply went from 1 to 6. Everything I had quoted before had understated engagement by about a third. The corrected numbers flattered the project, which is exactly why I published the correction alongside them.

Strategic outcome.

Design's contribution to code quality is now countable inside a system design does not run, rule by rule. The count is one I would trust from someone else.


Over ten days the reviewer posted 21 findings, 14 of them against my five rules. Of those 14, nine reached a decision. Three were fixed, one was accepted and tracked, one was declined with a stated reason and four merged with nobody acting on them. None were argued down as inaccurate. I also published a failure metric on purpose. Of the findings that got any answer at all, 44% merged unaddressed, which is four of the nine. A review process that only publishes its wins is not one anyone senior should believe.

The strongest evidence is not a number. On a finding about a picker that could render an empty, unexplained list, an engineer with no hand in building the reviewer replied unprompted, naming the commit that fixed it.

Product.

Happy-path-only changes get flagged on the diff, while the fix is still a comment rather than a ticket.

Engineering.

One reviewer rather than two on every front-end pull request, with each finding traceable to the rule behind it.

Design.

Five design rules run on every change to the front end without a designer in the room. The rules stay mine to maintain.

Organisational.

A design standard that lives in the engineering workflow rather than the wiki, with a published failure metric attached.


I cannot claim a measured improvement in shipped design quality, so I do not. Some of those outcomes are inferred from commits rather than stated by the author. A rule that never fires on a real problem still leaves no trace at all.


The principal signal is not that I built a bot. It is that I put a design standard where it would be argued with, then gave away everything that was not the standard so the standard would reach more code. The tool, the name, the trigger and my own metric all went. The rules stayed and I still own them.

This work shows that I can set a standard other teams run without me, back a proposal with evidence before asking for trust, take a direct challenge to my own number and rewrite it rather than defend it, design measurement that survives success and choose the outcome over the credit when the two pull apart.


Standards scale when they live in the workflow rather than the wiki. A rule that arrives on a diff gets argued with. The same rule on a documentation page gets nothing, not even disagreement. The lesson I would not have predicted is to check that your measurement survives success. Mine did not. The version where I keep the bot, the name and the dashboard is the version where the standard reaches fewer pull requests.

A cartoon illustration of a ghostly goblin rising from a mound of earth at night, surrounded by old interface screens like gravestones, with its goggles and clipboard left behind.
The design goblin after the merge. An illustration made for an internal presentation, not part of the product.

The principal move. I gave away the tool, its name and my own metric, because none of them was the standard. The rules were.

→ Full story and decisions in an interview