Algolia · 2026 · Design lead, author and owner of the design rules, built with an engineer
I built a design reviewer, then gave it away.
Algolia's dashboard is where its customers configure search. Every change to it arrives as a front end pull request and nobody was checking those against the design rules.
I turned five design rules into an automated reviewer that comments on every front end pull request, backtested it before asking anyone to trust it, then merged it into the always on bot another team already ran.
One always on reviewer now covers every front end pull request and every rule survived the merge. Four of the nine findings that got an answer merged with nobody acting on them, which I publish as the failure metric rather than hide.
The principal move was giving away the tool, its name and my own metric, because none of them was the standard. The rules were.
Every product has a gap between the design that was approved and the code that ships. Inside Algolia's dashboard that gap is where an accessibility affordance quietly disappears and a component gets rebuilt by hand instead of taken from the design system. Nobody reviews it, so those regressions ship in silence.
I turned five design rules into an automated reviewer that comments on every front-end pull request, then gave the tool away to the team already running a bot so the rules would reach everything. I backtested it before asking anyone to trust it, rewrote my own headline number when it was challenged and binned my own metric when I saw it could not detect its own silence. One always-on reviewer now covers every front-end pull request. The rules inside it are mine to maintain.
Design can be right and still ship wrong. Nobody reviews that gap. A pull request, the moment an engineer proposes a change, is the last point where a mistake is cheap to fix. Between the design handover and the customer, an accessibility affordance disappears, an error branch gets built to log to the console and show the user nothing, a component that already exists in the design system gets rebuilt by hand. None of that is carelessness. It is what happens when the last checkpoint on design quality is a designer's attention. There are far more pull requests than designers.
The cost lands on the customer, who meets the unhandled error. It lands on the design system, which decays one hand-built component at a time. And it lands on the design team, who meet the regression months later and argue for a fix that would have been a one-line comment at the time.
The obvious answer was more guidance. Guidance only works on people who go looking. Documentation, a checklist, a pattern page about error states. The person about to ship an unhandled error branch is not looking. So I put the standard where the work happens, as rules an automated reviewer applies to every front-end pull request while the code is still cheap to change. Not a document engineers should read. A comment on their diff.
I called this the floor, not the ceiling. I said so from the start. Catching mistakes at scale is not preventing them. The structural fix, a design system with fewer ways to get it wrong, is slower and worth more. I took the floor first because its findings show where the ceiling needs the investment.
I owned the pilot, chose its first intervention and wrote the rules it still runs on. The idea of a review gate came from more than one direction, so I do not claim I dreamt it up alone. What is mine is the call to make it the pilot's first move, the five rules, the backtest, the pilot groups, the measurement and, by explicit agreement at the merge, the rules' content going forward. An engineer I built it with tuned and landed the first version and later routed the merge.
The five rules are held deliberately to the judgment calls a linter cannot make.
Two further rules were designed and held back, because a reviewer nobody trusts yet should say less.
A reviewer that cries wolf gets muted. The accurate version was unaffordable. The obvious accuracy fix was a verification pass on every pull request. At roughly eight dollars a pull request it was ruled out on cost at this repository's volume, so accuracy had to come from scoping the rules better instead. The second tension was tone. A gate that reads like a compliance form gets ignored, so I gave the tool a name and a voice on purpose. Personality turned out to be the first thing a merge costs you.
The first conversation was about evidence, not an idea. Before proposing rollout I ran the reviewer read-only against eight already-merged dashboard pull requests, roughly seven thousand lines of diff. It produced ten findings. Two were genuinely useful. The one real bug was a setup wizard whose completion step failed silently on error, no message, no recovery, nothing, exactly the "only built the happy path" failure the exercise existed to catch. I published the backtest's limits alongside it. One manual pass rather than the live model, a small sample that never tested the destructive-action rule and full file context the live bot would not always have, so the real noise floor could run higher.
Outcome.
Rollout was argued from a results table rather than a slide, which is also what made the next challenge possible.
I had claimed a low false-positive rate. Read plainly against my own table, the figure was nearer 80% false positives. The number was challenged. I checked it against the table and the challenge was right. I had been counting in a way that flattered me. The assertive tier held, the single Blocker was correct, but precision across all ten findings was about 20%. So I rewrote the headline rather than defend it, then rewrote the measure too. A single pass or fail number was hiding the shape of the failure, so I split every outcome three ways. Wrong. Correct but not worth a comment. A fair question that still spends a reader's attention. With that split the cause fell out at once. The rules were firing on patterns a pull request merely touched rather than introduced. I scoped every rule to new code only and reported the 20% as a pre-mitigation baseline instead of quietly replacing it with a better number.
Outcome.
The reviewer got more accurate and so did I. The proposal went forward carrying its worst number on the front page.
When a second reviewer appeared on the same pull requests, one voice was the only workable answer. It was not going to be mine. Two bots on one thread get each other muted. Two copies of the same guidance drift apart the moment either is edited. The other bot already ran on every push, while mine ran on an opt-in label, only when someone remembered. So I argued against running both and folded mine into theirs. Every rule survived intact. I declined the easy version, which was to carry my opt-in trigger across, because the trigger was the problem. I gave up the tool, its name, its voice and my trigger. I kept the standard, with a split agreed at the merge. They own and run the bot. I own and maintain the design rules inside it.
Outcome.
One always-on reviewer on every front-end pull request instead of two competing ones, an order of magnitude more coverage in a fraction of the time. I verified that against the live repository with my own counting script rather than taking it on report.
My own measurement went in the bin with the tool, because it could not detect its own silence. The pilot metric counted comments actioned off a label I controlled, so it could only measure what the reviewer had already decided to say. A real problem the rules never fired on leaves no label and no trace, so the metric was blind by construction to the failures that matter most and would have looked healthiest when it was most wrong. So every finding now carries a machine-readable tag naming the rule that produced it. A script I wrote counts from those instead. Inside a system I do not own, my contribution would have disappeared the moment measurement moved off my labels, so I got that requirement built in before the new reporting shipped and checked it in the code myself.
Ten days after the merge I caught my own counting script understating the results. It capped the pull request list with no warning when the real population ran past the cap. Run deep enough to cover every pull request, findings went from 15 to 21 and findings that had drawn a human reply went from 1 to 6. Everything I had quoted before had understated engagement by about a third. The corrected numbers flattered the project, which is exactly why I published the correction alongside them.
Strategic outcome.
Design's contribution to code quality is now countable inside a system design does not run, rule by rule. The count is one I would trust from someone else.
Over ten days the reviewer posted 21 findings, 14 of them against my five rules. Of those 14, nine reached a decision. Three were fixed, one was accepted and tracked, one was declined with a stated reason and four merged with nobody acting on them. None were argued down as inaccurate. I also published a failure metric on purpose. Of the findings that got any answer at all, 44% merged unaddressed, which is four of the nine. A review process that only publishes its wins is not one anyone senior should believe.
The strongest evidence is not a number. On a finding about a picker that could render an empty, unexplained list, an engineer with no hand in building the reviewer replied unprompted, naming the commit that fixed it.
Product.
Happy-path-only changes get flagged on the diff, while the fix is still a comment rather than a ticket.
Engineering.
One reviewer rather than two on every front-end pull request, with each finding traceable to the rule behind it.
Design.
Five design rules run on every change to the front end without a designer in the room. The rules stay mine to maintain.
Organisational.
A design standard that lives in the engineering workflow rather than the wiki, with a published failure metric attached.
I cannot claim a measured improvement in shipped design quality, so I do not. Some of those outcomes are inferred from commits rather than stated by the author. A rule that never fires on a real problem still leaves no trace at all.
The principal signal is not that I built a bot. It is that I put a design standard where it would be argued with, then gave away everything that was not the standard so the standard would reach more code. The tool, the name, the trigger and my own metric all went. The rules stayed and I still own them.
This work shows that I can set a standard other teams run without me, back a proposal with evidence before asking for trust, take a direct challenge to my own number and rewrite it rather than defend it, design measurement that survives success and choose the outcome over the credit when the two pull apart.
Standards scale when they live in the workflow rather than the wiki. A rule that arrives on a diff gets argued with. The same rule on a documentation page gets nothing, not even disagreement. The lesson I would not have predicted is to check that your measurement survives success. Mine did not. The version where I keep the bot, the name and the dashboard is the version where the standard reaches fewer pull requests.
The principal move. I gave away the tool, its name and my own metric, because none of them was the standard. The rules were.
→ Full story and decisions in an interview