For the agent driving it

Wired by the repository, not by each machine

Adoption that depends on somebody remembering a per-machine install is adoption that stops at the person who set it up. winwright ships as a Claude Code plugin: two commands in the repository wire it, every clone is wired, and nothing is added to any path.

The friction this removes

What a hand-rolled harness costs, and what replaces it

None of this is a defect in the tools it names. Every one of them is what a general-purpose automation library does to somebody who has to answer, at the end of a run, whether the thing they were checking was actually checked.

A pass and a check that never ran are the same colour
the script you have
Assert.True(found is not null)
A precondition that was absent throws, is caught, and turns into a skip nobody counts — or worse, into a branch that returns early and reports nothing at all. The total moves and the colour does not.
here
DEGRADED (exit 2) - 1 unchecked (all the desk's)
The unevaluated reading is named in the summary, whose the absence was is named beside it, and the process exits 2 — a number CI can act on without anybody reading a word of the output.
The locator is a different string in every verb
the script you have
FindFirst(TreeScope.Children, new AndCondition(...))
A condition tree is written per call site, so the same control is addressed four different ways in four places, and the one that breaks is the one nobody recognises as the same control.
here
Window#main > Pane > Button#save
One grammar, read the same way by every verb, refused at parse time with the position and the reason rather than at run time with a null.
Every act is a click, so every act needs the screen
the script you have
SetForegroundWindow(h); MoveTo(x, y); Click()
Which means the run owns the desk, a notification that steals the foreground is a failure, and the same script cannot be run on the machine somebody is working on.
here
Act.Invoke(subject)
A pattern act asks the control through its own accessibility peer and needs no foreground. The verbs that do synthesise input are marked as such in the catalogue rather than discovered on a red run.
The expected string is typed into the test
the script you have
Assert.Equal("Relatório mensal", label)
A second copy of the truth, and it is the copy that goes stale — silently, on the day somebody edits the language file and nobody edits the test.
here
"covers": "report.labels"
The expected set is derived from the application's own language files, so switching the resolved language switches the expectation with it — and a key none of the files carries refuses the run rather than matching nothing. Every value in the set carries the file and line it came from.
A screenshot is written whatever it contains
the script you have
CopyFromScreen(bounds); bmp.Save(path)
A dialog over the window, a backdrop transmitting what is behind it, a page still computing, or nothing rendering at all — each writes a file, and each exits zero.
here
1 window(s) stand over 376x166 at 62,90, taking 15360 of its 62416 pixel(s): 'Update available' (pid 14820) over 240x64 at 194,92
Every way the picture can lie is checked, and the reading names the intruder, its process and the rectangle it covers — rather than cropping around it and writing the file anyway.
A case is two hundred lines that mostly repeat the previous case
the script you have
// twenty-seven copies of one runner
The loop, the waits and the verdict logic are re-authored per case, so a fix to any of them is applied to some of the copies and the rest keep the bug.
here
report.cases.json
Steps, locators, acts and expectations are fields; the loop, the waits, the attempts and the verdict belong to the engine. What is left in the file is the part that is about your application.

Four tools, arriving as schemas rather than as prose somebody has to read and retype

Which is the difference between a refusal and a guess: an agent that writes a case field by field is corrected at insertion, where the fix costs a retry, rather than at run time, where it costs a red run somebody has to read.

Reading — nothing is launched, nothing is pressed
winwright_formatreads
every field of a file, a case, a step and a fixture, whether it is required, and the closed list of what it accepts
winwright_vocabularyreads
every act, what each one needs said beside it, and whether the engine may repeat it
winwright_checkreads
a case read back before the file exists — either the loader's own refusal, addressed as cases[0].steps[1].act, or what a run of it would do
Running — it launches the application
winwright_rundrives
the cases a selection asks for: the verdict, a line per case that ran and per case it left alone, the exit code, and what outlived the run

The split is the whole reason the last one is a separate tool. Whether a file parses is a claim nothing about the machine can change; whether it passed is not — and a desk that cannot observe answers a hole, naming which condition is missing, rather than a red.

The two commands, and the one step they do not cover

Run once, in the repository that drives the application:

claude plugin marketplace add alegauss/winwright --scope project
claude plugin install winwright@alegauss --scope project

Both write into that repository's .claude/settings.json, so committing that file wires every clone. The server and the guard are .NET processes the plugin launches, so they are built once — dotnet build -c Release in the plugin's own clone. Skip it and you are told, not left guessing: each is wired through a launcher that looks for its assembly, and where there is none it writes the missing surface and the build command to stderr and exits 1. Never 2 — denying every write because a build is missing would put the guard in front of everything instead of in front of a harness script.

The hook that keeps the shortcut closed

A hand-written harness script is always available and always faster in the moment, and that is exactly how a 2,732-line one happens. So the plugin registers a PreToolUse hook: a write whose content names the engine's acting, locating or asserting namespaces is denied, and the refusal names the case file and the tool that replace it. It arrives before the work rather than after it — the difference between being asked to write the other thing and being asked to delete what you just wrote.

And it stays out of its own way in three places

  • ✓A scenario-file write is never denied — that is the verb it is pointing at.
  • ✓A project referencing the engine's source is never denied: a suite that drives windows on purpose is the one place a harness belongs, and a guard you turn off to work on the tool is a guard whose false denials nobody hears about.
  • ✓Anything it cannot read, it allows. A hook that denies what it did not understand is one that gets removed, after which nothing is guarded at all.

The skill that is not always loaded

It loads when a window is in play rather than on every turn, and what it says is which loop answers which question — which is the whole of what an agent needs in order to reach the right tool instead of the nearest one.

What it deliberately refuses

winwright is the substrate; the intelligence is the caller's. It reads trees, presses controls through their own patterns and assembles a verdict — it does not decide what is worth asserting.

✗No model

It calls no LLM.

✗No prompts

It stores none.

✗No recorder

Clicks never become a case.

✗No invented assertion

The tool never writes the test.

✗No second desk

It drives the one you are on.

✗No telemetry

Nothing is measured or sent.