The recorder watches a desktop application while you work in it and writes down the controls you touch. It is a reader: it asks Windows UI Automation which element is focused and reads that element's accessible Name, ControlType and AutomationId. It never clicks, never types and never sends synthetic input, which is banned across this product. The recorder looks where the Robot acts.
#What the recorder observes
Its engine is a poller. Every 200 milliseconds it asks for the focused element and prints one JSON object per line each time the focus changes. Reading focus instead of installing an input hook is why it needs no administrator rights and installs nothing. The same choice explains its blind spot: it cannot see a keystroke, so it records the field you typed into and never the text. That value is recovered in review, through the control's Value pattern.
The recorder also subscribes to the actions an application performs on itself. An Invoke raises a press event. A
Toggle or a Selection raises a note or a log line rather than a step. Every control is recorded with the window it
lives in, because a name is not unique across a desktop. Nine Notepad windows can each hold a document named
Text editor. A locator that names only the control cannot tell them apart.
#How a session starts
There are two ways to start one. The headless capture watches for a fixed number of seconds. The visual recorder opens a small Start and Stop window, which also offers Pause, Resume and a comment box:
# the visual recorder: press Start, demonstrate, press Stop
npx tsx orchestrator/cli.ts desktop record --ui
# a headless capture of one window, 60 seconds long
npx tsx orchestrator/cli.ts desktop record --seconds 60 --poll-ms 200 \
--window "Calculator" --out traces/recording.jsonl
Scope the capture to one window whenever you can: a recording is a demonstration inside one window and the capture
format says so. When you name a window the poller subscribes to that window's subtree and refuses
DESKTOP_WINDOW_NOT_FOUND rather than widening the search. A desktop wide capture of one toggle produced 758 events,
755 of them an unrelated progress bar. Scoped to the window, the same toggle produced one event.
#How you indicate an element
In the visual recorder you name the window first. The activity picker and the Indicate button stay disabled until one is chosen, because the window is the scope. You choose an activity, press Indicate and point at one control. The pick happens on a dwell, the pointer resting on the same element for about 900 milliseconds, or on a click. A green box follows the pointer and shows which element would be picked, so the pick is never a guess. Escape cancels. Point at the anchor names a stable element for finding a later one; the anchor rides on the next step rather than becoming a step of its own.
In a headless capture the indication is the work itself: each focus change is the signal.
#How the recorded sequence accumulates
Each capture goes into an append only stream, one JSON object per line. A focus change becomes a step. A summary line closes the run. A comment or a capture names the step it belongs to and is attached on the way in, so a trace reaches review with its notes already on their steps. Pause, resume and drop are written as lines too, because a gap in a recording that nobody explained is a gap nobody can trust.
Repeats fold rather than drop. A page that takes focus back between keystrokes produces one step per bounce, so a run
of consecutive focus changes on the same control becomes one step carrying repeat: N. The count stays in the
record. A genuine return to the same control later is its own step, because a run is only a run while it is
consecutive. Recording is the first of four passes: capture, prove, review and emit. The trace is evidence, not an
automation.
#Things that catch people out
The recorder cannot see what you typed. Recover the value in review rather than expecting it in the capture.
A step rides on a focus change, so a static text that cannot take focus is picked by pointing at it with an anchor, not by focusing it.
A recording is scoped to one window. An element from another window is refused at the point of capture rather than written under a window it does not belong to.
Closing the recorder window while a recording is running still writes the end line, so a capture is never left without its close. "You focused nothing" and "we heard nothing" are kept apart on purpose.