Why Claude Code Fixes the Wrong Element

Claude Code fixes the wrong element because it never saw the control you meant. It saw a screenshot, guessed at pixels, and edited the nearest plausible match.
When you write "make this button smaller" and three near-identical buttons sit in one row, a pixel-guessing agent has no reliable way to pick. It still reports success.
Most write-ups blame the model. The real gap is upstream: the agent was handed ambiguity and asked to resolve it where it cannot.
Why does Claude Code edit the wrong component so often?
The agent works from two weak signals: a flat image and your words. Neither names a specific control.
A screenshot is a grid of pixels. It contains a Save button visually, but nothing in the image says AXButton "Save" at a frame {x, y, w, h}.
The agent has to infer that. Inference over near-duplicates is where it drifts.
Then your prose adds fog. "This card", "the top one", "that dropdown" are deictic references that only resolve if the listener already knows where you were pointing.
Claude Code was not in the room. It fills the gap with a guess, and the guess lands on a sibling.
The failure is not comprehension. It is resolution. The agent understands "make the button smaller" — it just cannot tell which button, because nothing it received names one.
This is why the bug repeats across models. Swapping Sonnet for Opus does not add the missing fact.
The element was never identified in the first place, so a smarter reader of the same ambiguous input still guesses.
When Claude Code fixes the wrong element, is it the model or the input?
It is almost always the input. You can test this in one session.
Give Claude Code a screenshot and "fix the spacing on this row". Watch it edit a component that looks right in the image but sits three divs away in the DOM.
Now give it the exact selector or the exact React component name and file. The same model lands the edit first try.
Same weights, different outcome. The variable that changed was precision of reference, not intelligence.
"Add more context" is not the fix. A bigger screenshot, more of the file tree, or a longer description gives the agent more to read and no more certainty about the one element you mean.
Ambiguity does not shrink with volume.
So the useful question stops being "which model is best at UI". It becomes "how do I hand the agent a named target instead of a picture of one".
That is a solvable input problem. I go deeper on the mechanics in how AI agents know which UI element you mean.
What does naming the element actually mean?
On macOS, every on-screen control already has a name. The Accessibility API exposes the element under any point — role (AXButton), label ("Save draft"), value, frame, and parent chain.
VoiceOver has read this tree for years.
Naming the element means resolving that tree at the exact spot you pointed. You hand the agent the result: AXButton "Save draft" inside AXToolbar at a known frame.
That single change reframes the task. The agent stops asking "which of these did they mean" and starts matching a specific control to your code by label and role.
This is the design behind PinVari: you hold ⌥⌘A, circle or point at any on-screen element, and speak. It screenshots, transcribes on-device, and resolves the exact named accessibility element you circled — role, label, frame — before your agent ever reads it.
A few honest details, because the engineering is the point.
The capture overlay sits topmost, so a naive hit-test would resolve to PinVari's own window. chainExcludingSelf walks the on-screen window list and hit-tests the real app underneath.
Electron and Chromium build their accessibility tree lazily. PinVari sets AXManualAccessibility and retries for roughly 150ms until a labeled element appears, instead of reporting an empty group.
When a point lands on a bare AXGroup, labeledDescendant descends to the deepest labeled child so you get "Save draft", not "group".
Chromium identity attrs (AXDOMIdentifier, AXDOMClassList) give a title-less node a name.
None of that is exotic. It is the difference between shipping the pixel and shipping the name.
How do confidence and provenance stop the wrong-element guess?
Naming is necessary but not sufficient. Sometimes the resolution itself is uncertain — a title-less icon, a canvas region, an overlapping tooltip.
A system that pretends every resolution is perfect just moves the guess one layer down.
So each resolved element carries two extra fields when it reaches the agent over the local MCP server.
Confidence
A score for how sure the resolver is that this is the element you meant.
A circled element at high confidence is executed directly. Below the threshold the agent asks instead of guessing.
Provenance
Whether you deliberately circled it (trust it) or the cursor merely dwelled over it in passing (a hint, not a command).
Dwell detection is about 0.2s of hover. It resolves an element without a circle, but the agent should treat it as a candidate.
"Ask when unsure" is the whole trick. The wrong-element bug is a system silently resolving a coin-flip. Attach a confidence score and the coin-flip becomes a visible question.
Provenance also disambiguates the sentence itself. If you say "move this below that" while your pointer travels across two controls, PinVari records a timestamped pointer trail.
It binds each deictic word to the element the cursor was on at the instant you spoke it. The agent receives "this" → AXButton "Publish", "that" → AXButton "Save draft" — not one blurry region for both.
Does pointing beat writing a longer prompt?
For "which element", yes. Prose describes; pointing designates.
Here is how the common ways of telling an agent stack up. Skip the wide grid.
Read each as a card.
Screenshot plus "this button"
What the agent receives: pixels plus an ambiguous word.
Wrong-element risk: high. Near-duplicates collide.
Longer written description
What the agent receives: more prose, still no ID.
Wrong-element risk: high. Volume is not precision.
Hand-typed CSS or React selector
What the agent receives: a real ID, if you find it.
Wrong-element risk: low, but slow and easy to mistype.
Circle the element and speak
What the agent receives: AXButton "Save draft", frame, confidence, provenance.
Wrong-element risk: low. Resolved, named, scored.
Typing an exact selector works too. Precision is what fixes this.
Circling gets you the same precision in a second, without leaving the screen to hunt for the identifier.
There is a token angle as well. A named element and its window text are far cheaper to send than a full-resolution screenshot the agent re-parses on every turn.
I break that cost down in why screenshots waste Claude Code tokens.
What happens when the UI has no accessibility data?
Sometimes the tree is empty. A game canvas, a hand-rolled Electron surface, a WebGL view — the point lands somewhere with no labeled element.
Honesty matters more than a fake answer. When AX comes back blank, PinVari falls back to on-device Vision OCR on the exact region you circled.
It reads the visible text and hands that to the agent instead of inventing a label. On a browser it also reads the real URL from the AXWebArea, so the agent knows the page, not just the picture.
The rule holds in both paths: give the agent a grounded reference or admit uncertainty. Never manufacture confidence.
That is also why this runs on-device with no API keys and nothing uploaded by default. The resolver reads your screen locally and passes only the named result to the agent you already use.
If you are wondering how much any coding agent can see of your screen, can Claude Code see my screen covers the real answer.
Full focused-window text (up to 40,000 characters, including text scrolled out of view) plus an interactive-element map travels with the capture. Optional whole-screen OCR covers AX-blind surfaces.
When Claude Code fixes the wrong element, it is a reference problem, not a reasoning problem. Fix the reference — name the element, score it, mark how you pointed at it — and the "it fixed the wrong thing" reports mostly stop.
FAQ
Why does Claude Code keep editing the wrong component?
Because it works from a screenshot and your words, and neither names a specific control. When several components look alike, the agent infers which one you meant and often lands on a sibling.
Hand it the resolved, named element instead of a picture and the guessing stops.
Does a better prompt fix the wrong-element problem?
Rarely. A longer description adds words, not certainty about which element you mean.
The reliable fix is precise reference — an exact selector, or circling the element so it resolves to a named accessibility node with a frame and label.
How do I tell an AI agent exactly which UI element I mean?
Give it something that identifies one element unambiguously: a specific CSS or component selector, or a spatial capture that resolves the element under your pointer to its role, label, and frame.
Vague pointers like "this" or "the top one" only work if the agent already knows where you were looking.
What is a confidence score on a resolved UI element?
It is how sure the resolver is that it picked the element you meant. High confidence with a deliberate circle gets executed directly.
Low confidence gets flagged so the agent asks you rather than silently editing the nearest match.
Can Claude Code fix the right element in Electron or Chromium apps?
Yes, once the accessibility tree is awake. Chromium builds its AX tree lazily, so a resolver has to set AXManualAccessibility and retry until labeled elements appear.
If the surface is a true canvas with no elements, OCR on the circled region is the honest fallback.
Is pointing at an element faster than writing the selector by hand?
For identifying which element, yes. Circling resolves the named element in about a second without leaving the screen to search the DOM.
You get the same precision a hand-typed selector gives, with far less effort.
Hand your agent the exact element
PinVari resolves what you point at into a named, executable instruction — on-device, no keys, your own agent. One click inside PinVari connects Claude Code, Cursor, VS Code or Codex — or paste one CLI line from pinvari.com/connect.
PinVari → Connect → your agent (one click)Get PinVari — $39 →


