What Is Spatial Context for AI Coding Agents?

GuidesAugust 22, 20269 min readBy PinVari
What Is Spatial Context for AI Coding Agents?

Spatial context is everything on your screen that your repository cannot tell an AI coding agent: the screenshot, the exact element you are pointing at, the words you say about it, and the full text of the window it lives in. Delivered together, so the agent acts on the right thing.

It is the difference between "the toggle is broken" and AXSwitch "Email notifications" inside AXWindow "Settings", at 0.94 confidence, with your sentence attached. Most discussion of agent context stops at the repo — files, symbols, tests, a CLAUDE.md.

That is the half the agent can already read on its own. Spatial context is the half only you can see, and handing it over badly is why agents confidently patch the wrong element.

What is spatial context in plain terms?

Think about what actually happens when you find a bug in a running app. You are looking at a screen.

You know exactly which control is wrong, because your eyes are on it. You know what should happen instead.

And you have a repository open in another window that contains, somewhere, the code behind that control. The agent has the repository.

It does not have the screen, your gaze, or your finger. Everything on that side of the gap is spatial context.

It breaks into four things the agent is missing:

  1. What the screen looks like — the pixels.
  2. Which thing on it you mean — the region and the gesture.
  3. What you want done to it — your spoken or typed instruction.
  4. What surrounds it — the labels, the error text, the state of the window.

Give it one of those and it guesses. Give it all four, fused into a single named target, and it acts.

Key

Spatial context is not "let the AI see my screen." Seeing is the cheap part. The expensive part is resolution — turning a screenshot plus a gesture plus a sentence into one unambiguous, named thing the agent can act on.

Why is spatial context different from repository context?

Repository context is symbolic, stable, and already machine-readable. A function has a name, a file path, and a line number.

Two agents reading the same repo see the same thing. Spatial context is none of that.

It is transient, visual, and indexed by where your pointer happens to be at one moment. Scroll the page and it changes.

Resize the window and it changes. Switch to dark mode and half the pixels change while the meaning does not.

That instability is why the naive approaches fail. Pixel coordinates — "the element at (842, 391)" — are the most obvious encoding and the most fragile one.

They break on scroll, on a different display scale, on a window a few points wider than yours. The stable encoding is a name.

AXButton titled "Publish", inside AXGroup "Composer toolbar", inside AXWindow "Draft". That survives scrolling, resizing, and theming, because it describes identity rather than position.

I go deeper on the structure itself in what the accessibility tree is.

Where does the idea of spatial context come from?

It is not new. In 1980, Richard A. Bolt published "Put-That-There" at SIGGRAPH: a user pointed at a shape and said "put that there."

Speech supplied the verb and pointing supplied the nouns. Neither channel could have carried the command alone.

That is the founding experiment for spatial context, and its lesson is still the whole thing: the words and the pointing are one signal, not two. Forty-five years later, most tooling still treats them as two — a screenshot in one field, a text prompt in another, and a model left to align them.

The linguistics term for the phenomenon is deixis: words whose meaning depends entirely on the situation they are uttered in. "This", "that", "here", "there", "now".

A deictic word carries no content of its own. It is a pointer, and it resolves only against the context of the utterance.

When you say "make this blue" while your cursor is over a button, the sentence is complete for a human standing beside you and useless for an agent reading it an hour later in a log. The missing piece is not vocabulary.

It is the pointing, timestamped to the word.

Tip

If your bug reports are full of "this", "that", and "here", you are not being sloppy. You are speaking naturally. The fix is a tool that captures what you were pointing at when you said each word.

What are the channels of spatial context?

Four channels carry it in, and a fifth step decides whether any of it is usable.

Channel 1: The screenshot (pixels)

The screenshot is what the screen looked like. It is the easiest channel to produce and the weakest on its own.

A flat image has no element names, no roles, no frames. The model sees shapes and infers meaning.

A full-resolution capture is a large block of image tokens that carries no structure. The arithmetic is in why screenshots waste Claude Code tokens.

Channel 2: Pointing (region and gesture)

Pointing is the channel that turns "somewhere on this screen" into "here." A circle drawn around a control, a hover that lingers, a freehand loop around three related fields.

The gesture carries more than location. It carries intent strength: a deliberate circle around an element means something different from a pointer that merely passed over it.

A system that records which happened can weight them differently. Circle one row in a table and the agent gets that row, not the table, not the page.

Channel 3: Voice (deictic speech)

Voice is how the instruction arrives. Most tools throw it away by making you type instead.

Speech is where deixis lives. People say "make this blue" while pointing, so you need transcription with timing — which word at which millisecond, matched to the pointer.

Channel 4: The accessibility tree (named elements and window text)

This is the channel that converts everything above from a guess into a fact. On macOS, every on-screen control is exposed through the Accessibility API.

Ask the system for the element at a screen point — AXUIElementCopyElementAtPosition — and you get back a real node: role, title, value, frame, and the chain of parents it sits inside. That is not a description of a button.

It is the button. The same tree carries the surrounding text: the labels, the error message under the field, the value in the input, the heading of the section.

I unpack the resolver mechanics in how AI agents know which UI element you mean.

Heads up

The accessibility tree is not always populated. Canvas-rendered surfaces, games, and some custom renderers expose nothing useful. Any honest system needs an on-device OCR fallback for those and needs to say when it used it.

What is element grounding, and why do most tools miss it?

The four channels are inputs. On their own they are four blobs the model still has to align — the guessing problem moved one step later.

The fifth step is resolution, or element grounding: collapsing all four into one named element with a confidence score. Instead of an image, a coordinate, "make this blue," and a wall of window text, you hand the agent one card — role, label, parent chain, instruction, confidence, provenance.

A grounded element without a confidence score is still a guess, just a well-dressed one. If the resolver lands on a bare unlabeled group, or the tree came back thin, the number should drop and the agent should ask rather than silently pick something.

Provenance matters for the same reason. An element you deliberately circled deserves more trust than one your pointer happened to dwell on.

How does PinVari capture spatial context on macOS?

PinVari is a native macOS app built around exactly this pipeline. You hold ⌥⌘A, circle or point at anything on screen, and speak.

It captures all four channels and grounds them before your agent sees anything: named element (role, label, frame, confidence, circled-or-dwelled), up to 40,000 characters of window text including scrolled-out content, optional whole-screen on-device OCR, per-mark word buckets, deictic binding, ~0.2s dwell detection, and multi-display marks.

Transcription runs on-device. The delivery path is a local MCP server on 127.0.0.1, so the grounded capture goes to your agent — Claude Code, Cursor, Codex, Zed.

claude mcp add --scope user pinvari -- "$HOME/.pinvari/mcp/pinvari-mcp"

There is a reverse direction too. The agent can request spatial context mid-task with pinvari_request_capture.

It hits a point where it genuinely does not know which element you meant, and instead of guessing, it asks you to point. The notch island lights up, you circle the thing, and it continues.

How is spatial context different from computer use?

They sound similar and solve opposite problems. Computer use means the agent drives your machine: it captures frames, moves your cursor, clicks, and types.

Spatial context means the human drives and the agent receives. You point, you speak, the agent gets a grounded target and goes to work in the codebase.

Computer use puts the agent at the controls and streams frames. Spatial context keeps you at the controls and sends one resolved element card.

For the everyday loop of reviewing a build and pointing out what is broken, spatial context is the lighter and more accurate tool. You already know which element is wrong.

The job is transmitting that knowledge losslessly, not asking a model to rediscover it from a video feed.

Do you actually need spatial context, or is pasting a screenshot fine?

For one bug a week, paste the screenshot. Adding tooling to a rare task is not worth it.

The break-even arrives once UI review is a loop: repeated "which button?" rounds and full-resolution images crowding the context window. Browser-only tools are also blind on your IDE, native Mac app, and Electron client.

Spatial context through the system accessibility layer covers those surfaces. The honest baseline is in can Claude Code see my screen: no, not by itself.

FAQ

What is spatial context in AI?

Spatial context is the on-screen information an AI agent cannot get from your code: the screenshot, the specific element you are pointing at, the words you say about it, and the surrounding window text. It is delivered together so the agent acts on one unambiguous target.

How is spatial context different from computer use?

Computer use puts the agent at the controls — it captures live frames and drives your mouse and keyboard. Spatial context keeps you at the controls: you point and speak, and the agent receives a single grounded element plus your instruction.

Do AI coding agents need screen access?

Not raw screen access, no. What they need is the resolved identity of the element you care about, which can be produced on-device and sent as a small text payload.

What is deixis and why does it matter for AI agents?

Deixis is the linguistic term for words like "this", "that", "here", and "there" that only mean something in the situation they are spoken in. Agents fail on deictic instructions because the pointing that resolved the word was never captured.

What is element grounding?

Element grounding is the step that collapses the screenshot, the gesture, the speech, and the accessibility tree into one named element with a confidence score and a note about how it was selected. It is the difference between four raw signals and one actionable target.

Does spatial context work outside the browser?

Yes, when it is built on the operating system's accessibility layer rather than on a browser extension. macOS exposes the element under any point for every application, so native apps, IDEs, and Electron windows are all covered — with on-device OCR as the fallback for canvas-rendered surfaces.

Hand your agent the exact element

PinVari resolves what you point at into a named, executable instruction — on-device, no keys, your own agent. One click inside PinVari connects Claude Code, Cursor, VS Code or Codex — or paste one CLI line from pinvari.com/connect.

PinVari → Connect → your agent (one click)
Get PinVari — $39 →