Best Screenshot MCP Servers for Claude Code (2026)

The best screenshot MCP servers for Claude Code split cleanly by surface: Playwright MCP for anything in a browser, Peekaboo for native macOS windows, XcodeBuildMCP or ios-simulator-mcp for iOS simulators, Screenpipe for a searchable history of everything you looked at, and snap-happy when you want the smallest possible "just take a picture" server. If you want the agent to see the specific control you are pointing at rather than whatever it chose to photograph, that is a different job, and PinVari is the tool built for it.
Most roundups stop at "here are nine servers that take screenshots." The useful cut is directional: almost all of these are agent-to-screen tools, where the model decides what to look at.
Exactly one is you-to-agent. Those two directions solve different problems, and plenty of people end up wanting one of each.
Everything below was checked against each project's own README or documentation in August 2026. Where a capability is not documented, this post says so rather than guessing.
Which screenshot MCP servers are worth comparing?
Here is the field as cards. Read each one as a question about your workflow, not as a scoreboard.
Playwright MCP
- Screenshots: Yes (page + element)
- Who decides: The agent takes its own shot
- Named element: Page a11y snapshot, no confidence
- Beyond the browser: No
- Price: Free, open source
Peekaboo
- Screenshots: Yes (screen, window, app)
- Who decides: The agent
- Named element: AX UI map with element IDs
- Beyond the browser: Yes, any Mac app
- Price: Free, MIT
macos-automator-mcp
- Screenshots: No
- Who decides: N/A
- Named element: Whatever your script reads
- Beyond the browser: Yes, any Mac app
- Price: Free, MIT
ios-simulator-mcp
- Screenshots: Yes, plus video
- Who decides: The agent
- Named element:
ui_describe_pointon the simulator - Beyond the browser: Simulator only
- Price: Free, MIT
XcodeBuildMCP
- Screenshots: Yes (beta UI automation)
- Who decides: The agent
- Named element:
describe_uiwith frames - Beyond the browser: Simulator + device builds
- Price: Free, MIT
Screenpipe
- Screenshots: Continuous capture
- Who decides: Queries history
- Named element: AX tree, OCR fallback
- Beyond the browser: Yes, whole desktop
- Price: From $25/mo
snap-happy
- Screenshots: Yes (screen or window)
- Who decides: The agent
- Named element: No
- Beyond the browser: Yes, any window
- Price: Free, open source
PinVari
- Screenshots: Yes (region + per-mark)
- Who decides: You point; agent can also request
- Named element: Role, label, frame, confidence, provenance
- Beyond the browser: Yes, any Mac app
- Price: $39 one-time
The column that actually separates these tools is not "screenshots." Every row does that. It is who decides what gets captured, and whether the agent receives a named element or a rectangle of pixels it has to interpret.
If you are new to how any of this plugs in, what the Model Context Protocol is covers the protocol. Connecting MCP servers to Claude Code covers the claude mcp add mechanics that every tool below assumes.
See also the MCP server list if you want a wider catalog.
Is Playwright MCP the best screenshot MCP server for browsers?
microsoft/playwright-mcp is the one to install first if your bugs live on a web page. It runs the real Playwright stack, so the agent gets a browser it can drive, not just a camera pointed at one.
Two capture tools matter. browser_take_screenshot returns PNG, JPEG or WebP, full page or scoped to a single element.
browser_snapshot returns an accessibility snapshot of the page instead of pixels, which is usually the better thing to hand a model: it is text, it is structured, and it costs a fraction of the image tokens.
Install is npx @playwright/mcp@latest with Node 18+. It can also run standalone over HTTP with --port.
The limit is the obvious one. Playwright MCP sees inside a browser it controls, and nothing else.
Your Electron app, your IDE, your native menu bar, the Figma desktop client, and a Simulator window are all invisible.
It also drives a browser instance the agent launched. It is not looking at the tab you personally have open with your session cookies and your half-filled form.
What is the best screenshot MCP server for native macOS apps?
steipete/peekaboo describes itself as Mac automation that sees the screen and does the clicks, and that is a fair summary. It ships as a CLI plus a menu-bar app, and exposes the same observation and action toolset over MCP to Claude Code, Codex, Cursor and other clients.
What sets it apart from a plain screenshot server is the UI map. Peekaboo inspects running applications through the macOS accessibility APIs and produces a structured map of elements with IDs.
The agent then acts on those IDs — click element 7, type into element 3 — instead of guessing pixel coordinates from an image. That is the same insight that makes browser_snapshot better than a screenshot, applied to native apps.
Install is brew install steipete/tap/peekaboo. It is MIT licensed and free, and it requires macOS 15 or newer.
Multi-display behaviour is not documented in the README, and there is no OCR step. It reads semantic UI elements, so a <canvas> or a custom-drawn surface with no accessibility data gives it little to work with.
Peekaboo and Playwright are not rivals. Install both. Let Playwright own the web app and Peekaboo own everything else on the Mac, and your agent stops trying to solve native-window problems with a browser tool.
Does macos-automator-mcp take screenshots?
steipete/macos-automator-mcp shows up in every "macOS MCP" search, so it is worth stating plainly: it does not take screenshots. It is a Model Context Protocol server that lets clients discover and run AppleScript or JavaScript for Automation.
It exposes two tools. get_scripting_tips searches a bundled knowledge base of AppleScript and JXA recipes, and execute_script runs an inline script, a script file, or one of those knowledge-base entries.
I am including it because it is genuinely useful next to a screenshot server, not as one. If you want the agent to read a value out of a native app or toggle a system setting before something else photographs it, this is the right tool.
Ask it to see your screen and you will get a script that half-works.
Which MCP server screenshots the iOS Simulator?
joshuayoes/ios-simulator-mcp gives an agent tools against a booted simulator. Capture is screenshot (PNG, TIFF, BMP, GIF, JPEG) and record_video (H.264 or HEVC).
XcodeBuildMCP covers simulator and device builds, with a beta describe_ui that returns frames. Both are free and MIT licensed.
Neither one sees your Mac desktop. They are the right pick when the bug lives in a booted simulator, and the wrong pick when it lives in a native Mac window.
When do Screenpipe and snap-happy make sense?
Screenpipe continuously captures the desktop and lets the agent query history, with AX tree plus OCR fallback. It is the pick if you want "what was on screen ten minutes ago," not "what is this button right now."
It bills from $25/mo. That is a different product shape from a one-shot screenshot server.
snap-happy is the smallest "just take a picture" server. The agent can grab a screen or window.
There is no named element, no voice, and no confidence.
Use it when you truly only want pixels. Skip it when you already have Playwright or Peekaboo.
What if you want the agent to see the control you pointed at?
Can Claude Code see my screen? Not on its own.
Every server above except PinVari is still agent-initiated: the model decides what to photograph.
PinVari is you-to-agent. You hold ⌥⌘A, circle or point at a control, and speak.
It resolves the named accessibility element (role, label, frame, confidence, circled-versus-dwelled provenance) and hands that to Claude Code over a local MCP server. The agent can also call pinvari_request_capture mid-task so the notch island asks you to point.
On AX-blind surfaces it falls back to on-device Vision OCR. Below 0.8 confidence it asks rather than guessing.
It is a one-time $39 launch license, not a subscription. Nothing is uploaded by default.
That is a different job from Playwright or Peekaboo. Plenty of people run one agent-to-screen server and one you-to-agent server side by side.
A screenshot MCP server that only returns pixels still leaves the agent guessing which element you meant. Named accessibility data is what stops wrong fixes. If the README does not mention role, label, and frame, you are still in pixel land.
FAQ
What is the best screenshot MCP server for Claude Code?
Playwright MCP for browser work, Peekaboo for native Mac windows, and a human-initiated resolver if you need the agent to see the control you pointed at. There is no single winner because they cover different surfaces.
Can Claude Code take a screenshot of my desktop?
Not by itself. You install an MCP server that captures pixels or reads the accessibility tree, then the agent calls that tool. See how Claude Code connects to MCP servers for the install pattern.
Does Playwright MCP see native Mac apps?
No. It sees a browser instance it launched. Your IDE, Electron apps, and menu bar are invisible to it.
Is Peekaboo better than a screenshot?
For native apps, yes, because it returns a UI map with element IDs instead of only pixels. The agent can click a named node instead of guessing coordinates.
Do I need a paid screenshot MCP server?
No. Playwright, Peekaboo, snap-happy, and the simulator servers are free. Screenpipe is the paid continuous-history option. PinVari is a one-time purchase if you want point-and-speak.
What is the difference between a screenshot and a named element?
A screenshot is pixels the model has to interpret. A named element is role, label, and frame the model can act on directly. The second is smaller, cheaper, and more accurate.
Hand your agent the exact element
PinVari resolves what you point at into a named, executable instruction — on-device, no keys, your own agent. One click inside PinVari connects Claude Code, Cursor, VS Code or Codex — or paste one CLI line from pinvari.com/connect.
PinVari → Connect → your agent (one click)Get PinVari — $39 →


