Skip to content
AI Tool

Midscene Review

Midscene.js is a vision-based GUI agent for natural-language end-to-end testing and automation of web, PC, and mobile applications.

shipped Oct 2, 2026agentsfree
Domain rating57Monthly visits166/mo
agentscodeproductivity
Midscene — product screenshot

Why it matters

1Supports web, PC, and mobile application interfaces.
2Integrates with Playwright and Puppeteer.
3The open-source testing and automation tier is free.
4A stated model API cost example is $0.59 for 60 tasks; the precise usage unit is unclear.

About Midscene

Business Model
Open Source
Usage Pricing
$0.59 for 60 tasks per task
Platforms
Web, PC, Mobile
Target Audience
Developers and QA engineers looking for automated testing solutions

Pricing Plans

Open Source
Free
  • • End-to-end testing
  • • Natural language automation
  • • Support for multiple platforms

Cost Examples

  • • Total model API cost for 60 tasks: $0.59

Leadership

Zhang YimingFounder
API DocsOpen Source

Specs

API Available

Yes, public API

overview

What is Midscene?

Midscene is a vision-based GUI testing tool that enables developers and QA engineers to automate end-to-end application tests using natural language. It supports web, PC, and mobile interfaces through a unified API and test suite, and can be integrated with Playwright and Puppeteer.

features

Key Features of Midscene

Midscene uses vision and natural-language instructions to operate graphical interfaces for testing and automation across three stated platform categories.

  • Vision-based GUI agent for interacting with application interfaces.
  • Natural-language instructions for UI testing and automation.
  • Support for web, PC, and mobile platforms.
  • Unified API and test suite across supported platforms.
  • Integrations with Playwright and Puppeteer.
  • Text and vision multimodality.
  • Supports model options including Doubao Seed, Qwen-VL, OpenAI GPT, Google Gemini, and Zhipu GLM.
  • Testing and automation APIs are documented at https://midscenejs.com/api/.

use cases

Who Should Use Midscene?

Midscene is described for developers and QA engineers who need to automate application interfaces or include GUI interactions in end-to-end tests.

  • QA engineers writing end-to-end tests for web applications.
  • Developers automating UI checks with natural-language instructions.
  • Teams integrating GUI automation with Playwright test suites.
  • Teams integrating GUI automation with Puppeteer.
  • Developers testing interfaces across web, PC, and mobile platforms.

how to use

How to Use Midscene

Start with the Midscene.js project documentation and select a supported integration such as Playwright or Puppeteer. Configure the testing environment and model used for vision-based interaction before writing natural-language test actions.

  • 1Consult the documentation at https://midscenejs.com/ and API reference at https://midscenejs.com/api/.
  • 2Choose a supported platform: web, PC, or mobile.
  • 3Set up a Playwright or Puppeteer integration when using either test framework.
  • 4Configure a supported vision model for the test environment.
  • 5Write natural-language instructions for the UI actions to automate.
  • 6Run the test and evaluate the resulting application behavior.

pricing

Midscene Pricing & Plans

Midscene lists an open-source tier as free. The supplied pricing information also gives a $0.59 total model API cost for 60 tasks, but describes the usage price as "$0.59 for 60 tasks per task," so the exact per-task rate and what the figure includes are not clearly specified.

  • Open Source: Free.
  • Model API cost example: $0.59 total for 60 tasks, according to the supplied pricing data; exact billing unit is unclear.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Pros

  • +Supports GUI testing across web, PC, and mobile platforms.
  • +Accepts natural-language instructions for UI automation.
  • +Integrates with Playwright and Puppeteer.
  • +Provides a free open-source tier.
  • +Lists multiple vision-model options, including Doubao Seed, Qwen-VL, OpenAI GPT, Google Gemini, and Zhipu GLM.

Cons

  • −The supplied pricing description does not clearly establish the exact usage billing unit.
  • −The available information does not specify supported operating-system versions or device requirements.
  • −The supplied sources do not give performance benchmarks or test reliability measurements.
  • −The available information does not specify which features or platform integrations are included in the free tier.

Similar Tools

Midscene vs Competitors

Midscene is positioned around GUI testing across web, PC, and mobile, while the supplied descriptions of Stagehand, ZeroStep, Browser Use, and Skyvern focus on browser-based automation.

1
Stagehand↗

Focuses strictly on browser automation using an AI SDK built on Playwright with three core primitives: act, observe, and extract.

Stagehand is heavily tailored for browser-based automation and data extraction rather than cross-platform desktop or mobile native UI testing. If you need Midscene's multi-platform capabilities across mobile and desktop apps, Stagehand will not cover them.

2
ZeroStep↗

Plugs an ai() prompt function straight into your existing Playwright test suites without adding separate agent scaffolding.

ZeroStep relies on a hosted API backend that limits free runs to a monthly quota, unlike Midscene's self-directed model configuration. It is also strictly bound to web apps running inside Playwright.

3

Provides an autonomous multi-step agent loop in Python and TypeScript that navigates dynamic web interfaces using vision and raw CDP.

Browser Use is optimized for exploratory browsing and long-running autonomous web agents rather than structured, assertion-heavy E2E test suites. You trade away Midscene's deterministic test assertion helpers and cross-platform desktop drivers.

4

Combines computer vision with LLM reasoning over Playwright to navigate websites and complex workflows without pre-defined DOM selectors.

Skyvern is architected primarily for workflow automation and web-scraping agents rather than rapid assertion testing in software QA pipelines. It also focuses entirely on browser environments, leaving out mobile and desktop application testing.

More on Stork

Related AI Tools

Other tools in this category, matched by shared tags

One short daily email of tools worth shipping. No drip funnel.

one email a day · unsubscribe in two clicks · no third-party tracking

For builders

This page is doing a job for someone else’s tool.

AI agents read it. Buyers land on it. It answers in eight languages and over MCP. Your tool can have one like it — live in 24 hours.