Annual Defect Report 2026Download
Blog & Reports
ArticleOctober 8, 2026 · 7 min read · Mobot

Testing Is the Bottleneck in the Agentic Software Factory

Coding agents ship mobile code faster than QA can test it. What agents can do with Mobot's real-device testing today.

Coding agents now write, refactor, and open pull requests for mobile code at a pace no human team planned for. CI/CD pipelines keep up because they were built to scale with compute. Mobile QA does not scale that way. A change that touches checkout, login, or push handling still has to be proven on a physical iPhone or Android device, with a real OS build, real hardware state, and a real finger on the glass.

When code generation gets faster and real-device testing stays the same speed, testing becomes the constraint on the whole system. More PRs means a longer queue for the same device lab, or a growing pile of changes that ship having only been checked on emulators.

Mobot's direction is to make real-device testing something agents work with directly, the same way they work with a compiler or a CI job. This post walks through the architecture layer by layer.

ArchitectureHow the Agentic Testing Protocol fits together

Client

Agents and tools

  • Coding agents
  • CI/CD pipeline
  • Jira / Linear

Agent interface

Agent gateway

  • MCP Server (read-only)
  • Test Intent Schema
  • Agentic Testing Protocol

Test intelligence

Mobot AI

  • Test Planning and Coverage Selection
  • Test Case Generation
  • Tests that update as your app changes
  • Mobot AI Model

Robot fleet

Real devices, real robots

  • On Demand Robot Reservations
  • Automated Device Reset and Provisioning
  • Device Configuration Profiles
  • Mobot testing robots

Results to client agents

  1. Human QA validation
  2. Structured defect reports
  3. Jira / Linear ticket sync
  4. Agent verification re-test

Logs and exhaustive defect forensics

An architecture model, not a feature list. Available to agents today: read-only access through Mobot’s MCP server to test reports, test cases, observations, and device and network logs where captured.

Layer 1: The client — agents and tools

The architecture starts where work already happens. Three kinds of clients sit on the left side:

  • Coding agents that generate or modify app code and need to know whether the change works on a device before a human reviews it.
  • CI/CD pipelines that already gate merges and builds and need a real-device stage alongside unit and integration tests.
  • Jira / Linear, where defects are triaged and where work is assigned to humans and agents.

The design choice is that Mobot does not ask these clients to change how they work. A coding agent should not need a custom plugin per testing vendor, and a pipeline should not need a human to file a request for device time.

Layer 2: The agent interface — a gateway built for agents

The agent gateway is where agents connect to Mobot. It has three parts:

MCP Server. Agents that speak the Model Context Protocol can connect to Mobot and read their organization's testing data, covered in detail below.

Test Intent Schema. A way for agents and pipelines to say what should be tested, not a device-specific script, so the layers below can plan coverage and generate test cases. It maps to the Direct stage of the protocol below.

Agentic Testing Protocol. The protocol for the lifecycle of a test between an agent and Mobot, covered in the next section.

The point of the gateway is a stable contract. Clients integrate once; what happens behind the gateway can evolve without breaking them.

The Agentic Testing Protocol

The Agentic Testing Protocol defines the full lifecycle of a test between an agent and Mobot. In this architecture, the lifecycle has seven stages:

  1. Discover. An agent would learn which devices, OS versions, and slots are available.
  2. Reserve. An agent would book device time or join a queue.
  3. Provision. An agent would deliver the build. The device would be wiped and configured, and the app installed.
  4. Direct. An agent would state the objectives, scope, and acceptance criteria for the test.
  5. Execute. A Mobot testing robot and an operator run the test, and an agent would be able to observe progress.
  6. Report. Results and structured defect reports come back to the agent.
  7. Verify. An agent would deliver a fix, and Mobot would re-test that specific defect.

The stages form a loop: when a fix produces a new build, that build goes back to Verify.

Agents can already read through Mobot's MCP server — read-only — their organization's test reports, test cases, observations, and device and network logs where captured. A coding agent can pick up a defect Mobot found, read the report and logs, and start on a fix. The rest of the protocol — reserving devices, provisioning builds, directing tests, re-tests, ticket sync and review status — is coming next.

Layer 3: Test intelligence — Mobot AI

Behind the gateway, Mobot's AI layer turns intent into executable work. It has four components:

Test Planning and Coverage Selection. Not every change needs every test. This component decides what to run for a given request, which matters most when agents produce changes far more often than a human team would.

Test Case Generation. Mobot AI creates the test cases that exercise the intent the client described, so a team does not have to hand-write a case for every flow before it can be tested.

Tests that update as your app changes. When the app changes, test cases are kept up to date with it rather than left to break. That is the difference between a suite that keeps pace with agent-generated code and one that becomes its own source of noise.

Mobot AI Model. The model underneath the test intelligence layer that drives planning, generation, and maintenance.

Layer 4: The robot fleet — real devices, real robots

This is the part of the stack that cannot be emulated. Mobot runs tests on real iOS and Android devices with Mobot testing robots, so the test exercises the same hardware, OS build, and touch input that a user's phone does.

On Demand Robot Reservations. Booking robot and device time. For agents, this maps to the protocol's Reserve stage.

Automated Device Reset and Provisioning. Preparing the device and installing the build before a run. For agents, this is the protocol's Provision stage.

Device Configuration Profiles. Defined device configurations for a test. For agents, setting them is part of the Provision stage.

Mobot testing robots. The physical robots that execute the tests on the device.

Emulators and simulators are useful for fast feedback, but many mobile-only defects depend on the device itself: the hardware state, the OS build, a real interruption arriving mid-flow. A test layer for agents has to include the device, or agents will optimize against an environment users never touch.

The return path: results back to client agents

A test result is only useful if the agent that asked for it can act on it. The bottom band of the architecture is the return path:

Human QA validation. People review results as part of Mobot's service. In this architecture, a review state on each result would tell agents when a result is final.

Structured defect reports. Defects come back as structured data, not just prose and a screenshot. Agents read defects and issues as observations through the MCP server.

Jira / Linear ticket sync. In this architecture, defects would sync to the tracker the team already uses, Jira or Linear.

Agent verification re-test. When an agent delivers a fix, it would be able to request a re-test of that defect on a real device, so the loop does not stop at "PR opened."

Logs and exhaustive defect forensics. Agents can pull the device and network logs attached to a report, where captured, for root-cause work.

The full design closes the loop from code to verified fix.

What this means for engineering and QA leaders

If your team is adopting coding agents, the question is whether your testing capacity keeps up, on the devices your users actually hold.

Three things to look for in any test layer you put behind your agents:

  1. A stable, agent-native interface. Agents should connect to it directly, without a human relaying requests.
  2. AI that scales test creation and maintenance, so test coverage grows with code volume instead of falling behind it.
  3. Real devices and reviewed, structured results, so agents act on defects users would actually hit.

For background on the kinds of defects that reach production in mobile apps, see our Annual Defect Report.

See it run

Mobot is building the real-device test layer for agents, starting with read access your agents can use today. To see Mobot test your own app on real devices, start a free trial.

See how agents connect to Mobot on our agentic platform page.

Start a free trial →

Agentic testingAI agentsReal devicesMobile QA

See What Your Emulators Are Missing

Get a real, verified defect report from Mobot’s robots and QA analysts — on your app, on real devices.

or talk to sales