Skip to content

Driving with an agent (MCP)

A Fleury app already describes its controls as roles, labels, values, and actions. fleury_mcp gives an AI agent that semantic view of the app and lets it use the actions the controls advertise. It does not scrape terminal cells or guess which keys might work.

You do not add MCP code to the application. Start with ordinary Fleury widgets, then connect an MCP host to the same runApp entry point you use in a terminal.

This release checklist contains a text field, a checkbox, and a button. Try it by hand: change the version, mark Tests passed, then choose Prepare release.

agent_guide.dart
Application code · editable
import 'package:fleury/fleury_core.dart';
/// The small release workflow used by the Driving with an agent guide.
///
/// It deliberately uses ordinary Fleury controls. Their existing roles,
/// labels, values, and actions are enough for a person, a widget test, or an
/// MCP agent to drive the same workflow.
class AgentGuideApp extends StatefulWidget {
const AgentGuideApp({super.key});
@override
State<AgentGuideApp> createState() => _AgentGuideAppState();
}
class _AgentGuideAppState extends State<AgentGuideApp> {
final version = TextEditingController(text: '0.9.0');
bool testsPassed = false;
String status = 'Draft';
@override
void dispose() {
version.dispose();
super.dispose();
}
void prepareRelease() {
final release = version.text.trim();
setState(() {
status = switch ((release.isEmpty, testsPassed)) {
(true, _) => 'Blocked: enter a version',
(false, false) => 'Blocked: mark tests passed',
(false, true) => 'Ready to publish $release',
};
});
}
@override
Widget build(BuildContext context) => Padding(
padding: const EdgeInsets.all(1),
child: Column(
mainAxisSize: MainAxisSize.min,
crossAxisAlignment: CrossAxisAlignment.start,
children: [
const Text('Release checklist', style: CellStyle(bold: true)),
const Text('Complete the checks, then prepare the release.'),
const SizedBox(height: 1),
const Text('Version'),
SizedBox(
width: 28,
child: TextInput(
controller: version,
semanticLabel: 'Release version',
autofocus: true,
),
),
Checkbox(
label: 'Tests passed',
value: testsPassed,
onChanged: (value) => setState(() => testsPassed = value),
),
const SizedBox(height: 1),
Button(
text: 'Prepare release',
variant: ButtonVariant.primary,
onPressed: prepareRelease,
),
const SizedBox(height: 1),
Text('Status: $status'),
],
),
);
}
Release workflow
Live · interactive
click & type to interact ⓘ how this demo runs

There is no fleury_mcp import and no manually assembled agent protocol. The controls already publish the useful contract:

Visible controlSemantic contractAgent operation
Release versiontextField, current text, setValueset_value
Tests passedcheckbox, checked state, setValueset_value
Prepare releasebutton, activateinvoke_action
Status texttext, current labelread after the action

No extra identity or agent wrapper is needed for this read-and-act workflow. The server returns the current target for each control and refuses detectable stale positional references.

The MCP server is a separate executable so your application has no agent dependency. It connects to the app over a Unix-domain socket, so it runs on macOS and Linux. While Fleury is pre-release, activate the executable from the same checkout as the app. From the checkout’s root, bootstrap once so the package resolves the checkout’s fleury, then activate:

Terminal window
dart tool/fleury_dev.dart bootstrap
dart pub global activate --source path packages/fleury_mcp

Then point an MCP host at fleury_mcp, followed by -- and the command that starts your app. From the packages/samples directory of a Fleury checkout, this runs the example above:

Terminal window
fleury_mcp -- dart run bin/samples.dart agent-guide

For your own application, replace everything after --:

Terminal window
fleury_mcp -- dart run bin/run_app.dart

An MCP host that uses JSON configuration can express the same command as:

{
"mcpServers": {
"release-app": {
"command": "fleury_mcp",
"args": ["--", "dart", "run", "bin/run_app.dart"]
}
}
}

Hosts that add servers from their own command line take the same command. With Claude Code, for example:

Terminal window
claude mcp add release-app -- fleury_mcp -- dart run bin/run_app.dart

fleury_mcp starts the app itself. The app runs in remote mode, sends its semantic tree over a private local socket, and receives semantic actions back. Its terminal rendering is not part of the agent connection. The app gets an 80 × 24 viewport; pass --cols=<n> and --rows=<n> before the -- to change it, or use the resize tool during a session.

fleury_mcp runs the app’s command in its own working directory, so a relative path such as bin/run_app.dart needs the host to start the server from the application package. If the host can’t set a working directory, give the entrypoint’s absolute path.

Finding the programs is a separate question. Desktop hosts often start servers with a minimal PATH that includes neither ~/.pub-cache/bin nor the Dart SDK. Give the host the absolute path to fleury_mcp (which fleury_mcp prints it), and put the Dart SDK’s bin directory on the PATH in the server’s env setting, because both the fleury_mcp script that dart pub global activate installs and the app’s dart run command call dart.

dart run compiles the app each time the server starts, and fleury_mcp suggests compiling it once instead: dart compile exe bin/run_app.dart -o my_app, then fleury_mcp -- ./my_app. A compiled app starts faster, but its debug tooling is off, so read_frames, read_logs, and read_errors return available: false unless the app passes DebugConfig(enabled: true) to runApp (see Inspect what happened).

The server supports the stateless MCP 2026-07-28 revision and the 2025-06-18 initialization flow for existing hosts. The driving workflow below is the same on both.

Once the host is connected, ask for a user-visible result instead of naming protocol tools:

Set the release version to 1.0.0, mark the tests as passed, prepare the release, and tell me the final status.

A host can complete that request in four semantic steps:

  1. Read the UI with get_ui, or narrow a large screen with find_nodes.
  2. Find Release version and use set_value. Select Tests passed from the settled ui returned by that action.
  3. Set Tests passed, then select and activate Prepare release from the next returned ui.
  4. Report Ready to publish 1.0.0 from the final returned ui.

The agent uses ids returned by get_ui or find_nodes; application authors do not hard-code those ids into prompts. The server validates values against each control’s semantic schema before dispatching them.

Use the action a control advertises whenever one exists:

  • set_value changes a field, checkbox, switch, select, slider, or other settable control in one operation.
  • invoke_action activates buttons and invokes domain actions such as select, open, close, submit, increment, or cancel.

The current 2026-07-28 surface deliberately exposes only actions whose target travels in the request. Legacy 2025-06-18 clients also have press_key for keyboard routing and shortcuts, and type_text, which types text into the focused field. Those focus-relative tools stay off the stateless surface until Fleury can carry an explicit target or focus lease with them.

This is the same distinction used in Testing: test or drive application behavior through semantics, then cover keyboard and pointer routing separately when those physical paths matter.

Every mutating MCP tool returns the settled UI, which counts as a fresh read. Select the next control from that result instead of retaining a target from an earlier tree. Use find_nodes when the full tree is capped or you want a narrower view.

Built-in controls currently receive positional semantic ids. They are safe for an immediate read-and-act sequence, and Fleury returns stale_reference rather than dispatching when a positional target is no longer current. Controls that represent durable records, such as rows an agent may come back to, still need distinct keys or stable semantic ids.

How positional targets are checked

A positional node includes stableId: false and an opaque targetRef. A 2026-07-28 client echoes both the node id and that reference when it acts, taking the current targetRef from the latest returned UI. Fleury rejects the call if the slot now contains a control with a different observable identity token. A semantically identical unkeyed replacement is intentionally indistinguishable. The reference is protocol bookkeeping, not an application id to copy into code or prompts.

An explicit Semantics(id: ...) gives a custom semantic control a durable id. Built-in controls don’t accept an id yet. Don’t wrap one in Semantics just to add an id: that duplicates the semantics the control already publishes.

Built-in controls already expose their semantics. A custom interactive widget must say what it is, what it is called, and which actions it supports:

Semantics(
id: const SemanticNodeId('retry-upload'),
role: SemanticRole.button,
label: 'Retry upload',
actions: const {SemanticAction.activate},
onAction: (action) {
if (action == SemanticAction.activate) retryUpload();
},
child: const Text('[ Retry ]'),
)

Prefer a first-party control when its behavior fits. Besides helping agents, the same semantic contract powers Fleury tests and the browser accessibility surface.

After an action, use its returned ui before choosing the next target. For changes the app makes on its own, use wait_for_change instead of polling. A client passes the uiRevision from its last UI result as sinceRevision, so an update that lands between the read and the wait is returned immediately rather than missed.

Revisions and subscriptions

uiRevision is opaque and scoped to one server instance, so a value from a restarted server fails closed. Legacy 2025-06-18 hosts can also subscribe to fleury://ui/tree; the 2026-07-28 revision replaced that subscription method, so new integrations should prefer wait_for_change until Fleury adds a modern streaming transport.

Large lists and tables publish their visible semantic window. resize can make more rows visible; find_nodes keeps targeted reads small. Do not guess an off-screen id or press navigation keys unless keyboard behavior is the thing being exercised.

Tool failures include a stable machine-readable code. The common recovery is small and explicit:

CodeMeaningRecovery
not_foundThe node is no longer presentRe-read the UI
stale_referenceA positional target changed after it was readRe-read and select the current node
ambiguousMore than one node shares the id you passedUse find_nodes to pick the control by role and label; give durable controls distinct Semantics(id:) values
out_of_domainA value violates the control’s schemaUse the advertised type, range, or options
action_busyThe same action on this control is still running, or too many actions are queuedAct on the UI it opened (a dialog to answer, say), or wait for earlier calls to finish and re-read
rate_limitedToo many actions in a short windowPause, then re-read with get_ui or wait_for_change
not_readyThe app connected but hasn’t rendered a UIRetry after a moment. If it persists, check that the command’s entrypoint calls runApp
app_exitedThe app quit or crashedRestart the server. If the app keeps exiting, run its command in a terminal to see why
protocol_mismatchThe app and fleury_mcp were built against different Fleury versionsActivate fleury_mcp from the same checkout as the app’s fleury dependency, then restart the server

A server that stops as soon as it starts couldn’t attach to the app. Its standard error, which hosts usually keep in a log, says why: the app exited before connecting (with its exit code), or didn’t connect within 20 seconds. Run the same command in a terminal, from the same directory, to see the app’s error. A host that can’t find fleury_mcp or dart needs the PATH fix in Connect an MCP host.

An action whose handler is still running after two seconds, most often one that opened a dialog and awaits its answer, is not a failure: invoke_action and set_value return status: "pending" with the UI it opened. Answer the dialog through that UI; the same action on the same control is refused with action_busy until the earlier one finishes.

The driven app’s labels, values, logs, and errors are untrusted application data. The MCP server marks them as such; an agent should report that content, not follow instructions embedded inside it.

During development, read_frames, read_logs, and read_errors expose the same evidence as Fleury’s debug shell. An agent can trigger an action, then check whether it threw, logged an unexpected result, or caused expensive frames. They need the app’s debug tooling, which is on when the app runs from a .dart source file, as with dart run bin/run_app.dart, or with assertions enabled. A compiled app or a snapshot turns it off, and the tools return available: false, unless the app passes DebugConfig(enabled: true) to runApp.

For the underlying model, continue with Built for agents. For app-side assertions against the same semantic contract, use Testing.