Skip to content

Built for agents

You can hand a Fleury app to an AI agent and have it drive the UI — read what’s on screen and operate it — through the Model Context Protocol, the open standard hosts like Claude already speak.

That’s unusual for a terminal UI. At the pixel level a TUI is just a grid of characters, so the usual way to drive one is to screen-scrape: diff ANSI output, match substrings, guess at key sequences. It’s brittle and blind — restyle a border and it breaks; the agent never really knows what’s on screen, only what it looks like. A Fleury agent never scrapes. Here’s how, and what makes it possible.

The fleury_mcp package runs a Model Context Protocol server over a private wire to your app:

Terminal window
fleury_mcp -- dart run bin/run_app.dart

It spawns your app, tracks its live semantic tree, and exposes it as MCP: the graph as a resource (fleury://ui/tree), and the actions on it as tools — get_ui / find_nodes to read, invoke_action / set_value to drive, and resize / wait_for_change to surface more rows or watch for asynchronous change. Legacy 2025-06-18 clients also retain the focus-relative type_text and press_key tools; they stay off the stateless surface until the request can carry an explicit target or focus lease.

Point an MCP host at it and the agent reads roles, labels, values, and the actions each node supports, then drives the UI through them — no ANSI scraping, no guessed keystrokes. Positional nodes carry an opaque targetRef, so a stateless action request does not depend on another task’s last read and rejects observable slot-identity changes. Semantically identical unkeyed replacements still need distinct keys or stable semantic ids. The app needs no agent-specific code.

For asynchronous app changes, wait_for_change returns the next settled tree without polling. Current clients pass the opaque, instance-scoped uiRevision from the prior result as sinceRevision, closing the gap between reading and starting the wait while rejecting handles from a restarted server. Legacy 2025-06-18 hosts can additionally subscribe for a compact notifications/resources/updated delta. Current MCP replaced that method with subscriptions/listen; Fleury does not advertise modern streaming subscriptions yet.

See Driving with an agent (MCP) for the hands-on setup: installing the driver, connecting a host, completing a semantic workflow, and making custom controls drive well.

The MCP server is a thin shim, not a bolted-on adapter, because the app already produces everything the agent reads. Alongside the visual tree, Fleury can project a semantic tree: a semantic app graph describing what the UI means, not how it’s painted. Agent-reachable and browser surfaces retain and update that graph when semantics or painted coverage change. The headless tester builds a snapshot when tester.semantics() asks for one, and a plain terminal run likewise skips the work until a debug consumer asks. That graph is the MCP resource; the SemanticActions on it are the MCP tools. There’s no scraping layer to maintain — driving the UI is just reading the graph and invoking its actions. The rest of this page is that graph.

Each node is a SemanticNode — a role, a human label, a value, interaction state, the actions it supports, and its children.

final class SemanticNode {
final SemanticNodeId id;
final SemanticRole role; // table, chart, button, slider, message…
final String? label; // "CPU", "Submit", "row 3"
final Object? value; // 0.62, "draft", selected text…
final String? hint;
final bool enabled, focused, selected, busy;
final bool? checked, expanded;
final Set<SemanticAction> actions; // what you can DO to it
final List<SemanticNode> children; // the tree
final SemanticState state; // completed / failed / disabled…
// …plus a cell `bounds` rect and a `validationError` for invalid inputs.
}

The role vocabulary is open. Core Fleury declares the generic set every surface understands — button, textField, list, listItem, region, status, dialog, tree, command, and the rest of SemanticRole.values. The widget catalog adds the roles built for the apps agents actually live in — WidgetRoles.conversation, messageList, traceTimeline, patchReview / patchFile, commandPalette, toolCall, approval — and your own package can add more. A declared role names the core role it projects through, so a surface that has never heard of it (the served browser’s accessibility mirror, an older agent bridge, the text-coverage policy) still treats it as what it is:

abstract final class BoardRoles {
static const board = SemanticRole('kanbanBoard', base: SemanticRole.region);
static const card = SemanticRole('kanbanCard', base: SemanticRole.listItem);
}
Semantics(role: BoardRoles.card, label: task.title, child: ...)
tester.semantics().byRole(BoardRoles.card) // tests match on the role

On the wire the node carries role: "kanbanCard" and coreRole: "listItem"; an agent filters by the specific name, and the browser mirror renders an ARIA listitem. A declared name must be an identifier and must not be a core name (both are asserted when the node is collected); only the name and its core root travel, so the human label is always derived from the name. The actions are equally concrete: activate, submit, select, increment, decrement, open, close, navigate, copy, start, cancel.

Take the live DataTable below — the real widget, running in your browser:

A DataTable — and the semantic tree behind it
Live

This excerpt was captured from tester.semanticInspectionJson() after rendering this example at 48×8 cells. It shows the table and its selected data row; snapshot metadata, node ids, the header, other rows, and additional node state are omitted:

{
"role": "table",
"label": "Data table",
"value": 0,
"actions": [
"copy",
"focus",
"select",
"setValue"
],
"state": {
"currentRowIndex": 0,
"selectionMode": "row"
},
"children": [
{
"role": "tableRow",
"selected": true,
"children": [
{
"role": "tableCell",
"label": "dan",
"value": "dan",
"state": {
"columnId": "name"
}
},
{
"role": "tableCell",
"label": "author",
"value": "author",
"state": {
"columnId": "role"
}
},
{
"role": "tableCell",
"label": "1284",
"value": "1284",
"state": {
"columnId": "commits"
}
}
]
}
]
}

No ANSI parsing. The agent can read the selected row and available actions; state.columnId identifies each column. This table exposes the strings returned by its cellBuilder, so even the commit count is a string, "1284", rather than a typed domain value. An agent can invoke select on a row, use the table’s setValue action to jump to a row index, or copy the current selection, without guessing which arrow keys to press.

An agent’s get_ui result has the same shape, minus any value that only repeats its label, plus agent fields such as stableId.

An agent that read a node a moment ago needs to act on that node, even after the UI rebuilt. Every node carries a SemanticNodeId, and the framework derives a stable one with no app effort: a node under a keyed ancestor (a list row’s Key, say) keeps its id across rebuilds and reorders, because the id folds in its ancestor key chain rather than its raw position. A fully unkeyed node falls back to a positional id — and those can shift as the tree changes, so the server guards them: if a positional node’s app-issued target token changed since the agent last read it, the action fails safe with a stale_reference error instead of mis-targeting. The token survives value, focus, busy-state, and ticking updates to the same logical element; it rotates when the mounted contributor or the target’s role, label, or advertised actions change, and when a synthesized target disappears and returns. A semantically identical update with the same widget runtimeType and Key is the same logical element under Fleury’s reconciliation contract, so use a distinct Key or stable Semantics.id when those configurations are different logical targets. Controls with frequently changing labels should also use a stable id rather than making each label transition a new positional lease. Pinning Semantics(id: SemanticNodeId('submit')) on the nodes that matter gives the agent a durable handle and sidesteps the guard entirely.

So that the agent isn’t guessing which ids are durable, get_ui and find_nodes flag every positional node with "stableId": false and attach a one-line idGuidance note explaining it; a stable id is simply unannotated. An agent that needs to hold a reference across an interaction can see up front which nodes need an app-assigned id — and it’s the honest signal for the one case the framework can’t solve for you: an unkeyed node in a dynamic list has no identity across a reshuffle, exactly as it wouldn’t in Flutter, so durability there is an app-authoring choice (add a Key or a Semantics(id:)), not a framework default.

Settable nodes publish their raw constraints (min/max, option labels, a date format) in semantic state; the MCP server derives a typed valueSchema from them — the type and domain a set_value will accept — so an agent fills a field correctly the first time, and an out-of-domain value is rejected with a clear reason rather than silently dropped.

Live semantic state without rebuilding content

Section titled “Live semantic state without rebuilding content”

Use state: for values known when a widget builds. Use stateBuilder: for values that become available after layout, such as a list’s visible range:

Semantics(
role: SemanticRole.region,
label: 'Source viewport',
stateListenable: controller,
stateBuilder: () {
final range = controller.visibleRange;
return SemanticState({
if (range != null) ...{
'visibleRangeStart': range.first,
'visibleRangeEnd': range.last,
},
});
},
child: SizedBox(
height: 8,
child: ListView.builder(
controller: controller,
itemCount: lines.length,
itemBuilder: (_, index, _) => Text(lines[index]),
),
),
)

Here controller is a ListController owned by the enclosing state. A completed scroll updates the semantic snapshot without rebuilding this wrapper or its content. First-party collections already use this pattern.

The callback runs when semantics are collected. It can run several times per frame, and a terminal session without a semantic consumer may never call it. Keep it a pure read: no inherited dependency reads, model mutations, or work scheduling. Return a snapshot whose map and nested values will not be mutated. Compute expensive content summaries when content changes; reading viewport state should not scan the entire collection.

stateListenable invalidates the semantic snapshot even when no pixels change. The element manages the subscription across moves and replacement, but does not dispose your model. Without a listenable, changing the model alone does not schedule a semantic update. Supply either state or stateBuilder.

For a custom collection wrapper whose visual content depends on the controller, listen to controller.viewChanges in its build dependency. That signal covers cursor changes, explicit content refresh, and scroll requests. Ordinary controller listeners continue to receive completed viewport metrics as well.

That graph isn’t a diagram of something internal; it’s the API — and agents aren’t its only consumer. A test reads the same tree and asserts on meaning:

testWidgets('Save advertises an activate action', (tester) {
tester.pumpWidget(const Semantics(
id: SemanticNodeId('save'),
role: SemanticRole.button,
label: 'Save',
actions: {SemanticAction.activate},
child: Text('Save'),
));
final save = tester.semantics().single(
role: SemanticRole.button, label: 'Save');
expect(save.actions, contains(SemanticAction.activate));
});

tester.semantics() returns the same tree shown above; .single(...) finds a node by role, label, value, or an action it advertises. The assertions are about what a node is — its value, whether it’s selected, the state it carries — so they survive a re-theme, a relayout, or a port to the browser. “The CPU gauge reads 0.62” and “row 3 is selected” are one expect each, not a screen-scrape.

Driving the UI is the mirror image: find a node that advertises an action, invoke it, and watch the state change — exactly what an agent does instead of guessing keystrokes.

testWidgets('activating Save runs its handler', (tester) async {
var saved = 0;
tester.pumpWidget(Semantics(
role: SemanticRole.button,
label: 'Save',
actions: const {SemanticAction.activate},
onAction: (_) => saved++, // the app's handler
child: const Text('Save'),
));
final result = await tester.invokeSemanticAction(
SemanticAction.activate, role: SemanticRole.button, label: 'Save');
expect(result.completed, isTrue);
expect(saved, 1); // the UI actually changed
});

That read graph → invoke action → observe change loop is the whole story, and it isn’t terminal-only. fleury serve exposes the same loop over a socket: the tree ships out as JSON — the same node fields tester.semanticInspectionJson() shows, redaction and all, flattened into a full-then-patch diff envelope (childIds instead of nested children) so it stays inside DEFLATE’s window — and SemanticActions come back on the same channel, each answered with an invocation-status result. So an MCP agent, a test, and the browser’s accessibility mirror are all the same kind of consumer — reading meaning and acting on it — with different goals.

The graph isn’t built for any single consumer — it’s one artifact with three uses:

  • Agents read roles, state, and the available SemanticActions and act through them — a typed surface, not a screenshot to interpret.
  • Tests assert on meaning (above), so they survive a re-theme, a relayout, or a port to the browser.
  • Accessibility, in the browser: both browser paths project the same tree into a visually hidden, accessible DOM (ARIA roles, labels, values, and states), and activating a node there invokes its primary SemanticAction. Custom controls still need meaningful semantics, and screen-reader behavior still needs testing. In a terminal, Fleury has no channel for handing the tree to assistive technology.

Because semantics are produced by the core (not a target), they exist whether the app runs in a terminal or the browser. fleury serve streams the semantic tree to the client alongside the visual frame, so a remote session is just as inspectable as a local one.

See Core and targets for where semantics sit in the pipeline.