Mercury beta.21: better control of a running session
Beta.21 is out. Edit is less fussy about formatting, there’s a control for pausing a chat’s agents, and the model picker has been rebuilt.
The release includes fixes found while using Mercury to develop Mercury. Some concern what happens to the code being written. Others concern what the model and the screen are told about work already running, especially after an interruption or a model change.
Better matching, stricter writes
Edit now tolerates differences in indentation, trailing whitespace, typographic dashes and spaces when finding the text to replace. It reports how it matched. Different content still does not qualify, and differing candidate blocks are refused as ambiguous rather than choosing one.
That removes some avoidable retries without turning a close-looking match into permission to edit somewhere else. A file the session wrote itself also counts as read, and a successful edit tells the model that another read is unnecessary. These behaviours are recorded in the editing changes.
Writes containing omission placeholders such as “rest of code unchanged” are refused before anything lands. That phrase is not a replacement for the code it would overwrite.
Interrupted tool calls also give a more accurate account of what happened. If execution started before the interruption, the result can no longer claim that nothing changed. The model needs to check the resulting state before deciding what to retry.
Pausing the work
The new crew pause control is p. In a hosted chat’s crew view, it pauses the chat’s agents at their next safe point. Pressing it again resumes them from there.
The pause indicator reads the runner’s state. A keypress requests the pause; it does not establish that every agent has already reached it. Keeping that distinction visible matters when several agents are still finishing their current work.
Beta.21 also fixes sub-agents sharing a scratchpad folder and overwriting each other’s helper files. Each gets its own folder within the session’s scratchpad. This is particularly relevant when Mercury is working on itself: separate checkouts do not help much if the temporary scripts used to investigate them overwrite one another.
The loop guard addresses another kind of wasted work. It flags a tool call repeated with the same arguments and result, as well as repeating cycles of calls. By default, it warns rather than ending the turn. Automatic stopping on a repeated cycle is a separate setting.
The model picker, and the state behind it
The rebuilt picker groups models under provider headings that expand and collapse. Your signed-in account’s provider comes first, with the others ordered by recent use. Each model row keeps its alias, raw ID, state and context size together. / opens the filter.
Models returned by a live provider list remain selectable even when Mercury does not recognise their IDs. An unknown model is labelled accordingly, without inventing a context size.
Some of the UI work was below the picker itself. Updates published within the same millisecond could share a stamp, leaving a fresh model update absent from the live session state. On Linux, a missed directory-watch notification could leave that state waiting for the next heartbeat. Both are fixed.
These are the kinds of UI faults a clean TypeScript check does not settle. The values can have the right types while the interface is still showing an earlier state. Fixing the publication and notification handling is part of making the controls dependable.
There are more direct interaction fixes too. Clicking the usage or files header no longer consumes the draft in the composer. Tabbed code in transcript diffs and permission cards now wraps within the available width instead of running through the border.
More control over tool output
Bash and PowerShell now accept max_output_chars for each call. The model can ask for a smaller inline result while the complete output remains saved. The reference to that saved output is retained, including when the command fails.
The new ContextLeft tool reports context usage and the tokens remaining before compaction from the request’s measurements. It loads on demand through ToolSearch. Tool descriptions have also been shortened by removing repeated instructions, reducing the definitions sent with the first request. These changes are detailed in the tool notes.
Network outages now take their own bounded reconnect steps across providers instead of spending the API retry budget. The status row shows that Mercury is reconnecting. Rate-limited requests separately honour the provider’s retry delay first.
What passed
The release verdict records the local pool at 141 of 141 release suites green. The final hosted gate entry also records a successful release run.
The terminal-interaction checks are reported separately. Beta.21 did not receive a new full terminal-drive run. Individual reruns cleared some earlier failures, while others remain documented for follow-up. Narrow-width picker issues are also queued for beta.22. Passing the release pool does not mean every terminal interaction has been cleared.
I moved development onto a Mac mini during this release, after having to close applications on the laptop and split tests with Windows. That gave the development sessions and test pools room to run together.
The rule remains that a change moves on when the release pool is green. More development capacity helps keep those checks close to the changes, but it is not evidence that the application itself is faster. Performance still needs measuring, and the terminal still needs exercising under load.
The full release notes cover the remaining changes. To update an existing installation:
mercury update