Mercury CLI

Notes

Long chats, calmer view. How .28 builds on reliability

· by @dubsthedev

Long chats and compaction

The biggest fix in this update is automatic compaction. Previously, if a conversation failed to compact three times, Mercury could stop trying altogether. That left you stuck in a long chat even when another attempt could have worked.

Summaries now use the provider connection’s timeout instead of being cut off after two minutes. The recovery step also keeps track of the context it removes. There was a bug where it would clear some context, then skip the retry because it thought nothing had changed.

Conversations that are too large to summarise in one request are now handled in parts. If the model refuses the first summary request, Mercury tells you and retries once with the conversation supplied as text. When the size of the conversation is what makes requests fail, you get one warning saying so rather than a series of seemingly unrelated errors.

If the session does pause, you can run /compact manually. Sending another message also gives automatic compaction a fresh attempt.

Skill listings were getting lost during compaction, when resuming a child agent, and when a directory was discovered during collection. They are preserved now. There are also two smaller display fixes: /compact stops showing “held” once the runner has accepted it, and headless compaction no longer repeats an earlier failure as an extra output row.

Resuming sessions

A few things could go wrong when coming back to a saved session.

One bug left a session stuck opening when removed transcript rows referenced each other. Another caused the next request to fail on the OpenAI, Z.AI, Google and local routes when saved assistant replies had been stored as plain text. Both are fixed. Mercury now replays those replies in the format each route expects, and display-only notices and receipts are kept out of the assistant’s messages.

Saved titles and tags no longer get overwritten by similarly named fields inside tool inputs. Titles are cleaned up so that control-character leftovers from the first prompt do not appear in them, and the editor’s session list now shows tags.

A session created on an already-running runner now restores its own permission mode after a crash, rather than inheriting the runner pool’s. /config also shows the session’s permission mode and the saved default separately when they differ. Startup records its route before the first screen and input are ready, including plain-world launches.

Restarting a runner no longer repeats an agent’s completion notice or incorrectly marks a finished agent as stopped. That includes notices delivered alongside other input or attached partway through a turn.

Text-only models and images

On Z.AI, an image in the conversation could stop a session or agent with a provider error. This included screenshots and pictures read from disk.

Text-only models now receive an [image] placeholder instead. They still cannot see the picture, but its presence no longer prevents the conversation from continuing.

Mercury also recognises Z.AI’s image-rejection message and retries without the image. A crewmate gets one automatic retry; a second refusal still ends its turn.

Z.AI requests with thinking disabled now use a format the provider accepts. The pricing entries have also been updated.

Session stats

Several places were showing numbers from the whole repository when they should have been showing the session you were looking at. A new session’s TRACE panel could even open with a list of tool calls from across the machine.

TRACE now starts empty for a new session. /status, /deck and /trace count the focused session’s tool calls, matching the trace rail. The deck-chip explanation in the ? sheet has been updated too.

/monitor and the deck now read the session’s own spend and line counts. They were showing $0.00 and +0/-0 even after a session had edited files.

/usage also has a “This session” tally under each provider’s active billing slot. It shows tokens and cost by model, with scheduled and advisor usage on separate lines.

Health checks

The cockpit’s /health could report FAULT and suggest a fix for a problem that was not there, while mercury health reported CAUTION. A live catalogue that the process has not fetched yet is now treated as a timing gap. Edit outcomes also count only the session’s own edits.

For an OpenRouter-only sign-in, auth status and the health check now identify the correct provider and model. Auth status routes through OpenRouter and exits with code 0.

There is a new managed-policy row that warns about an invalid lock value, an unknown surface name or an unreadable policy file, with the relevant fix. The prompt row lists the session, coordinator and crewmate contracts with their sizes.

On Windows, a timed-out health check also cleans up its temporary gate-tree folder.

Terminal rendering

In Terminal.app on the Mac, a thin extra line was drawn above two-colour cells. You could see it as a second pupil on the critter, a line above the head on the enter screen, and the same artefact on images and sprites. Those are fixed.

The chat was also repainting rows that had not changed. Finished tool groups and the turn and session cost rows now stay put during new records and thinking updates. The launch animation looks the same but sends about a fifth of the bytes.

Slow terminal connections now trigger reduced idle animation, as slow hardware already did. While the connection is the reason, the status line says reduced · slow link. There is a delay before reducing motion and before restoring it, so it does not keep switching back and forth.

Text width is now measured against the terminal once per session instead of being guessed. That helps emoji, flags and East Asian text line up where the terminal actually draws them.

A missed startup reply no longer blocks terminal upgrades indefinitely. /exit also returns the shell prompt to the right place in console hosts, rather than drawing it over the first row of the old screen.

The chat view

The chat view should feel a bit calmer and easier to follow. Some of these smaller changes have taken a while to get to, but they should be noticeable when you are spending a lot of time in a session.

Tool rows now start with their tool’s mark. Once a call settles, the mark takes the tool’s colour; it turns red on an error and shows ✕ when the call was denied. Notices keep their dot.

During thinking, the frame shows how long the model has been thinking and how recently the stream sent something, for example ◐ thinking 4m · ↻3s. The status line also shows the time since the last byte and the number of thinking blocks received. A long think should not leave you staring at a frozen token count.

The thinking label now waits three seconds before shimmering on every thinking phase, rather than only on the first one.

Crewmates appear in the rail as soon as they launch. Their views no longer show the lead’s compaction message, and rows you have expanded stay open when switching between the lead and a crewmate.

There are fixes for stray characters appearing in the composer after view switches, split key sequences, mouse reports or terminal replies. Tabbing into the left rail shows the caret immediately, and the single-rail layout keeps the last-prompt card visible while there is still room for it.

On the Session Concourse, the NOW cell follows the session’s current activity. It stops showing a tool once its result arrives, rather than claiming that something like Glob is still running minutes later. Eval cells also show “running” with the elapsed time once the kernel accepts them, instead of staying on “Starting kernel…”.

Picking a model through the composer now saves it as the default, as the picker’s header promised. The confirmation says it was saved, and the effort wording matches the status line.

I have also added 2- and 5-minute advisor intervals to /config, alongside the existing 10, 20, 30, 45 and 60-minute options.

File tools

Changes made outside Mercury are now detected after a full file read. Editing a file with mixed line endings also preserves the edited line’s own ending, rather than replacing it with whichever ending is most common in the file.

Glob now handles patterns such as sub/*.txt under a supplied path when the session is working in a different directory. That combination could previously return “No files found” even when the files existed.

Grep’s context lines preserve the slashes actually present in the file, and their displayed paths use the same relative or absolute form as the matching lines.

The Edit schema no longer presents replace_all as required, and it now explains its four supported forms.

Hooks, headless runs and the background helper

Lifecycle hooks marked “run once” now stop after their first run. They were running again on every event. A hook that returns another event’s output format now produces an error naming both events, instead of being silently accepted. SessionStart also includes the model on startup and resume.

A headless run that ends blocked on user input now exits with code 1 and explains the blocker in its outcome row. A run that delivers its result still exits with code 0. Previously, a run could do nothing and still report success.

After an update or rollback, a background helper that has handed over to its replacement now exits on its first idle read. It no longer holds the daemon lock for another fifteen seconds.

That is .28. The chat changes are the part I am watching most closely, because they only really show themselves once a session has been running for a few hours. If you spend that kind of time in one, let me know how it feels. That is the bit I most want to hear about from you this time.