Tools that tell you what's happened
This update took a long time, and it means a lot to me. The tools were getting quite annoying to use, so I ended up rewriting them. A lot of the changes are about what the model gets told before and after a tool runs.
The tools a request carries
A session's request now carries full definitions only for the tools it reaches for. The rest are named and load when requested. The tools remain available without every definition going into every request.
Seeing what actually changed
An Edit result now shows the changed lines read back from the file, with line numbers just like a Read. The next edit can use those lines directly.
Previously, an edit that matched more than one place could leave the model trying again to find a unique match. The refusal now lists every match's lines in one answer, so it can see where the ambiguity is.
A partial Read says which lines it showed out of the file's total and how to continue. When a file is too large to show whole, the result gives its line count and two smaller sections ready to read.
Background commands also give a clearer account of what happened. Their start tells the model what will end them. A command moved to the background at its timeout returns what it printed so far and names the file holding its output.
Code cells now keep their state between calls. JavaScript cells also return final objects, arrays and awaited values that were previously being dropped.
Changing code by its shape
There are new tools for finding and rewriting code by its shape across files as one reviewed change. The language-service tools cover code questions, renames, moving a declaration or a whole file, code actions and formatting.
A small read-only tool handles the questions. The tools that change code load when needed.
Flow and permission prompts
Flow has changed too. File edits inside the project and read-only commands go ahead on their own. Everything else asks the way Default does, and destructive calls always ask.
A yes on a permission card is saved as a project rule where the rule can describe the call. If a card is left unanswered for five minutes, it is withdrawn and the session carries on with allowed work.
Shell commands are now judged on their actual structure. For example, wc -l < file reads a file, but it was asking for permission. It now runs as a read. Text inside a quoted heredoc and shell comments are also recognised for what they are.
The correction goes both ways. Docker commands such as docker run and docker exec were running without an ask even though they change state. They now ask; listing and inspecting remain reads. This applies in Bash and PowerShell.
Long chats
There is a little more compaction work following the previous update. The summary request now uses the session's own effort, thinking and output ceiling, and waits properly after a model switch.
The first compaction being refused is fixed. If a summary request does fail, Mercury reports the reason and keeps the conversation intact instead of calling the failure a success or saving the refusal as the summary. The next message can try again.
This feels like one more step closer to release. If you use these tools, let me know how the changes feel, especially whether the results are clearer and the permission prompts appear where you expect them.