Mercury CLI

Notes

Qwen, Mercury, and future plans

· by @dubsthedev

I've been playing around with local models in Mercury a lot lately. That led to 28 logged bugs, all fixed in beta.23. Ollama, LM Studio and llama.cpp have been tested live; vLLM still needs a Linux system with an NVIDIA card.

Setup now runs through /localsetup, from finding or installing a server to pulling a model, setting its context window and testing it. The model choice stays with the user; nothing is preselected. The tool list sent to local models is down from about 50k tokens to about 11k, while every tool remains accessible.

Testing and fixes

Tests covered Qwen 3.5 9B and 27B, Qwen 3.6 Coding 27B and 35B, Qwen 3.8 27B, and Gemma 4 31B. The hardware was a Mac mini M5 Pro with 15 cores, 48 GB of memory and models on an external 1 TB SSD. The 9B read 457 tokens per second and generated 32; the 27B generated 11.

One fault overestimated Qwen's memory requirement by a factor of four, rejecting the full context window despite sufficient memory. The calculation is fixed. Mercury also reads the server's reported window rather than assuming 4k, and selects the largest that fits, up to the model's full 262k.

Another fault reloaded the model with every message, costing about 15 minutes in one session. Models now warm up when selected and remain loaded.

Local models are already improving, and I hope that continues. I'd like to see more powerful models running locally through Mercury one day.

The other fixes

  • The new advisor provides advice from a second model, with its own memory. Off by default, with any pairing supported, including local models.
  • The output ceiling is up to 128k where the model allows it.
  • Turns that spend their whole budget thinking continue at the same effort and say so. This followed a review of 212,923 past requests.
  • The status row no longer shows a dead 0 while a request is running.
  • Crewmate chats are visible again in the crew view, with effort and idle status.
  • Notice rows are stamped in order, and each request records the effort it ran at.