On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from 5/24 to 23/24 functional and delivered successes across sequential output-budget and reasoning rounds. That is promising development evidence. It does not establish generalization or show that xhigh alone caused the difference.
The next experiment was less encouraging: lowering temperature to 0.8 on two selected difficult cases preserved 5/6 functional successes but reduced delivered successes to 4/6. The reused temperature-1.0 controls scored 5/6 on both metrics.
This is a follow-up to my local OpenCode experiments. The full article and public records provide the detailed methods and inclusion rules. This post is a self-contained summary of the development findings and their limits.
Response space and reasoning
The setup used Qwen3.8-27B Q4_K_M on one RTX 3090, llama.cpp b11146-7fe450e19, CUDA 12.8, and OpenCode 2.0.20. OpenCode operated on the workspace and ran tools; llama-serve
Discussion
Get the discussion rolling
A single comment can start something great.