Yeah, improving on 2024 performance was a pretty low bar to clear, by mid-2025 I saw the continuing improvement and decided that even if it was marginal at the time, learning how to use it was probably worthwhile given the improvements that seemed to be coming. Those improvements definitely did come from my perspective. Much of it was in the harnesses - many things I used to have to tell models explicitly, repeatedly in fall of 2025 they started doing without explicit prompting by spring of 2026. I also think I learned what they could be expected to do well and what was a waste of time trying which made me more productive with them as well.
I recently heard that Kimi v3 has made significant progress in code quality - with some people calling it “on par” with Claude. I primarily use Claude Opus - when I tried Fable during their free preview it “felt” even better, but not enough better to shell out a lot of extra cash for personal playtime projects. Kimi isn’t as accessible under the fixed monthly price model so I’m unlikely to try it anytime soon.
Progress has slowed at the frontier and open weight is right on their heels. The general advice is that Fable and the like are best for niche use cases and to make an open weight model your daily driver because they’re almost as good and are a lot cheaper. In my case for personal use, I’m perfectly happy with Qwen 3.8 27B running locally and don’t care to deal with data collection or huge bills to get a bit stronger of a model.
To switch from Claude to Kimi in service forms available to me would be a 5x price increase for how I use Claude (milking the subscription 7 day token limit dry after 5 days pretty consistently.) Decent self-hosting solutions seem to be running around $20K+ in capital equipment and more than $20 per month in electricity costs alone.
Yeah, improving on 2024 performance was a pretty low bar to clear, by mid-2025 I saw the continuing improvement and decided that even if it was marginal at the time, learning how to use it was probably worthwhile given the improvements that seemed to be coming. Those improvements definitely did come from my perspective. Much of it was in the harnesses - many things I used to have to tell models explicitly, repeatedly in fall of 2025 they started doing without explicit prompting by spring of 2026. I also think I learned what they could be expected to do well and what was a waste of time trying which made me more productive with them as well.
I recently heard that Kimi v3 has made significant progress in code quality - with some people calling it “on par” with Claude. I primarily use Claude Opus - when I tried Fable during their free preview it “felt” even better, but not enough better to shell out a lot of extra cash for personal playtime projects. Kimi isn’t as accessible under the fixed monthly price model so I’m unlikely to try it anytime soon.
I saw this article today summarizing a Mozilla report that open weight models are about 4 months behind the frontier: https://arstechnica.com/ai/2026/09/exclusive-open-chinese-models-close-gap-with-silicon-valleys-frontier-ai-models/
Progress has slowed at the frontier and open weight is right on their heels. The general advice is that Fable and the like are best for niche use cases and to make an open weight model your daily driver because they’re almost as good and are a lot cheaper. In my case for personal use, I’m perfectly happy with Qwen 3.8 27B running locally and don’t care to deal with data collection or huge bills to get a bit stronger of a model.
To switch from Claude to Kimi in service forms available to me would be a 5x price increase for how I use Claude (milking the subscription 7 day token limit dry after 5 days pretty consistently.) Decent self-hosting solutions seem to be running around $20K+ in capital equipment and more than $20 per month in electricity costs alone.