Fireworks has high variance amongst models and while some are served correctly; many of them are junk / broken and degraded and it seems like they don’t even know; because even running 1k MMLU Pro questions would flag it very quickly.
First day was a bit rough, first week was still a little rough, but it's been pretty smooth since then, even when learning how to fix things and trying new software.
I'm using Niri and Noctalia as my desktop setup, and it's been different than my Windows experience, but it feels fun and cool to just use a computer in a new way.
When I dropped my Claude sub a few months ago I first went to try pi. At first I was a little overwhelmed by having to browse extensions to get the experience I want (coupled with finding new model providers to use).
So I looked elsewhere and found Crush and Hermes. They're both very a e s t h e t i c, which I think can make using them fun, but ultimately if I had a nitpick, I would look back over at pi, and the grass looked greener. (And not to mention, both Hermes and Crush seem to have some drama/baggage.)
I'm back on pi, and happy with just a few packages I've downloaded for it.
I want to move away from claude but don't have the hardware for "good" local llms, and would rather use a ZDR api instead of one of the big apis for privacy reasons.
After my Claude subscription ended, I tried using Fireworks and OpenRouter to be less vendor-locked and access cheaper models. MiniMax, Kimi, and Qwen were the three I tried the most since they're fairly cheap, but I still burned thorugh credits way too fast. Now I'm using Codex.
I've heard that if you can deal with slower output, local LLMs can run on gaming graphics cards well enough. So they're not great for a live coding assistant, but running some tasks/agents overnight is an option.
It's been a few months since I looked around at this topic, but Fireworks and Openrouter were the two options I (briefly) tried.