Wow, you weren't kidding. I looked at their chart, and the cost-per-task for Fable is more than double Sol's. And DeepSeek absolutely stomps. Four cents per-task vs Sol's $1 and Fable's $3.
I might need to check out DeepSeek more. I had no idea the difference was this obscene. Makes me wonder if something's off with the benchmark. A 70x cost reduction vs. Fable seems too good to be true.
In my benchmarks of security auditing abilities of models, DeepSeek was roughly an order of magnitude cheaper than either GPT 5.5 or Opus 4.8, less than ten cents per task vs. roughly a buck each for the best American models at the time.
GPT 5.5 Pro was ~230x at almost $23 per task.
DeepSeek is my go-to when I need an API, and local Gemma 4 won't do because it's either too slow or not capable enough. DeepSeek isn't at the frontier but it's good enough for a lot of things, very cheap, and quite fast. Flash is even faster and cheaper, and still better than anything I can host locally.
> Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers to improve accuracy.
Dumb nit, but why not put your own press release through your model to prevent basic things like missing quote marks? Reminds me of that time an OAI released wildly inaccurate copy/pasted bar charts.
It does seem to raise fair questions about either the utility of these tools, or adoption inertia. If not even OpenAI feels compelled to integrate this kind of model-check into their pipeline, what's that say about the business world at-large? Is it that it's too onerous to set up, is it that it's too hard to get only true-positive corrections, is it that it's too low value for the effort?
Businesses do whatever’s cheap. AI labs will continue making their models smarter, more persuasive. Maybe the SWE profession will thrive/transform/get massacred. We don’t know.