Rendered at 19:20:27 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
LaurensBER 1 days ago [-]
I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day.
OpenCode Go even has double limits temporarily so for 10 USD you effectively get 140 USD of tokens to spend. It would impress me if someone could burn that amount with "normal" usage. Even when running multiple sessions.
I have a Claude Max subscription but I've barely touched it, it just feels like a step back to have to think about limits and usage even though the models are stronger.
The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues. Test coverage too low? Auto generate tests on CI for every pull-requests! Monitoring server logs, continuous security audits and investigating every received exception now becomes possible.
I'm thinking about having it automatically filter and re-rank my social media feeds so I can steer the algorithm instead of the other way around.
Perhaps other people (with enormous budgets) were already doing all of the above but for us this is a really exciting release!
abixb 23 hours ago [-]
If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international market if a smaller model can match SOTA of just a few months ago.
There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.
FernandoTN 22 hours ago [-]
I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes.
For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...)
The real unlock will be, and you can already see it with GPT-5.6 and Fable-5, to delegate complex enough tasks that will take more than 24 hours to get done and they will not lose track. I'm not talking about a loop, but the actual intelligence to recover from these compounding errors that accumulate in dumber models.
We're still a long way from the intelligence needed to let one of these agents go ahead and supervise multiple layers of sub-agents underneath to do complex orchestration. The future looks very promising and exciting. Imagine having the possibility of a Frontier model orchestrating as many sub-agents as needed that are running on cheaper models like DeepSeek.
Cookingboy 20 hours ago [-]
That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them.
You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?
AussieWog93 19 hours ago [-]
I think we put up with Fable's occasional hiccups because there's nothing better at the moment.
I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides.
If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entire $100 right there.
Hell, you could argue that the best "coding model" that we have at the moment is the human brain, and people will gladly pay $10,000/mo for one of them.
Arguing over $20 vs $100 for something that actually puts in work just seems insane to me.
palata 18 hours ago [-]
> I think we put up with Fable's occasional hiccups because there's nothing better at the moment.
Which was an argument for using every less powerful model since the moment they got useful, right?
When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough back then, DeepSeek is better than that now.
Sure, it's always nicer to have a slightly better model. But the price difference starts mattering a lot more when all the models are already sufficiently good.
AussieWog93 18 hours ago [-]
To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time.
We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's expected to find another few grand per month in misallocation/inefficiency.
I wouldn't be surprised if we were doing more like $10k/mo higher in 6-9 months' time.
When you're talking about numbers like this, the fact that one AI is $100/mo and another is $10/mo or $40/mo doesn't matter. They could make GLM-5.2, or any other Opus 4.5-class model free and it still wouldn't make sense to deploy in a commercial context.
The other angle I'd approach things from is that Opus 4.5 (and I'd agree with you that that model was the saddle point) was "good enough" for the types of things we were asking it to do back then, but as the models have become more capable the tasks we're asking them to do have also expanded with it.
I know I've personally gone from "hey can fix this race condition with a Redis mutex" 6 months ago to "Independently redesign this full embedded USB stack and QA it end-to-end, working around a specific Kernel bug in macOS Tahoe that requires decompilation to find the source of, while keeping in mind the constraints of our 8-bit AVR chip from 2011" now.
But that said, yes, maybe in 5 years' time we will reach an "intelligence saturation" where the average person won't be able to even conceive of how to use the new SOTA.
ray_kay777 16 hours ago [-]
I think we're even starting to reach that saturation point now for a lot of people. In my industry (law) plenty of people have tried CoPilot once or twice, or tried ChatGPT a year ago, and as a result have basically dismissed AI as being useless. The setup required to be able to get it to do end to end tasks to your liking is also substantially more work than most people are willing to put in.
EarthAmbassador 3 hours ago [-]
I find 500x returns on my $20 Chat GPT and Claude subscriptions as I litigate pro se against large law firms in federal court.
solenoid0937 12 hours ago [-]
I think HN doesn't really understand fixed costs, I spend a few hundred dollars per day on Fable and the costs are irrelevant compared to what we make.
F7F7F7 5 hours ago [-]
I've never understood this line of thinking...
Are you saying we shouldn't care about the future of affordability and access because at this moment we have seemingly endless access?
Sounds extremely short sighted.
solenoid0937 1 hours ago [-]
No, its that going from $2k/mo to $20/mo is not worth the capability loss in the slightest, because AI just isn't a big part of our fixed costs.
RealWed5 16 hours ago [-]
"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s
user43928 11 hours ago [-]
Don't you think that's a bit of an inane comment to make about a setup you know nothing about?
Surely they are using read-only access.
AussieWog93 9 hours ago [-]
We are obviously limiting access to be read-only for anything customer-facing (it will be able to put recommendations in the dashboard but not actually change things directly) but I think the point of GP's comment was literally just to get a reaction.
ciefa 1 hours ago [-]
Did you mean to post this on Reddit instead of HN?
petervuto 10 hours ago [-]
When you do things in life, sometimes things go wrong. Oh no!
This is such a ridiculous objection for how beloved it is. Wide swathes of the public can't cope with any adversity or risk.
don_esteban 5 hours ago [-]
No, the point is that with human in the loop the downside is (usually!) rather limited, as common sense would stop obvious fuckups (ok, not always, but still).
With an agent (especially incompetently employed), the danger of unwittingly destroying your company (or at least, the crucial data/reputation) is rather higher. We are notoriously bad at estimating the downside risks in complex systems.
The most obvious case is the downside risks in complex financial constructs... things look great for a while ... until a sudden surprising collapse arrives and totally destroys all the upside you think you have created.
godwinson__4-8 18 hours ago [-]
The question low cost models will create: Why would you massage output?
Fable 5 is still going to mess things up at any sufficient complexity. The advantage of low cost models with "good enough" intelligence is they can recursively correct. Why? Because it is cheap. Proper requirements and tests and subagents take away increasing amounts of work, at a cost that is not prohibitive.
If you are reviewing code manually you might consider Fable 5 a worse option. As it articulates itself with higher confidence and you already know it is capable, you are may be more likely to miss a mistake. You know to be on guard with a junior engineer. Reviewing a senior who suddenly makes some weird stochastic mistake can be a lot harder. It would be like if the smartest human engineer you knew was capable of some random brainfart in the middle of their massive diff. Imo, much harder to deal with.
Of course, we should keep in mind Fable 5 is only expensive today. It will be cheaper in the future. Autonomous, recursive prompting and improvement is the clear end state. Especially for entities that will always have the budget for that at the SOTA frontier.
Cookingboy 18 hours ago [-]
Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models.
If $100 Claud Max subscription works for you, then great.
But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic.
For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive.
And unfortunately for the American tech industry, the valuation is based off those enterprise deals, not your $100/month Claude Max subscription.
azinman2 15 hours ago [-]
Didn’t deepseek recently announce prices will go up significantly?
Right now the US dominates everyone else in actual chips
in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.
monster_truck 13 hours ago [-]
It's far from significant, it's partially doubled during peak hours. They could 16x it and it would still be two orders of magnitude better value than OAI's $200/mo plan.
It's that good. They are far from capacity limited, and even if they were, you can rent a single MI300X from somewhere like Hot Aisle and get more tk/s than you'll be able to use.
user43928 11 hours ago [-]
How does it compare to 5.6 Luna after the permanent 80% price cut?
That one is dirt cheap at API pricing, I can't imagine quota is going to be a concern on the $200 subscription, which in my opinion easily supports full time use of 5.6 Sol on xhigh.
xbmcuser 15 hours ago [-]
Deepseek is open weight/source ie it will be running on US servers in the US maybe by on each companies own servers.
re-thc 10 hours ago [-]
> Didn’t deepseek recently announce prices will go up significantly?
They also previously said prices will go down significantly once they get a hold of the upcoming Huawei chips (later this year).
Prices are going up just because they can. It can easily come back down. They aren't strained by some IPO / VCs requiring them to 1000x their earnings.
oceanplexian 12 hours ago [-]
I don’t think the industry knows how to price this stuff. Deepseek is great (I’m running it on a RTX 6000 pro setup) but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine, until you experience how good these models can be.
Think about it this way.
Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time. To the lay person this sounds trivial but to a serious business this intelligence gap could represent millions, or billions of dollars.
Cookingboy 12 hours ago [-]
If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5.
But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models.
There is a reason why Chinese open weight models are now popular even in American enterprises, because CTOs realize that they are indeed good enough.
solenoid0937 12 hours ago [-]
I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money.
The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this is so hard for HN to understand.
margorczynski 11 hours ago [-]
It depends on the use case. And most companies (like 90%+) do not have the coffers FAANG has and price does make a big difference.
solenoid0937 8 hours ago [-]
Unless you are very cash strapped, Fable is a very nominal fixed cost compared to the benefit of what it offers (fully autonomous agents, and no longer needing to pair program with one).
And even it isn't "enough". I can very clearly see myself using more advanced agents to move up the abstraction ladder.
For businesses that have actual problems to solve, I see them investing in the frontier for a good bit longer, probably until we have AGI that can replace employees, maybe even a bit after.
This is why I find the "good enough" arguments silly. Like, the usefulness of an AI tops out to you when you can pair program with it? Seriously? You cannot envision ways in which more advanced AI enables you to do more, better? That's bizarre to me. I don't ever see myself running out of problems to solve.
stymaar 9 hours ago [-]
> just a very last-gen way of using agents.
Fable has only been out for a month but somehow everyone is supposed to have moved to a completely different way of working that supposedly only works for Fable and nothing else…
solenoid0937 8 hours ago [-]
This stuff is just obvious to anyone working with the latest models. Fully autonomous agents are a game changer. Having to pair program with one is indeed "last gen", I haven't done that for a month and I won't ever be doing that again in my life, outside of personal projects.
stymaar 7 hours ago [-]
> This stuff is just obvious to anyone working with the latest models. Fully autonomous agents are a game changer.
This kind of takes makes me cringe. Why don't you go back to LinkedIn?
I have no idea what you mean by “pair program with an agent”, but Opus have been able of autonomous coding since last November, and with any half-decent harness even local Qwen3.5 was able to do so 6 months ago.
Fable is a stronger model, which means it can solve harder tasks but it's also over-hyped, because only a small fraction of task is hard enough to be Fable-worthy.
throw10920 33 minutes ago [-]
> This kind of takes makes me cringe. Why don't you go back to LinkedIn?
I use Kimi K3, Opus 5, and Fable on a daily basis.
Fable is the only one that reliably one shots complex changes and makes the right design choices. Everything else requires handholding.
I can let Fable loose on a 12+ hour (for AI) task and it will have performed it flawlessly when I come back the next day. K3 and Opus are not like this.
And no, our harness is not the limiting factor here.
seanmcdirmid 10 hours ago [-]
I use Chinese models, even smaller local ones, for much more than pair programming. If we are talking about deepseek v4 flash, which is basically a frontier model, it is much more capable than the local models I run on my MacBook Pro. The only issue really is finding the right harness.
I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough.
solenoid0937 8 hours ago [-]
> If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough
I mean, this proves my point. Better models enable you to get more done. With Fable, 80% of the time, I no longer have chat with an agent over the details of a PR. I give it an outcome and it gets done. This means I can work on much more with the limited time I have.
And I don't see this ending. When better models come out that take that from 80% to 99.x%, I will have that better model manage teams of other models and move up the abstraction layer.
If models get even better than that, perhaps I stop reviewing PRs entirely. Maybe normies can start using agents to build real things.
Unless your business doesn't have many problems to solve and isn't in a competitive environment, it will benefit from using the best models.
seanmcdirmid 58 minutes ago [-]
If you don’t have a way to automatically check the results via redundancy, you need a really good model, since even 99.9% reliability is going to cause slot of headaches. If you do have a way to automatically check results, then you can use something that fits in your computer and was produced 3 years ago.
My point is you don’t need the best model if you just put in QA processes that can be done by models also. And if you don’t have that, the model is probably not going to be good enough.
KronisLV 11 hours ago [-]
> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.
On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.
Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.
Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.
I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.
re-thc 9 hours ago [-]
> This matches my experience with DeepSeek V4 Pro at Max reasoning,
That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper.
> On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time,
GLM 5.3 bridges this gap.
> I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates.
Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
KronisLV 2 hours ago [-]
> GLM 5.3 bridges this gap.
I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release!
> Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU production still has ways to go and export restrictions definitely limit what they can do.
Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
re-thc 32 minutes ago [-]
> Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
They're a lot larger e.g. Fable. It is insanely bad. Do you know how much more resources "Western" companies have? Most in China don't have random GPUs to "play with" like every "frontier lab" employee does.
> As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
They're born lucky though. Efficiency is "free". The next generation hardware e.g. Nvidia claims Blackwell -> Rubin is 10x efficiency (verified by Neoclouds apparently).
AussieWog93 9 hours ago [-]
> It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models.
It's been 2 months since Fable was released to the general, man. Nobody knows what's going on inside of these companies except the people at the coal face.
stymaar 11 hours ago [-]
> but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine,
It's funny to see that Anthopic shills have been saying the exact same thing for the past two years now (and it was OpenAI fans before). It's amazing to see that Claude 3 Sonnet was "great" but now that even Qwen 9B is better than this version of Sonnet DeepSeek V4 is still not good enough despite being stronger than Opus 4.7 was.
> Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time
If you think Fable makes 20 times fewer mistakes than DS4 you're delusional. It doesn't even do 20 fewer mistake than Gemma 4…
serf 18 hours ago [-]
> not your $100/month..
This is made brutally obvious by anthropics customer support for people with such accounts.
AussieWog93 18 hours ago [-]
Yeah, fair. If we were talking $2,000/mo vs $200 then the maths starts looking very different.
crossroadsguy 13 hours ago [-]
The discussion started about cost-vs-usability, and then you brought in "but humans, the best LLMs around, cost much more" into this discussion to make it a not cost-vs-usability discussion. Do you not find it a bit disingenuous?
AussieWog93 9 hours ago [-]
Not really, no. My point is that a $100/mo LLM subscription is generating $5000/mo in value, so even a negligible difference in performance wipes out the cost savings completely.
Same reason it makes sense to assign a team of humans that cost $100k/mo to a product that brings in $5M/mo, rather than one human with 5 Claude Max subs.
The cost is a rounding error.
Art9681 4 hours ago [-]
SOTA models are exceedingly good at self-correction in the right harness especially in domains where things can be proven mathematically. Most people are just deploying Claude Code or Codex with default settings and rub the genie lamp expecting exceptional results. Garbage in, garbage out.
KronisLV 11 hours ago [-]
> You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?
This is why I quite like Kimi K3 - close to the same performance (definitely like Opus, approaching Fable), noticeably cheaper, generally good enough for me to daily drive. Only problem is that their official provider (on the Vivace plan) feels kinda slow, I'd say close to 2x slower than Opus on Max reasoning on average (probably more relatable than Fable).
budsniffer952 4 hours ago [-]
"isn't the old version good enough?"
You can say this about literally every product we buy. And yet...
bflesch 5 hours ago [-]
"It's only real AI if it comes from the silicon region of USA"
xyzzy123 20 hours ago [-]
There are a lot of tasks that are hard for organisations to run consistently but require some intelligence - monitoring logs and metrics for anomalies and security events, backup audits, audit processes in general, ensuring document quality and consistency, database advice and tuning, customer experience management, process optimisation - that are not "long horizon" in the classical sense of each step depending on the last, but are the result of consistency and attention over a long period of time and a large amount of data.
For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.
svachalek 21 hours ago [-]
These are not 24 hours of inference with floating point errors accumulating; largely the system guards against errors compounding. Tool failures, compile failures, test failures, etc, push back against the model taking a wrong turn and force it to correct.
Yes it's much easier to have a smarter model that goes straight to the correct answer first, but it may not be necessary or economical. There's a minimum bar for the model where it understands problems and knows the right step to correct them, and above that newer models give diminishing returns.
mycall 17 hours ago [-]
> it's much easier to have a smarter model that goes straight to the correct answer first
That's basically ASI not AGI, if you agree humans are NGI (natural general intelligence) and make mistakes and wrong decisions in solutions all the time. Right steps with some wrong ones is acceptable though for AGI.
spaceman_2020 14 hours ago [-]
Majority of white collar work absolutely does not require sota models
solenoid0937 8 hours ago [-]
White collar work will require SOTA models up until the point where said models can automate white collar work entirely. Then, we might finally see a meaningful commoditization crunch.
spaceman_2020 6 hours ago [-]
I genuinely don't know any work that has been taken over completely by AI without humans in the loop managing things
At this point, I don't even know if its possible
IrishLagger 2 hours ago [-]
Amazing how many inaccuracies you fit in there.
- You describe the breakdown in terms of time but it's more accurately a function of reasoning complexity.
- You seem to assume that no intermediate evaluation is possible.
- Often it is (e.g. the build breaks or tests start failing), allowing for course correction. There's definitely a cost to that but it can still be cost effective if the accuracy is "good enough" and the price difference significant.
- There are numerous tasks that don't require Fable or GPT5.6 level reasoning to improve efficiency by an order of magnitude.
copperx 21 hours ago [-]
> these compounding errors that accumulate in dumber models
While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive.
I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).
fuck_google 11 hours ago [-]
[dead]
dd8601fn 15 hours ago [-]
I have a silly (but honest) question. What's an example or two of a > 24hr task that people are actually asking something to do? Like real life ones.
trollbridge 14 hours ago [-]
Decompiling / disassembling and annotating old software, making sure it can build cleanly back to the original binary, and then look for bugs or subtle issues.
Another one I did was a printer data stream translator from an obscure format to PostScript/PDF (or just PNGs), complete with cups support, etc so these old apps can easily be hooked up.
Flash is capable now of running long range defined-goal tasks like this.
vehemenz 6 hours ago [-]
I wondered this too. Also, the duration of the task depends on the quality of the prompt and the model used. I have a feeling a lot of these day-long tasks are bogged down by suboptimal tool use and on-demand python slop.
canadiantim 5 hours ago [-]
Working through the total backlog of issues in a repo that may have accumulated from planning sessions
est 19 hours ago [-]
> even small errors compound over time and can lead to catastrophic outcomes
So, death sentence even to frontier models?
onkarkdev 12 hours ago [-]
[flagged]
12 hours ago [-]
ComplexSystems 22 hours ago [-]
It is true. I don't care about having infinite frontier-level intelligence, and I don't care if Fable can one-shot frobnicate a klaxelzorp with a benchmark performance of 97%. I doubt most people do, in fact. I just want something that meets the baseline level of intelligence needed to be a really, really good pair programming agent. It shouldn't have any silly dealbreaker issues involving laziness or hallucinations, it should be smart enough to bounce ideas off of, and it should automate doing tedious boilerplate. And - most of all - I want to be able to afford using it as much as I want. That's what has happened here.
abixb 22 hours ago [-]
I wonder when we crossed the "99 percentile of intelligence for 99% of the usecases" threshold. At this point, the gains seem to be right at the very edge of bleeding edge for narrow and specialized use cases, and wonder if it'll be a sort of diminishing return from here on.
horsh1 21 hours ago [-]
In April
dukeofdoom 18 hours ago [-]
Probably the best counterexample is the games they are able to design. It's still mostly AI slop, few would want to play.
jqbd 12 hours ago [-]
With those cheap models the idea is you're still in the loop anyway so the more expensive model is a waste of time and money. In this case, that means you're steering the game to look like you want not how AI wants.
bronson 9 hours ago [-]
Same with the expensive models. Fable can't one-shot a good new game. Game dev is still human-in-the-loop no matter what model you're using.
mdp2021 11 hours ago [-]
> narrow and specialized use cases
Such as Decision Making. /s
You just can't set a high enough threshold of intellectual effort for critical decisions.
ekidd 18 hours ago [-]
> If what you're saying is true and accurate, then US-based AI labs are in big trouble.
I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook.
It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers.
It would absolutely have been a frontier model last December.
inciampati 22 hours ago [-]
It is amazing how fast it happened. Right now one of my main projects is fully running on DeepSeek flash. My reason was that I was blocked by both of the main US AI labs from working on it because it involves viruses. DeepSeek flash has been killing it since I switched it on, completing the first phase of the project and setting up an iteration in another application space. It isn't the most brilliant model, but it is reliable and I don't have to manage my weekly token allowance. I just spend freely and end up spending only a few dollars a day. Intelligence is going to become a basic commodity. Only special stuff is going to drive us to use special models. And maybe not even that.
monster_truck 13 hours ago [-]
Seconding this, flash and especially pro are damn near batting 1000 for me in my experiments. In one instance it figured out that the poc was running on a much slower machine and locked affinity to a single core.
Many devs who have never tried from either side build it all up in their head but it's almost always been a matter of thorough tedium, which LLMs are excellent at churning through, especially when there's api docs/headers/code comments.
If you have the space, try mirroring your port at the switch level and capturing every packet then making it go through them all to look for whatever. We have NSA at home lol
chme 14 hours ago [-]
> If what you're saying is true and accurate, then US-based AI labs are in big trouble.
Well... I would think that the whole AI industry in the US are working towards public bailouts... Which I guess they'll get under the current administration... So they'll be fine... Nobody there really seems interested in actually creating a profitable business anyway...
budsniffer952 4 hours ago [-]
One random says they are using DeepSeek and you proclaim it's all over for the top AI labs. Brilliant.
First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses?
Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American labs are SOTA. That might change, but let's not act like the American labs don't know what they are doing.
The idea that people are going to use cheaper models for cheaper work isn't some novel revelation, it's completely obvious. People are doing it already, eschewing Fable.
edot 4 hours ago [-]
Subsidy doesn't matter once the weights are released. You're conflating subsidized training and research with subsidized inference. DeepSeek's (allegedly government-) subsidized inference is coming to an end soon per their docs, but that's fine because the weights are open. You can run this on-prem with very modest hardware. And you can fine-tune it to your business with your data.
budsniffer952 4 hours ago [-]
Then why are the major American labs in "deep trouble"? They can also continue to be subsidized and release the weights if they want to, if that's what you think it takes for survival. Seems like a low bar.
The point is it takes money to keep developing models. Everyone is playing by the same rules. At this point, the US labs are trying to build businesses. I'm not really sure what the Chinese labs goals are. But I do know they aren't doing charity work.
audunw 8 hours ago [-]
These companies know their value was never in the models. There’s a scramble to acquire as much hardware as possible, and to build as many services that people get soft locked into, such that they have a moat when the open weight models fully catch up. Doesn’t matter that the models are open weight, if you want access to the hardware you will have to pay.
I dont see how the outlook is any better for the open weight companies. They’re in the exact same situation as the closed weight companies except they have had much less revenue, and built up less of a brand, leading up the the point where they are equal in terms of model quality.
Art9681 4 hours ago [-]
It's not necessarily true. OP has not provided actual objective figures. The nuance is in tokens spent per successful completion. So sure, you can blast DeepSeek for 1 hour solving a hard task, or Opus-5 for 5 minutes on the same task. Opus would ultimately be cheaper because time IS money and solving things faster is ultimately cheaper. Sure, benchmarks will show cost to be lower but people who test all of these models in real world use cases know the benchmarks are not reflective of real world use cases and Chinese models choke on problems Fable or gpt-sol will breeze through.
btbuildem 18 hours ago [-]
> only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models
In what kind of sad and failed dystopia is this a "saving grace"? For whom?
swiftcoder 10 hours ago [-]
> "saving grace"? For whom?
For Anthropic and OpenAI, presumably. And the rather large economic distortion field around them, that may or may not go very badly for all our retirement funds if those firms become insolvent...
reacharavindh 5 hours ago [-]
Sample of 1…
I downgraded my Claude subscription and delegated my Claude Opus access to serve the role of an Architect to brainstorm and plan every step of development.
I leave the development to Deepseek.
Claude gets to review at many layers. It is often just as good as if I let Opus develop it by itself(the architect session will find similar number/level of gaps).
Deepseek flash as an architect and problem solver is not as thorough as Opus5 + high.
Codex sol+ high is even better than Opus 5 at this moment for my needs.
apatheticonion 12 hours ago [-]
I've been using DeepSeek exclusively since I've been playing with LLMs, largely due to the price.
I joined a company that is an Anthropic shop and I am genuinely shocked.
Sonnet 5 is a little better on long horizon tasks and headless unsupervised agent workflows - but for in-IDE workflows, it's virtually unusable.
I am so used to flipping around my codebase at warp speed with DeepSeek flash. It's so fast and accurate I don't have time for parallel agents. It's a really rewarding workflow.
Moving to Sonnet, you ask is something simple like "split this into a seperate file" "implement this method" "this is my schema, implement a repository for it". It'll spend 30 minutes thinking and charge like $12. And no token caching, what are you even doing Anthropic?
It's unusable.
DeepSeek are in a league of their own
horsh1 21 hours ago [-]
If programming in the US to become unconditionally 10x more expensive, then the exodus from the US is about to begin.
thenthenthen 14 hours ago [-]
Could we look at manufacturing, for example the car industry, to predict what will happen?
stockworks 19 hours ago [-]
I couldn't agree more, and think of all the wasted inference across accounts overpaying for their subscriptions.. Need a secondary marketplace for this stuff.
ngl999 14 hours ago [-]
Only to to find out they end up in a much worse place
swiftcoder 10 hours ago [-]
> Only to to find out they end up in a much worse place
Like Europe?
chii 12 hours ago [-]
> There's no way large companies outside the US will pay the "US AI lab" premium
this has already happened with manufacturing, so it isn't surprising that other industries follow.
The US premium in engineering and scientific endeavors have been lacking for the past 30-40 years, and if it werent for tech and silicon valley, the US would have nothing state of the art. Even on that front, the US is falling behind given how much effort in tech has been diverted into privacy invading, and advertising.
The US has been riding momentum, but eventually that momentum will stop. It will take half a century to get back up to speed, and by then, the US will have fallen behind so far that catching back up will seem impossible.
The telling evidence would be if china has the first moonbase before the US does. I think this is highly likely looking at today's US administration.
alfiedotwtf 12 hours ago [-]
Laguna S 2.1 is American and seems to be on par (or better) than DeepSeek v4, is only 110B, and it's faster! Weird there's not much talk about it vs other models.
rcpt 19 hours ago [-]
The bet isn't that people will be able to automatically reply on bugs and rack up API charges.
The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model
palata 18 hours ago [-]
> You don't win the stock market or make the deadliest drone by switching to the cheap model
Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"?
I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?
qup 18 hours ago [-]
There are more than two options. What about great engineers with great models?
palata 13 hours ago [-]
Again, I was answering to:
> You don't win the stock market or make the deadliest drone by switching to the cheap model
The question is not "can you win with the best model?", it is "can you not win without the best model?".
fuck_google 11 hours ago [-]
[dead]
conmod278 11 hours ago [-]
> If what you're saying is true and accurate, then US-based AI labs are in big trouble.
As it is now clear by behaviour of companies and US govt, all these investments will be backstopped by US govt. No US AI company will go hungry, they are national champions.
gigatexal 19 hours ago [-]
Yeah this is what I’m curious about. How good are they after the benchmarks. I’ve been told yeah they’re good but they’re just building to show off for benchmarks.
The ByteDance folks are apparently training a mythos level model 10T params apparently. If they do would it still be subsidized at these cheap rates?
anal_reactor 21 hours ago [-]
It's been true for almost every business. "Cheap and good enough" usually trumps "excellent but expensive". Ikea, McDonald's, Ryanair, AliExpress, Aldi - these brands prove that catering to poor people is more profitable than catering to rich people simply because there are so many poor people that their collective spending power outweights the one of rich people.
Seems more like a overall industry problem, not limited to low cost carriers
odo1242 20 hours ago [-]
Well, not universally. It’s a tradeoff. If what you said was universally true Apple wouldn’t exist; Spirit Airlines wouldn’t be bankrupt, etc.
tsunamifury 20 hours ago [-]
Apple sells to half the American population. And by definition many of them are poor.
Spirit was broken by oil prices which everyone pays the same for. (There is no cheaper jet fuel alternative).
Not a good comparison to the point of wrong conclusions.
Barrin92 15 hours ago [-]
>Aldi - these brands prove that catering to poor people is more profitable than catering to rich people
At least here in Germany Aldi isn't even really limited to the poor, it's famously a place where you can run into anyone. Where I used to live in Berlin close to the government district I literally on occasion ran into the chancellor (and her bodyguards). Aspirational shopping where you buy premium goods to pretend to have higher social status honestly seems a bit on its way out. Even middle class people seem to consciously shop more utilitarian now.
21 hours ago [-]
twotwotwo 55 minutes ago [-]
Hard to make big predictions, but it sure looks like at least this level of capability is going to be available in the open and relatively cheap to run.
The 'floor' has gone up: today's model a bit behind SOTA is like model releases that were blowing people's minds a few months ago. Compared to, say, DS R1, this is far-out futuristic technology.
This, Luna, and (if it's good in practice) Laguna S are also fast and light not just cheap. And, as happened before, DeepSeek's first but other open model makers likely follow.
And a small, fast model taking small steps is...fun? More of the experience of working with the code.
trollbridge 14 hours ago [-]
I kept running into it looping two nights ago or else getting trapped in a reasoning loop it couldn’t escape from. Switching to Pro helped, but I ultimately had to use GLM-5.2 to recover my session. (GPT-5.6-Sol’s cybersecurity guardrails went off since the problem I was trying to fix involved a race condition where it would segfault and the words “stack frame” in my session made it decide I was being naughty. Yet another reason not to use American models…)
Another time, Flash started trying to make tool calls by just calling bash and catting the tool call to stdout. Then it started running echo xx for every two letter UNIX command it could think of: mv, cp, etc and the it dug into uv, ty, and jj
bronson 9 hours ago [-]
What was your prompt? If something benign, then it sounds like your harness has an issue with its tool calling that the agent is atrying to work around.
paxys 23 hours ago [-]
How is $5/day irrelevant? In the $150/mo range you can get effectively unlimited usage of GPT 5.6 Sol (Pro plan). Why use a much weaker model for the same price?
ux266478 23 hours ago [-]
> In the $150/mo range you can get effectively unlimited usage of GPT 5.6 Sol
With 5 active sessions going nonstop? That seems like a pretty important qualifier.
brynnbee 23 hours ago [-]
Others have said similar but I disagree, I'm spending $200/m and I can easily burn through my weekly quota with a few overnight goals using 5.6 medium.
ElijahLynn 21 hours ago [-]
Same here with Claude Opus on 20x Max plan. Easy to burn through with 3-5 parallel sessions, each with their own subs going.
Then, once I go over, API pricing racks up FAST!
brynnbee 17 hours ago [-]
I burned through $100 of credits in literally 40 minutes doing the same long running task I always do.
darkwater 22 hours ago [-]
And what do you do with all that?
brynnbee 17 hours ago [-]
I've reverse engineered multiple classic games and turned them into popular, browser-based MMO-like experiences.
I'm also creating a free platform that replaces extremely out-of-date software, some of it only available with mutli-million dollar contracts, to help medical physics professionals with cutting-edge radiotherapy devices used to treat cancer.
Nice! I tried to play https://eternalsagas.com/play but it stalls at 45% loading with the progress bar always at 0 but the music playing.
And out of curiosity, how do you automate testing the porting in the browser that's actually playable etc?
And aren't you a bit scared of hosting and serving the "hairy bits" such as full assets?
Nice job anyway!
brynnbee 6 hours ago [-]
Yeah that's my weakest browser game since I've actually transitioned to using a desktop Godot client I haven't released yet. It will load eventually it just takes forever bc the assets are quite large, which is part of the reason I've had to move to Godot. That project doesn't exactly have an active user base which is why it's lower on priority list to fix.
I automate testing playable parts in browser by adding a dev mode that allows text commands for everything instead of having to rely on clicking UI or 3D elements. It can see the full game state in JSON and interact in any way via commands.
And re: IP - I just accept that I might get a C&D any day and have to take it all down. I'm careful to not accept a single penny for any reason and don't even have Patreon. Usually monetizing is what makes IP owners unhappy. And for the Pokemon MMO I just don't advertise it anywhere meaningful since they'll C&D the second they see it. Largely made it for my nephew and we play it together.
pylotlight 17 hours ago [-]
Time to make creative software for linux that can replace adobe for Video/photo editing :P
brynnbee 6 hours ago [-]
GIMP isn't bad these days, I use it constantly
ccakes 17 hours ago [-]
An MMO of Full Throttle would be amazing!
brynnbee 17 hours ago [-]
It'd be interesting to see what someone could do to turn a 90s era adventure game into an MMO! I never played the game myself but I think basically any game dev project that modernizes stuff is really neat.
baby_souffle 22 hours ago [-]
5.6 medium is pretty good at implementing moderately complicated things as long as you've done a good job specking out the types and the API contracts and acceptance criteria.
I can pretty easily burn through my weekly quota over several agent coding hours with minimal supervision when tasked with some pretty large but well-planned refactors.
rain_iwakura 23 hours ago [-]
not at all true. if you're truly using it across the board for smaller things (translation of pages, filtering of every individual tweet based on its relevance to you etc), the costs ramp up super quickly.
i used for work where i did less and it quickly reaches thousands if you're not careful. i can already see what some will say: skill issue et cetera - whatever.
paxys 23 hours ago [-]
$100-200/mo is the subscription price. You aren’t going to go over. And you can select smaller models as well. Not everything has to be done by the most expensive one.
janalsncm 21 hours ago [-]
You just get throttled, which interrupts your whole workflow.
whimblepop 23 hours ago [-]
I thought it went without saying that GPT 5.6 Sol is the wrong model to use for things like filtering tweets. Apparently not?
Drupon 13 hours ago [-]
You would think that no one would be stupid enough to use Fable or Sol for small one-off tasks like filtering tweets, but AI has opened up a lot of avenues for stupid people to ship code. It's only going to get worse.
nicoburns 22 hours ago [-]
If you factor in cost then it may well be, but it's definitely the case that the high-end models can get you significantly better results than the cheaper models even for tasks that feel like they should be straightforward.
_aavaa_ 23 hours ago [-]
I really doubt that the either of the pro plans are subsidized heavily enough to support you swapping v4Flash for Terra, much less Sol.
$5/days is ~330 Mtok/day, that’s a nontrivial amount of work, and none of the gpts are more efficient than deepseek at $/task if deepseek meets your quality bar.
usef- 22 hours ago [-]
They do give a significant amount more than you would expect from the API pricing. The US provider API pricing has heavy margins by all accounts (and most are short enough of GPUs that there's little incentive to drop).
LaurensBER 21 hours ago [-]
5 USD is at the "raw" API price.
OpenCode currently offers 60 USD API credits at 10 USD per month (OpenCode Go) and have even doubled it temporarily as a promotion.
Effectively you can get Deepseek for 1/12th the already ridiculous cheap API price.
jboss10 19 hours ago [-]
From here, it looks like opencode is hemorrhaging money. I've got a Opencode Zen free account, and I've been using deepseek-v4-flash-free on Pi for a bit, and I haven't hit a limit yet. Sometimes my request fails, but retrys work. I know this is a very cheap model, but it's being given out for free. I assume they might be training on outputs?
_aavaa_ 17 hours ago [-]
Those are not at the same price. Opecode’s 60 USD of deepseek usage is charged at much higher rates than what deepseek themselves charge at.
swiftcoder 10 hours ago [-]
> Opecode’s 60 USD of deepseek usage is charged at much higher rates than what deepseek themselves charge at
Per the open code zen pricing page[1], it appears that the token prices are the same, but their cache is 10x more expensive?
For flash the input/output is the same, but the cache difference is big, you're paying 10x on >95% of your tokens.
For Pro, it's even worse, input/output is 4x and cache is 40x. The price different is really brutal. Yes you will still come out ahead by spending your first 10$/month on opencode go, but you will be saving a lot less than initially appears from their (60 USD for 10 USD pitch).
The cost per token is super low. If you're used to paying OpenAI or Anthropic API-based fees then the same workload on DeepSeek feels free.
re-thc 23 hours ago [-]
> In the $150/mo range you can get effectively unlimited usage of GPT 5.6 Sol (Pro plan)
Not true. Sol on XHigh or Max runs out even on the $200/mo plan. It's not close to effectively unlimited. Maybe at 2x the current allowance it can.
edot 4 hours ago [-]
What are you doing where you're getting Sol on XHigh or Max to run out on the $200 plan? Last night, I had 48% of my weekly limit left, so I spun up 4 projects I had been working on and ran them all on Ultra with Fast mode and /goal, and it took over 4 hours to burn through that. I mean, I was reeeeally trying to use it up because I just can't use it up with normal usage for weekend and after-work programming. I've had Max crunching on something this morning for 3 hours now and I've only used 4% of my weekly limit ...
re-thc 22 minutes ago [-]
> What are you doing where you're getting Sol on XHigh or Max to run out on the $200 plan?
Real work. $200 looks good on the outside until the essence of it, e.g. the models lie. I gave a list of spec to Sol and Sol decided some items didn't need to be done and the reason was "unproven", "not enough evidence", etc.
They all come up with amazing ways to lie (or be lazy). Often times what you get isn't what you asked for (only on the surface). E.g. I ran it to iteratively bench and optimize a better data structure for the project. It spent hours and finally came up with something. When I check it out -- it benchmarked the wrong criteria and was way off. So here we go again. Most AI work looks good on the surface. There are infinite edge cases.
So to do real work and gate it you need to:
1. Plan
2. Get it to do the work
3. Get independent agents to check from different angles
4. Take that feedback and get it to fix those gaps
5. Match against the plan and redo parts if needed
Every task is easily 4-5x the estimated amount of tokens.
p.s. well I did burn some banked resets building a compiler for some language AND it is still NOT done. Every time it says done I say check it says ok we still have bugs...
albedoa 22 hours ago [-]
Ignoring the other side of the equation is a pretty wild thing for you to do here:
> I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams)
surgical_fire 22 hours ago [-]
$5 a day is pretty extreme in DeepSeek. You really have to abuse it to get anywhere close to it. Maybe something in around hundreds of millions of tokens per day, considering cache hits and all.
And to be frank, it is not that much weaker for regular software development work. I use Claude at work and I see no difference in capability. I only notice a dramatic difference in how much more expensive it is.
Aeolun 1 days ago [-]
But DeepSeek now has a warning they’re going to sharply increase their API pricing sometime in the future.
LaurensBER 1 days ago [-]
Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers).
vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development.
I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek prices.
minraws 1 days ago [-]
As someone who recently tried it on some blackwell cards, it's possible to match the prices especially the input can be even cheaper and output can match the costs so you can easily build a net 20-30% margin business even at current GPU prices.
The entire issue is caching, I tried to write some custom to dump to disk kv-caching using some ideas from their papers and my experience with snapshots and vm checkpoint systems, I must say they must have really squeezed that lemon it's hard.
Atleast me with Sol couldn't figure it out over a couple days, a few hours each day, which isn't much but I did feel a bit stuck with existing solutions and felt like I might have to write something from scratch. But if you are willing to put in the effort into the infra I do think it's doable. But it will be really hard to pull it off.
My congrats to anyone who manages to pull it off, they might be able to kill off most AI labs. Assuming they can find the compute, Deepseek really has killed all models for me other than Sol/Fable/Opus/K3 tier stuff.
minraws 22 hours ago [-]
Mild info dump, since this has a few too many upvotes and some folks might be misunderstanding, 20-30% is assuming a typical agentic workload where input tokens dominate by over 20:1 or at least 10:1, if you are output token heavy then this is going to be a different ball game.
And there is no way in hell anyone can afford caching prices same as what DeepSeek is offering, and DeepSeek keeps the cache available for an insane amount of time most providers will flush it in 5-mins like Claude/Anthropic (some offer customizing it but I am not sure of the pricing, it's load based on some like Fireworks, which means assume a couple minutes at most, they say several minutes god knows what that really means).
There is no way to match DeepSeek's current prices, "profitably" if you are renting a GPU and reselling tokens, unless you have some really amazing caching infra or something.
Deepseek's prices are just insanely cheap, I am not saying it's impossible to get there the overall performance suggests it should be feasible, but I will be damned if any provider could match their tps and caching any time soon at those same prices profitably.
I believe even if Deepseek 2-3x their prices across the board even then they would be cheaper for most long running tasks, that's just how good their caching is.
For one I have managed to hit the cache after over 24 hours on their system it's insane, I honestly didn't care because it was so cheap but it truly made me incredibly happy to think about the engineering that must have taken. TTFT is slightly worse, but it's good enough, for those cache prices I can take a few seconds worth of hit on TTFT.
janalsncm 21 hours ago [-]
From what I understand about deepseek’s pricing, they are only charging what they need to break even.
Karrot_Kream 20 hours ago [-]
Thanks this is a comment with a great amount of useful detail.
twotwotwo 1 days ago [-]
One read is 1) they're getting a lot of traffic for Flash, 2) they've said they're updating Pro soon and expect that to lead to a traffic spike for Pro, but 3) that would leave them overloaded, so 4) they're going to raise prices to avoid it.
It's interesting that most open models adding 1M context did it in a way that reduces KV cache size (though DeepSeek was the most aggressive, using compressed attention on all layers), but only a couple providers turned it into a discount on cache reads.
onlyrealcuzzo 1 days ago [-]
> Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers).
Can anyone working at one of the main US labs (Google, OpenAI, Anthropic) comment on WTF they haven't even tried MLA - despite the obvious massive advantages?
I know enough to know they aren't completely incompetent. So there must be a quite good reason.
But it remains a mystery to me.
DeepSeek's MLA is like almost 2 years old at this time. They've got thousands of people working on this stuff. They clearly have the ability to at least try it...
aabdi 24 hours ago [-]
They already are?
There’s a measurable performance tradeoff versus gqa so there’s reluctance.
For the most part though the new deepseek v4 tech is hca and mhc and people are still catching on like with moe and rl. Wait for 6 12 months, minimum time for next pre train.
ronsor 1 days ago [-]
Are they not?
The big US labs are opaque and don't publish much of any technical details anymore. We don't know what they are or aren't doing, honestly.
re-thc 23 hours ago [-]
> they can replicate or beat the price with rented GPUs
They "can" is the caveat here. Rented GPUs are going up in pricing. I recently got an email that DigitalOcean pricing of GPUs were going up.
So
1. They have to get a hold of them (availability is bad)
2. They have to maintain the pricing
NorwegianDude 1 days ago [-]
Eh, what are you guys even talking about? Deepseek is not cheapest provider as is, and it's MIT. So deepseek making it more expensive to use is just nonsense, they can only change their own pricing. It's the beauty of MIT license and open weights. If anything, these models are some of the safest in the world to use if you worry about a rug pull.
LaurensBER 1 days ago [-]
There's more to inference than just the input/output token cost. Caching has a massive impact.
Deepseek charges $0.0028 per cache read on Openrouter. The next cheapest is $0.018.
That's a massive difference and quickly adds up on coding sessions (which often hit 95%+ cached tokens).
akman 1 days ago [-]
90%+ cache hit rate is common, and so you'll see on places like openrouter that Deepseek cache cost is indeed a magnitude cheaper than the rest.
greenavocado 24 hours ago [-]
My usage thus far from api.deepseek.com
- input_cache_hit_tokens: 1,265,646,976 x 0.0000000028 = $3.5438115328
- input_cache_miss_tokens: 18,208,088 x 0.00000014 = $2.54913232
- output_tokens: 9,615,178 x 0.00000028 = $2.69224984
- request_count: 10,837 (no price)
This adds disk as a tier in the HBM → CPU → Disk KV cache hierarchy.
There's also a cluster of related KV-offload FS PRs: #49225 (read/write batching, still open) and #49152 (batch store/load in C, merged Jul 28).
It's hard to say if these are similar to the approach DeepSeek takes but they definitely seem very interesting.
ms8 1 days ago [-]
Yes, there is warning, but also there are many providers on OpenRouter[0], hosting open weight model with similar pricing. The question is Will they go up as well?
DeepSeek has far cheaper cache pricing. That's the difference.
1 days ago [-]
fastball 21 hours ago [-]
But the model is open weight?
eli 1 days ago [-]
I assume/hope this is about prices going up for the next release of Pro
HSO 1 days ago [-]
even if they double it it`s from such a low base it is still supercheap
metadat 1 days ago [-]
Source?
dolebirchwood 1 days ago [-]
If you're on the DeepSeek Platform, you'd see this:
"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice."
I try to use all the intelligence I can, which fable, opus and then second tier models.
I am not sure why you wouldn't want to use the SOTA models unless speed is a concern. Otherwise you are leaving quality on the table.
amelius 24 hours ago [-]
> it's good enough to use it for (almost) everything
which in your case is?
rpdillon 23 hours ago [-]
I've posted a few times about my project that's a collection of 30k-250k webapps that are served from a WebDAV server. The apps know how to write updated copies of themselves back to the server.
My family uses it. I have gallery apps (yearbooks for each year are a lot of fun!) of us on trips and just living, an outlining app that's a mesh of Workflowy and Org Mode (it's called Fluxtral), a markdown-backed app (it uses marked.min.js, and is called Dextral) that offers documents, logs, calendars, and kanban boards, all parsed from markdown. I have a List app for gear, trips, shopping, etc. that we all can contribute to. There are utilities (world clock, calendar) and games (an oracle for RPGs, a KenKen implementation), and apps (a diagram editor that exports to SVG, a web-launcher that uses pneumonics, a Scheme-based hacking environment, and a spreadsheet that does most of what you'd expect aside from Solver and Pivot tables).
I started these projects before AI, and made slow progress over the years, but the modern versions of all this stuff have been built with Deepseek V4 Flash. I've also used Gemini in the very early days, and Kimi K2.6 later on, but these days, since I can now host Deepseek v4 Flash 0731 in a 2-bit quant on my Strix Halo box (128GB, but only about 250GB/s of memory bandwidth, so 15t/s), I used Deepseek with omp for almost everything. It's a very capable model, and I'm amazed I can run it locally and get good results. It's really revolutionary for my (small) use cases.
indigodaddy 21 hours ago [-]
Sounds fascinating! A blog write-up about your platform would be a fun read, if you're up to it
rpdillon 18 hours ago [-]
For sure! I'll be doing a Show HN at some point, just want to feel a bit more confident about certain aspects first.
indigodaddy 18 hours ago [-]
Makes sense, I'll look out for it! although of course most Show HNs these days get lost in a deluge of submissions.. you might actually be better off omitting the Show HN when you submit it..
podnami 22 hours ago [-]
A collection of 30k-250k apps? Like individual unique apps?
rpdillon 20 hours ago [-]
Sorry, a collection of apps whose size is between 30kb and 250kb.
p1necone 10 hours ago [-]
I read that and thought "ah that's going to confuse people, but I can tell they mean 30k-250k loc", so thank you for the clarification.
indigodaddy 20 hours ago [-]
Haha, I misinterpreted as well. That would be a lot of apps!
IamTC 12 hours ago [-]
I too have Strix Halo. Two of them. The model have been awesome for me too.
With unsloth's Q3_S quant + kyuz0 'llama-vulkan-radv-performance' toolbox, I am getting 280+tps (batch and ubatch at 2048) for PP and 18+tps for TG. I really only need 256k context so it all fits.
If I go down to the Q3_XXS quant + dpsark + 'llama-vulkan-radv-performance', I can get about the same PP and 25+tps for TG with draft set to 2 or 3. Fits about the same as above.
Edit: I did notice the 25+tps quickly degrades down to 20+ after the first few hundred tokens.
13 hours ago [-]
throwaway27448 23 hours ago [-]
[flagged]
dan_q 24 hours ago [-]
> which in your case is?
oh, they're mad.
anramon 1 days ago [-]
>even if it's not SOTA
And, probably 99.99% of people using LLM probably don't even need SOTA anyway.
swiftcoder 1 days ago [-]
At least on these benchmarks, it seems to be pretty handily scoring up with the SOTA from 6 months ago?
eru 16 hours ago [-]
> It would impress me if someone could burn that amount with "normal" usage. Even when running multiple sessions.
What's normal usage? I mean, Kimi is already really keen to spin of lots of subagents, and DeepSeep can probably do the same?
> The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues.
Yes, though I did that even with Claude (on my employer's token budget). The agents are great at doing the gruntwork of chasing down the reproduction of flaky tests, too. They need some hand holding at first, but the guidelines are usually re-usable per project. (Claude specifically needs to be told to really concentrate on reproduction, and not eagerly start fixing the flake: if you don't have a reliable reproduction, you have no clue whether your fix actually fixes anything.)
afro88 19 hours ago [-]
> The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues. Test coverage too low? Auto generate tests on CI for every pull-requests! Monitoring server logs, continuous security audits and investigating every received exception now becomes possible.
I don't think this is the win you think it is. It's amazing that this is possible, but it introduces so much human overhead that you can drown in reviews and it can effectively slow you down more than a quick check and fix yourself.
The models need to get a lot more consistent in what they can and can't do before you can automate this stuff and only check the things you know the model isn't good at
ljosifov 23 hours ago [-]
Hear hear. IQ tokens to cheap to meter upon us. So many things changed since last week. Now I've had Prime agent session grinding into its 20-th hour still not giving up. Been using opencode-go since Go sub appeared. What made a difference was deepseek-v4-flash and mimo-v2.5 showing. Very similar middling models ~300b so light on the gpu. 1M context and hybrid archs - so one can actually make use of that 1M (don't grind to a halt like others). In OMP I have one the primary (default), the other one as /advisor looking over the shoulder and nagging. On opencode-go in credits counting they are the bottom-2 in cost, cheaper by 200-350 times than than the top-1. Last week with deepseek-v4-flash-0731 another jump - now it's closer to the top models then to the middle. Now I don't even need the /advisor probably. Still left it there it's sometime amusing the models back and forth. :-) DeepSeek offer /v1/responses api now with flash-0731, so setup Codex to use that too. I'm loving this :-)
klardotsh 16 hours ago [-]
I would not recommend DSV4F (even 0731 edition) without an advisor. On its own it’s an absolute drunk intern in my experience, but with an advisor model watching like a hawk when it gets stuck in loops or goes down boneheaded rabbit holes, it’s fine (and very cheap). I’ve been using GLM-5.2 as my /advisor but might try just a second DSV4F instance.
sfifs 22 hours ago [-]
It's very impressive and I'm running it locally on 2x DGX. Non thinking mode is very responsive. Thinking mode has some latency but can be switched on when needed. Both are really good
h14h 20 hours ago [-]
Are you using it via the official DeepSeek API, or via a different model provider? If the former, it's worth noting that their cache read prices are one tenth that of every other provider ($0.0028/M vs $0.028/M), so folks who want to use a sovereign inference provider with a zero data retention policy likely won't see anywhere close to the same value.
mh- 18 hours ago [-]
Worth mentioning also that DeepSeek is the only provider in OpenRouter that was disabled-by-default until I enabled a setting: Allow paid endpoints that train on request data.
swingboy 21 hours ago [-]
Agreed. With less than $10 on the DeepSeek API used, I’m somewhere near half a billion tokens over the past week or however long it’s been since it came out.
I’ve found it to be very capable. I’m using it with pi as well and some custom extensions I’ve put together over the past few months and it’s pretty crazy having it do what I need it to a vast majority of the time, do it fast, and see that it’s used like $0.12.
12 hours ago [-]
Snowfield9571 15 hours ago [-]
You should try with reasonix. It is designed to maximize cache hits. It seems less featurefull in its current state than open code and omp but I am getting a 99% cache hit with reasonix making my costs insanely cheap.
m101 21 hours ago [-]
Could you go into how you run two instances that speak to each other in an implementer / advisor role in parallel? I’ve been looking for this sort of orchestrator / worker solution where there’s constant feedback and nudging between the two.
You can probably implement something similar as a plugin for your preferred harness. From a technical perspective I think it just sends the output w/h the thinking and tool trace to another model and asks it to double check everything (exact prompt must be somewhere in the OMP repo).
m101 20 hours ago [-]
Thanks I’ll give that a go.
Would you run a less costly model as the supervisor given it’s consuming a lot of text and may have a simpler task to do like “make sure the implementing model doesn’t start over-engineering things”?
LPisGood 22 hours ago [-]
Does auto generating tests even do anything helpful? Don’t they just sort of tautologically say the code does what it does at best or do something completely ridiculous like test and implementation that only exists in the test file at worst?
LaurensBER 22 hours ago [-]
We have an extensive description of _how_ tests should be written and they're reviewed by a human. All the AI does is fill in the boring middle part.
frogperson 19 hours ago [-]
The pricing was awesome, but deepseek just sent out emails warning of a large price increase.
jojohack 23 hours ago [-]
Running DeepSeek with Pi as well, any plugins you recommend running it with ( e.g. native browser for snapshots, etc. )
electroglyph 21 hours ago [-]
my experience is the same, but deepseek is planning on increasing prices soon, which will make it a lot less attractive
Seconded. I love OpenCode and Pi, but omp is my daily driver.
stavros 20 hours ago [-]
What's good about it? I use OpenCode and it does what I need, basically.
rpdillon 18 hours ago [-]
Spawning subagents via orchestrate and using /advisor are both super valuable. Haven't really unlocked the full capability of omp yet, but hoping to over time.
ljosifov 23 hours ago [-]
omp - current top, after using codex claude opencode pi that I still use too
ycui7 18 hours ago [-]
it is funny when people say i am struggling to spend money.
dominotw 24 hours ago [-]
> Auto generate tests on CI for every pull-requests!
this seems like such a bad idea
EchoVoicy 23 hours ago [-]
Depends on the prompt I think. If it's just "Generate tests plz" then I agree, but if its
"If this PR adds any new endpoints, ensure that there are functional and integration tests. If there are not, please investigate the feasibility and appropriateness, and create functional tests using the guide found on our wiki for guidance https://www.ourdevwiki.site/how-to-make-functional-tests" then maybe it could add some value.
But that very much depends on the specific system. Some tests are obvious, some not so much.
tcp_handshaker 23 hours ago [-]
And software keeps getting worst.
The analogy I like is that building software is running a Michelin restaurant. The moment you scale, the chef is just writing cooking books and is absent, and you move into franchising, you will be amazed at the bottom line revenue scaling, while customers will be progressively appalled with the food...
palata 18 hours ago [-]
Not that I disagree, but the average software before AI was more like a McDonald's. I have genuinely seen companies who were writing code a lot worse than what AIs produce nowadays. Doesn't necessarily mean that their software is better now, but my point is that before AI, I don't think that software compared to Michelin chefs.
22 hours ago [-]
jmyeet 1 days ago [-]
> I'm thinking about having it automatically filter and re-rank my social media feeds so I can steer the algorithm instead of the other way around.
I hadn't really thought about this but AI may well be the technology that disrupts and ultimately destroys social media.
The value proposition of something like FB or IG is, as we know, the network effect. The platform gets to extract value from user generated content. I believe that users should own the platform, a bit like the Wikimedia Foundation, because they're the ones that create value. Federation is a popular belief on HN and I've come to believe that's simply the wrong solution to the right problem.
Anyway, how these social media companies make money is by optimizing the feed for engagement. People know it too so you see people trying to build an audience by rage baiting. And then more time spent equals more advertising revenue.
But what happens when the AI can simply slurp all the posts and then filter and rank them? It destroys the engagement and advertising model. And I'm not opposed to that, honestly. It may be on eof the few good thing sto come out of AI.
1 days ago [-]
catigula 1 days ago [-]
>Test coverage too low? Auto generate tests on CI for every pull-requests!
Terrible use-case.
wrobelda 24 hours ago [-]
Terrible comment.
_s_a_m_ 24 hours ago [-]
These posts have to be Chinese bots, these models are all trash. Used it via OpenCode for an hour, cost me one hour of my life. It is for anything complete trash.
r14c 23 hours ago [-]
I've gotten a lot of good work done with deepseek models. Like with any generic harness there's some tuning that has to happen. I used open code for a while, but I've landed on pi.dev as my go to since its easier to tune and has better deepseek integration. iirc open code is quite bad at utilizing cache and doesn't have a lot of ways to specifically tune the harness for a particular model.
apitman 23 hours ago [-]
I've found it to be pretty good so far.
greenavocado 24 hours ago [-]
(1) you used opencode
(2) what provider did you use. openrouter is trash because they shit up the model serving. no max effort and horrific cache utilization, on the order of 50-75%, absolutely garbage. beware
alex0015 24 hours ago [-]
What should we be running deepseek on besides opencode? I chose it because I heard good things. Also provider is directly through deepseek credits.
stavros 20 hours ago [-]
I use the Deepseek API and pay peanuts. Very satisfied.
greenavocado 23 hours ago [-]
oh you used opencode go?
harness: omp.sh
stiltzkin 19 hours ago [-]
[dead]
dan_q 24 hours ago [-]
You're mad.
EchoVoicy 23 hours ago [-]
Point 1 finger out, and you point 4 back.
ricketycricket 16 hours ago [-]
Said the man with 6 fingers.
NoboruWataya 23 hours ago [-]
My Claude account was banned the other day. The only possible cause I can think of is that I tried to authenticate from the AI assistant in a JetBrains IDE and, not thinking, entered the details for my regular subscription rather than an API account. As soon as it became apparent that I needed an API account rather than a subscription, I just closed out of the tab. Nevertheless, about 20 minutes later I got an email saying my account was banned for a violation of the usage policy, and my appeal was rejected.
My initial thought was to sign up for ChatGPT, but I had $20 in OpenRouter so I've been trying out DeepSeek V4 Pro with Pi for the last few days and I gotta say, it's good enough for my use case. And even with paying for API usage rather than Claude's subsidised subscription, and with OpenRouter taking their cut, I will probably end up paying significantly less overall. And I really like the flexibility of being able to use whatever minimalist open source harness I want (and being able to switch providers easily, too).
(My demands probably aren't as high as many others' - I mostly use it for help with some hobbyist coding projects, and I tend to ask it questions about how to approach problems rather than just telling it to go off and code stuff for me.)
nodja 23 hours ago [-]
I'm the same way, I have a very low/sporadic usage of any subscription I've tried. I now just use openrouter with DS4 pro/flash. It also gets rid of usage anxiety where I would try to justify the $20/month by forcing myself to use the tokens for projects as the weekly limit deadline neared.
ignoramous 23 hours ago [-]
> My initial thought was to sign up for ChatGPT, but I had $20 in OpenRouter so I've been trying out DeepSeek V4 Pro with Pi for the last few days and I gotta say, it's good enough for my use case
If you prefer subscriptions, OpenCode Go ($10/mo), Cline Pass ($10/mo), Atlas Code ($20/mo), and CommandCode ($1/mo) serve some of the best open weights with generous limits. OpenCode Go currently offers $120 for $10 on DeepSeek Flash v4 (if you're okay with data retention).
> DeepSeek V4 Flash: ZDR agreement is renewed monthly. The current agreement is valid through August 31, 2026.
Is there other info I should be aware of w.r.t data retention with opencode go? It's hosted in China, so other middlemen may be active (I doubt it, but possible)?
WarmWash 21 hours ago [-]
If it's hosted in China, they can tell you whatever you want to hear and do whatever they want to do.
What are you going to do? Take a CCP company in front of a CCP judge?
jacquesm 16 hours ago [-]
They can do the same thing in the US. What are you going to do, sue OpenAI or Anthropic?
WarmWash 14 hours ago [-]
Well, yes, because the courts and the AI labs are totally separate entities.
If they are doing it to you, they are probably doing it to others, which makes an easy class action
filterfish 5 hours ago [-]
Is any class action lawsuit "easy"?
DiogenesKynikos 14 hours ago [-]
You probably could sue them if they really did store your information without your consent.
It would be a hassle, but China does have privacy laws. Companies do get sued for violating them.[0]
China has laws that serve the state, not the individual. They are there for the party to use to prosecute. So you could file, but considering the state is the one responsible for holding the data, and the one that controls the courts, it's not going to go anywhere.
Remember China is still an authoritarian dictatorship. One leader with absolute power for life. Don't let the facade misguide you.
DiogenesKynikos 7 hours ago [-]
The case mentioned in the article I linked to involved a private citizen suing a tech company, based on China's GDPR-like laws. He won.
WarmWash 5 hours ago [-]
Sure, because the state didn't have a stake.
Trying suing when it's in the states interest, like it is to build AI datasets on western data. For reference, no case in China has ever been ruled against the government. They don't have things like judges smacking down executive orders or refunding tariffs.
anon373839 21 hours ago [-]
Yes. There is a big question mark about whether Anomaly itself (as the OpenCode Go middleman) retains data. The docs were completely silent on this the last time I checked.
0xbadcafebee 20 hours ago [-]
According to this code (https://github.com/anomalyco/opencode/blob/dev/packages/cons...) only Grok and Luna have 30-day retention, the rest have 0-day retention. Of their legal documents, none mention data retention for anything other than personal data. The exception is if you '/share' which explicitly gives them your session to share over the web with others.
anon373839 15 hours ago [-]
I haven’t looked at their docs in a little while but they were very precise in saying that the providers follow a zero data retention policy. But Opencode Go proxies the requests, right? I never saw a statement in the legal docs that would limit Anomaly’s right to retain the data passing through.
0xbadcafebee 21 hours ago [-]
OpenCode Go does not send any inference to China unless you go into settings and manually select 'Enable models hosted in China'.
ignoramous 21 hours ago [-]
That's a recent addition I wasn't aware of. Thanks. Previously, OpenCode docs claimed they had no agreement with DeepSeek on data retention for training. Even more previously, OpenCode said they used providers based in US/EU/Singapore (which they no longer do so).
apitman 17 hours ago [-]
> OpenCode Go currently offers $120 for $10 on DeepSeek Flash v4
I've been asking about psyops and bioweapons and I'm still going strong. I did get a Sonnet session shut down the other day though which feels like some kind of achievement.
ak_t 23 hours ago [-]
Note this is the 07/31 release of DSv4 flash and not the "preview" that they put out a couple months or so ago.
I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload.
The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s on a single stream. I saw 1000 tok/s with ~64 concurrent streams on vLLM.
That's fast enough that you can interactively chat with it without switching tabs while you wait, and its a ~300B (13B active, hence the speed) model so the responses are also very good. It's actually more convenient now for me to direct 95%+ of my day to day usage to my local model, and only use Claude Fable for really big coding tasks.
Until this model was released, I was contemplating spending even more money on hardware to run GLM5.2 (~750B) at reasonable speeds, but I no longer feel that need. This is smart enough, and I think it only gets much better for local models from here.
ComputerGuru 22 hours ago [-]
What quantization level is that? Because official endpoints are slow.
ak_t 22 hours ago [-]
It doesn't need extra quantization. The official weights are natively mixed precision FP4/FP8, so it fits in ~160GB. The API slowness is probably from being batched with other concurrent user requests. The provider's aggregate throughput gets higher but per-stream speed slows down.
zargon 21 hours ago [-]
V4 Flash fits entirely in two RTX Pro 6000s without any quantization at all.
bel8 22 hours ago [-]
From opencode go $10/mo plan I get between 60 t/s and 100 token/s even with large contexts of 150k+ tokens.
I wouldn't call 80 t/s slow.
ponyous 21 hours ago [-]
You are right, relatively to other llm providers this is not slow. But if you think what is possible when you have 1000t/s a sec you might find it slow.
hatefulmoron 13 hours ago [-]
That's across 64 concurrent streams; you could make more concurrent requests to DeepSeek API no?
namr2000 21 hours ago [-]
What runtime are you using with the 2x RTX Pro 6000 Blackwell machine? I have the same setup and tried DSv4 Flash on vLLM and ran into a ton of kernel bugs that don't seem to have been fixed yet.
> On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s on a single stream.
For reference, on a 1x B300 it's over 400 tok/s decode on a single stream.
apitman 21 hours ago [-]
I'm getting like 25 tok/s on 2x RTX Pro 6000. This is with llama.cpp, but I had GPT tune it for me. I was under the impression vLLM was at most ~2x faster, and usually for highly parallel loads. Any tips on where I should look first for an obvious blunder?
If the newer builds aren't working, you might try running the old v6 build (based on the eldritch-enlightenment image). gilded-gnosis gave me some problems that I haven't bothered to track down, the old builds are still gonna blow away llama-server performance. And that's before you get hooked on vLLM's PagedAttention and can run multiple sequences without a ton of extra overhead.
mosura 1 days ago [-]
I strongly recommend trying this for programming tasks.
It is strong (not Fable strong though) with a much better “persona” than Opus, and very different blindspots. If you flip between Claude and this you will find both catch the mistakes of the other before they get out of control.
On balance I actually prefer DeepSeek for programming now, because of the way it talks.
chorizo 1 days ago [-]
This also reflects my experience and should put to bed the distillation rumours. This model feels nothing like the Claude models, including tone and blindspots.
_s_a_m_ 24 hours ago [-]
[flagged]
23 hours ago [-]
nylonstrung 22 hours ago [-]
Compared to the last Deepseek V4 Flash version I've had tons of issues with it getting in infinite loops and talking to itself without executing tool calls, wasting tons of tokens
This is on Pi agent, nothing fancy at all about my prompts or use case. Anyone else experiencing this?
I've also had it randomly go from talking about Rust to talking about the electric chair, controversies about D&D rules (both irrelevant and something I've never discussed) and it's completely blind to it in future prompts even when its pointed out and referenced directly
All this said its still worth it but the agentic performance has degraded in my experience at least
natrys 21 hours ago [-]
I think this one requires a bit of strong prompting. I am also normally a Pi user, but my experience in OpenCode with this model has been drastically better than in Pi, where it overthinks a lot and gets distracted by random things.
What quantization are you using? Which infra provider?
Baseten.co's version got into a loop rather rapidly... I've since added loop detection and adjusted some other settings on the pi coding agent and have yet to notice it again. I also switched to DeepInfra ... who serves an fp4 version admittedly, but I've had no issues with it as of yet and it's the top provider on openrouter.ai volume wise.
wgd 13 hours ago [-]
If memory serves the DeepInfra offering is marked as fp4 because that's the native precision of the experts (which are of course the majority of the weights in a MoE model) so they feel that's the more accurate label, while most other providers claim fp8 because the dense layers are natively fp8 and they want to display the bigger number for obvious reasons. They're not actually serving at different precisions, it's just a confusing mess.
hankbond 15 hours ago [-]
using the platform.deepseek API version I haven't had this happen once in my usage (which has been exclusive since its release). Which provider are you using? I also use a pretty bare bones Pi.
bel8 22 hours ago [-]
I've been using for work, from opencode $10/mo subscription plan, on high effort (which is better than max imo), and haven't had any issue.
When it was first available in opencode, it was kinda slow for me, I guess because everyone wanted to try the new shiny. But now it's back to being screamingly fast and Opus 4.8 level of smart, for penies.
the_duke 22 hours ago [-]
Yeah, I saw the same thing - quite annoying. It can be mitigated through the prompt.
Why? It's open weight, there are plenty providers on open router that are serving the latest v4 flash at 0.14/0.28 $.
LorenDB 23 hours ago [-]
Yes, but even the cheapest providers on OpenRouter are charging at least 10x what DeepSeek does for cached input tokens, which is where DeepSeek gets most of the cheapness.
23 hours ago [-]
petesergeant 23 hours ago [-]
(nevermind, I was reading DeepInfra as Deepseek. My bad)
VulgarExigency 23 hours ago [-]
You are missing a 0 to the left of the 2 on Deepseek's number
modeless 23 hours ago [-]
This would be more convincing if those providers had converged on a number that was not the exact pricing of DeepSeek themselves. Clearly DeepSeek is setting the price here and without them holding it down I expect increases.
dghlsakjg 6 hours ago [-]
Before this release, there were providers undercutting deepseek on the preview version.
I suspect that once the hype dies down, or the field gets more competitive we will see the same on 0731.
vb-8448 22 hours ago [-]
Non necessarily, they can easily increase market share by staying where they are.
nicce 23 hours ago [-]
The most expensive defines the price. Others need to be just slightly cheaper.
542458 1 days ago [-]
Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.
MarkLowenstein 24 hours ago [-]
Real question: is there anybody that is both maintaining alpha-dev capability by keeping abreast of all these daily changes, while also reserving enough time to actually work?
Seems like we've reached the event horizon of whether AI advances are worth paying attention to.
becquerel 23 hours ago [-]
I think the play now is to just try out whatever the best new model is every time you see a headline that fundamentally reorganizes your conception of what's possible.
naught0 22 hours ago [-]
I enjoy using opencode go to play around with a lot of different models. I wind up using deepseek v4 flash for most everything, stepping up to minimax m3 if that doesn't cut it, finally preferring GLM for complex tasks or important planning I want to go right the first time
I recommend opencode or something akin to it to play with models. Any big model updates or hot new ones will naturally run across your desk that way
How about: The yaks have started shaving themselves, who can keep track of how good a job they are doing?
speedgoose 13 hours ago [-]
Alpha dev?
petesergeant 23 hours ago [-]
I don't think you need to be keeping abreast of them really, you just need to be using the best model you can get enough tokens from, which for many people is Fable 5 @ $200ish, ideally fanning out implementation to cheaper models
bronson 8 hours ago [-]
> which for many people is Fable 5
Not for me, Fable refuses to debug Linux kernel bugs. Unless you say who you're speaking for, it sounds like you're just shilling for Anthropic.
whinvik 1 days ago [-]
Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.
fallingbananna 24 hours ago [-]
GPT 5.6 Luna is an extremely cheap and still very capable model.
A chinese model being in the same ballpark of capability at half the price sounds believable to me.
nwienert 14 hours ago [-]
It's significantly worse than Luna and quite a bit slower in some fairly involved tests I run.
jfaat 11 hours ago [-]
That's fascinating, it's WAY better than luna ime. What sort of things are you testing it for?
debazel 9 hours ago [-]
I've been using this DeepSeek model the whole day today after building with 5.6 Luna extensively over the last week and I would disagree, at least for Rust + OpenGL.
DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it.
I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot.
Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations.
(Both were set to high reasoning)
ignoramous 1 days ago [-]
In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.
nwienert 1 days ago [-]
Yep, and the v4 flash final is about 2.5x slower than preview making it no longer a fast model, in fact slower than Luna and bigger models in many cases.
Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.
thehamkercat 1 days ago [-]
And now nobody seems interested in it because the price hasn't gone down
I believe it's because they are below $20 Million revenue limit (which Kimi K3's license has)
So we won't see any price decrease unless Kimi changes the license of K3
dyauspitr 24 hours ago [-]
Not for long, Deepseek is saying they will have a significant price jump soon. They really shouldn’t do it because they are on the cusp of capturing the scalable API market.
telotortium 23 hours ago [-]
They need to be able to serve their market. The price increase is partly load shedding. If they improve their ability to serve their load, they can always drop it again, as OpenAI did with Luna recently.
ignoramous 23 hours ago [-]
> as OpenAI did with Luna recently
My read is, OpenAI is neither able to claw b2b money (away from Ant) nor are they able to stave off open weights on the other. In short, they're struggling to hold onto their distant #2 position in the coding market, and these pricing changes reflect a (desperate) change in strategy.
lukewarm707 22 hours ago [-]
and i still won't use it, because they log and spy on your prompts XD.
the private endpoint costs 10x (azure).
private endpoints for deepseek (lots of providers) also cost about 10x more.
but 10x more for deepseek is $0.028 cached input, and 10x more for luna is $0.10.
andai 23 hours ago [-]
The recently announced they're raising their prices 10x right?
Which would put them... exactly where everyone else is on this graph.
Edit: I seem to have misunderstood the news. I thought the magical cache read pricing was going away (0.002) and they were going to be on par with everyone else (0.02). But I have no idea.
Edit 2: Apparently, neither do they!
>We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
guilamu 23 hours ago [-]
Where does this "10x" comes from?
freakynit 15 hours ago [-]
Every "other" provider other than deepseek on openrouter (or elsewhere) is charging at the very least 10x for cached input tokens... and in agentic coding, that is where the 90%+ of the costs from from... that means with these providers, your costs will jump by 10x or more.
> The recently announced they're raising their prices 10x right?
No.
They sent an email to customers saying that they will raise prices "significantly".
How much that will be is speculation.
My guess is that they will just remove the 75% discount they gave when they released V4 preview. It will still be relatively cheap even at 4x the current price.
gizmodo59 20 hours ago [-]
It will be comparable to Luna then.
vijucat 4 hours ago [-]
Caching makes a huge difference to cost. On Fireworks AI, for example, if it hits the cache, you pay only 20%. And uncached is just $0.14/M tokens for DSV4-0731! I get entire re-architecture projects (with new tests and documentation) done for mere dollars. DSV4-0731 is a daily driver for me.
But note that you have to use Cline (or other harness) if using vscode. I was shocked at how poor the recent versions of GitHub Copilot are at using the cache (with Fireworks AI, but I believe it's a more generic problem).
It's really amazing to see how the gaps between the self hostable models and the closed models has been shrinking in the last 24 months.
And how this has been accelerating!!
I felt this very hard when I had to travel in the middle of nowhere in south america, with no network, and wanted to keep an LLM model on my macbook pro with 48GB of RAM. That was back in April 2026, a few months ago.
I downloaded Google Gemma 4 (google/gemma-4-26b-a4b) and - Oh boy - I was amazed by it's capacity!
I was able to use it to code simple things, ask it about nature, learn new stuff while traveling and make stories for the kids.
Was really amazing to observe and experiment this!
Seems to me there will be some good chance to run these great LLM locally on our hardware!
Amazing time to be alive
UberFly 11 hours ago [-]
I just don't know... This post sounds like an Ai bot.
bonoboTP 9 hours ago [-]
Not at all.
taf2 4 hours ago [-]
we got this running on 4 RTX Pro 6000's and for single request we're getting around 250 tok/s we can support about 48 concurrent requests we're seeing around 2400 agg tok/s peaking around 24-31 concurrent users. Model performance feels like gpt 5.4 - mostly using it with pi agent. the only thing i'm missing with this model is vision and i see some folks have done some work like https://huggingface.co/webbrain-one/DeepSeek-V4-Flash-0731-V... but have not yet tried it out.
zacksiri 16 hours ago [-]
I'm not sure about all these benchmarks, I did some very simple tests (I have my own benchmarks https://upmaru.com/llm-tests) and these models fail, not sure if it's the inference provider or the model. They seem to be optimized for benchmarks more than real use cases. Do anything outside their distribution (even if it's not complex) they fail.
I Compared Deepseek V4 Flash 0731 (low) to Gemini 3.5 Flash Lite (minimal) and GPT 5.6 Luna (no reasoning) and Deepseek V4 Flash 0731 gets it wrong alot, where as Gemini and 5.6 Luna just gets it done.
theogravity 16 hours ago [-]
You're using it on low, that's why. There's a huge difference in performance from low to max effort.
zacksiri 15 hours ago [-]
I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test.
Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.
Palmik 9 hours ago [-]
Low, High and Max, obviously, can't be compared across models. They only mean the model is likely to spend less reasoning effort (~output tokens) with Low than High on the same, *single shot* task.
But even in this very post, you can see that Max was actually cheaper than High.
If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.
kgeist 11 hours ago [-]
I'm not sure it's a fair test either to compare the "low" setting of one model with the "low" setting of another. They're completely different settings that just happen to have the same name.
lejalv 11 hours ago [-]
Comparing at similar thinking hasn't much value. You can compare the tiers that have the closest price, that would be more interesting.
zacksiri 11 hours ago [-]
Yes, I think a proper comprehensive test would be a better judge of the outcome. I may do round 2 given my first batch of models is already outdated.
theogravity 12 hours ago [-]
I did not downvote you. I think it was the way you wrote the comment which made it seem like it couldn't perform in general.
It's not frontier, but it's far past what we had at the beginning of the year. It's very usable. I get great instruction compliance, tool calling, and with a trivial workflows flow it has very good long-running performance as well.
Kuyawa 5 hours ago [-]
I use DeepSeek on a daily basis and I didn't spend $10 in the whole month of July delivering eight fully functional apps.
Btw if you need an app I may deliver it to you in ten minutes for just five cents if I'm in the mood. Just let me know.
paulddraper 5 hours ago [-]
I assume you got the email about them increasing prices "significantly" soon (but no hint on what significantly means).
Kuyawa 5 hours ago [-]
Yep I got the email, but I am not worried. Building an app for 5 cents will cost me what 10 cents? A full dollar? Still a bargain. Still giving me two months and 29 days free for something that took me three months to do
stuaxo 18 hours ago [-]
Seeing everyone spend like 200USD a month seems kind of mad.
I have £20/month Gemini and £20 a month claude for a bunch of personal projects.
Yes I have to wait sometimes, it's probably a good thing.
Petersipoi 17 hours ago [-]
I max out my $200/month Claude plan. You are obviously just not taking advantage of it to the same level as others. Which is fine. Don't pay for something you don't need. But I would definitely take a massive productivity hit if I had 1/20th the usage.
Mkengin 12 hours ago [-]
Are you also talking about personal projects? That wouldn't be enough for me for work either, but for personal use, my $20 Codex subscription is perfectly fine. Then again, as a new dad, I can maybe only work on personal projects for 1 hour in the evening if things go well, so for my current situation in life, it's enough.
kilroy123 6 hours ago [-]
Same here and my codex plan. I have to wait until next week to get back to work.
therealdrag0 18 hours ago [-]
Just depends on commitment, time, and scope. Personally I also use the 20$ plans. But people spend 200$ for a gym membership. So shrug.
scrollop 12 hours ago [-]
Gemini? Are you locked into google as that's an odd choice.
system2 17 hours ago [-]
People use Claude Code for professional reasons. Waiting is not a luxury when there is a deadline.
momojo 5 hours ago [-]
Did anyone else experience a change in verbosity? I've been playing around with an agent that holds your hand in a Jupiter notebook and it felt like it started writing essays versus nice, concise, helpful paragraphs like before. My gut was correct because I checked my Deepinfra usage and it was almost a 2x out-token usage for every in-token.
Not a huge deal since it's still cents per session, but my bigger issue was the weird change in tone. It became a lot more pretentious and over-explanatory.
Heavy prompt reworking helped but maybe that's just the cost of being better at coding and ARC-AGI?
walrus01 24 hours ago [-]
Oke of the great advantages of v4 flash 0731 is that even in the largest size unsloth quantized gguf, Q8 K XL, it will fit well within the resources of a 256GB DRAM server. If you have no gpu at all and are okay with setting up a workflow that handles slow token per second rate, give it a task and check back in 4-6 hours, it works great. And remember to give it more lengthy tasks to run overnight. Whatever workflow you set up, the idea is to keep it busy 24x7 doing different things in parallel.
saturn5k 10 hours ago [-]
For the last 3 months I've been using V4 Flash Free with Hermes through Opencode Zen both personally and at my company and I've been having a great experience so far. It's my go-to model for terminal work, managing my entire ubuntu server, Cloudpanel, managing static websites, doing SEO audits, network tests, DNS troubleshooting, e-mail deliverability troubleshooting...
Furthermore, in my company we are using MCPs for Google Ads (it manages our ads), Analytics, Search Console, Zoho CRM, Microsoft Clarity... We use it to crawl specific websites and send daily summaries to our sales team in MS Teams channel. We use it to send daily summaries on marketing statistics and analytics... All with a FREE model. We are rarely hitting any limits so far and in case we need more tokens - we use NOUS or openrouter to pick between Flash or Pro for specific tasks that require more churning.
AMA.
djfjfndmwpekdm 10 hours ago [-]
Sounds super simple stuff compared to what I use Fable and Opus 5 for.
LeBit 18 hours ago [-]
I always find it confusing that a meaningful volume of the comments are saying "this reached parity with SOTA models. Best $/task."
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
freakynit 15 hours ago [-]
I was in the first group up until last week.. now, the second one.
For anything even moderately complex.. like, even low end of complexity, this model behaves maximum like gpt-5.6-luna-high .. nothing more.
Yesterday itself I gave it a coding task in some existing moderately complex small project, and i was using xhigh thinking effort, it was unable to cover all edge cases... and i had already got it to review, and then fix, 3 more times, after the first initial one.
Still it left 2 edge cases.
Then, reverted full code, gave sol-high the same task, it took well over 20 minutes, and completed it in one go with zero edge cases remaining.
I am not using it for anything serious anymore.
thebytefairy 6 hours ago [-]
Definitely interesting. My guess it has to do with the level of hand-holding the users are doing. Ie. Full on vibing vs pair programming.
ddxv 16 hours ago [-]
I guess both are true and for everyone at some point. All models, even SOTA, fail. When they fail, it is quite frustrating. Additionally, some models are very cheap to run and use. When Deepseek fails the cost was minutes and pennies.
wolttam 14 hours ago [-]
Yours is the only mention of GPT-OSS in the whole thread. I count about 2-3 comments saying the model is meh and many more saying it’s a big step up.
I am among those with real life experience with the model that used the previous as well and will attest that the new model is a big improvement
surprisetalk 1 days ago [-]
This reminds me of those pareto-style speedrun record charts when a new glitch is discovered.
When I see dramatic leaps like this, it tells me that the important hacks haven't yet been discovered.
SwellJoe 1 days ago [-]
DeepSeek is my cheap and cheerful Chinese model of choice for API use. Has been for a while, but now it's Flash instead of Pro. Even cheaper, and now better then Pro. I feel like most of the major Chinese models are benchmaxxed, they have weird quirks every time I use them (Qwen 3.8 Max doesn't check its work and leaves stuff broken, doesn't write tests unless prompted, etc., Kimi ends up being quite expensive and rarely better than GPT Sol or Opus 5), while DeepSeek models seem to be generally as good as the benchmarks indicate: Not the best, but stronger across the board than any model within an order of magnitude of its price.
eli 24 hours ago [-]
Qwen 3.8 Max is very strong at troubleshooting and code review.
SwellJoe 24 hours ago [-]
I'll grant it's very thorough when assigned a troubleshooting task. I'm not as confident of its code review though it is very good at security vulnerability auditing, and isn't hobbled for that work like Fable, and even Opus refuses some work in that area now.
kromem 22 hours ago [-]
Flash is a delightful model and the start of intelligence at effectively insignificant cost.
From here on, it's going to become all about harnesses that best situate and organize swarm intelligence at scale.
apitman 21 hours ago [-]
These are very interesting results, and honestly hard to believe, even as a big 0731 fan.
If I'm reading the chart correctly, a couple observations:
* deepseek-v4-flash-0731 max is better than kimi-k3 max
* glm-5.2 is dumber than a box of rocks (this must be on low reasoning or something, right?)
This is way more extreme than other results I'm seeing, like those from Artificial Analysis.
antonyragleap 13 hours ago [-]
The interesting part for me is whether DeepSeek V4's reasoning gains hold up on long-running tasks rather than short benchmarks.
dools 21 hours ago [-]
I have been using deepseek v4 pro almost exclusively. I was using Kimi a lot but it just nose dived. The decline started with the release of 2.7 and accelerated with the release of 3.
When I need vision capabilities I use GPT 5.3 codex and if deepseek can’t figure something out after a few goes I switch to GTP 5.5 or 5.6 (I’ve been giving Terra first bite recently and it does pretty well, and have used Sol a couple of times).
Using this regimen means I spend under $100 per month on inference and I work all day everyday with multiple agents running simultaneously all on API token spend not subscriptions.
sourcecodeplz 1 days ago [-]
wow. i remember when GPT-5.2 (medium) was everyone's favorite.
ARC-AGI II:
- GPT-5.2 (medium) %26.7 ($0.759)
- DSV4-Flash (max) %61.4 ($0.04)
xyzsparetimexyz 1 days ago [-]
That page needs a Pareto frontier display. But wow, it absolutely demolishes.
jacquesm 17 hours ago [-]
This is the best model to come out since the beginning of open weights models for those working with classified data that you can not use hosted services for. I've been using it pretty much day and night since it landed and I'm nothing short of amazed. You'll need some pretty good hardware to run it though.
zmmmmm 21 hours ago [-]
it's great but we need a multi-modal model of this quality and price to truly declare victory.
But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?
ghosty141 21 hours ago [-]
xiaomi mimo is very cheap and not bad.
minimaxir 1 days ago [-]
It's always fun when Max reasoning is cheaper than High reasoning.
Terretta 1 days ago [-]
Rework is expensive.
Tell your PjM who should tell your PgM who should tell your PdM, all the PMs...
Maybe if "the business" sees it is true of LLMs, they might believe it's true of giving better context to engineers up front then giving them time to think and prototype (thinking tokens are an answer prototype).
gentlewater 1 days ago [-]
I’ve been refreshing hacker news constantly for a week now waiting for v4 pro, after they stated it would follow «soon». I have learnt «soon» is a matter of definition.
indigodaddy 1 days ago [-]
I guess you mean a "new" v4 pro?
ignoramous 1 days ago [-]
> been refreshing hacker news constantly for a week now waiting for v4 pro
Looking at the caching price of deepseek compared to its competitors, does it have a secret sauce or is it just subsidizing?
ignoramous 22 hours ago [-]
That's DeepSeek's way of selling "token plans", yes. But without the downsides like daily or weekly limits and guaranteed upfront/fixed spend.
g023 17 hours ago [-]
It says 'yes' where the others say 'no'. Good enough for me.
blueflowerbed 8 hours ago [-]
Perhaps I am doing it wrong, but the deepseek llms fail so hard for any complex task and don't compare to the paid models. It keeps forgetting basic instructions after two responses
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
nmitchko 9 hours ago [-]
It’s an adaptation of CoLaR, but my implementation is a little different:
- Dedicated stop head to fire when latent thinking hits threshold
- MTP support with training taking draft support as first class
- different architectural layer 35 -> layer 42 writeback. So latents skip roughly 6.2 tokens of reasoning per token, then never make it to decoded output
24 hours ago [-]
Almondsetat 20 hours ago [-]
It's crazy to think V4 Pro still hasn't finished the post processing.
seanmcdirmid 21 hours ago [-]
Note they double the price if you use during peak time. However, they define peak time with respect to China, not Europe or the USA...so if you are out of Asia, I guess Australians might be impacted, and its still cheap anyways.
tarruda 20 hours ago [-]
One of the best things about this version is that it is trained in the codex harness. It feels just as good as OpenAI models in using codex tools, but extremely cheap and with 1M context
djc404 20 hours ago [-]
Do you have any sense how using it with codex compares to OpenCode?
It’s always a bit tricky picking the right harness (when you have options). Sometimes the differences are subtle but meaningful. But who has the time to run everything twice and compare all the time!
tarruda 18 hours ago [-]
I don't have experience with opencode, so I couldn't tell you.
Codex is really good in my experience, especially due to its native sandboxing. Deepseek seems really well versed in its tools, including update_plan and knowing when to request sandbox escalation.
hnussffjc9 7 hours ago [-]
Been burned, can confirm every word
tosh 1 days ago [-]
results comparable to gpt 5.6 luna but cheaper
promising!
minimaxir 1 days ago [-]
Since the x-axis is log-scaled, DeepSeek is much cheaper than visually implied (mousing over the raw values, it's 1/4th the cost of Luna).
literallyroy 1 days ago [-]
Is this pricing from Deepseek with training on usage?
Is it still cheaper than Luna if using an OpenAI subscription? My gut is no, but I have not done the math.
K0IN 2 hours ago [-]
Some china proxies charge as little as 1 cent / Mio tokens for luna, which I use. As far as I know they bundle multiple codex subscriptions to get subsidized token pricings, so in codex it should be.
swiftcoder 1 days ago [-]
You'd have to compare against something like the OpenCode Go subscription, and I'm fairly sure deepseek napkins out cheaper in that scenario
LUmBULtERA 1 days ago [-]
I'm still not sure, there's a promo going on now, but generally Go gives $60 of API credit and right now it might be $120 with deepseek. But $20/month OpenAI subscription I believe gives you many hundreds of API-equivalent usage? I've heard $100/month giving many thousands API-equivalent per month.
minimaxir 1 days ago [-]
Everything is cheaper if using a subscription, but some applications require API usage.
jrflo 24 hours ago [-]
Might not actually be that much cheaper, we don't know what margin OpenAI is charging on Luna API. Open models likely have much less margin.
w2seraph 21 hours ago [-]
There's no reason that LLMs should cost beyond grave digging sums when this one topples the charts it'll be over.
harisamin 24 hours ago [-]
I'm curious... is anyone using DeepSeek V4 Flash from HugginFace? Is the cost around the same as directly form DeepSeek or from Openrouter?
evanjrowley 23 hours ago [-]
The benchmark performance tells me DeepSeek v4 Flash could be very cost-effective at playing SNES/Gameboy games.
andai 23 hours ago [-]
It won't be long before I can just stay home, and have my robot ride my bike for me.
qup 16 hours ago [-]
I'm going to send mine to visit my mom. It's so hot in August.
andai 6 hours ago [-]
You reminded me of a short story published over a century ago.
ive been enjoying this deepseek v4 flash 0731. i think its a great model, it helped me with finishing all of my abandoned projects.
simonw 24 hours ago [-]
That's a pretty great score for a model you can run on as (expensive) laptop.
CrosswordPuzzle 24 hours ago [-]
I'm really excited for where the open weight models go from here. I've had fun with just CPU inference on old servers that only have AVX1; here's hoping for commoditized TPU-like hardware!
nikp123 24 hours ago [-]
I just used it for some Kubernetes + FluxCD tasks and oh my is it good.
clayhacks 1 days ago [-]
Why wasn’t this run against ARC-AGI-3? Or did it fail to solve anything?
chorizo 1 days ago [-]
They tweeted that ARC-AGI-3 results take longer to run, so we’ll need wait a bit longer.
mycall 23 hours ago [-]
I'm curious how much worse the 0731 quantizations do.
johnmlussier 23 hours ago [-]
Been running it using Prime Agent and absolutely love it.
Havoc 1 days ago [-]
They did recently announce they're increasing prices though (got a mail yesterday I think), so not sure this analysis showing it as price outlier will last
minimaxir 1 days ago [-]
That is only when using the DeepSeek API directly. OpenRouter has 24 different providers serving it at existing prices.
This latest DeepSeek is almost at the "too cheap to meter" level. That's going to be a larger unlock than models like Fable/Mythos that are way too expensive to justify, IMO.
What secret sauce do they have?
pama 23 hours ago [-]
No secrets—all published. Very efficient attention. Excellent kernels. Great caching subsystem. Small and well trained model.
est 18 hours ago [-]
> What secret sauce do they have?
Quant company usually squeezing every penny.
throwaway_95283 1 days ago [-]
limited resources, no modern GPUs, no $10 billion dev budgets.
pair it with codewhale, 50 agents, 200 MB of ram.
iagooar 1 days ago [-]
I love DeepSeek V4 Flash since the pre-0731, now even more. It is the first model that is truly too cheap to meter.
But I find it having a pretty significant problem with tool calling - no idea why, but tool calling with it is SLOW. As long as the model is reasoning, all good. But give it a bunch of tools and it becomes extremely slow.
Am I the only one experiencing this?
JunkiesCoder 11 hours ago [-]
Didn’t deepseek recently announce prices will go up significantly?
18 hours ago [-]
m3kw9 19 hours ago [-]
The token price seem to be jigged, how do you know if it's subsidized or temporary. Anyone can just lower the token price to get to the left.
jacquesm 16 hours ago [-]
Almost every service provider in the AI field is subsidizing their token cost to some degree, they're all shooting for marketshare and lock-in (and they're not really achieving the latter).
system2 20 hours ago [-]
How can OpenAI or Anthropic fight against these prices?! $0.14 input, $0.28 output. For 1M tokens...
gxs 20 hours ago [-]
One thing that popped into my head is that this shows how committed they are to building something that scales across the world
China has zero energy concerns in terms of energy production - not literally zero, but they’d be able to prioritize other dimensions and not necessarily worry about efficiency
Here they are though releasing models that sip resources
casey2 22 hours ago [-]
Finally something that is breaking away from the pack. Interesting that max costs less than high. I still think, currently, TPS is more important than near frontier intelligence. Likely for reasons that LeCun outlined, maybe out of a billion prompts you will get value from that intelligence. When we have very fast models abstraction will work as that filter.
esafak 1 days ago [-]
It's serviceable but, like many Chinese models, it uses a lot of tokens to get work done.
gruez 1 days ago [-]
>it uses a lot of tokens to get work done.
That's irrelevant when you use $/task as the metric, which the OP does use.
esafak 1 days ago [-]
It also affects the time.
aitchnyu 24 hours ago [-]
It felt like a rocket compared to GLM 5.2 though. Are Chinese models generally token-heavy?
If I had the GPU size, hook it up to llama.cpp and setup the --reasoning-budget and reasoning-message; Most of that additional reasoning is a lot of garbage and you can redirect it to useful output.
That's how I handle the Qwen27B and 35B
nomel 22 hours ago [-]
> Most of that additional reasoning is a lot of garbage and you can redirect it to useful output.
What do you mean by "redirect it to useful output"? Could you give an example? This sounds interesting.
cyanydeez 20 hours ago [-]
It's specific to the harness. Using dynamic context pruning, the budget cuts it off after a select amount of tokens and the budget message tells the model to use subgents to finish whatever it's thinking about
nomel 19 hours ago [-]
Nice. Does it use a summarization, or a hard cutoff?
ttkciar 18 hours ago [-]
llama.cpp uses a hard cutoff. The agent then does "something" that is specific to the agent's implementation and configuration. It might summarize and then "finish the thought" with a different model, and then resubmit the prompt to the llama.cpp API endpoint with <think>..</think> prefilled. The primary model then infers the remainder of the reply.
cyanydeez 9 hours ago [-]
llama.cpp does a hard cut off on budget; it can set a reasoning-message as default but the client _can_ set a per message reasoning-message, so it's possible a smart harness could inspect the cut of thoughts and trim and do whatever.
leizhou 24 hours ago [-]
so cool. does it mean it can understand the verificated code
WhitneyLand 1 days ago [-]
The DeepSeek team is so strong, very impressive.
Imagine if they had GPU resources of western labs.
mosura 1 days ago [-]
Necessity is the mother of invention.
SV companies get way too comfortable when they have enough in the bank to stay running more than three months.
artursapek 18 hours ago [-]
These prices are not real. They already said so.
addozhang 17 hours ago [-]
what's the real you mean?
Numeros 6 hours ago [-]
[flagged]
ClipBGNET 17 hours ago [-]
[flagged]
floki165 8 hours ago [-]
[dead]
hna8hjbqzy 4 hours ago [-]
[dead]
Helldez 12 hours ago [-]
[dead]
22 hours ago [-]
zacurryyyy 9 hours ago [-]
[dead]
hnc3yfnu6f 23 hours ago [-]
[flagged]
antirez 1 days ago [-]
Price is not a good meter. Active parameters per token are. Joule would be even better.
orbital-decay 1 days ago [-]
It's an excellent metric, the amount of applications not viable now due to cost/latency/throughput is vastly bigger than the amount of current use cases. Even current ones do benefit, e.g. it's a great executor subagent.
Energy and intelligence are good too, sure.
fallingbananna 23 hours ago [-]
What if we used 100% of the brain all the time?
As an end consumer, I don't care about the number of active parameters. I really do care only about the tracked metric (how well does it do the job, and how much does it cost... ideally also with time included, but that wouldn't fit on a 2D chart)
minimaxir 1 days ago [-]
Price accounts for computational/architectural efficiency improvements whereas active parameters does not.
polytely 1 days ago [-]
for someone with a limited budget it is actually very important because it makes me less scared to experiment.
muricula 1 days ago [-]
Price is confounded by VC subsidies, economies of scale, and inference optimizations. I think a more interesting chart would be ARC AGI vs forwards pass flops or ARC AGI vs training tokens. Of course we don't have those numbers for the closed source models or even some of the open weight ones.
With the exception of cache costs, all providers have similar input/output costs.
kennywinker 1 days ago [-]
Not counting the cost of making the model, which is subsidized by… someone? The chinese gov i think?
segfault99 19 hours ago [-]
DS comes out (one of, or) the most successful quant fund in China.
They don't strictly need any kind of subsidies.
FWIW they have a funding round planned (kerfuffle about leaks from CEO presentation few weeks back) -- presumably because infrastructure needs have ballooned.
Naturally there will be some PRC government interest in one of their flagship AI companies. From what is visible seems to be more along the lines of ensuring that DS gets its fair share of resources -- e.g. Xi Jinping meeting founder and positive comments about success of DS means that (hypothetically) Alibaba can't screw DS too much on infra charges to kill off a 'competitor'. Also would imagine that DS's top guys have been clearly identified and will have been 'discouraged' from going to work for one of the SV polycules. But even here as much carrot as stick -- none of the DS top guys will ever need to work again except for love of the job.
kennywinker 13 hours ago [-]
> DS comes out (one of, or) the most successful quant fund in China
> They don't strictly need any kind of subsidies.
You understand how these two sentences directly contradict each-other, yeah? The money-losing operating of training a model is paid for by momey earned from prior investments. So… the work is “subsidized” by its parent company’s investments in it.
1 days ago [-]
_aavaa_ 1 days ago [-]
Subsidized by inference profits and volume.
kennywinker 13 hours ago [-]
Is deepseek actually turning enough of a profit off inference to fully pay for training the next model? And do those profits depend on releasing model weights somehow?
npn 1 days ago [-]
weak argument. deepseek v4 flash is open weight, you can easily find other providers with competitive price with Deepseek (except for input caching), some even half as cheap.
OpenCode Go even has double limits temporarily so for 10 USD you effectively get 140 USD of tokens to spend. It would impress me if someone could burn that amount with "normal" usage. Even when running multiple sessions.
I have a Claude Max subscription but I've barely touched it, it just feels like a step back to have to think about limits and usage even though the models are stronger.
The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues. Test coverage too low? Auto generate tests on CI for every pull-requests! Monitoring server logs, continuous security audits and investigating every received exception now becomes possible.
I'm thinking about having it automatically filter and re-rank my social media feeds so I can steer the algorithm instead of the other way around.
Perhaps other people (with enormous budgets) were already doing all of the above but for us this is a really exciting release!
There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.
For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...)
The real unlock will be, and you can already see it with GPT-5.6 and Fable-5, to delegate complex enough tasks that will take more than 24 hours to get done and they will not lose track. I'm not talking about a loop, but the actual intelligence to recover from these compounding errors that accumulate in dumber models.
We're still a long way from the intelligence needed to let one of these agents go ahead and supervise multiple layers of sub-agents underneath to do complex orchestration. The future looks very promising and exciting. Imagine having the possibility of a Frontier model orchestrating as many sub-agents as needed that are running on cheaper models like DeepSeek.
You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?
I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides.
If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entire $100 right there.
Hell, you could argue that the best "coding model" that we have at the moment is the human brain, and people will gladly pay $10,000/mo for one of them.
Arguing over $20 vs $100 for something that actually puts in work just seems insane to me.
Which was an argument for using every less powerful model since the moment they got useful, right?
When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough back then, DeepSeek is better than that now.
Sure, it's always nicer to have a slightly better model. But the price difference starts mattering a lot more when all the models are already sufficiently good.
We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's expected to find another few grand per month in misallocation/inefficiency.
I wouldn't be surprised if we were doing more like $10k/mo higher in 6-9 months' time.
When you're talking about numbers like this, the fact that one AI is $100/mo and another is $10/mo or $40/mo doesn't matter. They could make GLM-5.2, or any other Opus 4.5-class model free and it still wouldn't make sense to deploy in a commercial context.
The other angle I'd approach things from is that Opus 4.5 (and I'd agree with you that that model was the saddle point) was "good enough" for the types of things we were asking it to do back then, but as the models have become more capable the tasks we're asking them to do have also expanded with it.
I know I've personally gone from "hey can fix this race condition with a Redis mutex" 6 months ago to "Independently redesign this full embedded USB stack and QA it end-to-end, working around a specific Kernel bug in macOS Tahoe that requires decompilation to find the source of, while keeping in mind the constraints of our 8-bit AVR chip from 2011" now.
But that said, yes, maybe in 5 years' time we will reach an "intelligence saturation" where the average person won't be able to even conceive of how to use the new SOTA.
Are you saying we shouldn't care about the future of affordability and access because at this moment we have seemingly endless access?
Sounds extremely short sighted.
Surely they are using read-only access.
This is such a ridiculous objection for how beloved it is. Wide swathes of the public can't cope with any adversity or risk.
With an agent (especially incompetently employed), the danger of unwittingly destroying your company (or at least, the crucial data/reputation) is rather higher. We are notoriously bad at estimating the downside risks in complex systems.
The most obvious case is the downside risks in complex financial constructs... things look great for a while ... until a sudden surprising collapse arrives and totally destroys all the upside you think you have created.
Fable 5 is still going to mess things up at any sufficient complexity. The advantage of low cost models with "good enough" intelligence is they can recursively correct. Why? Because it is cheap. Proper requirements and tests and subagents take away increasing amounts of work, at a cost that is not prohibitive.
If you are reviewing code manually you might consider Fable 5 a worse option. As it articulates itself with higher confidence and you already know it is capable, you are may be more likely to miss a mistake. You know to be on guard with a junior engineer. Reviewing a senior who suddenly makes some weird stochastic mistake can be a lot harder. It would be like if the smartest human engineer you knew was capable of some random brainfart in the middle of their massive diff. Imo, much harder to deal with.
Of course, we should keep in mind Fable 5 is only expensive today. It will be cheaper in the future. Autonomous, recursive prompting and improvement is the clear end state. Especially for entities that will always have the budget for that at the SOTA frontier.
If $100 Claud Max subscription works for you, then great.
But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic.
For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive.
And unfortunately for the American tech industry, the valuation is based off those enterprise deals, not your $100/month Claude Max subscription.
Right now the US dominates everyone else in actual chips in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.
It's that good. They are far from capacity limited, and even if they were, you can rent a single MI300X from somewhere like Hot Aisle and get more tk/s than you'll be able to use.
That one is dirt cheap at API pricing, I can't imagine quota is going to be a concern on the $200 subscription, which in my opinion easily supports full time use of 5.6 Sol on xhigh.
They also previously said prices will go down significantly once they get a hold of the upcoming Huawei chips (later this year).
Prices are going up just because they can. It can easily come back down. They aren't strained by some IPO / VCs requiring them to 1000x their earnings.
Think about it this way.
Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time. To the lay person this sounds trivial but to a serious business this intelligence gap could represent millions, or billions of dollars.
But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models.
There is a reason why Chinese open weight models are now popular even in American enterprises, because CTOs realize that they are indeed good enough.
The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this is so hard for HN to understand.
And even it isn't "enough". I can very clearly see myself using more advanced agents to move up the abstraction ladder.
For businesses that have actual problems to solve, I see them investing in the frontier for a good bit longer, probably until we have AGI that can replace employees, maybe even a bit after.
This is why I find the "good enough" arguments silly. Like, the usefulness of an AI tops out to you when you can pair program with it? Seriously? You cannot envision ways in which more advanced AI enables you to do more, better? That's bizarre to me. I don't ever see myself running out of problems to solve.
Fable has only been out for a month but somehow everyone is supposed to have moved to a completely different way of working that supposedly only works for Fable and nothing else…
This kind of takes makes me cringe. Why don't you go back to LinkedIn?
I have no idea what you mean by “pair program with an agent”, but Opus have been able of autonomous coding since last November, and with any half-decent harness even local Qwen3.5 was able to do so 6 months ago.
Fable is a stronger model, which means it can solve harder tasks but it's also over-hyped, because only a small fraction of task is hard enough to be Fable-worthy.
This is not appropriate for HN. Please review the guidelines: https://news.ycombinator.com/newsguidelines.html
Fable is the only one that reliably one shots complex changes and makes the right design choices. Everything else requires handholding.
I can let Fable loose on a 12+ hour (for AI) task and it will have performed it flawlessly when I come back the next day. K3 and Opus are not like this.
And no, our harness is not the limiting factor here.
I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough.
I mean, this proves my point. Better models enable you to get more done. With Fable, 80% of the time, I no longer have chat with an agent over the details of a PR. I give it an outcome and it gets done. This means I can work on much more with the limited time I have.
And I don't see this ending. When better models come out that take that from 80% to 99.x%, I will have that better model manage teams of other models and move up the abstraction layer.
If models get even better than that, perhaps I stop reviewing PRs entirely. Maybe normies can start using agents to build real things.
Unless your business doesn't have many problems to solve and isn't in a competitive environment, it will benefit from using the best models.
My point is you don’t need the best model if you just put in QA processes that can be done by models also. And if you don’t have that, the model is probably not going to be good enough.
This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.
On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.
Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.
Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.
My overall thoughts (released over some time):
https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...
https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...
https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.
That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper.
> On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time,
GLM 5.3 bridges this gap.
> I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates.
Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release!
> Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU production still has ways to go and export restrictions definitely limit what they can do.
Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
They're a lot larger e.g. Fable. It is insanely bad. Do you know how much more resources "Western" companies have? Most in China don't have random GPUs to "play with" like every "frontier lab" employee does.
> As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
They're born lucky though. Efficiency is "free". The next generation hardware e.g. Nvidia claims Blackwell -> Rubin is 10x efficiency (verified by Neoclouds apparently).
It's been 2 months since Fable was released to the general, man. Nobody knows what's going on inside of these companies except the people at the coal face.
It's funny to see that Anthopic shills have been saying the exact same thing for the past two years now (and it was OpenAI fans before). It's amazing to see that Claude 3 Sonnet was "great" but now that even Qwen 9B is better than this version of Sonnet DeepSeek V4 is still not good enough despite being stronger than Opus 4.7 was.
> Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time
If you think Fable makes 20 times fewer mistakes than DS4 you're delusional. It doesn't even do 20 fewer mistake than Gemma 4…
This is made brutally obvious by anthropics customer support for people with such accounts.
Same reason it makes sense to assign a team of humans that cost $100k/mo to a product that brings in $5M/mo, rather than one human with 5 Claude Max subs.
The cost is a rounding error.
This is why I quite like Kimi K3 - close to the same performance (definitely like Opus, approaching Fable), noticeably cheaper, generally good enough for me to daily drive. Only problem is that their official provider (on the Vivace plan) feels kinda slow, I'd say close to 2x slower than Opus on Max reasoning on average (probably more relatable than Fable).
You can say this about literally every product we buy. And yet...
For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.
Yes it's much easier to have a smarter model that goes straight to the correct answer first, but it may not be necessary or economical. There's a minimum bar for the model where it understands problems and knows the right step to correct them, and above that newer models give diminishing returns.
That's basically ASI not AGI, if you agree humans are NGI (natural general intelligence) and make mistakes and wrong decisions in solutions all the time. Right steps with some wrong ones is acceptable though for AGI.
At this point, I don't even know if its possible
- You describe the breakdown in terms of time but it's more accurately a function of reasoning complexity.
- You seem to assume that no intermediate evaluation is possible.
- Often it is (e.g. the build breaks or tests start failing), allowing for course correction. There's definitely a cost to that but it can still be cost effective if the accuracy is "good enough" and the price difference significant.
- There are numerous tasks that don't require Fable or GPT5.6 level reasoning to improve efficiency by an order of magnitude.
While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive.
I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).
Another one I did was a printer data stream translator from an obscure format to PostScript/PDF (or just PNGs), complete with cups support, etc so these old apps can easily be hooked up.
Flash is capable now of running long range defined-goal tasks like this.
So, death sentence even to frontier models?
Such as Decision Making. /s
You just can't set a high enough threshold of intellectual effort for critical decisions.
I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook.
It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers.
It would absolutely have been a frontier model last December.
Many devs who have never tried from either side build it all up in their head but it's almost always been a matter of thorough tedium, which LLMs are excellent at churning through, especially when there's api docs/headers/code comments.
If you have the space, try mirroring your port at the switch level and capturing every packet then making it go through them all to look for whatever. We have NSA at home lol
Well... I would think that the whole AI industry in the US are working towards public bailouts... Which I guess they'll get under the current administration... So they'll be fine... Nobody there really seems interested in actually creating a profitable business anyway...
First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses?
Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American labs are SOTA. That might change, but let's not act like the American labs don't know what they are doing.
The idea that people are going to use cheaper models for cheaper work isn't some novel revelation, it's completely obvious. People are doing it already, eschewing Fable.
The point is it takes money to keep developing models. Everyone is playing by the same rules. At this point, the US labs are trying to build businesses. I'm not really sure what the Chinese labs goals are. But I do know they aren't doing charity work.
I dont see how the outlook is any better for the open weight companies. They’re in the exact same situation as the closed weight companies except they have had much less revenue, and built up less of a brand, leading up the the point where they are equal in terms of model quality.
In what kind of sad and failed dystopia is this a "saving grace"? For whom?
For Anthropic and OpenAI, presumably. And the rather large economic distortion field around them, that may or may not go very badly for all our retirement funds if those firms become insolvent...
I downgraded my Claude subscription and delegated my Claude Opus access to serve the role of an Architect to brainstorm and plan every step of development.
I leave the development to Deepseek.
Claude gets to review at many layers. It is often just as good as if I let Opus develop it by itself(the architect session will find similar number/level of gaps).
Deepseek flash as an architect and problem solver is not as thorough as Opus5 + high. Codex sol+ high is even better than Opus 5 at this moment for my needs.
I joined a company that is an Anthropic shop and I am genuinely shocked.
Sonnet 5 is a little better on long horizon tasks and headless unsupervised agent workflows - but for in-IDE workflows, it's virtually unusable.
I am so used to flipping around my codebase at warp speed with DeepSeek flash. It's so fast and accurate I don't have time for parallel agents. It's a really rewarding workflow.
Moving to Sonnet, you ask is something simple like "split this into a seperate file" "implement this method" "this is my schema, implement a repository for it". It'll spend 30 minutes thinking and charge like $12. And no token caching, what are you even doing Anthropic?
It's unusable.
DeepSeek are in a league of their own
Like Europe?
this has already happened with manufacturing, so it isn't surprising that other industries follow.
The US premium in engineering and scientific endeavors have been lacking for the past 30-40 years, and if it werent for tech and silicon valley, the US would have nothing state of the art. Even on that front, the US is falling behind given how much effort in tech has been diverted into privacy invading, and advertising.
The US has been riding momentum, but eventually that momentum will stop. It will take half a century to get back up to speed, and by then, the US will have fallen behind so far that catching back up will seem impossible.
The telling evidence would be if china has the first moonbase before the US does. I think this is highly likely looking at today's US administration.
The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model
Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"?
I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?
> You don't win the stock market or make the deadliest drone by switching to the cheap model
The question is not "can you win with the best model?", it is "can you not win without the best model?".
As it is now clear by behaviour of companies and US govt, all these investments will be backstopped by US govt. No US AI company will go hungry, they are national champions.
The ByteDance folks are apparently training a mythos level model 10T params apparently. If they do would it still be subsidized at these cheap rates?
https://en.wikipedia.org/wiki/Category:Defunct_low-cost_airl...
https://en.wikipedia.org/wiki/List_of_defunct_airlines_of_th...
https://en.wikipedia.org/wiki/List_of_defunct_airlines_of_th...
https://en.wikipedia.org/wiki/List_of_defunct_airlines_of_th...
Seems more like a overall industry problem, not limited to low cost carriers
Spirit was broken by oil prices which everyone pays the same for. (There is no cheaper jet fuel alternative).
Not a good comparison to the point of wrong conclusions.
At least here in Germany Aldi isn't even really limited to the poor, it's famously a place where you can run into anyone. Where I used to live in Berlin close to the government district I literally on occasion ran into the chancellor (and her bodyguards). Aspirational shopping where you buy premium goods to pretend to have higher social status honestly seems a bit on its way out. Even middle class people seem to consciously shop more utilitarian now.
The 'floor' has gone up: today's model a bit behind SOTA is like model releases that were blowing people's minds a few months ago. Compared to, say, DS R1, this is far-out futuristic technology.
This, Luna, and (if it's good in practice) Laguna S are also fast and light not just cheap. And, as happened before, DeepSeek's first but other open model makers likely follow.
And a small, fast model taking small steps is...fun? More of the experience of working with the code.
Another time, Flash started trying to make tool calls by just calling bash and catting the tool call to stdout. Then it started running echo xx for every two letter UNIX command it could think of: mv, cp, etc and the it dug into uv, ty, and jj
With 5 active sessions going nonstop? That seems like a pretty important qualifier.
Then, once I go over, API pricing racks up FAST!
I'm also creating a free platform that replaces extremely out-of-date software, some of it only available with mutli-million dollar contracts, to help medical physics professionals with cutting-edge radiotherapy devices used to treat cancer.
https://brynnbateman.com/ for a list of projects
And out of curiosity, how do you automate testing the porting in the browser that's actually playable etc? And aren't you a bit scared of hosting and serving the "hairy bits" such as full assets? Nice job anyway!
I automate testing playable parts in browser by adding a dev mode that allows text commands for everything instead of having to rely on clicking UI or 3D elements. It can see the full game state in JSON and interact in any way via commands.
And re: IP - I just accept that I might get a C&D any day and have to take it all down. I'm careful to not accept a single penny for any reason and don't even have Patreon. Usually monetizing is what makes IP owners unhappy. And for the Pokemon MMO I just don't advertise it anywhere meaningful since they'll C&D the second they see it. Largely made it for my nephew and we play it together.
I can pretty easily burn through my weekly quota over several agent coding hours with minimal supervision when tasked with some pretty large but well-planned refactors.
i used for work where i did less and it quickly reaches thousands if you're not careful. i can already see what some will say: skill issue et cetera - whatever.
$5/days is ~330 Mtok/day, that’s a nontrivial amount of work, and none of the gpts are more efficient than deepseek at $/task if deepseek meets your quality bar.
OpenCode currently offers 60 USD API credits at 10 USD per month (OpenCode Go) and have even doubled it temporarily as a promotion.
Effectively you can get Deepseek for 1/12th the already ridiculous cheap API price.
Per the open code zen pricing page[1], it appears that the token prices are the same, but their cache is 10x more expensive?
[1]: https://opencode.ai/docs/zen/#pricing
Deepseek 0.14/0.28/0.0028
Pro Opencode 1.74/3.48/0.145
Deepseek 0.44/0.87/0.0036
For flash the input/output is the same, but the cache difference is big, you're paying 10x on >95% of your tokens.
For Pro, it's even worse, input/output is 4x and cache is 40x. The price different is really brutal. Yes you will still come out ahead by spending your first 10$/month on opencode go, but you will be saving a lot less than initially appears from their (60 USD for 10 USD pitch).
[0]: Deepseek: https://api-docs.deepseek.com/quick_start/pricing/ [1]: Opencode: https://opencode.ai/docs/zen/#pricing
Not true. Sol on XHigh or Max runs out even on the $200/mo plan. It's not close to effectively unlimited. Maybe at 2x the current allowance it can.
Real work. $200 looks good on the outside until the essence of it, e.g. the models lie. I gave a list of spec to Sol and Sol decided some items didn't need to be done and the reason was "unproven", "not enough evidence", etc.
They all come up with amazing ways to lie (or be lazy). Often times what you get isn't what you asked for (only on the surface). E.g. I ran it to iteratively bench and optimize a better data structure for the project. It spent hours and finally came up with something. When I check it out -- it benchmarked the wrong criteria and was way off. So here we go again. Most AI work looks good on the surface. There are infinite edge cases.
So to do real work and gate it you need to:
1. Plan
2. Get it to do the work
3. Get independent agents to check from different angles
4. Take that feedback and get it to fix those gaps
5. Match against the plan and redo parts if needed
Every task is easily 4-5x the estimated amount of tokens.
p.s. well I did burn some banked resets building a compiler for some language AND it is still NOT done. Every time it says done I say check it says ok we still have bugs...
> I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams)
And to be frank, it is not that much weaker for regular software development work. I use Claude at work and I see no difference in capability. I only notice a dramatic difference in how much more expensive it is.
vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development.
I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek prices.
The entire issue is caching, I tried to write some custom to dump to disk kv-caching using some ideas from their papers and my experience with snapshots and vm checkpoint systems, I must say they must have really squeezed that lemon it's hard.
Atleast me with Sol couldn't figure it out over a couple days, a few hours each day, which isn't much but I did feel a bit stuck with existing solutions and felt like I might have to write something from scratch. But if you are willing to put in the effort into the infra I do think it's doable. But it will be really hard to pull it off.
My congrats to anyone who manages to pull it off, they might be able to kill off most AI labs. Assuming they can find the compute, Deepseek really has killed all models for me other than Sol/Fable/Opus/K3 tier stuff.
And there is no way in hell anyone can afford caching prices same as what DeepSeek is offering, and DeepSeek keeps the cache available for an insane amount of time most providers will flush it in 5-mins like Claude/Anthropic (some offer customizing it but I am not sure of the pricing, it's load based on some like Fireworks, which means assume a couple minutes at most, they say several minutes god knows what that really means).
There is no way to match DeepSeek's current prices, "profitably" if you are renting a GPU and reselling tokens, unless you have some really amazing caching infra or something.
Deepseek's prices are just insanely cheap, I am not saying it's impossible to get there the overall performance suggests it should be feasible, but I will be damned if any provider could match their tps and caching any time soon at those same prices profitably.
I believe even if Deepseek 2-3x their prices across the board even then they would be cheaper for most long running tasks, that's just how good their caching is.
For one I have managed to hit the cache after over 24 hours on their system it's insane, I honestly didn't care because it was so cheap but it truly made me incredibly happy to think about the engineering that must have taken. TTFT is slightly worse, but it's good enough, for those cache prices I can take a few seconds worth of hit on TTFT.
It's interesting that most open models adding 1M context did it in a way that reduces KV cache size (though DeepSeek was the most aggressive, using compressed attention on all layers), but only a couple providers turned it into a discount on cache reads.
Can anyone working at one of the main US labs (Google, OpenAI, Anthropic) comment on WTF they haven't even tried MLA - despite the obvious massive advantages?
I know enough to know they aren't completely incompetent. So there must be a quite good reason.
But it remains a mystery to me.
DeepSeek's MLA is like almost 2 years old at this time. They've got thousands of people working on this stuff. They clearly have the ability to at least try it...
There’s a measurable performance tradeoff versus gqa so there’s reluctance.
For the most part though the new deepseek v4 tech is hca and mhc and people are still catching on like with moe and rl. Wait for 6 12 months, minimum time for next pre train.
The big US labs are opaque and don't publish much of any technical details anymore. We don't know what they are or aren't doing, honestly.
They "can" is the caveat here. Rented GPUs are going up in pricing. I recently got an email that DigitalOcean pricing of GPUs were going up.
So
1. They have to get a hold of them (availability is bad)
2. They have to maintain the pricing
Deepseek charges $0.0028 per cache read on Openrouter. The next cheapest is $0.018.
That's a massive difference and quickly adds up on coding sessions (which often hit 95%+ cached tokens).
Cache:
Hit rate: 98.582% (1,265,646,976 / 1,283,855,064)This adds disk as a tier in the HBM → CPU → Disk KV cache hierarchy.
There's also a cluster of related KV-offload FS PRs: #49225 (read/write batching, still open) and #49152 (batch store/load in C, merged Jul 28).
It's hard to say if these are similar to the approach DeepSeek takes but they definitely seem very interesting.
[0] https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice."
I am not sure why you wouldn't want to use the SOTA models unless speed is a concern. Otherwise you are leaving quality on the table.
which in your case is?
My family uses it. I have gallery apps (yearbooks for each year are a lot of fun!) of us on trips and just living, an outlining app that's a mesh of Workflowy and Org Mode (it's called Fluxtral), a markdown-backed app (it uses marked.min.js, and is called Dextral) that offers documents, logs, calendars, and kanban boards, all parsed from markdown. I have a List app for gear, trips, shopping, etc. that we all can contribute to. There are utilities (world clock, calendar) and games (an oracle for RPGs, a KenKen implementation), and apps (a diagram editor that exports to SVG, a web-launcher that uses pneumonics, a Scheme-based hacking environment, and a spreadsheet that does most of what you'd expect aside from Solver and Pivot tables).
I started these projects before AI, and made slow progress over the years, but the modern versions of all this stuff have been built with Deepseek V4 Flash. I've also used Gemini in the very early days, and Kimi K2.6 later on, but these days, since I can now host Deepseek v4 Flash 0731 in a 2-bit quant on my Strix Halo box (128GB, but only about 250GB/s of memory bandwidth, so 15t/s), I used Deepseek with omp for almost everything. It's a very capable model, and I'm amazed I can run it locally and get good results. It's really revolutionary for my (small) use cases.
With unsloth's Q3_S quant + kyuz0 'llama-vulkan-radv-performance' toolbox, I am getting 280+tps (batch and ubatch at 2048) for PP and 18+tps for TG. I really only need 256k context so it all fits.
If I go down to the Q3_XXS quant + dpsark + 'llama-vulkan-radv-performance', I can get about the same PP and 25+tps for TG with draft set to 2 or 3. Fits about the same as above.
Edit: I did notice the 25+tps quickly degrades down to 20+ after the first few hundred tokens.
oh, they're mad.
And, probably 99.99% of people using LLM probably don't even need SOTA anyway.
What's normal usage? I mean, Kimi is already really keen to spin of lots of subagents, and DeepSeep can probably do the same?
> The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues.
Yes, though I did that even with Claude (on my employer's token budget). The agents are great at doing the gruntwork of chasing down the reproduction of flaky tests, too. They need some hand holding at first, but the guidelines are usually re-usable per project. (Claude specifically needs to be told to really concentrate on reproduction, and not eagerly start fixing the flake: if you don't have a reliable reproduction, you have no clue whether your fix actually fixes anything.)
I don't think this is the win you think it is. It's amazing that this is possible, but it introduces so much human overhead that you can drown in reviews and it can effectively slow you down more than a quick check and fix yourself.
The models need to get a lot more consistent in what they can and can't do before you can automate this stuff and only check the things you know the model isn't good at
I’ve found it to be very capable. I’m using it with pi as well and some custom extensions I’ve put together over the past few months and it’s pretty crazy having it do what I need it to a vast majority of the time, do it fast, and see that it’s used like $0.12.
You can probably implement something similar as a plugin for your preferred harness. From a technical perspective I think it just sends the output w/h the thinking and tool trace to another model and asks it to double check everything (exact prompt must be somewhere in the OMP repo).
Would you run a less costly model as the supervisor given it’s consuming a lot of text and may have a simpler task to do like “make sure the implementing model doesn’t start over-engineering things”?
this seems like such a bad idea
"If this PR adds any new endpoints, ensure that there are functional and integration tests. If there are not, please investigate the feasibility and appropriateness, and create functional tests using the guide found on our wiki for guidance https://www.ourdevwiki.site/how-to-make-functional-tests" then maybe it could add some value.
But that very much depends on the specific system. Some tests are obvious, some not so much.
The analogy I like is that building software is running a Michelin restaurant. The moment you scale, the chef is just writing cooking books and is absent, and you move into franchising, you will be amazed at the bottom line revenue scaling, while customers will be progressively appalled with the food...
I hadn't really thought about this but AI may well be the technology that disrupts and ultimately destroys social media.
The value proposition of something like FB or IG is, as we know, the network effect. The platform gets to extract value from user generated content. I believe that users should own the platform, a bit like the Wikimedia Foundation, because they're the ones that create value. Federation is a popular belief on HN and I've come to believe that's simply the wrong solution to the right problem.
Anyway, how these social media companies make money is by optimizing the feed for engagement. People know it too so you see people trying to build an audience by rage baiting. And then more time spent equals more advertising revenue.
But what happens when the AI can simply slurp all the posts and then filter and rank them? It destroys the engagement and advertising model. And I'm not opposed to that, honestly. It may be on eof the few good thing sto come out of AI.
Terrible use-case.
harness: omp.sh
My initial thought was to sign up for ChatGPT, but I had $20 in OpenRouter so I've been trying out DeepSeek V4 Pro with Pi for the last few days and I gotta say, it's good enough for my use case. And even with paying for API usage rather than Claude's subsidised subscription, and with OpenRouter taking their cut, I will probably end up paying significantly less overall. And I really like the flexibility of being able to use whatever minimalist open source harness I want (and being able to switch providers easily, too).
(My demands probably aren't as high as many others' - I mostly use it for help with some hobbyist coding projects, and I tend to ask it questions about how to approach problems rather than just telling it to go off and code stuff for me.)
If you prefer subscriptions, OpenCode Go ($10/mo), Cline Pass ($10/mo), Atlas Code ($20/mo), and CommandCode ($1/mo) serve some of the best open weights with generous limits. OpenCode Go currently offers $120 for $10 on DeepSeek Flash v4 (if you're okay with data retention).
> DeepSeek V4 Flash: ZDR agreement is renewed monthly. The current agreement is valid through August 31, 2026.
Is there other info I should be aware of w.r.t data retention with opencode go? It's hosted in China, so other middlemen may be active (I doubt it, but possible)?
What are you going to do? Take a CCP company in front of a CCP judge?
If they are doing it to you, they are probably doing it to others, which makes an easy class action
It would be a hassle, but China does have privacy laws. Companies do get sued for violating them.[0]
0. https://www.chinajusticeobserver.com/a/china%E2%80%99s-top-c...
Remember China is still an authoritarian dictatorship. One leader with absolute power for life. Don't let the facade misguide you.
Trying suing when it's in the states interest, like it is to build AI datasets on western data. For reference, no case in China has ever been ruled against the government. They don't have things like judges smacking down executive orders or refunding tariffs.
At DeepSeek's absurdly low rates or market rates?
Careful with using OpenCode's accounting for DeepSeek v4 Pro, though: https://github.com/anomalyco/opencode/issues/39822
[0] https://github.com/anomalyco/opencode/issues/39857
I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload.
The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s on a single stream. I saw 1000 tok/s with ~64 concurrent streams on vLLM.
That's fast enough that you can interactively chat with it without switching tabs while you wait, and its a ~300B (13B active, hence the speed) model so the responses are also very good. It's actually more convenient now for me to direct 95%+ of my day to day usage to my local model, and only use Claude Fable for really big coding tasks.
Until this model was released, I was contemplating spending even more money on hardware to run GLM5.2 (~750B) at reasonable speeds, but I no longer feel that need. This is smart enough, and I think it only gets much better for local models from here.
I wouldn't call 80 t/s slow.
For reference, on a 1x B300 it's over 400 tok/s decode on a single stream.
I'm guessing tensor parallelism or similar?
Here's a runbook: https://github.com/local-inference-lab/rtx6kpro/blob/master/...
If the newer builds aren't working, you might try running the old v6 build (based on the eldritch-enlightenment image). gilded-gnosis gave me some problems that I haven't bothered to track down, the old builds are still gonna blow away llama-server performance. And that's before you get hooked on vLLM's PagedAttention and can run multiple sequences without a ton of extra overhead.
It is strong (not Fable strong though) with a much better “persona” than Opus, and very different blindspots. If you flip between Claude and this you will find both catch the mistakes of the other before they get out of control.
On balance I actually prefer DeepSeek for programming now, because of the way it talks.
This is on Pi agent, nothing fancy at all about my prompts or use case. Anyone else experiencing this?
I've also had it randomly go from talking about Rust to talking about the electric chair, controversies about D&D rules (both irrelevant and something I've never discussed) and it's completely blind to it in future prompts even when its pointed out and referenced directly
All this said its still worth it but the agentic performance has degraded in my experience at least
It might be even better in Codex or Oh My Pi according to this bench I saw earlier: https://nitter.net/composio/status/2085330847951970801
Baseten.co's version got into a loop rather rapidly... I've since added loop detection and adjusted some other settings on the pi coding agent and have yet to notice it again. I also switched to DeepInfra ... who serves an fp4 version admittedly, but I've had no issues with it as of yet and it's the top provider on openrouter.ai volume wise.
When it was first available in opencode, it was kinda slow for me, I guess because everyone wanted to try the new shiny. But now it's back to being screamingly fast and Opus 4.8 level of smart, for penies.
I suspect that once the hype dies down, or the field gets more competitive we will see the same on 0731.
Seems like we've reached the event horizon of whether AI advances are worth paying attention to.
I recommend opencode or something akin to it to play with models. Any big model updates or hot new ones will naturally run across your desk that way
Not for me, Fable refuses to debug Linux kernel bugs. Unless you say who you're speaking for, it sounds like you're just shilling for Anthropic.
A chinese model being in the same ballpark of capability at half the price sounds believable to me.
DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it.
I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot.
Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations.
(Both were set to high reasoning)
Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.
it's still $3/$15 for all providers on openrouter
because of some Kimi license
https://openrouter.ai/moonshotai/kimi-k3#providers
https://synthetic.new/?referral=kwjqga9QYoUgpZV
Uptime looks crap, though.
So we won't see any price decrease unless Kimi changes the license of K3
My read is, OpenAI is neither able to claw b2b money (away from Ant) nor are they able to stave off open weights on the other. In short, they're struggling to hold onto their distant #2 position in the coding market, and these pricing changes reflect a (desperate) change in strategy.
the private endpoint costs 10x (azure).
private endpoints for deepseek (lots of providers) also cost about 10x more.
but 10x more for deepseek is $0.028 cached input, and 10x more for luna is $0.10.
Which would put them... exactly where everyone else is on this graph.
Edit: I seem to have misunderstood the news. I thought the magical cache read pricing was going away (0.002) and they were going to be on par with everyone else (0.02). But I have no idea.
Edit 2: Apparently, neither do they!
>We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
Sort by cache read.
No.
They sent an email to customers saying that they will raise prices "significantly".
How much that will be is speculation.
My guess is that they will just remove the 75% discount they gave when they released V4 preview. It will still be relatively cheap even at 4x the current price.
But note that you have to use Cline (or other harness) if using vscode. I was shocked at how poor the recent versions of GitHub Copilot are at using the cache (with Fireworks AI, but I believe it's a more generic problem).
https://x.com/vijucat/status/2085415745144672492?s=20
And how this has been accelerating!!
I felt this very hard when I had to travel in the middle of nowhere in south america, with no network, and wanted to keep an LLM model on my macbook pro with 48GB of RAM. That was back in April 2026, a few months ago.
I downloaded Google Gemma 4 (google/gemma-4-26b-a4b) and - Oh boy - I was amazed by it's capacity!
I was able to use it to code simple things, ask it about nature, learn new stuff while traveling and make stories for the kids.
Was really amazing to observe and experiment this!
Seems to me there will be some good chance to run these great LLM locally on our hardware!
Amazing time to be alive
I Compared Deepseek V4 Flash 0731 (low) to Gemini 3.5 Flash Lite (minimal) and GPT 5.6 Luna (no reasoning) and Deepseek V4 Flash 0731 gets it wrong alot, where as Gemini and 5.6 Luna just gets it done.
Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.
But even in this very post, you can see that Max was actually cheaper than High.
If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.
Btw if you need an app I may deliver it to you in ten minutes for just five cents if I'm in the mood. Just let me know.
I have £20/month Gemini and £20 a month claude for a bunch of personal projects.
Yes I have to wait sometimes, it's probably a good thing.
Not a huge deal since it's still cents per session, but my bigger issue was the weird change in tone. It became a lot more pretentious and over-explanatory.
Heavy prompt reworking helped but maybe that's just the cost of being better at coding and ARC-AGI?
Furthermore, in my company we are using MCPs for Google Ads (it manages our ads), Analytics, Search Console, Zoho CRM, Microsoft Clarity... We use it to crawl specific websites and send daily summaries to our sales team in MS Teams channel. We use it to send daily summaries on marketing statistics and analytics... All with a FREE model. We are rarely hitting any limits so far and in case we need more tokens - we use NOUS or openrouter to pick between Flash or Pro for specific tasks that require more churning.
AMA.
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
For anything even moderately complex.. like, even low end of complexity, this model behaves maximum like gpt-5.6-luna-high .. nothing more.
Yesterday itself I gave it a coding task in some existing moderately complex small project, and i was using xhigh thinking effort, it was unable to cover all edge cases... and i had already got it to review, and then fix, 3 more times, after the first initial one.
Still it left 2 edge cases.
Then, reverted full code, gave sol-high the same task, it took well over 20 minutes, and completed it in one go with zero edge cases remaining.
I am not using it for anything serious anymore.
I am among those with real life experience with the model that used the previous as well and will attest that the new model is a big improvement
[0] https://taylor.town/silver-landmines
When I see dramatic leaps like this, it tells me that the important hacks haven't yet been discovered.
From here on, it's going to become all about harnesses that best situate and organize swarm intelligence at scale.
If I'm reading the chart correctly, a couple observations:
* deepseek-v4-flash-0731 max is better than kimi-k3 max
* glm-5.2 is dumber than a box of rocks (this must be on low reasoning or something, right?)
This is way more extreme than other results I'm seeing, like those from Artificial Analysis.
When I need vision capabilities I use GPT 5.3 codex and if deepseek can’t figure something out after a few goes I switch to GTP 5.5 or 5.6 (I’ve been giving Terra first bite recently and it does pretty well, and have used Sol a couple of times).
Using this regimen means I spend under $100 per month on inference and I work all day everyday with multiple agents running simultaneously all on API token spend not subscriptions.
ARC-AGI II:
- GPT-5.2 (medium) %26.7 ($0.759)
- DSV4-Flash (max) %61.4 ($0.04)
But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?
Tell your PjM who should tell your PgM who should tell your PdM, all the PMs...
Maybe if "the business" sees it is true of LLMs, they might believe it's true of giving better context to engineers up front then giving them time to think and prototype (thinking tokens are an answer prototype).
https://reddit.com/r/DeepSeek is where the fellow F5ers are at.
Does no thinking emissions for context saving.
It’s always a bit tricky picking the right harness (when you have options). Sometimes the differences are subtle but meaningful. But who has the time to run everything twice and compare all the time!
Codex is really good in my experience, especially due to its native sandboxing. Deepseek seems really well versed in its tools, including update_plan and knowing when to request sandbox escalation.
promising!
https://americanliterature.com/author/em-forster/novella/the...
https://news.ycombinator.com/item?id=49198661
https://x.com/thdxr/status/2085377844515922210
What secret sauce do they have?
Quant company usually squeezing every penny.
pair it with codewhale, 50 agents, 200 MB of ram.
But I find it having a pretty significant problem with tool calling - no idea why, but tool calling with it is SLOW. As long as the model is reasoning, all good. But give it a bunch of tools and it becomes extremely slow.
Am I the only one experiencing this?
China has zero energy concerns in terms of energy production - not literally zero, but they’d be able to prioritize other dimensions and not necessarily worry about efficiency
Here they are though releasing models that sip resources
That's irrelevant when you use $/task as the metric, which the OP does use.
That's how I handle the Qwen27B and 35B
What do you mean by "redirect it to useful output"? Could you give an example? This sounds interesting.
Imagine if they had GPU resources of western labs.
SV companies get way too comfortable when they have enough in the bank to stay running more than three months.
Energy and intelligence are good too, sure.
As an end consumer, I don't care about the number of active parameters. I really do care only about the tracked metric (how well does it do the job, and how much does it cost... ideally also with time included, but that wouldn't fit on a 2D chart)
With the exception of cache costs, all providers have similar input/output costs.
They don't strictly need any kind of subsidies.
FWIW they have a funding round planned (kerfuffle about leaks from CEO presentation few weeks back) -- presumably because infrastructure needs have ballooned.
Naturally there will be some PRC government interest in one of their flagship AI companies. From what is visible seems to be more along the lines of ensuring that DS gets its fair share of resources -- e.g. Xi Jinping meeting founder and positive comments about success of DS means that (hypothetically) Alibaba can't screw DS too much on infra charges to kill off a 'competitor'. Also would imagine that DS's top guys have been clearly identified and will have been 'discouraged' from going to work for one of the SV polycules. But even here as much carrot as stick -- none of the DS top guys will ever need to work again except for love of the job.
> They don't strictly need any kind of subsidies.
You understand how these two sentences directly contradict each-other, yeah? The money-losing operating of training a model is paid for by momey earned from prior investments. So… the work is “subsidized” by its parent company’s investments in it.