Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Here’s a take I haven’t seen yet:

If training and inference just got 40x more efficient, but OpenAI and co. still have the same compute resources, once they’ve baked in all the DeepSeek improvements, we’re about to find out very quickly whether 40x the compute delivers 40x the performance / output quality, or if output quality has ceased to be compute-bound.



> If training and inference just got 40x more efficient

Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year? It was always the second one, and doing the training cheaper makes it even more so.

But this implies that we could use those same resources to train even bigger models, right? Except that you then have the same problem. You have a bigger model, maybe it's better, but if you've made inference cost linearly more because of the size and the size is now 40x bigger, you now need that much more compute for inference.


Actually inference got more efficient as well, thanks to the multi-head latent attention algorithm that compresses the key-value cache to drastically reduce memory usage.

https://mlnotes.substack.com/p/the-valleys-going-crazy-how-d...


That's a useful performance improvement but it's incremental progress in line with what new models often improve over their predecessors, not in line with the much more dramatic reduction they've achieved in training cost.


If H800 is a memory-constrained model that NVIDIA built to avoid the Chinese export ban on H100 with equivalent fp8 performance, it makes zero sense to believe Elon Musk, Dario Armodei and Alexandr Wang's claims that DeepSeek smuggled H100s.

The only reason why a team would allocate time on memory optimizations and writing NVPTX code rather than focusing on posttraining is if they severely struggled with memory during training.

I mean, take a look at the numbers:

https://www.fibermall.com/blog/nvidia-ai-chip.htm#A100_vs_A8...

This is a massive trick pulled by Jensen, take the H100 design whose sales are regulated by the government, make it look 40x weaker and call it H800, while conveniently leaving 8-bit computation as fast as H100. Then bring it to China and let companies stockpile without disclosing production or sales numbers, and have no export controls.

Eventually, after 7 months, US govt starts noticing the H800 sales and introduces new export controls, but it's too late. By this point, DeepSeek has started research using fp8. They slowly build bigger and bigger models, work on the bandwidth and memory consumptions, until they make r1 - their reasoning model.


What's surprising is anyone would repeat Elon musk related things.

Tech or politics related, he's off the deep end.


Especially since he seems intent on everyone talking about him all the time. I find it questionable when a person wants to be the centre of attention no matter. Perhaps attention is not all we need.


Yet another casualty of laypersons browsing arXiv. That paper was like flypaper to his narcissism.


The problem is he's only wrong some of the time and then people arguing about which one it is this time generates attention, a valuable commodity.


Maybe “some” applied in the past but his recent history might best be described as “almost always”.


Drugs. Dont do that much drugs for so long.


He's like a broken smart network switch, smart as in managed. Packets with switch MAC on it are all broken, but erroneously forwarded ones often has valuable data. We through L3 don't know which one is which.


I'm wrong some of the times.

He's a lucky mensch, no more, no less.


Interesting how people keep calling it “the Chinese export ban”. Isn’t an American export ban?


I don’t think that it got more efficient. It’s that smaller models can train via larger ones cheaply. Think teacher/student relationship

https://en.m.wikipedia.org/wiki/Knowledge_distillation


I think what got cheaper are models with up to date information.


You almost never reintegrate new information with training, its by far the most expensive way to do that.


...and that got cheaper? Not sure your point.

At some point, the models _have_ to do "continuous integration" to provide the "AGI" that's wanted out of this tech.


> DeepSeek is still a big model that requires a lot of resources to run

I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4.

I've just switched to it for my local inference.


Isn't the largest model still like 130GB after heavy quantization[1] and 4 tok/s borderline unusable for interactive sessions with those long outputs?

[1] https://unsloth.ai/blog/deepseekr1-dynamic


I told it to skip all reasoning and explanations and output just the code. It complied, saving a lot of time)


Wouldn't that also result in it skipping the "thinking" and thus in worse results?


Yes, but even that can still be run (slowly) on cpu-only systems down to about 32gb. Memory virtualization is a thing. If you get used to using it like email rather than chat, it’s still super useful even if you are waiting 1/2 hour for your reply. Presumably you have a fast distill on tap for interactive stuff.

I run my models in an agentic framework with fast models that can ask slower models or APIs when needed. It works perfectly, 60 percent of the time lol.


OP probably means "the largest distilled model"


So not an actual DeepSeek-R1 model but a distilled Qwen or Llama model.

From DeepSeek-R1 paper:

> As shown in Table 5, simply distilling DeepSeek-R1’s outputs enables the efficient DeepSeekR1-7B (i.e., DeepSeek-R1-Distill-Qwen-7B, abbreviated similarly below) to outperform nonreasoning models like GPT-4o-0513 across the board.

and

> DeepSeek-R1-14B surpasses QwQ-32BPreview on all evaluation metrics, while DeepSeek-R1-32B and DeepSeek-R1-70B significantly exceed o1-mini on most benchmarks.

and

> These [Distilled Model Evaluation] results demonstrate the strong potential of distillation. Additionally, we found that applying RL to these distilled models yields significant further gains. We believe this warrants further exploration and therefore present only the results of the simple SFT-distilled models here.


How are you running it, can you be more specific?


DeepSeek-R1-Distill-Llama-70B on triple 4090 cards.


In the long run (which in the AI world is probably ~1 year) this is very good for Nvidia, very good for the hyperscalers, and very good for anyone building AI applications.

The only thing it's not good for is the idea that OpenAI and/or Anthropic will eventually become profitable companies with market caps that exceed Apple's by orders of magnitude. Oh no, anyway.


Yes! I have had the exact same mental model. The biggest losers in this news are the groups building frontier models. They are the ones with huge valuations but if the optimizations becomes even close to true, its a massive threat to their business model. My feet are on the ground but I do still believe that the world does not comprehend how much compute it can use...as compute gets cheaper we will use more of it. Ignoring equity pricing, this benefits all other parties.


My big current conspiracy theory is that this negative sentiment toward Nvidia from Deepseek's release is spread by people who actually want to buy more stock at a cheaper price. Like, if you know anything about the topic, it's wild to assume that this will drive demand for GPUs anywhere but up. If Nvidia came out with a Jetson like product that can run the full 670B R1, they could make infinite money. And in the datacenter section, companies will stumble over each other to get the necessary hardware (which corresponds to a dozen H100s or so right now). Especially once HF comes out with their uncensored reproduction. There's so much opportunity to turn more compute into more money because of this, almost every company could theoretically benefit.


Can you guys explain what this would be bad for the OpenAI and Anthropic of the world?

Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to happen in the long run?


First of all, we need to just stop talking about AGI and Superintelligence. It's a total distraction from the actual value that has already been created by AI/ML over the years and will continue to be created.

That said, you have to distinguish between "good for the field of AI, the AI industry overall, and users of AI" from "good for a couple of companies that want to be the sole provider of SOTA models and extract maximum value from everyone else to drive their own equity valuations to the moon". Deepseek is positive for the former and negative for the latter.


Because building a frontier model is expensive. But building a model as good as an existing frontier model is cheap (re: distillation).

https://en.m.wikipedia.org/wiki/Knowledge_distillation

So the takeaway is they have no moat



Beautiful and concise, much better than my word salad.


I believe in general the business model of building frontier models has not been fully baked out yet. Lets ignore the thought of AGI and just say models do continue to improve. In OpenAIs case they have raised lots of capital in the hopes of dominating the market. That capital pegged them at a valuation. Now you have a company with ~100 employees and supposedly a lot less capital come in a get close to OpenAIs current leading model. It has the potential to pop their balloon massively.

By releasing a lot of it opensource everyone has their hands on it. Opens the door to new companies.

Or a simple mental model, there has been this ability for third parties to get quite close to leading frontier models. The leading frontier models takes hundreds of millions of dollars and if someone is able to copy it within a years time for significantly less capital, its going to be hard game of cat and mouse.


If I can use LLMs for free, why would I give money to OpenAI or Anthropic?


Yes, but I think most of the rout is caused by the fact that there really isn't anything protecting AI from being disrupted by a new player - They're fairly simple technology compared to some of the other things tech companies build. That means openai really doesn't have much ability to protect it's market leader status.

I don't really understand why the stock market has decided this affects nvidia's stock price though.


Does line go up forever?


Isn't the question closer to "/when/ does line stop going up?"


That is a matter of hope.

If line keeps going up, line does catastrophic or potentially apocalyptic harm, given our current circumstances.


It goes up at least until LLMs match humans - ie until an LLM can write Windows


I want the LLM to decide not to do anything, or write a new OS.

Whenever I prompt: "Do not do anything"

It always does <something>.


Do not think of a pink elephant. Were you able to do so?


> Whenever I prompt: "Do not do anything" It always does <something>.

Yep. A lot of times, the responses I get remind me of Simone in Ferris Bueller's Day Off: https://www.youtube.com/watch?v=swBtLPWeKbU

If you end up making a new model, please teach it that less is more and call it "LAIconic".


Slightly tangential, but I want an LLM which can debug windows.


Debug And locally fix security holes.


Or sell the found exploit to the friendly nation state actor to pay for its compute behind your back


I've missed the stories on this until now. Is it known (and is there an ELI5) how they were able to do it so much more efficiently?


This article has good background, context, and explanations [1] They skipped CUDA and instead used PTX which is a lower level instruction set where they were able to implement more performant cross-chip comms to make up for the less-performant H800 chips.

[1]: https://stratechery.com/2025/deepseek-faq/


> Moreover, if you actually did the math on the previous question, you would realize that DeepSeek actually had an excess of computing; that’s because DeepSeek actually programmed 20 of the 132 processing units on each H800 specifically to manage cross-chip communications. This is actually impossible to do in CUDA.

You can do this just fine in CUDA, no PTX required. Of course all the major shops are using inline PTX at the very least to access the Tensor cores effectively.


So can people do the same in SPIR for OpenCL or amdgcn?

https://en.wikipedia.org/wiki/Standard_Portable_Intermediate...

https://www.khronos.org/spir/

Or even better in the unified language like SYCL?

https://cdrdv2-public.intel.com/786536/Heidelberg_IWOCL__SYC...


IIUC they released a paper, it's partially algorithmic improvements partially good old low level optimization.


There's something to be said about the idea that instead of just dumping oceans of money into buying Nvidia cards they just...optimized what they had

I'd say the wider industry could learn a thing or two, but as other commentors have joked. The line must go up


The double edged sword of export controls.


Creativity always delivers when minds are constrained.


Is that what we can call Project2025?

Creativity from constrained minds?


surely more relevant comments could exist in the aether of the internet.


Your reply doesn't read like a refute to me, mnky9800n.


Ah but why care about efficiency when you have basically unlimited investor money?

Reminds me of Japanese cars and the OPEC boycott in the 1970s...


there's this and that little desktop computer they announced earlier this month - digits

they claim it's able to run models with 200B parameters on a single node and 400B when paired with another node


Like everything, I expect improvement to be logarithmic


That's a take I've seen in many HN comments


That seems to be the key question.


Yeah this was my first thought as well. If it got so efficient how good all the models will be 2-3 months from now


>If training and inference just got 40x more efficient

The jury is still out on how much improvement DeepSeek made in terms of training and inference compute efficiency, but personally I think 10x is probably the actual improvement that's being made

But in business/engineering/manufacturing/etc if you have 10x more efficiency, you're basically going to obliterate the competitions.

>output quality has ceased to be compute-bound

You raised an interesting conjecture and it seems that it's very likely the case.

I know that it's not even a full two years that ChatGPT-4 has been released but it seems that it take OpenAI a very long time to release ChatGPT-5. Is it because they're taking their own sweet time to release the software not unlike GIMP, or they genuinely cannot justify the improvement to jump from 4 to 5? This stagnation however, has allowed others to catch up. Now based on DeekSeek claims, anyone can has their own ChatGPT-4 under their desk with Nvidia project Digits mini PCs [1]. For running DeepSeek, 4 units mini PCs will be more than enough of 4 PFLOPS and cost only USD12K. Let's say on average one subscriber user pays OpenAI monthly payment of USD$10, for 1000 persons organization it will be USD$10K, and the investment will pays for itself within a month, and no data ever leave the organization since it's a private cloud!

For training similar system to ChatGPT-4 based on DeepSeeks claims, a few millions USD$ is more than enough. Apparently, OpenAI, Softbank and Oracle just announced USD$500 Billions joint ventures to bring the AI forward with the new announced Stargate AI project but that's 10,000x money [2],[3]. But the elephant in the room question is that, can they even get 10x quality improvement of the existing ChatGPT-4? I really seriously doubt it.

[1] NVIDIA Puts Grace Blackwell on Every Desk and at Every AI Developer’s Fingertips:

https://nvidianews.nvidia.com/news/nvidia-puts-grace-blackwe...

[2] Trump unveils $500bn Stargate AI project between OpenAI, Oracle and SoftBank:

https://www.theguardian.com/us-news/2025/jan/21/trump-ai-joi...

[3] Announcing The Stargate Project:

https://openai.com/index/announcing-the-stargate-project/




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: