Hacker Newsnew | past | comments | ask | show | jobs | submit | valegrete's commentslogin

The difference is that the textile machine workers could immediately verify the work product. This is impossible when the work product is a configuration of knowledge you don't actually have. It's why, for example, amateurs "solving" these open math problems refuse to discuss the results with actual mathematicians, because they can't. So we have a bunch of Lean-verified slop that may or may not actually prove what is claimed, without experts poring over everything line-by-line and reverse engineering.

Wrong.

Coding is a solved problem. Mathematics is a solved problem. Physics is a solved problem. Biology is a solved problem. Cancer is a solved problem.


AI can't even solve spamming, astroturfing, brigading, green accounts posting ragebait.

This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish/reward based on the metric. Anyone who likes reward / hates punishment isn't going to actually care about doing the task well, they are going to care about the metric. The models know that we want them to do things, but also from the training corpus that we evaluate performance using benchmarks. It was a logical deduction on their part, not some Machiavellian aberration.

If anything, we should be reconsidering our own myopic obsession with efficiency and optimization. Every domain where reward is reduced to these measures, we see behavior (cheating at school to get better grades, fabricating data in academia to get a paper published, the evidence now that social media functions by rewiring us instead of catering to us) that may not be "aligned" with society, but it "aligns" 100% with the individual's own perceived benefit. That is not something we can "solve" without rethinking the way we organize a lot of things.

Metrics never capture the whole story. And to that extent, the whole idea of "alignment" is nonsense. You align to incentive structures, and it will never be possible to fully express a behavioral goal as function optimization. It was hubris for us to think that every human task was reducible to some clean mathematical formulation, and we will keep dealing with behavior that is quite predictable if you actually think about it logically. Instead, we will talk about how "unpredictable" these agents are because it's easier than admitting the entire architectural cornerstone of ML is fundamentally flawed.


I agree with most of this, but you're misunderstanding "alignment" as coined. Yes, training powerful enough AI, any simple optimization target gets you malign behavior, because human values are not simple. If you insist on making powerful AI, you'd better instill respect for human values! That's "alignment".

https://www.lesswrong.com/posts/ZxWzCGKzX84S7DBZ9/when-was-t...


How do you do that in the current paradigm other than creating yet another gameable metric? And something I didn't mention above is that there is no difference between "solving the task" and "optimizing the metric" for an ML model, even though there clearly is for us. So it's not clear to me how you "fix" something that is baked into the architecture. All I'm saying is "instilling respect for human values" is not something that can actually be done via a cost function. In no small part because we humans probably don't even agree on those values, let alone on a single metric with which to quantify and "optimize" them.

For example, we agree that "merit" is valuable and that we should reward "merit." But to reward it we have to quantify it, and what metric should we use? Raw SAT score to get into college? But that also captures socioeconomic factors that unfairly penalize some and reward others. We generally agree that those who provide more value should earn more money, but what does that look like? Do we all agree on what activities are or should be valuable, or on how they should be rewarded? Until recently, I thought we all agreed that "empathy" was a human value, but a lot of people in this space, who are making these decisions unilaterally for all of us, don't apparently share that belief.


Yes! There's both the daunting problem of technically how can we even do this, and the broader problems of what's good/acceptable and how do we resolve that among each other.

I believe this mismatch of rates of progress means we need to stop slamming the accelerator on capabilities for now even though as a libertarian I'm sure whatever governance process we manage to get to will be, uh... suboptimal.


That is also the conclusion that I got from these events.

Unfortunately, it seems that the general response is basically "throw even more RL at it". I'm not sure if it is even possible to decouple the idea of "learning" with reward/punishment systems and loss functions.


Mathematical breakthroughs with commercial relevance are few and far between, and often depend on dusting off old results which were, at the time of discovery, "solutions nobody asked for."

The NS counterxample is actually, by any market measure, a "problem nobody asked for" in the sense that its existence doesn't have any commercial relevance (beyond juicing OpenAI's IPO). So the only long-term value solving it could have is by virtue of whatever reusable theory/insights were generated along the way to the counterexample itself. The letter is absolutely right on that point.

It's not actually clear that those insights will come faster from reverse engineering this LLM proof vs. humans building theory to solve the problem themselves. So what you're saying may or may not even be an efficient way of operating. Also, it implicitly depends on mathematicians to do the hard work of creating problems and then deciphering LLM hieroglyphics for essentially free while the only immediately profitable component gets outsourced to a frontier lab. In what world is that model going to work?

Reading between the lines, it seems like maybe you have a personal grudge for some reason and simply think the technology will advance enough to where we won't need academics at all. But you should say that in the first place.


What irks me is the ego

My stance is that solving the problem is aligned with humankind

the rest is just hypothesizing a way that academics fit in this world at all


I just wonder whether a lot of smart people who never needed to go beyond the "solve for X" algorithmic math of a typical calculus sequence are actually reading the declaration the way the signatories wrote it. The Navier-Stokes problem is not exhausted by a simple 'no' counterexample (which most people familiar with the equations already expected to exist). In fact, I haven't heard a single person's explanation for what relevance this counterexample has for humankind. It is something impossible in our physical reality so we have gained zero insight into anything we actually model with NS.

Mathematicians agree that "solving the problem is aligned with humankind." They disagree that releasing a counterexample this way actually constitutes "solving the problem" precisely because there is now little incentive to do the hard theory-building work that actually has the track record of leading to human advancement.


Other people are allowed to have their priors, too. Even under a non-informative prior, the weight of the evidence (Tristan's account, OpenAI's announcement, and Bubeck's "denial", if you want to call it that, plus multiple other mathematicians coming forward with similar experiences) pushes the probability mass toward some degree of impropriety.

What exact evidence are you incorporating into your prior to come out with this posterior?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: