Hacker Newsnew | past | comments | ask | show | jobs | submit | more EncomLab's commentslogin

No one says that a thermostat is "thinking" of turning on the furnace, or that a nightlight is "thinking it is dark enough to turn the light on". You are just being obtuse.


Yes. A thermostat involves a change of state from A to B. A computer is the same: its state at t causes its state at t+1, which causes its state at t+2, and so on. Nothing else is going on. An LLM is no different: an LLM is simply a computer that is going through particular states.

Thought is not the same as a change of (brain) state. Thought is certainly associated with change of state, but can't be reduced to it. If thought could be reduced to change of state, then the validity/correctness/truth of a thought could be judged with reference to its associated brain state. Since this is impossible (you don't judge whether someone is right about a math problem or an empirical question by referring to the state of his neurology at a given point in time), it follows that an LLM can't think.


>Thought is certainly associated with change of state, but can't be reduced to it.

You can effectively reduce continuously dynamic systems to discreet steps. Sure, you can always say that the "magic" exists between the arbitrarily small steps, but from a practical POV there is no difference.

A transistor has a binary on or off. A neuron might have ~infinite~ levels of activation.

But in reality the ~infinite~ activation level can be perfectly modeled (for all intents and purposes), and computers have been doing this for decades now (maybe not with neurons, but equivalent systems). It might seem like an obvious answer, that there is special magic in analog systems that binary machines cannot access, but that is wholly untrue. Science and engineering have been extremely successful interfacing with the analog reality we live in, precisely because the digital/analog barrier isn't too big of a deal. Digital systems can do math, and math is capable of modeling analog systems, no problem.


It's not a question of discrete vs continuous, or digital vs analog. Everything I've said could also apply if a transistor could have infinite states.

Rather, the point is that the state of our brain is not the same as the content of our thoughts. They are associated with one another, but they're not the same. And the correctness of a thought can be judged only by reference to its content, not to its associated state. 2+2=4 is correct, and 2+2=5 is wrong; but we know this through looking at the content of these thoughts, not through looking at the neurological state.

But the state of the transistors (and other components) is all a computer has. There are no thoughts, no content, associated with these states.


It seems that the only barrier between brain state and thought contents is a proper measurement tool and decoder, no?

We can already do this at an extremely basic level, mapping brain states to thoughts. The paraplegic patient using their thoughts to move the mouse cursor or the neuroscientist mapping stress to brain patterns.

If I am understanding your position correctly, it seems that the differentiation between thoughts and brain states is a practical problem not a fundamental one. Ironically, LLMs have a very similar problem with it being very difficult to correlate model states with model outputs. [1]

[1]https://www.anthropic.com/research/mapping-mind-language-mod...


There is undoubtedly correlation between neurological state and thought content. But they are not the same thing. Even if, theoretically, one could map them perfectly (which I doubt is possible but it doesn't affect my point), they would remain entirely different things.

The thought that "2+2=4", or the thought "tiger", are not the same thing as the brain states that makes them up. A tiger, or the thought of a tiger, is different from the neurological state of a brain that is thinking about a tiger. And as stated before, we can't say that "2+2=4" is correct by referring to the brain state associated with it. We need to refer to the thought itself to do this. It is not a practical problem of mapping; it is that brain states and thoughts are two entirely different things, however much they may correlate, and whatever causal links may exist between them.

This is not the case for LLMs. Whatever problems we may have in recording the state of the CPUs/GPUs are entirely practical. There is no 'thought' in an LLM, just a state (or plurality of states). An LLM can't think about a tiger. It can only switch on LEDs on a screen in such a way that we associate the image/word with a tiger.


> The thought that "2+2=4", or the thought "tiger", are not the same thing as the brain states that makes them up.

Asserted without evidence. Yes, this does represent a long and occasionally distinguished line of thinking in cognitive science/philosophy of mind, but it is certainly not the only one, and some of the others categorically refute this.


Is it your contention that a tiger may be the same thing as a brain state?

It would seem to me that any coherent philosophy of mind must accept their being different as a datum; or conversely, any that implied their not being different would have to be false.

EDIT: my position has been held -- even taken as axiomatic -- by the vast majority of philosophers, from the pre-Socratics onwards, and into the 20th century. So it's not some idiosyncratic minority position.


Clearly there is a thing in the world that is a tiger independently of any brain state anywhere.

But the thought of a tiger may in fact be identical to a brain state (or it might not; at this point we do not know).


Given that a tiger is different from a brain state:

If I am thinking about a tiger, then what I am thinking about is not my brain state. So that which I am thinking about is different from (as in, cannot be identified with) my brain state.


> What I am thinking about is not my brain state

Obviously the thing you are thinking about is not the same as your thinking about it, nor the same as your brain state when thinking about it. Thinking about a thing is necessarily and definitionally distinct from the thing.

The question however is whether there is anything to "thinking about thing" other than the brain state you have when doing so. This is unknown at this time.


Earlier upthread, I said

>> the thought "tiger" [is] not the same thing as the brain state that makes [it] up.

To which you said

> Asserted without evidence.

This was in the context of my saying

>> There is undoubtedly correlation between neurological state and thought content. But they are not the same thing.

Now you say

> the thing you are thinking about is not the same as your thinking about it, nor the same as your brain state when thinking about it.

Are we at least agreed that the content of the thought "tiger" is not the same thing as the brain state that makes it up?

> The question however is whether there is anything to "thinking about thing" other than the brain state you have when doing so. This is unknown at this time.

If a tiger is distinct from a brain state, which I think we agree on, and if our thoughts are about real things such as tigers, which I assume we agree on, then how can there not be more to thought than the associated brain state?


> Are we at least agreed that the content of the thought "tiger" is not the same thing as the brain state that makes it up?

No. I don't agree that "the content of [a] thought" is something we can usefully talk about in this context.

Thoughts are subjective experiences, more or less identical to qualia. Thinking about a tiger is actually having the experience of thinking about a tiger, and this is purely subjective, like all qualia. The only question I can see worth asking about it is whether the experience of thinking about a tiger has some component to it that is not part of a fully described brain state.

> If a tiger is distinct from a brain state, which I think we agree on, and if our thoughts are about real things such as tigers,

We also have thoughts about unreal things. I don't see why such thoughts should be any different than the ones we have about real things.


>> If a tiger is distinct from a brain state, which I think we agree on, and if our thoughts are about real things such as tigers, which I assume we agree on, then how can there not be more to thought than the associated brain state?

> We also have thoughts about unreal things. I don't see why such thoughts should be any different than the ones we have about real things.

Let me rephrase then:

If a tiger is distinct from a brain state, which I think we agree on, and if our thoughts can be about real things such as tigers, which I assume we agree on, then how can there not be more to thought than the associated brain state?

A brain state does not refer to a tiger.


I realize I'm butting in on an old debate, but thinking about this caused me to come to conclusions which were interesting enough that I had to write them down somewhere.

I'd argue that rather than thoughts containing extra contents which don't exist in brain states, its more the case that brain states contain extra content which doesn't exist in thoughts. Specifically, I think that "thoughts" are a lossy abstraction that we use to reason about brain states and their resulting behaviors, since we can't directly observe brain states and reasoning about them would be very computationally intensive.

As far as I've seen, you have argued that thoughts "refer" to real things, and that thoughts can be "correct" or "incorrect" in some objective sense. I'll argue against the existence of a singular coherent concept of "referring", and also that thoughts can be useful without needing to be "correct" in some sense which brain states cannot participate in. I'll be assuming that something only exists if we can (at least in theory if not in practice) tie it back to observable behavior.

First, I'll argue that the "refers" relation is a pretty incoherent concept which sometimes happens to work. Let us think of a particular person who has a thought/brain state about a particular tiger in mind/brain. If the person has accurate enough information about the tiger, then they will recognize the tiger on sight, and may behave differently around that tiger than other tigers. I would say in this case that the person's thoughts refer to the tiger. This is the happy case where the "refers" relation is a useful aid to predicting other people's behavior.

Now let us say that the person believes that the tiger ate their mother, and that the tiger has distinctive red stripes. However, let it be the case that the person's mother was eaten by a tiger, but that tiger did not have red stripes. Separately, there does exist a singular tiger in the world which does have red stripes. Which tiger does the thought "a tiger with red stripes ate my mother" refer to?

I think it's obvious that this thought doesn't coherently refer to any tiger. However, that doesn't prevent the thought from affecting the person's behavior. Perhaps the person's next thought is to "take revenge on the tiger that killed my mother". The person then hunts down and kills the tiger with the red stripes. We might be tempted to believe that this thought refers to the mother killing tiger, but the person has acted as though it referred to the red striped tiger. However, it would be difficult to say that the thought refers to the red striped tiger either, since the person might not kill the red striped tiger if they happen to learn said tiger has an alibi. Hopefully this is sufficient to show that the "refers" relationship isn't particularly connected to observable behavior in many cases where it seems like it should be. The connection would exist if everyone had accurate and complete information about everything, but that is certainly not the world we live in.

I can't prove that the world is fully mechanical, but if we assume that it is, then all of the above behavior could in theory be predicted by just knowing the state of the world (including brain states but not thoughts) and stepping a simulation forward. Thus the concept of a brain state is more helpful to predicting their behavior than thoughts with a singular concept of "refers". We might be able to split the concept of "referring" up into other concepts for greater predictive accuracy, but I don't see how this accuracy could ever be greater than just knowing the brain state. Thus if we could directly observe brain states and had unlimited computational power, we probably wouldn't bother with the concept of a "thought".

Now then, on to the subject of correctness. I'd argue that thoughts can be useful without needing a central concept of correctness. The mechanism is the very category theory like concept of considering all things only in terms of how they relate to other things, and then finding other (possibly abstract) objects which have the same set of relationships.

For concreteness, let us say that we have piles of apples and are trying to figure out how many people we can feed. Let us say that today we have two piles each consisting of two apples. Yesterday we had a pile of four apples and could feed two people. The field of appleology is quite new, so we might want to find some abstract objects in the field of math which have the same relationship. Cutting edge appleology research shows that as far as hungry people are concerned, apple piles can be represented with natural numbers, and taking two apple piles and combining them results in a pile equivalent to adding the natural numbers associated with the piles being combined. We are short on time, so rather than combining the piles, we just think about the associated natural numbers (2 and 2), and add them (4) to figure out that we can feed two people today. Thus the equation (2+2=4) was useful because pile 1 combined with pile 2 is related to yesterday pile in the same way that 2 + 2 relates to 4.

Math is "correct" only in so far as it is consistent. That is, if you can arrive at a result using two different methods, you should find that the result is the same regardless of the method chosen. Similarly, reality is always consistent, because assuming that your behavior hasn't affected the situation, (and what is considered the situation doesn't include your brain state) it doesn't matter how or even if you reason about the situation, the situation just is what it is. So the reason math is useful is because you can find abstract objects (like numbers) which relate to each other in the same way as parts of reality (like piles of apples). By choosing a conventional math, we save ourselves the trouble of having to reason about some set of relationships all over again every time that set of relationships occurs. Instead we simply map the objects to objects in the conventional math which are related in the same manner. However, there is no singular "correct" math, as can be shown by the fact that mathematics can be defined in terms of set theory + first order logic, type theory, or category theory. Even an inconsistent math such as set theory before Russell's Paradox can still often produce useful results as long one's line of reasoning doesn't happen to trip on the inconsistency. However, tripping on an inconsistency will produce a set of relationships which cannot exist in the real world, which gives us a reason to think of consistent maths as being "correct". Consistent maths certainly are more useful.

Brain states can also participate in this model of correctness though. Brain states are related to each other, and if these relationships are the same as the relationships between external objects, then the relationships can be used to predict events occurring in the world. One can think of math and logic as mechanisms to form brain states with the consistent relationships needed to accurately model the world. As with math though, even inconsistent relationships can be fine as long as those inconsistencies aren't involved in reasoning about a thing, or predicting a thing isn't the point (take scapegoating for instance).

Sorry for the ramble. I'll summarize:

TL;DR: Thoughts don't contain "refers" and "correctness" relationships in any sense that brain states can't. The concept of "refers" is only usable to predict behavior if people have accurate and complete information about the things they are thinking about. However, brain states predict behavior regardless of how accurate or complete the information the person has is. The concept of "correctness" in math/logic really just means that the relationship between mathematical objects is consistent. We want this because the relationships between parts of reality seem to be consistent, and so if we desire the ability to predict things using abstract objects, the relationships between abstract objects must be consistent as well. However, brain states can also have consistent patterns of relationships, and so can be correct in the same sense.


Thanks for the response. I don't know if I'll have time to respond, I may, but in any case it's always good to write one's thoughts down.


Does a picture of a tiger or a tiger (to follow your sleight of hand) on a hard drive then count as a thought?


No. One is paint on canvas, and the other is part of a causal chain that makes LEDs light up in a certain way. Neither the painting nor the computer have thoughts about a tiger in the way we do. It is the human mind that makes the link between picture and real tiger (whether on canvas or on a screen).


>Rather, the point is that the state of our brain is not the same as the content of our thoughts.

Based on what exactly ? This is just an assertion. One that doesn't seem to have much in the way of evidence. 'It's not the same trust me bro' is the thesis of your argument. Not very compelling.


It's not difficult. When you think about a tiger, you are not thinking about the brain state associated with said thought. A tiger is different from a brain state.

We can safely generalize, and say the content of a thought is different from its associated brain state.

Also, as I said

>> The correctness of a thought can be judged only by reference to its content, not to its associated state. 2+2=4 is correct, and 2+2=5 is wrong; but we know this through looking at the content of these thoughts, not through looking at the neurological state.

This implies that state != content.


>It's not difficult. When you think about a tiger, you are not thinking about the brain state associated with said thought. A tiger is different from a brain state. We can safely generalize, and say the content of a thought is different from its associated brain state.

Just because you are not thinking about a brain state when you think about a tiger does not mean that your thought is not a brain state.

Just because the experience of thinking about X doesn't feel like the experience of thinking about Y (or doesn't feel like the physical process Z), it doesn't logically follow that the mental event of thinking about X isn't identical to or constituted by the physical process Z. For example, seeing the color red doesn't feel like processing photons of a specific wavelength with cone cells and neural pathways, but that doesn't mean the latter isn't the physical basis of the former.

>> The correctness of a thought can be judged only by reference to its content, not to its associated state. 2+2=4 is correct, and 2+2=5 is wrong; but we know this through looking at the content of these thoughts, not through looking at the neurological state. This implies that state != content.

Just because our current method of verification focuses on content doesn't logically prove that the content isn't ultimately realized by or identical to a physical state. It only proves that analyzing the state is not our current practical method for judging mathematical correctness.

We judge if a computer program produced the correct output by looking at the output on the screen (content), not usually by analyzing the exact pattern of voltages in the transistors (state). This doesn't mean the output isn't ultimately produced by, and dependent upon, those physical states. Our method of verification doesn't negate the underlying physical reality.

When you evaluate "2+2=4", your brain is undergoing a sequence of states that correspond to accessing the representations of "2", "+", "=", applying the learned rule (also represented physically), and arriving at the representation of "4". The process of evaluation operates on the represented content, but the entire process, including the representation of content and rules, is a physical neural process (a sequence of brain states).


> Just because you are not thinking about a brain state when you think about a tiger does not mean that your thought is not a brain state.

> It doesn't logically follow that the mental event of thinking about X isn't identical to or constituted by the physical process Z.

That's logically sound insofar as it goes. But firstly, the existence of a brain state for a given thought is, obviously, not proof that a thought is a brain state. Secondly, if you say that a thought about a tiger is a brain state, and nothing more than a brain state, then you have the problem of explaining how it is that your thought is about a tiger at all. It is the content of a thought that makes it be about reality; it is the content of a thought about a tiger that makes it be about a tiger. If you declare that a thought is its state, then it can't be about a tiger.

You can't equate content with state, and nor can you make content be reducible to state, without absurdity. The first implies that a tiger is the same as a brain state; the second implies that you're not really thinking about a tiger at all.

Similarly for arithmetic. It is only the content of a thought about arithmetic that makes it be right or wrong. It is our ideas of "2", "+", and so on, that make the sum right or wrong. The brain states have nothing to do with it. If you want to declare that content is state, and nothing more than state, then you have no way of saying the one sum is right, and the other is wrong.


Please, take the pencil and draw the line between thinking and non-thinking systems. Hell I'll even take a line drawn between thinking and non-thinking organisms if you have some kind of bias towards sodium channel logic over silicon trace logic. Good luck.


Even if you can't define the exact point that A becomes not-A, it doesn't follow that there is no distinction between the two. Nor does it follow that we can't know the difference. That's a pretty classic fallacy.

For example, you can't name the exact time that day becomes night, but it doesn't follow that there is no distinction.

A bunch of transistors being switched on and off, no matter how many there are, is no more an example of thinking than a single thermostat being switched on and off. OTOH, if we can't think, then this conversation and everything you're saying and "thinking" is meaningless.

So even without a complete definition of thought, we can see that there is a distinction.


> For example, you can't name the exact time that day becomes night, but it doesn't follow that there is no distinction.

There is actually a very detailed set of definitions of the multiple stages of twilight, including the last one which defines the onset of what everyone would agree is "night".

The fact that a phenomena shows a continuum by some metric does not mean that it is not possible to identify and label points along that continuum and attach meaning to them.


Looks like we replied to each others comments at the same time, haha


Your assertion that sodium channel logic and silicon trace logic are 100% identical is the primary problem. It's like claiming that a hydraulic cylinder and a bicep are 100% equivalent because they both lift things - they are not the same in any way.


People chronically get stuck in this pit. Math is substrate independent. If the process is physical (i.e. doesn't draw on magic) then it can be expressed with mathematics. If it can be expressed with mathematics, anything that does math can compute it.

The math is putting the crate up on the rack. The crate doesn't act any different based on how it got up there.


Or submarines swim ;)


think about it more


The entire paper is riddled with anthropomorphic terms - it's part of AI culture unfortunately. When they start talking about "planning", "choosing", "reasoning" it biases the perception of their analysis. One could certainly talk about a night light equipped with a photoresistor as "planning to turn on the light when it is dark", "choosing to turn on the light because it is dark, and "reasoning that since it is dark, it turned on the light"- but is that accurate?


I agree. "Planning" means we come up with alternative sets of steps or tasks which we then order into sequences or acyclic directed graphs and then pick the plan we think is the best. We can also create "Plan B" and "Plan C" for the cases that the main plan fails to execute successfully.

But as far as we know does AI internally assemeble subtasks into graphs and then evaluate them and pick the best one?

Is there any evidence in the memory traces of the executing AI that there are tasks and sub-tasks and ordering and evaluating of them, then taking a decision to choose and EXECUTE the best plan?

Where is the evidence that AI-programs do "planning"?


I love this analogy.


This is very interesting - but like all of these discussions it sidesteps the issues of abstractions, compilation, and execution. It's fine to say things like "aren't programmed directly by humans", but the abstracted code is not the program that is running - the compiled code is - and that is code is executing within the tightly bounded constraints of the ISA it is being executed in.

Really this is all so much slight of hand - as an esolang fanatic this all feels very familiar. Most people can't look a program written in Whitespace and figure it out either, but once compiled it is just like every other program as far as the processor is concerned. LLM's are no different.


And DNA? You are running on an instruction set of four symbols at the end of the day but that's the wrong level of abstraction to talk about your humanity, isn't it?


DNA is the instruction set for protein composition- not thinking.


Why? For most games (especially a "2d hide and seek") you drop a zip file and press a single button on your dev console to publish.


> For most games (especially a "2d hide and seek") you drop a zip file and press a single button on your dev console to publish

If you read through the actual pipeline, you'll see it's not just "drop a zip file and press a single button". Firstly, you're missing everything that goes before even having a ZIP file, making that process reproducible is valuable regardless of what you do later. Secondly, doing this for three platforms would mean repeating the same thing three times, in slightly different ways. Automating it just makes sense at that point.

Overall, building and publishing a game via automation just brings about every benefit from CI/CD, just to a "game development" context instead, so you'd do it for the same reason you'd automate any software release process.


If you automate your release process:

1. You won't forget how to ship a release if you go months between releases

2. You won't make mistakes when you release code - forgetting a crucial step along the way for example

3. Related: you can add tests to your release process - so your release doesn't go out if you made some last-minute mistake that broke the build

4. You'll release more often: the automation has de-risked your release process and reduced friction around it, which means you can ship with more confidence and less ceremony

5. You can reliably share that release process with other collaborators - not end up in a situation where only one person's laptop is able to ship

6. If your laptop breaks or gets stolen it won't harm your ability to release software

All of the above are true for non-game projects, I don't see why they shouldn't apply to games as well.


And that zip file might be not the version you thought it was, if it's not from a build server it might include any number of uncommitted or gitignored changes, more transfer steps mean more things that can go wrong. These automations are not about reducing keypresses, they are about reducing the number of trivialities to mess up. People tend to have other things on their mind when going through that kind of routine.


I recently heard some salty chap here on HN use the term "ClickOps" as opposed to "DevOps" or "GitOps" to describe the "just click on the button in the dev console web page to do that" approach.

https://www.wiechtig.com/blog/clickops-is-the-worst/

https://www.lastweekinaws.com/blog/clickops/

https://blog.equinix.com/blog/2022/12/01/what-is-clickops-an...


The implication that any software is "mysterious" is problematic - there is no "woo" here - the exact state of the machine running the software may be determined at every cycle. The exact instruction and the data it executed with may be precisely determined, as can the next instruction. The entire mythos of any software being a "black box" is just so much advertising jargon, perpetuated by tech bros who want to believe they are part of some Mr. Robot self-styled priestly class.


You're misunderstanding. A level of abstraction is necessary for operation of modern systems. There is no human alive who, given an intermediate step in the middle of some running learning algorithm, is able to understand and mentally model the full system at full man-made resolution, that is, down to the transistor level, on a modern CPU. Someone wishing to understand a piece of software in 2025 is forced to, at some point, accept that something somewhere "does what it says on the tin" and model it thusly rather than having a full understanding.


It's not misunderstanding at all - but your response is certainly an attempt to obfuscate the point being made. The moment you represent anything in code, you are abstracting a real thing into it's digital representation. That digital representation if fully formed at every cycle of the digital system processing it, and the state of the system - all the way down to the transistor level may be precisely determined. To say otherwise is to make the same error as those who claim that consciousness or understanding are indefinable "extra-ordinary" things that we have to just accept exist without any justification or evidence.


Okay, then, you're just using your own personal definition of "black box" instead of the one everyone else uses.

Something that's a black box is unknown to the speaker. It's not understood to be unknowable to anyone.


So your claim is that there are instructions, data, or both that are unable to be determined in what, is by definition, a fully deterministic machine?


By an individual person, yes. I claim that there exists no single human capable of fully understanding the totality of the software and hardware down to the individual transistor level.


That's a very wrong statement. Pretty sure I could explain all the maths, all the physics, all the electronics, all the operating systems and all the user space of a single high level language operation, when I was a fresh graduate. Now, I have forgotten most of the physics and electronics, since the university was quite some time ago, but feel free to ask any decent student of an IT bachelor, they should be able to pretty much build the PC from scratch. Sure, modern processors and whatnot add a bunch of optimizations, but you seem to really overstate the complexity of the computer.


We're talking about two separate things.

I'm talking about understanding, fully, the state of the CPU. Not just the conceptual operation of the CPU. Like, given a specific, modern AMD or Intel CPU, understand fully all states of all transistors.


I agree and never claimed that "a single person" could - but just because something is too complex for a single person to fully understand does not make it "mysterious" or a "black box". So what is the claim you are making? Anything beyond the complexity of a single person to understand = magic?


We're just using different definitions for "black box".

My definition is that it's something unknown, yours is that it's something unknowable.


The mystery was never in the "how do computers calculate the probabilities of next tokens" but rather in the "why is it able to work so well" and "what does this individual neuron contribute to the whole model"


The mystery is in how the data is encoded in the parameters and why LLMs performance scales so well with parameters. The key seems to be almost orthogonal vectors that allow neural networks to store so much data. They allow 2^(cn) vectors to be learned in an n-dimensional space with c being a constant.Since almost orthogonal vectors have very small dot products, they minimally interfere with each other, allowing many concepts to coexist with limited cross talk which enables superposition


I don't know any serious programmer who thinks that, just because each operation is simple, the operation of the whole thing can't be mysterious.


But the weights trained from machine learning are a black box, in the sense that no human designed e.g. the image processing kernels that those weights represent.

That is one reason people are skeptical of them, not only is training a large model at home expensive, not only is the data too big to trivially store, but the weights are not trivial to debug either


uLisp is amazing - capable of leveraging a lot of work in a small package.


This is the "Wozniak Standard" (sometimes called the Coffee Test) - Drop an AI enabled robot in front of a random house and ask it to bring you a cup of coffee. The robot would need to enter the house, locate the kitchen, locate the coffee machine, locate the coffee, locate the filters, locate the coffee mugs, locate a measuring spoon - then add the correct amount of water, the filter, the correct amount of coffee, start the brew cycle, wait for the brew cycle to finish, pour your coffee, then exit the house and deliver the mug to you. Extra points for adding cream and sugar.


I like that test. If the AI couldn't find a measuring spoon it would need to grab any spoon it could and just "eyeball it". Also, if there wasn't an actual mug then maybe a glass will work (but not a pint glass) and certainly not a plastic cup. When delivered it would have to know to say "couldn't find a mug so i grabbed a glass". there's other things too, can't find regular coffee but it found some instant coffee? The AI would need to decide if that will work or should it ask first. All of those things are petty easy for a human.


That’s a very high standard. I’d fail repeatedly.


It is not a high standard, I am sure you could train a chimp to pass this test[1]. If you know how to use a standard coffee maker and live in a typical American home, and the test is done in an typical American home with a standard coffee maker, you can definitely pass this test 100% of the time.

I understand that many people don't live in America and don't know how to use a coffee maker. That is 100% irrelevant. There is a frustrating tendency in AI circles to conflate domain knowledge with intelligence, in a way that invariably elevates AI and crushes human intelligence into something tiny.

[1] The hard part would be psychological (e.g. keeping the chimp focused), not cognitive. And of course the chimp would need to bring a chimp-sized ladder... It would be an unlawful experiment, but I suspect if you trained a chimp to use a specific coffee maker in another kitchen, forced the chimp to become addicted to coffee, and then put the animal in a totally different kitchen with a different coffee maker (but similar, i.e. not a French press), it figure would figure out what to do.


"locate the filters, locate the coffee mugs, locate a measuring spoon" in a random house in America is a very high standard. We’ll have to agree to disagree on that. If you teleport me into a random house, I’ll likely spend at least an hour trying and failing at that task, and most of their cabinets and drawers will be open by the end of it.

It also excludes corner cases like "what if they don’t have any filters"? Should the robot go tearing through the house till they find one, or do nothing? But what if there were some in the pantry — does that fail the test? There’s all kinds of implicit assumptions here that make it quite hard.


and what if there's only a Nespresso machine, a Keurig machine, instant, a french press, a moka pot, or a cappuccino machine (we can argue if an americano is actually coffee, but if that's what the house has, and no drip machine + accoutrements, you're not getting anything else)? Human or bot, that's a lot of possibilities to deal with, but for a bold human unfamiliar with those, they're just a YouTube video away (multiple ones if it's a fancy cappuccino machine). Until AI can learn to make coffee or change an oil filter on a 1997 GMC from watching a YouTube video, it'd be hard to consider it human-grade, even if it has been trained on all of YouTube, which assumedly Google has done. There are certainly things people do on YouTube that I couldn't do after a lot of intense practice, though, so I'm not totally convinced that's the right standard. It doesn't cost millions of hours and dollars of training and fine tuning time for me to, say, be able to tie a bow tie from a YouTube video though, even if it does take me a couple of tries.


It probably shouldn't continue to surprise me how often people's "AI benchmarks" exclude a significant fraction of actual, living, humans from being "human-grade".


You can't honestly claim that it would take you an hour to accomplish such a high probability task - have you never visited the house of a friend or family and had to open a few cabinets to find a water glass or a bowl or a spoon?

As for the point of corner cases being hard - I mean that's the point here, isn't it?


You might refuse to do it, but I doubt you'd ever actually completely fail it. If someone offered to pay $10 million if someone could go into a house, make a cup of coffee, and come back out with it, I imagine just about any functional adult would figure out a way to return with a cup of some sort of liquid resembling coffee. I don't see anyone saying, "Sorry, making a cup of coffee is too difficult, I'm going to forfeit the $10 million."

But sure, without proper compensation a lot of people would probably just say "I can't do it" as a way of avoiding the task.


Repeatedly? As in you would come back and tell whoever you're with "I gave up"? Like I can understand wanting to ask for e.g. "where do you keep the coffee", but if that wasn't possible -- say the host is asleep, and I'm there taking care of them -- I would certainly be able to figure it out. Just open cabinets and peek / carefully rummage around until you find what you need.


It would just scan for the MAC of the WIFI-enabled Keurig, home in to that, and then ping the RFIDs of the capsules, and grab one of those.

Wasn't that easy?


Sounds like you rediscovered the long held practice of "Rubber Ducking".


Haha, that’s a good point. I think the main change from using these (reasoning) models is that I’m more cognizant of my thinking process, rather than there being a novel technique.


eh, writing it out is closer to "proto design document" or just plain ol "whiteboarding"


the only reason I haven't written a design doc or busted out the felt tip for my rubber duck is because it can't read.


Have you tried? Maybe it can.


I mean, I'm sure there's some founder out there pitching an AI-powered rubber duck dev productivity tool.

At the very least, someone at Copilot must have pitched a rubber duck avatar as the new Clippy by now...


But now the rubber duck can talk back and also on occasion hallucinate and lie to you to confirm your delusions.


I guess I'll stop using this really useful tool because someone else is incompetent.


One of my co-workers joked at the time that "sure AlphaGO beat Lee Sedol at GO, but Lee has a much better self-driving algorithm."

I thought this was funny at the time, but I think as more time passes it does highlight the stark gulf that exists between the capability of the most advanced AI systems and what we expect as "normal competency" from the most average person.


> it does highlight the stark gulf that exists between the capability of the most advanced AI systems and what we expect as "normal competency" from the most average person

Yes, but now we're at the point where we can compare AI to a person, whereas five years ago the gap was so big that that was just unthinkable.


I mean people thought ELIZA was AI back in the 1960's. Everyone always thinks "this is it!!".


> people thought ELIZA

But which people? Those people which show that a supplement of extra intelligence, also synthetic, is sought.


It was. The definition of "AI" keeps shifting.


Love me some good old whataboutism (sure, LLMs are now super-intelligent at writing software, but can they clean my kitchen? No? Ha!)


The computer beat me at chess, but it was no match for me at kickboxing. - Emo Phillips

Tale as old as time. We can make nice software systems but general purpose AI / Agents isn't here yet.


Worse than that: it seems that it's much easier to make computer achieve superhuman feats in cognitive work, than it is to make it do even most basic physical interactions with the real world.

In short: the natural order of things is that computers are better at thinking, and people are better at manual labor. Which is the opposite of what we wanted.


AI is just hydraulics for the mind. Or should be.

I choose a direction and apply force.


i.e. the "geohot method". Why this guy still pulls so much traction is a mystery.


lotta good cracks back in the day

some of his "watch me program" vids are also pretty popular. I learned a lot from watching his workflow.

but I'd never trust him when talking about economics or politics, or even where to go for dinner.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: