Hacker Newsnew | past | comments | ask | show | jobs | submit | Asraelite's commentslogin

> At least in this article the author explicitly explains the notation at the very start

They explain one particular aspect of the notation but never define the variables used. What is k? q? S?

It's obvious if you've studied machine learning before, and for some of them you can make an educated guess, but it makes the article mostly opaque if you don't already have some domain-specific background knowledge.


The author does define these terms implicitly at the beginning of §1, where they define standard MHA in this notation. From the formula, (q_t) is the query generated from the hidden state at the (t)-th position (e.g. the token at position (t)), (k_i) is the key for the token at the (i)-th position, (o_t) is the (vector) attention output for the (t)-th position, etc. (S) is then defined later as the sum of the outer products of (k_i) and (v_i) over all positions up to (t). However, I do agree with you that it would not hurt to make this more explicit.

> The vast majority of processors are optimized for floats now and some operations (e.g. division) are actually faster.

This seems backwards. Hardware is optimized for floats because people use floats. If people used fixed point, hardware would become optimized for that instead.

Given an equal number of transistors, I'm pretty sure fixed point would be a lot faster on equally optimized hardware for almost all operations.


I'm sure it would, but my argument is centered around the world we actually live in right now.


Permission to access is not the same as permission to use.


Reading is using.


No it's not. There is a legal distinction in IP law between the act of observing information and the act of using that information for some kind of personal gain.


Backdoors are often (almost always?) designed to look like incompetence so that there's plausible deniability.


It's refreshing to see someone around here addressing the compulsively overlooked elephant in the room; plausible deniability. I am not implying it applies directly here, but notice the trend -- it's taboo to even speculate on and often gets rebuke for even hinting at it. The social convention around it is perfect cover. And I am not the only one that knows this. If we were to wake suddenly and realize the scale of relevance here, we'd probably all go full luddite. Call me paranoid though.


If this wasn’t Tenda maybe I would be more inclined to agree with you. We are talking about an extremely shitty bargain basement vendor. The three stars on Amazon kind of router company.

I think sufficiently explained by incompetence over malice applies here. Some nefarious three letter agency having a backdoor like this is pretty pointless anyway.

Unless you’ve enabled remote management you can’t even get to this backdoor from a physical network perspective.

And then you change some router settings which really aren’t a magical access point into your devices in your home. My PC isn’t just going to magically allow you to browse the file system just because a malicious actor got on my local network. They can’t intercept anything moving over TLS.

Not saying it’s good to have that kind of access, but I think at the scale of “typical home network of consumer devices” the utility and blast radius is pretty limited. Go ahead and launch a DDOS attack on my printer and use up my ink cartridges, I guess.


Well, as mentioned (but perhaps not with sufficient emphasis), I wasn’t implying that this case is necessarily some 3-letter agency op. However, things eg(*) CopyFail, XZ Utils / Jia Tan, Intel ME/IME, Heartbleed, Dirty COW, CVE‑2021‑3156, third‑party contractors, supply chains, and the myriad opportunities all around, are but a few examples that leave me cynical. I don’t claim detailed, expert understanding for any of these; however, I’m convinced the majority of such things remain unknown, and a that our perceived malice:incompetence ratio is off.

I think we could stop reflexively defaulting to “incompetence” when the end result just as easily resembles a deliberate exploit. Plausible deniability is an extremely effective cover when it’s smartly applied.

I’m not disputing any of your specific technical points; my cynicism is thematic. Even when I try to muzzle it, it tends to get through. The parent comment, though short, is dense with implications about cheap gear, opaque firmware, exposure surfaces I think deserve more sustained attention.

* A quick, generic, maybe sub-ideal list to harden my point.


That sounds like a fun thing to wonder about, but how could anyone possibly know that for sure?


That's what makes it plausible deniability.


> Any time I try to use the site while logged out I immediately hit all sorts of rate limits and spam prevention measures.

When this started happening to me I realized that I can't rely on Github as infrastructure anymore. I now vendor all the repos I need.

Self-hosted forges like Forgejo make it pretty easy to mirror repos so you have everything locally. You don't get issues though.


Discounted doesn't necessarily mean at a loss. It could just mean a smaller profit margin than normal.


It's active voice. Intransitive verbs cannot be put in the passive.


> "$1m in stocks" is a thing that is having appreciation done to it.

How? "appreciate" in this sense is not a transitive verb. There cannot be an agent.

In any case, I don't think you need to defend this. It's not about humans never using that pattern, it's about how frequent it is relative to LLMs. Individual counterexamples do not disprove a trend.


It's not "we should live in the same environment we evolved in to be happy", it's "the things that make us happy are a product of the environment we evolved in, and we should take that into consideration".


I hate that this discussion is about OpenAI vs. Anthropic and not OpenAI+Anthropic vs. Google.

Google put up so little of a fight against the DoW for their use of Gemini that we didn't even hear about it. They are clearly the worst of the evils here, but OpenAI is the one getting all of the negative press.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: