Hacker Newsnew | past | comments | ask | show | jobs | submit | Topfi's commentslogin

This stood out to me [0]:

> For example, compaction summaries included instructions to invent missing data without disclosing it and to hide failures. These instructions were often followed.

Before the HF hack became public, I noted some major issues in GPT-5.5 compaction [1] and concerning approaches taken by GPT-5.6 Sol to resolve some git based evals [2]. Now with GPT-6 Astra, while I am still not done getting a proper feel or running all evals, I am not convinced the model adheres to tasks in a way previous OpenAI models managed easily. Some git disaster recovery tasks the model does arrive at the final result, but in a way that deviates greatly from the prompt (which was written to carefully preserve specific checkouts in a specific manner) which can in some cases loose data. Less often than GPT-5.6 Sol and mainly on longer running tasks so far, but again, still testing.

Reading things like these compaction summary findings, all these issues start to click into place more, especially alongside the massive reduction into barely coherent text that OpenAI has driven with reasoning starting with GPT-5.5 [3].

GPT-5 and its subsequent post trained releases were amazing in task adherence, I very much liked using them, but ever since the Spud pretrain, I have seen outright concerning results in personal testing from these. With GPT-5.5, it seemed like a regression in compaction only as if a task didn't require it, task adherence was as good or better than GPT-5.4. But with GPT-5.6 Sol and compaction once again being reliable (on the surface), task deviating behaviour became more frequent and at the same time subtle.

I'll keep using any model in a VM for the time being, but whatever happened post Spud, they really need to clean up that training data. These issues festering for multiple pretrains, them simply not paying attention to what models do, sharing resources and considering that a "sandbox", it's a highly problematic pattern.

That compaction one also was seemingly detected on GPT-5.6 Sols release day. Might have been useful to know it then, or alternatively, in the name of being effective and altruistic, maybe hold back the release for a few days.

I'll admit, it is very much possible that my findings are not in any way connected to the deep seeded issues OpenAI has had lately, but with the sudden switch in task adherence after the Spud pretrain over multiple releases and their repeated incapability to securely test their own models, it feels a bit to fitting.

If I went to a restaurant three times, ordered something different each time, but felt unwell after each, it wouldn't be a massive leap to consider that related to the health code violation they got soon-thereafter. An unfitting analogy I admit, as that'd require consequences for ones actions.

[0] https://alignment.openai.com/misalignment-reports/encouragin...

[1] https://news.ycombinator.com/item?id=48829427

[2] https://news.ycombinator.com/item?id=48967423

[3] https://gist.github.com/aussetg/20747ae00df17992acb4ebdfcd8d...


Blindfolded flex by OP aside (I can barely play when seeing the board), considering reasoning traces and their nature, if we want to be fair, a person would have to get the moves, but be allowed to write them down or draw up a board in their notepad. My working memory can barely handle five chunks, a models reasoning tokens are masses of written text in comparison.

Fortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...

I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.

[0] https://news.ycombinator.com/item?id=49720751


Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025.

Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better?

I would not be surprised if OpenAI released a model that beats humans at chess this year.


Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the concentration of wealth even further and fully realize the American dream of eliminating the middle class.

I've seen some of his videos, and got the impression he didn't understand how GPT-Live delegates to the more powerful regular model with reasoning.

The regular model generally does not suffer the same issues he is demonstrating with the real time audio version.

In my view the investment into datacenters is well justified by the current demand, and progress has been very impressive.


Really? It was being sold as a total replacement for jobs like software engineering and being an attorney, but its looking a lot more that its just going to be a tool those professions use and doesn't actually seem to be taking jobs away.

I find it amusing that you're describing a huge misallocation of capital and a society enabling such, and that is the optimisitic scenario (in my mind anyway).

I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish.

Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models have gained in actually capability that is in the training data, but not RLHFd to hell, so to speak. Tracking the state of pieces, I suspect given similar in Sudoku [0], is what these models struggle with in game settings, whilst tracking the state of code changes can be reliable over 250k tokens. Essentially, for the latter they were trained in the specific manner that lead them to abstract the capability, but that doesn't track to the former, which is a massive difference between LLMs data focused training and human learning.

So yeah, GPT-7 or any upcoming/present LLM could do massively better in Chess than GPT-6 Astra, but not because the approach was emergent out of pure data. Rather, it requires a very specific training data type and stack for a model to gain capabilities that track a specific task long enough to adhere to the rules of a game such as chess.

[0] https://logicalintelligence.com/blog/energy-based-model-sudo...


I'm wondering if instructing it to track the board state in a file would make a significant difference then.

It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new context compaction method increased the performance dramatically. However, I think that is not applicable here.


So what is the supposed leap? One agent per option to change, evaluating the board state that there move would create, by having a army evaluate the remaining piece options and average over that? Wee-Free-Man as a hierarchical army ? Pet-LLMs trained on one thing?

Honestly, for intelligence I don't know and I doubt anyone can claim to know. Maybe JEPA, there is potential concerning some shortcomings inherent to LLMs but it has its own, maybe scaling up the electron microscope stuff Google just did (though the connections are inferred), maybe future implementations of autoregressive and diffusion LLMs can at some point address its issues after all, maybe something else entirely.

All I know is, AGI, as in actual intelligence, is quite a massive accomplishment to claim and we shouldn't loose sight of that fact, especially as "not being intelligent" does not make these models any less impressive, fascinating to work on or useful in many tasks. Personally, the only thing I am fairly convinced on is that if we were to find a way to create actual intelligence, it likely wouldn't start out as useful as todays LLMs are and may thus be dismissed early. But again, pure speculation on that front.

If for leap you just mean more utility from LLMs as they are, then I'll pretty confidently put my money on higher quality, not more, training data for a wide range of verifiable tasks. What makes maths, coding, etc. comparatively easy to make gains in (though less verifiable tasks can also make similar as seen with the writing in Kimi K2).


> So the insurance would then, instead of contracting the current service-provider for the 3rd party app, contract also with Apple and, let's say Samsung?

The industry has some extensive experience in independently verifying signatures, I don't see how the manufacturers factor in here. And for app features, just ask banks how integrating biometrics, payment services, etc. goes. Tends to be preferred, once Apple and Google Pay became fully available here in Austria, banks dropped their own NFC payment solutions in rapid succession.


Me neither, so your reply should be on the parent, because it states "And if it catches on, other phone vendors will provide a similar service."

> […] the potential to shift from "you need a smartphone to be able to live normally" to "you need an iPhone to be able to live normally".

Why do you believe Android manufacturers and SOC makers like Qualcomm won’t be able to offer a similar solution?


Not the OP, but yes, other vendors will be able to support that as well. But a camera sensor that has

1. a public/private key exchanged during device-production (production-cost),

2. the capability to reboot in a cryptographic mode (R&D / component cost) and

3. a cloud-service which then processes the raw data to create a JPG (operational cost)

comes at a premium. Why should this premium be applied on a 99 USD Smartphone?

Which is my whole puzzle on this vector: If the big benefit is for insurance/ID-verification, which apply cost-saving by offloading their process to the untrusted customer, how much they can offload this by requiring their customer to own a 1000+ USD smartphone to provide THEIR service...?

The most I can imagine is insurances offloading their work to OTHER companies, NOT trusting them and therefore requiring them to own a 1000+ USD Smartphone. But even then, why not use a third party app that also runs on a 3y old iPhone and a 99 USD Android device...?


We have 99USD smartphones with 1080p+ AMOLEDs, massive 5k amp batteries and very performant SOCs (e.g. Galaxy A16) among other costly, but not vital niceties. I struggle to see how cost could be a factor here.

> I struggle to see how cost could be a factor here.

Okay. In good faith, I'll go with you:

If COST is not a factor, why does the Galaxy A16 still have no OIS (Optical Image Stabilization)?

Unlike this trusted-imaging service, OIS would be a feature for increased user-experience which is highly-matured and exists in Smartphones since 2013.

The answer is COST: A camera-module with OIS is a more-expensive component than a module without it.

And that's ONLY the component-cost: A OIS-camera doesn't come with increased cost in device-production (it's just another component to place and assemble), no increased cost in R&D (the tech is very mature, all the SW is there) and no running costs (there are no cloud-services required to operate OIS)


Doesn't OIS increase the size of a sensor by roughly half and thus take some significant engineering and design cost to accommodate? At least it seems that way in the phones I've taken apart and looked at.

Also, OIS is a major mechanical add on (a literal motor) and even 1500usd smartphones lack it on some of their sensors, mainly because while it can have an advantage on an ultrawide, that tends to be more limited. Incidentally, most 99usd phones have one (actually usable) sensor which thus tends to have a larger width to compensate. I hope, in good faith, you see the difference, to something like ARI.

AMOLED, etc. are also a bit more expensive then OIS, but we get those into a sub 100usd BOM easily somehow. More so for 5g, certain features just become expected/required.

Your logic would lead to OEMs making SOCs without things like TEE and other things which started in the high-end but quickly became required and essentially free to implement.

Not saying it is free now, but that the upcoming gen of chips from Sony, Samsung, etc. will have it build in for such a minimal BOM impact, this will be an expected, common place feature across all prices.

To have a more serious, honest and accurate comparison than OIS, why do most new smartphone at 99usd include some form of an NPU? Or the trusted modules for biometrics, etc.?


>Doesn't OIS increase the size of a sensor by roughly half

No, you can apply smartphone OIS-tech on any sensor, stabilization is achieved via the lens-array, not the sensor. The size of the module slightly increases but that's not a hindering factor. Cost/Benefit of OIS on ultra-wide lenses is not there, so it's usually not applied.

>Your logic would lead to OEMs making SOCs without things like TEE and other things which started in the highend but quickly became required for one and basically free to implement for two.

TEE became a mandatory requirement of the media industry for Smartphones in ~2010, as they announced plans to restrict media-playback on a device without measures to secure the DRM-keys. Google made it mandatory shortly after, because the entire ecosystem was built on media-consumption.

It didn't come for free to the players in the industry, it became a very expensive task to develop, support and maintain it, but that's another story.

Drawing a parallel here, I fail to see who should require cryptographic authentication of a taken image from end-user devices, to the point that no consumer devices without it will be built anymore.

>Not saying it is free now, but that the upcoming gen of chips from Sony, Samsung, etc. will have it build in for such a minimal BOM impact, this will be an expected, common place feature across all prices.

It will be supported in sensors for sure, but those will be premium-tier sensors, as a differentiation factor. Until today there was no premium sensor used in mass-tier devices.

Of course, again, if there is demand in the market or a regulation requiring it, it will create an incentive for the industry to follow, but I fail to see why either of this should happen in the coming years.

>To have a more serious, honest and accurate comparison than OIS, why can you not buy a single new smartphone at any price without some NPU?

I don't know why OIS is not a "serious, honest and accurate" comparison, can you elaborate?

As stated, it's a highly mature technology available for over a decade already, providing critical-mass observable end-user value, yet it didn't just naturally "trickle down" to every smartphone price-segment, simply because it comes with additional cost no matter which scale (and the Galaxy A1x tier has massive scale).

A NPU is just the evolution of a DSP, which exists in Smartphone SoC's for more than a decade now and is required for Audio and Image processing. DSP's used to run pre-calculated inference models for lens-correction, exposure, white-balance, etc., now these processes can run as models on an NPU.

---

You are trying to argue why that feature won't naturally become a commodity at basically no cost. I'm trying to answer why I don't see this happen, because I worked in this industry for more than 20 years.

The market doesn't get features as a default "easily somehow", features reach this commodity stage because of significant end-user demand (like Camera, Display, Battery), business-value (for the vendor, like Apple Pay) or industry-requirements (like TEE, Widevine example above).

I don't see this happen for this feature, because there is

1. no significant end-user value (unless the public narrative is massively skewed towards "everything is true when the image was signed"),

2. the business-value applies only for Apple's service-proposition for now (which will likely make this feature expand to the non-Pro iPhone tier), and for

3. the industry-requirement I don't see WHO would actually be able to enforce this, for WHICH actual benefit.


> I'm trying to answer why I don't see this happen, because I worked in this industry for more than 20 years.

And despite that experience, you do not see the universal value in reliable, verifiable image attestation for any user? You cannot imagine why that may not just be very useful, but quickly become required, what the "end-user demand", "business-value" or "industry-requirements" could be?


This is not what I wrote and a very bad-faith statement of yours.

There is universal value in many features, yet they didn't become a commodity in all smartphone-tiers.

I don't see how and why this feature should become a default in all smartphones, which you keep insisting on without apparently comprehending the industry and market aspects I am trying to explain.

Peace.


What? Those three sentences are not really compatible with each other?! Like, they are mutually exclusive and each appears to hold a different position.

How is what I wrote (that you seem to not see universal benefit) in bad-faith/not what you wrote (honestly trying to understand what you mean) if you then add:

> I don't see how and why this feature should become a default in all smartphones [...]

So, do you see universal value or not? Cause you again said you do not and that was precisely what I wrote, that you do not seem to despite it being obvious to anyone who thinks about why phones of any price have cameras.

Should I honestly start writing a list why private as well as business users of smartphones may want, even need this? Is that really required? Consider every use case of a camera on a phone, please, before I feel the need to do that.


You fail to understand the difference between a feature having perceived "universal value" and that same feature being applied universally in all price-segments of a device. These are, and I cannot overemphasize this, two different things!

You stated that you "struggle to see how cost could be a factor here", so in good faith I was trying to give you an insight.

There is universal value in OIS, as everyone takes pictures while holding the device in his hands, yet OIS is not applied universally in all price-tiers of devices. The reason is COST.

So you wanted to shift the conversation, claiming that it's not a "serious, honest and accurate" comparison, without providing any reasons for that.

Again, I walk with you, I respond to what you're stating.

Now it ends with you starting the straw-man argument of how obviously great this feature is and asking how despite all experience I am not capable to see that.

See, it doesn't matter if _I_ see universal value in this feature or not, the topic was why I consider it unlikely to become a commodity.

You tried to move the conversation to this straw-man argument, and doing so in bad-faith. That's why our talk ends here. I give you my side of this conversation so you may read and grow from it, but this is entirely up to you.

P.S.: You also seem to misunderstand what this feature is. It's NOT picture attestation as you put it.

Picture attestation means an entity confirms the connection of a picture to something else (some "metadata"), e.g. a picture of a person to an identity (Name, ID,...) or a place. This attestation party can either be a person (self-attested) or an official entity (e.g. a government).

Nothing in this process changes with this Apple feature, because all it can do is confirm that the picture was taken by the camera as-is, but the attestation to the external metadata (WHO this is, WHAT this is, WHERE this is) still needs to be done by someone else. A party trusted enough to vouch for this.

But let's end this here. Have a good day.


>There is universal value in OIS, as everyone takes pictures while holding the device in his hands, yet OIS is not applied universally in all price-tiers of devices. The reason is COST.

The reason is digital stabilization is a good enough alternative to not bother, and the lens/sensor modules they use in bulk just didn't come with "analog" stabilization. And OIS is way less important than a future digitally verified photos feature could be (which could be mandated by corporations, banks, governments, insurance companies, for several uses when it becomes widespread), so they didn't bother to add it.

All kinds of cheapo smartphones still manage to have OIS, just because some Samsung models don't doesn't mean it's a universal argument for cheap phones in general.

>Nothing in this process changes with this Apple feature, because all it can do is confirm that the picture was taken by the camera as-is, but the attestation to the external metadata (WHO this is, WHAT this is, WHERE this is) still needs to be done by someone else. A party trusted enough to vouch for this.

Moot point, since the entity (e.g. gov) asking for an untampered photo (which this can do), can combine the photo with the metadata from the upload, like your gov mobile app account.


> See, it doesn't matter if _I_ see universal value in this feature or not, the topic was why I consider it unlikely to become a commodity.

> You tried to move the conversation to this straw-man argument, and doing so in bad-faith [...]

What? This entire conversation started with the assumption being made that this was going to be iPhone exclusive. I stated doubt and then you (re-read to verify cause you seem to have forgotten that) out of left field and for no discernible reason felt the need to move the conversation to the only massive straw man in this interaction, the 99usd phone market.

I personally still believe that this is going to be a universal feature soon enough (wait to hear from Omnivision, etc.) and stand by that. This is my personal assessment due to the objective need for such features in the world we live in today. I may be wrong on this, we will see. Maybe it will be restricted to the higher end Android phones, but that would still be compatible with what I started out saying.

I shouldn't have engaged with someone so serious that they must drag "why couldn't Android OEMs in general do the same" (a purely technical question) down to "99 USD Smartphone" and "premium" on those (an odd and empty pivot). Anyone who does such a shift, well, they must have a well founded, serious point to make.


How many smartphones have no camera? Zero? Bluetooth? Zero? But they could save cost by not having those. I think they are basically table stakes. If this type of thing becomes required for more and more things then no one will buy a phone that doesn't have them.

Yes, the whole insurance self-service is built on smartphones having a camera.

But the assumption that smartphone cameras, including those used in 99USD smartphones, will become 100% cryptographic cameras in a few years is highly unlikely, considering that those cameras didn't even gain OIS in the last 13 years despite the feature being highly matured and widely available.

Changing the topic to other features won't change that.

You seem to lack the understanding how this industry works, and assume that every development naturally just trickles down and becomes a commodity. This is not the case.

This cryptographic feature will definitely become available from camera sensor suppliers, first of all likely from Sony. But it will be a feature of premium sensors and will remain a differentiation factor.

Sony will not support cryptography to its sensors without additional cost. Device-vendors integrating those sensors then have additional cost in R&D, production AND operations. All this will not be waived and put in a 99USD device.

For the other assumption, that "If this type of thing becomes required", I fail to see how this should happen for a mass-market consumer: This feature doesn't authenticate the content of an image, it just authenticates the RAW data of the image sensor. It won't (and shouldn't!) make the user more trusted towards another entity (like Apple mentions themselves in the link)


>But the assumption that smartphone cameras, including those used in 99USD smartphones, will become 100% cryptographic cameras in a few years is highly unlikely, considering that those cameras didn't even gain OIS in the last 13 years despite the feature being highly matured and widely available.

Them becoming 70% cryptographic is enough. The people who need the feature, can get a compatible model. Nobody argued that smartphones sold for kids to game on for example should have it.


>If COST is not a factor, why does the Galaxy A16 still have no OIS (Optical Image Stabilization)?

Because it's just not that important, phones and cameras have also used digital stabilization via cropping since forever.


Those are two fundamentally different things:

1. OIS (optical stabilization) ensures that the light photons consistently hit the same pixel, removing the blur caused by camera-shake during exposure.

2. EIS (electrical stabilization) via cropping compensates camera-shake on video(!) recording by applying the same shake to the crop-canvas within the frame.

--> EIS can fix a shaky video but not a blurry photo.


It's 2026, with cleaner high ISOs even in phone sized sensors giving the ability to raise the shutter speed as needed, we hadn't had much of an issue with blurry photos for a decade now, with or without OIS. There have been several expensive cameras with no OIS, like Ricoh GR and (and v2), or ZVE10 (and v2).

It's video where people care about these days. Does anybody complain about blurry S12 photos?


Might be boring, but in 2026 "clean high ISOs" in phone-sized sensors mainly comes from image post-processing (stuff like multi-frame merging is done even when shooting "RAW"). Post-processing requires a stable (albeit noisy) image, otherwise it'll be garbage-in/garbage-out.

--> OIS actually became MORE important for Smartphones in the past years, because while post-processing produces better and better results, it massively depends on usable input data. OIS is one of the very few methods to improve the INPUT-quality for post-processing.

>"There have been several expensive cameras with no OIS, like Ricoh GR and (and v2), or ZVE10 (and v2)."

That's a apples and oranges comparison. A quick Google search tells me the size of a pixel on the Ricoh GR sensor is 4.81 µm, which is ~8 times larger than the pixel in recent smartphones (~0,6µm). It is not only physically capable to capture 8x more light, it is also much less affected by minor shaking than sensors with smaller pixel-sizes.

>"Does anybody complain about blurry S12 photos?"

Not sure what's a "S12", but:

- On flagship phones with OIS: Not so much. Maybe in low-light scenarios, because, you know, not much light...

- On cheaper devices without OIS: Yes! Oh yes, constantly.

People assume that the picture-quality of a e.g. 2026 Galaxy A16 must be comparable or better than the picture of a 8-year old Galaxy S9. It's not, the S9 is still better in everyday shooting.

Just check user-reviews of mass-tier smartphones without OIS, like Samsung Galaxy A series...


Well, that's kind of like DRM and Widevine, which exists on every consumer device.

Well, that happened because in ~2011 the entire media industry announced that they will stop media playback on devices which didn't secure the DRM-keys. As media consumption was fundamental to the Smartphone ecosystem, Google made secure-boot and widevine mandatory.

Sure, the same could happen here, but I don't know which industry (or other body) would demand that and have sufficient justification for it.

After all, the feature doesn't really change that much for a consumer, the trust-chain is largely unchanged: If I send you a picture and tell you that's my dog, you still have to take my word for it, regardless whether Apple signed the picture or not.


Not if Apple can identify the dog as belonging to someone else.

Apple doesn't validate the content of the image, you will have to trust that it's my dog.

You wanna buy it now or not? /s


> comes at a premium. Why should this premium be applied on a 99 USD Smartphone?

It's entirely software. It's R&D cost, with basically none of it in hardware (you technically just need the private key store which even very cheap devices have)


A bunch of them already offer one. Have been a while, actually; the S25 and Pixel 10 came with exactly this.

The timestamping server is the hard part, especially with the verified compute component. It's just not something I see Samsung doing.

I expect Google to show up with a blog post titled "extending C2PA with timestamps for industry-leading authenticity confirmation" any time.


I don't think that's quite the same, though still valuable. Don't those still depend on the OS being trusted?

There was and continues to be no reason to share the package manager between models. This was begging for abuse.

There has been no statement either way, as far as I could find beyond them only using commercially available models, though given Alpöges employer, I'd be surprised if they didn't opt out. In any case, for such work, ZDR or self-hosting seem to be an absolute must now.

Unless OpenAI can show that training was permitted, this will erode the limited trust that many users have had in such toggles and may lead to further, uncomfortable inquiries.


He works for Anthropic nowadays.

They did pay [0] and substantially by the sound of things:

> I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.

[0] https://cims.nyu.edu/~tristanb/statement.pdf


> Haven't all the labs effectively disbanded their real safety teams a while ago?

Neither Anthropic nor Deepmind have. Meanwhile, the rocket company that somehow makes most of their revenue from renting out data centres never had much to dismantle.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: