Hacker Newsnew | past | comments | ask | show | jobs | submit | blfr's commentslogin

How do the users know they're getting what they're paying for?

I disabled automatic downgrading/rerouting because it sometimes takes me a second to tell when the answer came from a different model than I wanted. You could easily sell Opus as Fable for a good while.


You can't, it's all reputation based. Similar to whether drug users don't really know what they got were diluted or not.

At most big festivals there is free drug testing by harm reduction charities. Not that they'll test you for drugs, but they'll test your drugs to make sure they are what you think they are.

At Fusion Festival, I saw a big bulletin board completely covered in notices of "we tested this pill, here's a photo, here's what they thought was in it, here's what was actually in it"

They also spelled the name of the charity wrong on all the maps, so that's nice.


I think the drug users will know before the Opus users.

i doubt the users really care as long as they get the job done, AI in 90% of use cases is just a race to the bottom, nobody cares as long as it's cheap and gets the job done.

So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.

They can sleep just fine being the only player in town actually not loosing subsidized money.

They've also been way behind before, and caught up to being only a little behind. We'll see how things shake out. We're deep in the present but who knows how things will look a year or two in the future.

I don’t think Google cares about being the most intelligence AI as much as it cares about monetizing it with all its products, which requires speed.

Google has long said that this is what it cares about most, the fastest at giving the correct answer to questions.


Gemini models are at or near the top in several categories, though, so I'm not sure the takeaway is that they're shamefully far behind.

Google still has several enormous advantages here:

1. Google Books, Youtube and the Google Search index all provide vast amounts of legally acquired training data.

2. They can easy people into AI using the info box. I think this strategy is working even if it does cannibalize their main revenue source. Better than just withering and leaving all of the money to OpenAI/Anthropic. I would not be surprised if Google has significant layoffs due to reduced ad revenue at some point, but I think they'll still be on top.

3. They already have their hooks into people's lives through Gmail, Google Calendar, Android, etc. The only other companies that come close are Apple (but for a much smaller number of people), and Microsoft (but only for business).

The fact that Google's models might be 20% worse, or a few months behind Anthropic's is completely insignificant in comparison to those things.


They are tied for first using the Google-proof question and answer benchmark: https://artificialanalysis.ai/evaluations/gpqa-diamond

Maybe that's their only goal?


Is Google trying to compete with OpenAI and Anthropic re: maximally intelligent models? Google seems to be the only one of the three that doesn't pray and self flagellate at the altar of AGI.

No. Demis is busy creating another documentary about how great of a human being he is . and giving interviews to fawning journalists projecting profundity over his every word.

If you're Demis, at least, you sleep fine because you were personally an early investor in Anthropic.

The YouTube golden age is right now. The content is the most varied, from super niche 2k views a video vloggers to massive terrestrial channels alternatives. To the point where I think creative poverty of established fields is caused by younger creatives being on YouTube and TikTok instead of writing rooms at old media.

Windows has fallen off hard but Linux has never been better. The personal computing in general has never been better: OLED screens, dark modes, everything runs on virtually every platform...

The only thing we've really lost is the Internet community. Eternal September turned out to be truly eternal and then went on steroids. But the tools? Never better.


> Linux has never been better

I replace Windows with Linux on my last Windows computer this week. A mixed bag of old motherboard, new-ish graphics card etc. I think I would have had a hard time to get an acceptable usage experience without using an LLM to guide me through the hiccups. (Odd window manager boot problem, sound device issue). But I agree, the Linux ecosystem seem to have something for everyone and on modern hardware it behaves relatively well. LLM as your local tech support works surprisingly well too.


> LLM as your local tech support works surprisingly well too.

Isn't that partly why I would imagine Linux is better than ever? I'm a relatively technical person compared to some of my peers and without LLMs I just wouldn't have the time to sit and toy around with search up what random errors mean, why things aren't working, why things are crashing, making a forum post and waiting for a response or searching through stack overflow.

But now, it's no longer an issue!


LLMs as tech support is truly the singular new reason causing Linux usage to jump atm. All the other reasons are just slowly accumulative.

LLMs as tech support is the feather breaking the back of the Windows stranglehold.


dunno.

If you gave two people with no other knowledge of computer access: Kubuntu and Windows 11

I'm sure they'd have an easier time with Kubuntu, and this gap has been widening for quite some time in the favour of KDE.

The largest issues seem to be either power users doing esoteric things to their systems, or people who want to work the Windows way.

Drivers (for as much as I can tell) are mostly solved, with some minor touch buttons on some laptops being an issue in some cases.

I'm not comparing Linux of 2007 to now, I'm comparing Windows now to Linux of now- the only real edge Windows has going for it is that it's already installed when you buy your computer.


I generally agree with you (I'm typing this on a laptop running Void Linux), but I do still keep a Windows laptop for gaming and music production - I've tried running Bitwig Studio on Linux, and it's easier on Windows. While I've heard good things about gaming on Linux/Steam on Linux, I've never bothered to try it myself, being content to just use Windows PC for gaming :)

For me the main thing is Windows has gotten so bad that I was finally willing to spend time getting Linux running. LLMs are nice, but not necessary. Claude helped me with a driver issue at least knowing where to look but the answer was ultimately on a Linux forum. I think LLMs help but it isn't that much better than just searching old forum posts.

At least for me, the availability of LLM help had nothing to do with my decision to switch to Linux, it was Proton making it to where games run great on Linux and Windows being intolerably bad.


The actual content yes, but the recommendation engine switching from "Here is a video that is kind of like this one, you might like it" to "Here is a video that we predict will keep you on this website for as long as possible" definitely soured the experience.

PocketTube. I organize all my subscriptions and only watch those now, I rarely actually go to the home page with recommendations.

I like that sentiment and fully agree regarding YouTube. Unlike Facebook, it's still possible to tune your feed to your exact interests. I have never had a better experience with any video medium, neither linear TV, nor streaming, than I have with YouTube right now.

Instagram is now much better for this with its "customise your algorithm" feature. I was able to strip out all the crap it thinks I like and focus on one specific topic.

> The content is the most varied

If no one can find it, you basically have a modern digital version of the "if a tree falls and no one hears it does it make a sound".

Discoverability in youtube is currently absurdly shit. The algo just decides you like something and shows you the exact same videos you have already watched 30 times. Its so so bad, specially with how good the front page and side recommendations used to be.

Now its a fight of shorts, repeated videos, AI garbage, and search ignoring explicitly your request.


Maybe it's your algorithm then because I found lots of great stuff that it surfaced to me that were not just repeats but adjacent topics. For example if I watch Veritasium I get videos about other math and physics topics and creators.

> Maybe it's your algorithm

Nope, its just the set up of Youtube.

The front page is by design around 60% videos you have already watched. Best case scenario.

It used to be 9 videos in 3 rows of 3. now its 2 videos, then 5 shorts.

In the side bar the default was "Related" which was almost always new videos, directly related to the video you just watched. Now its a tab called "All" which aggregates related tabs such as "Watched" which again is videos you already watched. Content from the same channel (not necessary related to the current video)

All of this hides possibly related videos and narrows down discoverability.

it also ignores the push of youtube to reduce visibility of things (the front page being mostly shorts by volume is pretty telling)


Most companies just get you a Claude team sub and maybe a couple of skills.

We only get Copilot. I’m not very happy.

What do you find worse about it? I've been switching between it, Codex, and Claude Code to try to compare them, and my only conclusion so far has been that it's nice that Copilot has both OpenAI and Anthropic models as options.

The parent poster is almost certainly talking about inference within workflows and not for interactive coding agents.

As much as I dislike 'em, this sounds mean spirited. And Alibaba admits in this very tweet that Fable is next level (it is).

I dislike their practices, but the main motivation for hoping they'll crash is that I think their immense overvaluation posses too much economic risk.

I want as much misfortune as possible to befall OpenAI and Sam Altman after what they did to the memory market.

How dare they buy things

But they didn't. The deal to buy 40% of the world's memory never happened after the price increases the news generated.

How dare they not buy things

How dare they claim to buy stuff to boost their standing while actually just faking the entire thing.

> Not eager to ruin the evening by starting a debugging session

Use a clanker. Nothing has brought back so much joy into homelabbing for me as having LLMs handle the boring stuff.

I launch it in opencode on the system or elsewhere to ssh in (I think giving it KVM is an overkill but an option nonetheless) and tell it to fix things. Not just ask questions and generate configs. I have backups, let it rip. They have gotten shockingly good at it even when they write bash wrappers just to catch some logs.


Not really, not for 1¢/day. Most massive platforms that I use for free (Google Workspace, Cloudflare Pages, even Oracle Cloud always free tier) have outlived many low cost solutions I tried.

"Why is it true?" is usually an even more cancellable/fireable offence (which is also why you see it discussed less).

I admonish Gemini and demand explanation nearly every day of how it's possible that Google invented the thing, has the best infrastructure for inference, and somehow falls behind Anthropic and even OpenAI.

NotebookLM is pretty cool since it can hold a ton of context but this is so far below my (and frankly just reasonable) expectations of Google.

I downgraded my Gemini subscription and got Claude. Still can't believe how much better it is. Fable is way better, that's a given. But Claude even has a real .deb repo. Something Antigravity had and managed to lose.


My response is going to be about Gemini generally and less about NotebookLM.

Google's last frontier model release was Gemini 3.1 Pro, which was in February of this year[1]. At the time, it was ahead of the (at the time) flagship models of Opus 4.6 and GPT 5.2/5.3. From my recollection of the time, it was the best model in the world.

Anthropic released Opus 4.5 Nov '25, 4.6 in Feb '26, 4.7 in April, 4.8 in (late) May. Then Fable in June. 4.7 beat 3.1 Pro on multiple metrics. Fable eats it for breakfast. However, I want to note the 3 month gap between those first two Opus versions.

OpenAI released 5.2 Dec '25, 5.3 Codex Feb '26, 5.3 Instant Mar, 5.4 Mar, 5.5 (late) May, 5.6 July. 5.4 beats 3.1 Pro on agentic benchmarks[2], seems to be similar/losing on non-agentic. 5.5 seems stronger than 3.1 Pro[3].

Gemini 3.5 Pro is alleged to be launching within the week. Why do I type this all out? Because I think Google is getting a bad rap. They are delayed on a frontier release by a month or two and are being regarded as if they cannot release frontier models. I think their last release demonstrates strength and we need to see a weak release before we call them "behind" (in any reasonable sense). These companies swap back and forth constantly. I recall a multi-month span where 2.5 Pro was just the best thing out there by a large margin (in my opinion).

[1]: https://blog.google/innovation-and-ai/models-and-research/ge...

[2]: https://www.anthropic.com/news/claude-opus-4-7

[3]: https://www.anthropic.com/news/claude-opus-4-8


In my experience, Gemini 3.x wasn’t just getting a bad rap, it was significantly worse in practice. It could analyze codebases and report back from a one-shot prompt as good as Claude or Codex but any slightly complex task that carried on for more than a few minutes led to hanging, seemingly infinite loops, and bizarre and nonsensical hallucinations, to the point of being unusable for serious work. The Claude and Codex counterpart models at the time rarely had such issues for the same type and duration of complex work, if at all. To be fair, later Claude especially started having hanging issues as many people noticed but that’s been better recently.

I think you have rosy eyed glasses (or never inspected the output too closely), Gemini 3.1 Pro was very bad at hallucinating.

Antigravity sucks so bad that I have started to feel that google really doesn't wanna compete, they just wanna hang in there at number 2 or 3, to just annoy the number 1 and 2.

I tried Antigravity recently with Flash 3.5, and it got stuck in a loop saying the same sentence over and over. I haven't seen this pathological behavior from other LLMs in months.

Fable in Cowork can get stuck in loops, I've had a single prompt use 83% of a session quota on a single prompt before I realized something was amiss.

Its explanation: "the wasted tokens came from re-rendering the document to verify layout after a page-orientation bug."


Fable is very bad at wasting tokens when it comes to doing these renderings and diffs - I had the exact same experience as you.

I recently got a trial for the Google AI Pro subscription and it burned through a week's in a single prompt. Quite impressive for sure.

For coding, I don't think they're even number 3, anymore. Seems more like 4th or 5th (unbelievably, even Mecha Hitler seems to do better, though I'm hopeful Gemini 3.5 Pro will turn things around).

I share your sentiment. I'm still paying for Gemini but it's almost useless to me. Google has some serious internal problems. Perhaps Gemini is merely being plagued by aggressive cost controls, or perhaps there are deeper flaws in Google’s approach. Either way, I can't trust NotebookLM with serious work and have stopped using it.

Apple is placing a major bet on Google and Gemini for iOS 27. If Gemini's decline is any indication of what's to come, Apple could be in serious trouble in six month's time.


> Google has some serious internal problems.

I suspect it is part leadership change (sundar/kurain) and over indexing on Ai for doing the job on top of a model that is just not as good (esp flash 3.5). Google Cloud / Gemini Enterprise sent me the greatest Slop Deck of all time. It was quite obvious it was Ai generated and the rep had only read a few slides of the 30+. I wonder if they are even aware after losing the sale


I was still using Gemini CLI even though Claude is better, just cause of inertia. Then one day it started refusing to work, saying I need to install Antigravity instead. Idk if that's an IDE or has a leaner CLI, but doesn't matter, I'm gone. I don't care how the sausage is made, don't randomly break the thing I'm using.

My wild guess is Google isn’t willing to play as dirty as the others wrt data acquisition.

Also enshittification is slowly starting. Yesterday I asked Gemini (in the Android app) for a recommendation for an app for sound recording. Instead of it answering I got a popup to allow Gemini access to open the app store (or something like that, I didn't allow it). When I declined it just stopped the conversation. It was actually hard to get it to just reply with a list of apps. And I have an AI subscription with Google!

Yeah, it grinds my gears. They could probably have lightning speed inference and the best model if they were interested in doing so.

Some good takes so I hate to be glib but no one because 1. the US was founded long after the invention of writing and even the printing press plus 2. bronze age morality and worldview are completely alien to us in ways that no subculture in the US is.

The article references "England has Shakespeare (1600s), Spain has Certantes (1600s) Russia has Pushkin (1800s)" whom I assume are intended as their nations' versions.

You going to stand on your statement with that as a reference? We're closer to Pushkin than Pushkin was to Homer.


Heh, we're closer to the Beowulf poet than the Beowulf poet was to Homer.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: