This is a dimension I hadn't considered the impact of; if METR isn't an impartial actor, then "AI Doom" is directly aligned with Anthropic's interests, and this whole thing starts to smell like a different and much scarier variety of bull.
Ah! That's what that is! I am so pleased, I've maintained a shell script forever to make sure MacOS is prioritizing the Ethernet adapter, I can delete that damn thing now.
The "assumed it was a policy" thing resonates with me. My agents have such trouble distinguishing a note from a law, and they seem to looooove following laws. I've had similar fuckups; a stray constraint can turn into some wild tradeoff decisions which compound if you aren't paying attention.
Interesting. For example, GitHub’s built in Copilot review for example will checkout the whole repo and do a broad review. I found it to be WAY better than our prior CI reviews that just looked at diff and had TONs of false positives. (We have also built our own review tools like Copilot that check out repo and they are similarly league ahead IMO than diff reviewers.)
I found diff reviewers so heavy in false positives, all my coworkers just ignored them.
This is only true if you only run the code reviewer in the final PR.
The way we set it up was that the code review responded to two signals:
1. GH PR webhook
2. The exact same review agents running as an MCP tool that the local agent can invoke before pushing.
Practically, what this means is that it's OK to have a false positive because the local agent can make the check with the full context.
This would be the same as if the entire team used Codex and, for example, had a sub-agent configured to run code reviews using a smaller model. In this case, the benefit to this tool-based approach is that the exact same agent works for all harnesses across a team and also works in the PR itself.
They won't need these cameras soon, AI-based home inspection tools are on the rise and insurers are using them in lieu of humans. The fine-print is absolutely dreadful, as you might expect.
I was forced to use one in order to renew my home insurance policy. Using a supplied app (from a LexisNexis company), I was required to take video of almost my entire home and property. There was no right of refusal; if I wanted to be insured with $Company, this is what I would do. I have little choice, given my area is losing insurers in droves because of definitely-not-global-climate-change.
Because I'm a prick, I made sure one of my bachelor party gags were in each room. Hope somebody got a laugh.
Anyway, this is what AI looks like - cost reduction, failure minimization, which is all well and good when there's choice in the market, and those who desire privacy have options.
For persistence, I think the most glaring gap is the inference - if these rogue agents need to call Claude's API, then they aren't persistent - Anthropic can turn them off. A memory resident program needs disk to be persistent, but above all it needs CPU; without it there are only bits.
And so the implication is the swarm has access to enough local compute to perform its own inference, using a large enough model to provide the capabilities to be dangerous.
This is impractical today.
It might be practical tomorrow - if the Chinese are allowed to continue to develop large open weight models that we can quantize and ablate and make small enough to run on a million standard PCs.
Or if *anyone* is allowed to continue to develop AI in the way every other technology has developed - improving, shrinking, optimizing.
And so the argument REALLY isn't against the Chinese - it's that we can never allow this technology to advance outside of trusted labs. If they get their way, they will need to keep this tech locked down forever, which will require much much more than what they are asking.
> It's that we can never allow this technology to advance outside of trusted labs.
Impossible. In 10 years you'll be able to train today's models for ~$10M from Moore's law and training innovation. In 20 years it'll be ~$200k. The genie won't go back in the bottle.
Depending on what your threat model is, and how much you trust OpenRouter and their upstream providers, they offer zero data retention (googleable: ZDR) APIs which come at a higher cost. You can find other zdr providers as well.
Is this just posturing (I won't speculate on the intended audience), or does the admin believe this will actually net more jobs for US citizens?
TFA doesn't go into great detail about how this actually pushes offshoring - the variety that I'm familiar with at least (X/Y% blend of US/"near-shore" aka anywhere-but-west-asia roles). I have no doubt however that Y is just going to get larger while X stays the same or decreases (not even factoring aipocalypse).
What constraint specifically makes them believe this is going to net in the way they want us to believe?
reply