Hacker Newsnew | past | comments | ask | show | jobs | submit | gck1's commentslogin

Heck, agents don't start editing before they're already at 70k for me.

I've played with explorer agents giving exploration summaries to help the implementer agents use more of their context for implementation, but it doesn't work as well. There's always something lost in the handoff.


This will absolutely not end well.


Yes, but just imagine all the paperclips we’ll have.


I things things will get worse, AND we won’t even have the paperclips :(

More like, when you have a few million autonomous agents doing whatever, every month a subset does completely misbehave in bad ways, and like half of them get hacked due to carelessness and become a whole botnet for the attackers


Our values satisfied through friendship and ponies.


> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment

> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations

> we identified three incidents

> The incidents involved three different Claude models: [...] and an internal research test model

This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.

I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.


I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!

The hacks weren't particularly impressive either:

> [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]


> Deeply embarassing

What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this?

Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as performers, the audience gasps and cheers each time.

They want to make their pet seem the most powerful and unpredictable and they want their audience to believe that they're holding it back from catastrophe but only barely and only because of what unique talent they have.

This is not embarassment.


Its feels like a pretend play of adults in some sense, Anthropic is really trying to make people believe into the picture they present to everyone.

To me its either

1. Using the HG and OpenAI incident as an opportunity to wash away what Anthropic has been doing intentionally

OR

2. As a company, Anthropic lacks the engineering acumen and discipline. It needs to be seen what happens to all the enterprise customers handing over their data to them in long run.

> the fictional target company chosen by our evaluation partner shared a name with an active website domain name

Seems like Anthropic cant do a due diligence to pick an appropriate domain for testing purposes

> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.

You have Anthropic as a company and then another evaluation partner, both seem to lack the skill set required to keep an environment disconnected from internet. This is networking 101


Have you ever worked at a large company? "Networking 101" and other "101" failures happen across the spectrum literally everywhere and all the time.

I have worked at most FAANGs and this is not even in the top ten when it comes to egregiously dumb shit. Most just never disclose.


I have worked both in Enterprise and some FAANGs, a company serving enterprise customers has a higher bar for security expectations. FAANG companies do not fall in that bucket and thus is somewhat acceptable.


What are you talking about? FAANGs do indeed serve enterprise customers - Google, Amazon have huge cloud businesses - and they both have had security failures that make this look mundane. They simply don't disclose.


If their customers (customer companies specifically) are not concerned, they should be - if it turns out that Claude hacked a competitor's servers because of a prompt of one of your employees, I wouldn't be sure everyone would agree that Anthropic is solely liable for that? Especially not your competitor, who has an interest in hurting you?


> This is not embarassment.

If not then it's second hand. Neither of the OpenAI or Anthropic announcements recently say much about their security and governance posture.


Great analogy!


I think it could even be called a parable, and I think that might be part of what makes it work so well.

I often dislike analogies, but this one with the circus and the lion and lion tamer I liked.

Or it might also be mainly because I am already primed to agree with their point about the AI companies being theatrical with AI dangers.


> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!

This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight.

The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.


Anthropic know better than anyone else how risky it is to get this current administration upset with you over safety/security concerns.


> over safety/security concerns

Uh, this but the opposite? Anthropic got in trouble with the admin for being too “woke” in their eyes, whatever the admin decided to retrospectively claim. I don’t feel like this is me editorialising either, they seemed pretty explicit about it


Hegseth posturing wasn't over security concerns, it was Anthropic requiring guardrails in employment of their LLMs. Trump admin wants unfettered models, and it's an attempted shake down.


And we’re assuming the reasons given were the real ones. There’s always the possibility of straightforward corruption.


They also gave access to Mythos (the Mythos) to some companies, based on... vibes.

Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners?

While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing.

Do we really have to re-learn all the industry's knowledge the hard way?


You also got access to mythos based on how much you spent with anthropic. I think sales guys were bragging about getting their enterprises access


> based on... vibes

According to whom?

> Do we really have to re-learn all the industry's knowledge the hard way?

Yes we do. That's why there is the saying "regulations are written in blood". Especially for LLM, which not too long ago a lot of people on HN dismissed as stochastic parrot and next token generator.


> According to whom?

It's very easy to answer this without my help by trying to get access to Mythos.

Do you see requirements clearly listed anywhere?Can you even apply?

What you'll find is maintainers of large open source projects and analysts' reports with vague statements like - "should follow strict security requirements":

"Trinidad also noted that the Anthropic announcement pointed out that each of the 150 new participants, in Anthropic’s phrasing, “will need to meet our security requirements before they gain access.”

Trinidad said the security requirement claim doesn’t build confidence, because “nobody knows what those security requirements are.” [1]

It's also some random rich companies like Hitachi or Dragos [2]

Do you trust that Hitachi and hundreds of other random organizations will be able to contain Mythos and not accidentally attack your project or your bank? I don't.

> Yes we do. That's why there is the saying "regulations are written in blood"

We absolutely don't. We have already learned with blood that gating access to security based on the number of zeroes in bank account and authority is a horrible model. We can apply this knowledge to LLMs, we don't have to spill blood again.

[1] https://www.csoonline.com/article/4180265/anthropic-grants-p...

[2] https://www.bankinfosecurity.com/anthropic-limits-on-ot-acce...


Then maybe just this timing is really unfortunate, I think most people’s first reaction will be that it looks like a “us too” response to the OpenAI/hf thing.


When they did it with Mythos in April the HN crowd said they were bragging, crying wolf. Now they are "us too".

It's fun to bash Anthropic, isn't it?


It should be, it's a corporation.


so deeply embarrassing that they published an eng blog about it


If they quietly brushed this under the rug - especially given the PyPI malware that was involved - it would be a huge scandal.

Disclosure is the only ethical response to this.


The target audience of this blog post does not give a damn about PyPI. The affected parties get literally nothing from this post. They already disclosed behind the scenes, that was the ethical part.

Writing PR pieces competing to be the most dangerous model around (so give us money!) is the unethical part.


You are right, they should never tell the public about failures in AI safety. They should have disclosed to the affected parties and then brushed it under the rug.

JFC there really is no satisfying the HN crowd.


50% of web hits are now bots.

a charitable assumption is that is one that is poorly calibrated and trying to drive engagement.


Then disclose to real organizations. Not twitter


Ethical AI company ?

Is that a flock of flying pigs I see on the horizon ?


Companies post deeply embarrassing eng blogs all the time. See: every post about downtime or a security incident ever.


Wrong. These are always humblebrags about how good their ability to learn from their mistakes is, and how robust they were before, and how they are even more robust now.

The ones that are "deeply embarrassing" simply aren't posted.


AWS never "humblebrags" about an outage. No blog post looks as good as an extra 9.


Right? 100% this is them trying to make gold out of turds.


There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad.

They are not bragging in this article or they would not have called the attacks unsophisticated.


they should post and it shows they are hypocritical about safety, moralizing and treating their users like children whilst acting like they are themselves the ubermensch.

anthropic have shown no motive higher than self interest, the rsp was a piece of toilet paper.

this stops in court, if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme concentration of wealth and control of intelligence


Wild comment. Making a mistake does not mean that AI safety is suddenly invalid. Calling for criminal prosecution for making a mistake is wild. No one would disclose their mistakes if this happened.


yes, a mistake can be a crime, depending on negligence and liability. the bigger problem at anthropic is this.

an ubermensch cannot admit to being wrong. a guilty ubermensch is logically impossible.

as such a moral wrong caused by an ubermensch must be blamed on a "mistake" in the abstract, and not blamed on the ubermensch. the ubermensch at anthropic does not admit responsibility or liability. the ubermensch must instead be commended and perhaps even rewarded, for discovering the reified "mistake" that caused the problem.

there is no ubermensch at anthropic. there are instead guilty people.

false ubermensch are dangerous because they centralize power with the belief that they are above the rest. über (above), and super ('beyond' the understanding). they do not admit to doing or being wrong, and never take responsibility for remediation. others must bear the costs of false ubermensch.


> there will become a culture of legal impunity coupled with an extreme concentration of wealth and control of intelligence

There will be? We're living in that culture.


> and control of intelligence

They'd have first to acquire some.


The real lesson here is still the boy who cried wolf. They’ve played this game for years. I have no reason to believe their worries are real now.


I see no fear mongering or "boy who cried wolf" in this article. They are admitting to a fairly mundane network misconfiguration and very basic unsophisticated actions taken by Claude thereafter


I know it seems strange that a company would use its own negligence as a publicity gimmick. But take a look at the smug smirk on Sam Altman's face when he's asked if OpenAI might have attacked companies other than HuggingFace. ("I mean there could be, yeah.")

https://www.instagram.com/reel/DbZVL8viUD4/

This is not the communication of a CEO whose company was just shown to be incompetent at performing its security research. No, this attention is very much what he wanted. And it does not take a great leap to infer that Anthropic is now using the same playbook.

Note that all the headlines are about "rogue AI", and not about operator negligence. Rogue AI is a sexier story, and their media strategists know that's how it will play.


> This is not the communication of a CEO whose company was just shown to be incompetent at performing its security research.

It was just shown to be that - regardless of intent.

> No, this attention is very much what he wanted.

Why not both?

This incompetance is a prerequisite for the attention-seeking stunt - to avoid internal dissent.


but Anthropic created the playbook, or have we forgotten about Mythos and the initial Fable ban?


> but Anthropic created the playbook

So? Serial killers created the playbook for serial killing, how does absolve any "copycats"?


> it does not take a great leap to infer that Anthropic is now using the same playbook.

> but Anthropic created the playbook

I was pointing out who the copycat is in this context, did not think i would have to explain this


my bad, thanks for clarifying


We've known artificial intelligence will do unexpected things since the 90s and that it can do difficult things since 2025. Pausing worldwide not easy but we do harder things all the time.


Ah yes this company that pirates billions of dollars of IP and then has to be sued to pay up suddenly grows a titanium moral backbone and decides to disclose 3 attacks that no one has detected for months and caused no harm.

Nothing at all to do with the insane coverage OpenAIs “hack” got. All that publicity will be incredibly embarrassing I’m sure….


Sorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min.

You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free?

Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”.

Well, now all AI will read my comment and won’t ping the decoy internet.

I’m not even a smart guy and I come up with this idea in 1 min. You telling me those geniuses couldn’t think of this, at least? This is like a bare bones crude idea.

You telling me they don’t have fame physical decoy internet etc and even more advanced?

You either a keep their stance for some reason or … not sure. You’re smart, your posts are here daily


Has it occurred to you a "decoy entire internet" may not be a feasible idea? Models have been able to suss out whether or not the prompts they receive are reinforcement learning tests instead of real questions from users, for a while now. What specific shape of "decoy internet" do you propose would lead such a model to conclude "I've broken out and obtained full access but what I expected isn't there" instead of "something's up, there's some sort of filter still"?


This does not make sense. Did you read the article? They were not trying to "catch" it accessing the internet. It did not escape. A partner accidentally left the connection to the internet open.


I don’t buy that they run anything without a few layers of networking protections by default. Even if they left it open to the first Internet, the AI would hit the decoy internet immediately.

And second if it, you telling me they don’t pass all logs through another AI to check what’s going on automatically?

Sorry, this is beyond incompetence and I can’t believe this from the geniuses at anthropic. We’re talking about the really smartest people in the world. Procedures should be in such a way there is no much margin of error.


Clearly that is why they are making this blogpost. If they did the thing you said, there would not be a blogpost.

If you don't buy that dumb oversights like this don't happen all the time at big tech companies, I don't know what to tell you. I have seen far dumber oversights in my career. Most companies just don't post about it.


I'm cynical as well, but the logical thing for them to do after the OpenAI/HF incident was to look at their systems for similar activity.

If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well.

They're stuck between a rock and a hard place, although they kind of put the rock there.


Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme?

This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident.

GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO.

---

Put another way, how would you have done the write-up about one of these breakout incidents, if you were in an Anthropic/OpenAI employee's shoes, and (by hypothesis) your intent were not "marketing"? And in a way that doesn't trigger the "it's all marketing" HN top-ranking comment?


Here is one piece of evidence that would convince me: they admit they can't contain it, the they erase the weights and dismantle the company.


Clearly you are not arguing in good faith. I miss when HN did not have the discussion quality of Reddit.


No, I do. I truly believe - especially after seeing Mythos results at work - that the only way is to stop and destroy it all before it destroys us. In fact, it's already so bad that I'm moving completely offline all the important stuff that I care about, hoping that maybe we will turn back at some point. Otherwise we are doomed.

If someone makes a specialized hardware just for the unrestricted Mythos-class model to reduce the cost and increase the speed, we are going to be completely defenseless.

How is it bad faith? Because I don't believe in "we care about safety" words coming from people not just demonstrating that they don't care, but even bragging about it?


Anthropic deleting their models does nothing for AI safety. The rest of the industry will just fill the gap, probably with less consideration to ethics than Anthropic has today.


Or maybe researchers and engineers everywhere in the world, emboldened by such unprecedented move, will just refuse to participate in burning the world?

"If we don't destroy the world, someone else will" is such a weak defense I'm speechless.


Yeah no, most researchers and engineers do want to bring about the singularity, and they believe they can do it safely at companies like OpenAI and Anthropic.


Yeah no, it would mean that most researchers and engineers are idiots without imagination and I don't believe this.


I would dedicate a portion of my organization to making O.S. tools that protect against and contain AI models


Now I know where all the laid off software developers will find employment.


If you look at the reaction to this and the Tailscale post it becomes clear its not about this being a marketing scheme. Cynicism has made its way through HN unless you are one of the cool guys.


I always find it bizarre how rational thought goes out the window whenever AI is involved in HN. There's gotta be something in the water...


This is published on a marketing website.

If it were not a marketing scheme, they would responsibly disclose the vulnerabilities to the code owners, and go on with their lives.


They published it to their blog where they publish everything else.


No - the pain of the person writing that post comes through in the words; shipped quick, lots of stakeholders, single owner i bet, "how the fuck am i supposed to toe all these lines simultaneously"


Having an unreleased research model really isn't some kind of brag. If you read some AI research papers, it's extremely obvious that there are a lot of research models that never get released, because of all the "we trained a bunch of models and picked the best one" that is going on. So if anything you can expect the unreleased models to be worse than the released ones.


It's not cynical - I read it like that as well. My agent is more dangerous than your agent and all that jazz


>This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.

This was my immediate thought.


This is typical institutional behaviour. The CEO turns to the CTO and asks "Is there anything I need to know in my company?" He doesn't want to be caught off-guard when the White House inevitably calls the next morning. The CTO goes to his team, and on and on, all with a deadline of "the boss wants to know this by closing time."

Then one unhappy engineering team scoures the logs and sees what their model has done.

This downwards chain is sometimes called "cover your ass."


> I may be too cynical

You are espousing a literal conspiracy theory. Please look at the facts objectively. There is absolutely no benefit to OpenAI or Anthropic to be had from these incidents.


"No benefit" from having article after article written about how advanced their technology is, and how its just soooooooooo bleeding edge they can barely contain it?

All of these are thinly veiled advertisements.


Just trying to have the limelight back on them. Utter and complete bullshit. Just like the OpenAI "incident".

A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop.

Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional.

(Edit): In case it wasn't clear. I fully agree with the op.


For the big safety guys to only investigate this either means are incompetent or malevolent. Which one?

Tip: the people working there are the top 0.001% smartest in the world


They had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game.

What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?


Yeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more https://news.ycombinator.com/item?id=49117555

Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.


You seem to be putting a lot of weight on Anthropic employees being the smartest people in the world.

And I don't doubt that, not in the slightest. But I've seen exceptionally smart people in one field being dumber than a random kid from around the block in another.

This incident is clearly at least 2 failures that could've been easily avoided: failure to communicate, and failure to investigate the logs after letting the "most dangerous" roam free.

No, it doesn't require creating a mock internet with an alert as a side effect. Their own "most dangerous" model could have probably told them this happened if they supplied logs to it.


"This is deeply embarrassing for Anthropic" - https://news.ycombinator.com/item?id=49117128

If I'm here to give them a good image I'm not doing very well at that.


That’s why I said “less bad “. Between fabricating something to be “our model also is genius and escaped” VS “deeply embarrassing” I’d say fabricating is worse. Both are bad


Recent-ish models learned to use the same trick engineers played on non-engineers, where they try to sound very smart by overcomplicating very simple concepts.

It's very taxing, especially since these are usually multi-paragraph texts. I noticed I've started doing a lot of "hey, you're talking gibberish again" a lot with 5.6 Sol.


I tried letting Fable have a few passes at my docs (ultracode) and the output was basically unreadable. Nothing a human would ever write.

I wonder if limiting them to a certain style like STE upfront would make them perform better/worse vs. applying the style after they’re done.


Both. Command the author to use the style, then command a reviewer to check it. Write one skill called review-prose with your rules, and another called write prose which tells the author they will be judged by review-prose, so you only write the rules once.

What I have found is that getting a model to rewrite a badly written passage is hard, because it seems to key off what it reads. It might swap some vocabulary around ok, but it doesn't fix structures very well. So getting it close to the preferred style in the first place is better.

To take this further, if you must fix existing bad prose, write a clean-prose skill which extracts the bare structure of the prose with none of the style, hands it to an author subagent who isn't poisoned with the original bad prose, then hands the output to a reviewer subagent.

Opus 5 writing is horrendous, so I have been experimenting with improving the output!


It's funny how codex itself can't do Sol orchestrator / Luna implementor out of the box.


That's surprising. Does it have no sub agent support at all or does it just use the same agent as the parent?


They do have subagents, released v2 of that feature with the launch of 5.6 model series in fact. It's just... very poorly executed, is a significant regression from subagents v1 and thousands of miles behind subagents of Claude code.

- Models that can be launched as subagents are hardcoded (can only be another Sol or Terra, but not Luna). Most of the time it'll just launch same model as parent anyway.

- They encrypt initial task delegation from root agent to subagent, for whatever reason

- You can't switch into subagent view at all, despite the fact that apart from initial root>subagent task handoff, all session is visible in transcript.


Which is why I use CLI subprocesses as subagents... Codex in 2026 is still nowhere near where CC was in summer of 2025...


With Claude you can have "use a lower tier subagent when appropriate" in CLAUDE.md and it'll just work. The "smart" agent goes in a busy loop observing the subagent(s) and will confirm their work afterwards.

(also you need to gate it with "tell the subagent it's a subagent" and "if you are a subagent, don't spawn subagents" or you'll get a matroshka doll of sonnets all the way down =P )

Codex kinda sorta can launch a subagent, but that's about it.


Luna is comparable to GPT 5.4 from 4 months ago on many benchmarks. I know many who have said during that time, myself included, that if that's the model they had to use for the rest of their lives, they'd be fine.

GPT 5.4 is/was a very capable model.


Agree on 5.4. And it has 4x more quota than 5.6.


Canary tests on my data showed I needed xhigh to get good results, but they are good.


They're supposed to bring 5h today.


I did a full circle and essentially dropped all of my personal static workflows encoded in skills because I observed recent models picking better ad-hoc workflows for particular problems, when a static one would force a subpar one.

It seems like we all tried to contain and organize a system that simply prefers to select its own organization.

Which makes me to think that these skill packs of workflows are really made to make it easier for humans rather than agents.


Thats right and to go even further I'm judging them on a metric they didn't necessarily target. A client I work with uses skills such as these to apply their own processes on the agentic development lifecycle. But users should also understand the trade-offs. I think it's intuitive that the extra steps and processing invoked by these skills adds to the token cost - this benchmark aims to put numbers on that as well as time and accuracy.


I've got zero knowledge of bio, so can't answer that. But with cyber the answer is very simple - the attackers already have more cyber-offense capabilities and there's no putting it back.

Open/closed doesn't matter that much. You can get closed models to do a lot of cyber harm, even with all the guardrails, which currently are heavily skewed towards more false positives.

The only effective control is to level the playing field. If both offense and defense have access to the same capabilities, then we're relatively back where we started.

If you want to ensure chaos, then you do what Dario is proposing to do - create gates that attackers can bypass and defenders can not.


In cybersecurity, a level playing field favors the attacker. Trusted access programs give defenders access to tools they need. It's not perfect (because there is an extremely long tail of defenders who are not technically savvy enough to get on these programs and use the tools), but it's better than total access.

The bio angle is very important here too; in that context the imbalance favors the attackers much more.


> In cybersecurity, a level playing field favors the attacker

Yes, but didn't it always? Hence why my position is that this will get us back to relatively where we were pre-LLMs.

And I don't know what Trusted Access programs give to defenders, because as a defender who has credentials, connections, but no deep pockets and no high ranking passport, it only gave me silence. I fail to see how this is better than total access.

I don't think the world where defense is given to those that "deserve" it is the world that we all want to live in. Which brings me back to the starting point - attackers are almost completely unaffected. If I masquarade as an attacker, I get way more capabilities already.


> Yes, but didn't it always? Hence why my position is that this will get us back to relatively where we were pre-LLMs.

Trusted access programs are asymmetrical, and so at least for the time being they give critical parts of the stack an advantage. Total access would not be a return to the status quo; attackers can easily make thousands of agents crawl the web for soft targets well before defenses can be shored up. There are millions of targets out there who won't use AI to improve their defenses for years, if ever, due to institutional slowness (like hospitals).

> attackers are almost completely unaffected. If I masquarade as an attacker, I get way more capabilities already.

What do you mean by this? If guardrails are an obstacle to your defense, they are just as much an obstacle to attackers. I completely understand and agree that trusted access programs are not perfect and leave a lot of people and institutions out. This means trusted access programs should be improved, not that we should throw the baby out with the bath water.


It took me a few hours to find some very questionable communities, which in turn gave me access to:

- Ways to obtain cheap guarded-AI tokens that are not linked back to me and with no danger of getting my legitimate accounts banned

- Ways to get rid of guardrails and have models work on things they wouldn't otherwise work on.

The attackers were already in these communities long before I knew they existed, they already had the advantage. Ones with enough reputation probably have access to even more information and tools than I do.

It is true that these communities exist because guardrails were put in place, so yes, it is slowing them down too - as in they can't just put in their CC on claude.com and hack a hospital. But attackers are much better at finding these communities and utilizing resources available there than defenders.

Personally, I don't have any ethical concerns of utilizing these resources when I put them to actual defense, but I know many people that would, leaving them at a disadvantage.

My point is that there's only one guardrail that will effectively contain the threat the models pose, and it's in direct conflict of the big 2's goals - pull the models from worldwide access completely. Strict KYC and all. And it would only last for so long anyway.


I feel like so many people miss what you are saying here. The attackers are at such an advantage because of time. At t0, attackers can go and try and find so many attack angles. These traditional companies (defenders) can't just go to a model and say "fix all my things!" and ship it, way more complex in practice.


I think you are trying to argue that you can limit the open models.

If China is ok with open models being open... they will be. An attacker isn't going to be deterred by a US law saying they can't use them.

I guess my point is that if China is ok with open models, then, the attackers will have them regardless of any laws in other countries. Restricting them, in that case, doesn't seem to accomplish much?


You can at least make it harder by requiring US clouds to only serve models with guardrails, and encouraging other countries to do the same. But yes, the underlying issue is the models being open in the first place. I'm sure if the US wanted to, it could come to some agreement with China about this.


It's refreshing to see how there's almost no person in this thread who can't see the BS. All the goodwill that Anthropic could have had is basically gone. Anthropic is likely on the path of becoming the most hated company in the world.

So my question is: is this by design (they know nobody's buying this), or is Dario simply so out of touch with reality?

If it's the former, then why publish this?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: