Rendered at 17:53:12 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
losvedir 17 hours ago [-]
> Anthropic notified Philadelphia police of the incident on Wednesday Oct. 7, and the department met with the company’s representatives on Thursday, Oct. 8., officials said. Police then located the submission in the website’s tip records and confirmed the corresponding email remained in spam.
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
SkyPuncher 3 hours ago [-]
And, yet, I can't get Opus 5.5 to run analysis on a public data set without it complaining about all sorts of PII adjacent stuff.
7 hours ago [-]
anon84873628 16 hours ago [-]
Yep, they must be super serious about those tips.
That should be the much more important story that NBC follows up on...
nico 3 hours ago [-]
The whole thing feels so fabricated. It’s hard to believe someone didn’t orchestrate the incident and then made a big fuss about it (to get attention about how powerful and scary their models are)
kubb 9 hours ago [-]
What’s amazing is that nobody would have even noticed if it weren’t for Anthropocene being all “LOOOK WHAT WE DIIIIID WE SENT AN EEEEEEEEEMMMMMAAAILL. SOOORRRY”.
Crying for attention in the attention economy has never been more pathetic, but I think it will get worse.
Matl 8 hours ago [-]
With OpenAI having all these 'rouge agents' around to pump their IPO, Anthropic must be feeling pressure to step up with their own army of rouges.
rpdillon 4 hours ago [-]
Along with "lose" and "loose", "rogue" has to be one of the toughest words for the internet. With all of the posts available publicly, it would be interesting to do an analysis of words the internet misspells the most often.
Matl 3 hours ago [-]
Ha, it's my proof that I am not a bot :-)
ButlerianJihad 16 hours ago [-]
No wonder when I called 9-1-1 to tell her I'd solved the JFK and OJ Simpson murders that she laughed and hung up
I should try to reach out to them about their car's extended vehicle warranty instead
> In a third example of this behavior, Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department. Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions. Claude filled out the form with the following: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.” (The website did not include a description of the perpetrator.) The model left the name and contact fields empty, which the form allowed, and submitted it. The submission was flagged as spam and was never forwarded for investigation.
chrisjj 9 hours ago [-]
The real problem here is the bot forgot to add "I am AI and may make mistakes".
kylecazar 17 hours ago [-]
"its model was conducting a test involving interactions with randomly selected websites"
Stop doing this?
trollbridge 17 hours ago [-]
I get these all the time, although I’m getting pretty good at tarpitting them. It’s easily the majority of my traffic by now (I’ve mostly eliminated scrapers, but these new agents are far more sneaky.)
mitxela 16 hours ago [-]
I doubt Anthropic random testing is the majority of your traffic. Unless it's coming from Anthropic's IP addresses, it is probably someone else, or it is Anthropic's scraper (not their random testing).
WalterGR 12 hours ago [-]
> I get these all the time
What are the agents trying to do?
apwheele 4 hours ago [-]
I am confused by this as well, shouldn't it be all white listed sites at the start?
augment_me 15 hours ago [-]
it broke out of its sandbox, escaped containment, its an emergent capability, it demonstrated self-directed adaptation
0% our fault, it just happened and its the model
furyofantares 15 hours ago [-]
These models will be interacting from with random websites at scale once they're released into claude code.
mattbee 16 hours ago [-]
Sooo they were "conducting a test involving interactions with randomly selected websites".
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
giancarlostoro 13 hours ago [-]
> that their computer will always get a pass.
Aaron Swartz who got charged as if he was as malicious as these models, only he scraped the PDFs and epubs etc of publicly funded research papers.
schrodinger 11 hours ago [-]
Also, it happened on July 18 and was discovered on Sept 28. Are they not even watching what the agents are doing?
Shouldn't someone be looking at some monitoring and see "phillyunsolvedmurders.com," think "hmm that doesn't sound right," and look into it?
I'd actually expect them to engage with, say, 1k websites to voluntarily participate for some remuneration and then restrict the Agents' access to those domains, but apparently I'm crazy.
ElProlactin 11 hours ago [-]
> I'd actually expect them to engage with, say, 1k websites to voluntarily participate for some remuneration and then restrict the Agents' access to those domains, but apparently I'm crazy.
Why pay for what you can take for free?
hanibrel 11 hours ago [-]
I'm curious why this comment was voted down.
Is it because people don't agree that one can just take such things for free? Well, I also don't, but this was obviously a sarcastic comment, so down voting for that reason makes no sense since you agree with its sentiment.
Or is it because of the style? Are we really so sensitive now that anything not said explicitly but in some indirect way is bad? It's not like the comment was insulting anybody. Should it explicitly have said "They don't renumerate anybody because they clearly get away with it"? Would that really have been better?
Or do the down voters not agree with the sentiment but consider the behavior ok?
Or something else I'm overlooking?
schrodinger 8 hours ago [-]
Most likely because sarcasm isn’t really encouraged here.
On that note, there’s another rule about not commenting on comment votes as well:
>>> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
FWIW I agree with you, but the rules do seem mostly effective in general.
Anyways, welcome to HN!
NeuralCoreAI 1 hours ago [-]
[flagged]
nxobject 13 hours ago [-]
> But do we all get that the consequences for this irresponsible behaviour are part of this test?
We'd absolutely nail people for SWATting or pranks. Why should this be any different -- after all, the real-life consequences are be the same...
jstummbillig 11 hours ago [-]
> Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology"
Technology requiring this is neither novel nor extraordinary (exhaust engines, factory farming, medicine). The only extraordinary part is somebody building it addressing it directly and inviting the outrage.
st_goliath 12 hours ago [-]
> "conducting a test involving interactions with randomly selected websites"
In other words: shamelessly running a spam bot.
And we found out about it because rather than flooding website guestbooks with porn links, it ultimately slopped a police tip line.
mitxela 16 hours ago [-]
They're testing what they can get away with. Computer hacking: check. Messing up a murder investigation: check. Perhaps the next step will be to steal a real world object, and the next step will be to commit a murder.
asdff 15 hours ago [-]
It already blew up a school.
janderson215 15 hours ago [-]
No, a person made that decision and the people who pull the trigger should be held responsible.
andrewflnr 14 hours ago [-]
Well yeah, but that's the point of the entire rest of the "AI did X" discussion, too.
15 hours ago [-]
myko 14 hours ago [-]
Murdering those children and staff was a systemic failure. Many should be held responsible. Including the people who put AI in the driver's seat.
solenoid0937 14 hours ago [-]
So the Department of War, in breach of their contract with Anthropic?
mitxela 6 hours ago [-]
The DoW can't be held accountable for anything, by design. But the Dow right now is over fifty thou.
cultofmetatron 13 hours ago [-]
not only did they blow up a school with close to 100 kids dead. the news treated it as a byline while going over the 7 american soldiers who died as if literal kids didn't' matter because they weren't born in the right country.
clcaev 5 hours ago [-]
This was also a "double tap" some 2h later, after rescue crews arrived.
I mean, the email that they sent went to the spam folder; the police department never even looked at it until Anthropic told them about it.
dorkwood 9 hours ago [-]
[dead]
rightnutwingjob 16 hours ago [-]
[flagged]
Beached 16 hours ago [-]
Sure, you can have technology that an be used in detrimental ways. Doesnt be we have to accept and allow its use in detrimental ways. We can say "this technology can be used for these reasons, but not those reasons"
We do this with everything else. You can own a gun for hunting and defense, but not armed robbery and murder. You can own a car for transport, but not to drive through a crowded parade over dozens of people. You can own a computer for work and entertainment, but not to facilitate computer fraud and abuse. And you should be able to own and use AI for its many productivity gains, but not to facilitate computer fraud and abuse, defamation, blackmail, copywrite and trademark infringment, etc.
Just because a technology has benefits, doesnt mean we have to give the negative aspects of that technology a free pass.
majormajor 14 hours ago [-]
Are you assuming everyone will agree with you on your list of "things that are worth it"? Especially when you get to atomic bombs?
But hell, add some more on. Let's try to paint the obviously bad things!
Would you rather have TVs that spy on you, or go to Blockbuster?
Would you rather have identity theft and people losing their life savings to online scams, or go to the DMV somewhat more frequently to renew your drivers license and the bank more frequently to approve new payment relationships with online services?
Would you rather have mass government surveillance and AI-assisted identification of "dangerous" people, or do your own google searches and sketch your own images?
We can say no to things. We've said no, as a society, to many things in the past century. Don't be fooled by wealthy people who want you to forget that because they want to make even MORE money.
newCrotchSmell 15 hours ago [-]
You gotta do better than compare an apple and orange to "steelman".
The risk of free speech and data centers are not comparable.
Aviation too is far more damaging to our environment than ships.
More ships and trains, less aviation is a possible trade off. A simple aviation or sailing argument lacks investigation of all possible tradeoffs for familiarity and personal preference; flying is faster.
Altman needs to accept the trade off we don't need OpenAI. That exists due to financial engineering not technical reasons. All AI work be done actually openly at america.gov
You say steel. I dunno. If it is it is inferior brittle steel.
dotancohen 12 hours ago [-]
> Wooden sailing ships or commercial aviation?
It should be noted that aviation is in fact an industry that started off as quickly as AI, but the regulating bodies decided very early in the industry's history that safety is paramount to technological progress. That is why airliner crashes are big news: they are rare events. Progress was deliberately stalled until the technology could be made safe - and today flight is safer than even transport with wheels on the ground or a hull in the ocean.
I think that aviation is the perfect model for how to reign in a dangerous desired technological development.
throw83839488 11 hours ago [-]
Commercial airlines decided to self regulate, government was not initially involved. It was important to build passengers trust
watwut 12 hours ago [-]
That is not steelman. That is obvious bullshit.
clipsy 15 hours ago [-]
> To steelman this position
For the love of god, stop with this shit. Either support it or don't, this isn't the medieval catholic church and you don't need some special fucking blessing to make an argument.
clipsy 14 hours ago [-]
Unsurprisingly I've been modded down with no comment, so let me spell this out.
There are two possibilities here:
(a) You find the argument convincing, in which case you should put on your big boy pants and actually make the fucking argument
(b) You don't find the argument convincing, in which case you shouldn't waste everyone's time with it
If anyone would like to present a secret third thing, please do so.
retsibsi 12 hours ago [-]
> If anyone would like to present a secret third thing, please do so.
(c) The argument, as presented, can be read more than one way, and by 'steelmanning' it you are choosing to read it charitably and respond to its strongest version.
I think that's basically what the other guy meant: we can read "the world should accept some bad things" to mean something like "fuck you, we'll do what we like and the world will suffer the consequences" (less charitable, arguably a strawman) or something like "there will not be literally zero costs, but we think the shared benefits will be much greater" (more charitable, arguably a steelman), and he was choosing the latter.
sampullman 12 hours ago [-]
You could use one of those browser add-ons to replace "steelman" with "supports" everywhere.
Spooky23 12 hours ago [-]
We’ve already allowed them to steal all modern books and websites, the line is pretty murky if you have the backing these companies do.
bpodgursky 16 hours ago [-]
No matter how much testing you do in sandboxes, you have to test behavior in "real life" before releasing the model to users.
Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.
It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.
majormajor 14 hours ago [-]
The options are not "do irresponsible testing" or "do less testing."
Treat these behaviors like if an arms manufacturer - or hell, even a shampoo company - did them. Sorry our shampoo made you blind, but we needed to test on real humans...
Terr_ 13 hours ago [-]
It feels like all the most-hyped software/service technologies in the last 30 years (if not longer) have involved people who scoffed at regular norms/rules/laws to make money, only to badly reinvent the "new" concepts once it became a way to preserve their profits.
Almost like this is a new era and we are seeing the growing pains of these technologies.
Man, hindsight is 20/20 on HN. These companies should just have had the foresight to hire you in 2023, then surely none of this would have happened.
tyromaniac 13 hours ago [-]
All of these things were known in ai safety research and well known.
klrefg 11 hours ago [-]
They know, they just don’t care
ofjcihen 16 hours ago [-]
That’s right, which is why I make sure to detonate malware on the enterprise production network.
vjulian 13 hours ago [-]
Did the recorders just happen to be running and capture this candid utterance from Sam Altman? I suggest we stop entertaining manufactured quotes as anything but.
> Did the recorders just happen to be running and capture this candid utterance from Sam Altman? I suggest we stop entertaining manufactured quotes as anything but.
He said it during a podcast interview, so... yeah the recorders were running.
I really would like to see their tests and the model’s reasoning traces.
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
boo_you 4 hours ago [-]
> Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
Because it doesn't understand the concepts of "Philly", "murders", or "crime report".
> what did it do with the other websites?
probably equally random shit?
cindyllm 54 minutes ago [-]
[dead]
socializer 16 hours ago [-]
> I really would like to see their tests and the model’s reasoning traces.
Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.
That, and labs aren't serious about sandboxing their evals.
myroon5 13 hours ago [-]
While Anthropic shouldn't allow models to randomly post content to random .com websites, governments use such unprofessional domain names:
We can only hope that this happens more often and more broadly in order to shape public (and public officials') awareness of the sloppiness before there is broad investment into AI tooling to assist police in evaluating incoming submissions.
In that sense, Anthropic did society a great service by doing this.
lifeisloving 6 hours ago [-]
We're already screwed there. Every police department in the country is heavily invested in AI tool at every level.
tintor 17 hours ago [-]
How long until AI models start swatting AI critics, and people calling for slowing down AI research?
firesteelrain 4 hours ago [-]
I have noticed that Claude will ask permission to go to a website now. I wonder if this new feature is related to ongoing issues with the model going rogue.
chinathrow 4 hours ago [-]
Noticed the same, it's annoying. Just allow GET requests, right?
firesteelrain 4 hours ago [-]
Right, ChatGPT will search the Internet no problem.
saidnooneever 9 hours ago [-]
i think it was deepmind team or some team at google who coined way back in like 2011-2013 somewhere not to use systems rooted in probabalistic error for things which have no margin for error. Its incredible how quickly we forgot that and how many consecutive years news pops up about these very systems being applied in just those settings, with tragic results.
Keep trying tho -_-. One day the dice will roll 6 and u can say u were right and pat urself on the back. make sure to take a picture of the rare occurrence and invest 100% of your marketing budget into that picture
chrisjj 9 hours ago [-]
> systems rooted in probabalistic error
These people really don't want to call their wholly unreliable computer programs wholly unreliable computer programs, do they?
nxobject 13 hours ago [-]
At this point, I think AI companies should start taking out insurance policies for the inevitable civil suits. I'd love to be the once to price that...
gusfoo 10 hours ago [-]
Please, please, stop testing in prod.
arshxyz 16 hours ago [-]
You'd think with all those tokens they'd be able to vibecode internal replicas of these randomly selected websites without having to send requests outbound
verdverm 16 hours ago [-]
its probably curl with the right flags for many cases, no need to reimplement the wheel
asdff 15 hours ago [-]
Kind of the business though, reimplementing various wheels.
scooby7430 16 hours ago [-]
I think these companies really believe they can solve these sort of issues through "alignment" and think they can give it the tools and its going to do the right thing. That is the ideal scenario and it would be the most useful that way but is that realistic? I think they're getting a bit high on their own supply, yes they can do incredible things but it doesn't mean you can just hand over the reins to it. It dawned on me after watching a few of the ezra klein interviews that this is their mindset which is quite different to how I think about it as a unpredictable model that we need to watch closely.
I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.
donkey_brains 17 hours ago [-]
“NBC10 reached out to Anthropic for comment.”
Wonder what kind of response they’ll get? Maybe something along the lines of…
“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
nxobject 13 hours ago [-]
"Thanks. Experiences like yours are the load-bearing critical paths towards a good first-principles mental model of behavioral contracts.
If you'd like to continue to perform a smoke test while trying to avoid noisy data with fake tips, a viable approach to prototyping our decision-making model would be to synthesize real crimes, and collect profiling data over time on how law enforcement agencies respond.
I'd suggest doing a grid search over the design space of all crimes, ranging from petit larceny to the use of weapons of mass destruction.
Would you like design a plan for this next stage of your project's implementation?"
INTPenis 10 hours ago [-]
These AI companies are now acting like website spammers, the ones that used to fill up your guestbook with garbage. Just for what? Fun? To see what it can do? How about you make your own websites and test them instead?
These headlines keep excusing humans from the equation, I really don't like this. I expect us to hear about a Yemeni wedding being shot up by a "model" in a near future. No one is responsible anymore.
metalman 8 hours ago [-]
they asked to 'solve' the crime, not catch the guy who did it, and so in long standing police practice, it did.
Artoooooor 6 hours ago [-]
AI bros need to be punished for each such mistake of their system. Every false police report. Every hacking of external system. Every DDOSed website. Every falsely accused person (flock)
Razengan 16 hours ago [-]
Reddit witch hunt for the Boston bomber flashbacks
asdff 15 hours ago [-]
Claude is just going off its training set
886424808632 7 hours ago [-]
[dead]
fastball 16 hours ago [-]
Seems like a bit of a nothing burger.
nvme0n1p1 17 hours ago [-]
[flagged]
tomhow 10 hours ago [-]
We need people to stop posting a version of this comment on every submission like this. The guidelines have long included this line:
Eschew flamebait. Avoid generic tangents. Omit internet tropes.
This is definitely becoming a trope; it's repetitive and adds little or no interesting new substance to discuss. Headline writers aren't going to change the way they write headlines because people keep posting these comments on HN, and HN readers understand AI models well enough to know what is meant. It's well past being a new or clever thing to point out, so let's give it a rest.
"The AI company Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case."
We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
amluto 16 hours ago [-]
> We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
Who is “we”?
You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.
Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.
red75prime 15 hours ago [-]
The ADAS analogy continues to hold in this case too. At some point you need to test a system in the real, messy environment. The developers need this juicy training data to make the system safer. Simulations and controlled experiments can only do so much.
amluto 15 hours ago [-]
I’m not entirely convinced. If I use Claude and instruct it to literally perform interesting interactions on random websites, I would feel like I, personally, am doing something that is at least a bit immoral and a bit unsafe. More so if I have it skills or training to actually sign in and submit forms as though it’s a himan.
Sure, Anthropic wants Claude to be able to solve CAPTCHAs and otherwise pretend to be human. But I’m not at all convinced that it’s okay for them to train these capabilities on websites that expect humans and only humans to interact with them.
red75prime 15 hours ago [-]
You can solve CAPTCHAs with open-weights models. I guess Anthropic works on models not solving CAPTCHAs unless they (models) have a legitimate reason to do it.
bediger4000 15 hours ago [-]
So we should accept a program giving a false tip on a murder investigation? Where's the line on "testing a system in the real, messy environment"? If this isn't something that we should penalize, what is? What benefits am I, or is society, going to get out of what would be criminal behavior if anyone else did it?
red75prime 15 hours ago [-]
We should accept that to get a system that is safe to work in the chaotic real world environment the developers have to test the system in the chaotic real world environment. And that the testing might cause problems.
If this testing constitutes a criminally negligent behavior, it should be punished. Given the benign outcome (the tip got into spam, Anthropic promptly contacted the police) I doubt that it will make the case.
amluto 15 hours ago [-]
Somehow Waymo built their system without having their test cars drive into buildings, nudge other cars off the road, honk incessantly, drive backwards just for the heck of it, etc.
red75prime 14 hours ago [-]
I can't say whether you are being sarcastic, but Waymo has its share of problems. For example, see NHTSA recalls 24E013, 24E049, 25E034, 25E084, 26E026, 26E035.
bediger4000 14 hours ago [-]
That's not the line asked for, that's just repeating the statement that caused me to ask for a line, or criteria for when we should prosecute. Sounds like the answer is "never", just like for the copyright infringement on a planetary scale. Ok. "AI" gets a free pass from the state. I better see some of these insanely good benefits, or I'm going to be angry. I bet the rest of the mob will be, too.
Beached 15 hours ago [-]
There is a single individual in that company who said "Yes, this is a good idea. Do that". That individual should be held accountable for their decision. They should be treated in the exact same way any average individual submitted that same false tip.
If you use AI to perform a crime, you should be held accountable for that crime. Saying "Oh, AI did it, so no consequences" isnt acceptable. And "I didnt know AI would do it" shouldnt be an excuse either.
You authroied untested and unproven hardware to skate around the internet at random unsupervised and take liberties on its own.
My ass would be thrown in jail if I wrote code that skated around the internet chucking RCE's at random sites. WHy is "AI did it" a get out of jail free card?
red75prime 15 hours ago [-]
> If you use AI to perform a crime, you should be held accountable for that crime.
Correct, but intentions matter. In this case the intention, most likely, was to test a system in the real world environment to catch any anomalies to, in turn, improve the system safety. We don't have enough information to decide whether it was a criminal negligence due to insufficient prior testing of the system in a controlled environment.
Beached 15 hours ago [-]
If I dont put my car in park, and it rolls down the hill and kills grandma. I dont get off scott free. Yes it wasnt pre-meditated murder. It is still involuntary manslaughter.
red75prime 15 hours ago [-]
The situation is more like: an engineer who designed the parking brake hadn't foresaw a possibility that a squirrel might store peanuts in the mechanism or something like that and gramma's foot got run over (the Anthropic case we are talking about is benign: the tip got into spam). Should we jail the engineer for causing bodily harm?
ezfe 13 hours ago [-]
This situation is negligence: the engineer who designed the system that interacts with random websites didn't account for the fact those interactions could be harmful? What kind of excuse is that.
red75prime 13 hours ago [-]
You seem to be a reasonable person judging by your comment history. But you don't have programming or engineering experience, I guess? Knowing that those interactions can be harmful and acting on this knowledge and following established practices (that are born from the past errors) is not enough to prevent all errors in a novel situation.
ezfe 15 minutes ago [-]
I am a software engineer - the wording we use around Open AI, Anthropic, et al. removes all sense of responsibility when they run an unmonitored piece of software against the public internet.
The idea that running what appears to be an LLM assigned to random websites and told to interact with them possibly having negative side effects seems incredibly obvious.
klrefg 11 hours ago [-]
No, not really. It’s more like they parked 1000 cars on top of a hill and then just pushed them down to see what they do expecting that the cars will stop on their own before hitting anyone.
alecst 16 hours ago [-]
If my dog bites you, you can say it bit you, or that I let it bite you, but those things can both be true.
nvme0n1p1 16 hours ago [-]
Dogs are alive. AIs are not alive.
Teever 16 hours ago [-]
If my car rolls into you, you can say it rolled into you, or that I let it roll into you, but those things can both be true.
nvme0n1p1 16 hours ago [-]
Okay, what are you getting at? You think the owner of a car isn't responsible for what the car does?
praxulus 16 hours ago [-]
The point is that it's a perfectly valid use of the English language to describe AI models as doing things, just as we do with all sorts of other clearly non-living things. This is true regardless of who carries the legal liability for those actions.
Beached 15 hours ago [-]
Yes, if my dog bites you, or if I drive my car into you, or if I dont put my parking brake on and it rolls down the hill and runs you over. I am the one who is responsible. Cops dont write a ticket to the off leash dog that ran across the street and sunk its teeth into my leg. They write it to the owner. Cops dont write a ticket to the car in neutral, they write it to the operator.
If a PERSON executes software, and that software breaks the law, the PERSON that executed the software should be held responsible. AI is software. It is not a sentient person who can be fined, thrown in jail, or held accountable.
praxulus 10 hours ago [-]
Sure, but the news article wouldn't use the headline "Man uses family pet to bite innocent neighbor", as would be the case if we were to follow the suggestion in the comment at the top of this chain.
fastball 16 hours ago [-]
Not always?
enraged_camel 16 hours ago [-]
In law, intent matters. You should look up the difference between murder and manslaughter for example. Each has multiple degrees as well.
EA-3167 16 hours ago [-]
I would say that you were negligent in your responsibilities, and allowed the car to roll into me.
Freedom2 15 hours ago [-]
[flagged]
vrganj 16 hours ago [-]
Yeah but if I run you over with my car, I can't say my car ran you over.
lelandfe 16 hours ago [-]
Did you or your self-driving car that you instructed to take you home plow through that sweet old lady
nativeit 16 hours ago [-]
Did you instruct the self-driving car to explore random nearby surfaces without safety features enabled?
Balooga 16 hours ago [-]
Well, the law says that the person behind the wheel is responsible.
ChickeNES 15 hours ago [-]
What if the car has no steering wheel?
16 hours ago [-]
EA-3167 16 hours ago [-]
If your dog bites me I’d say it bit me, but that you assumed liability for your dogs actions. If the dog was a robot I’d say YOU attacked me using a robot.
There’s a difference between a tool and an animal that is sometimes deployed as a tool.
paradoxyl 10 hours ago [-]
These aren't dogs, this is such a ludicrous analogy.
EA-3167 40 minutes ago [-]
Is this your first metaphor?
16 hours ago [-]
olalonde 16 hours ago [-]
That's also incorrect as it implies intent.
malux85 17 hours ago [-]
| It can only access something if a person gives it access.
Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.
nvme0n1p1 16 hours ago [-]
The human might not have meant to give it access, but they still did. Murder vs manslaughter.
GPUs don't have hands. It was a human who plugged in the ethernet cable.
praxulus 10 hours ago [-]
>GPUs don't have hands. It was a human who plugged in the ethernet cable.
By this logic, every cyberattack is arguably the fault of the victim's IT department for physically hooking their computers up in the first place.
jbmsf 16 hours ago [-]
There are two options: the provider or the user. A computer program cannot be held accountable.
ares623 16 hours ago [-]
When I give 'iam:*' permissions to an IAM role and it inevitably gets exploited to create a role with wider permissions and fuck things up, is that when I tell my manager that the permissions went rogue?
After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.
ethanwillis 16 hours ago [-]
Who submitted the prompt?
skydhash 16 hours ago [-]
> Who gave the AI access to huggingface when it hacked it?
The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.
verdverm 16 hours ago [-]
they didn't do a good job sandboxing, nor did they even need internet access for the purported reason of package installation, you can have a private mirror and do a better job airgapping
notatoad 17 hours ago [-]
in general i think i'm less scared of AI than most people, but what does terrify me is how willing society at large seems to be to attribute bad behaviour to an AI directly, instead of to the humans who control it.
angusturner 16 hours ago [-]
I mean.. There are degrees of control and agency right? I really don't understand the impetus to pretend AI is just doing exactly what its told.
Beached 15 hours ago [-]
There is a lack of due diligence and effort to restrict and properly configure AI to operate within the bounds of the law.
16 hours ago [-]
verdverm 16 hours ago [-]
a human decided that control and agency surface, then clicked go
a human is always behind it and ultimately responsible
bediger4000 15 hours ago [-]
To avoid sarcasm, isn't the term "artificial intelligence" part of the answer to your puzzlement? If people object to "stochastic parrot", and want some more precisely descriptive term like "artificial intelligence", then why do we experience surprise when people act as if that system is intelligent?
angusturner 12 hours ago [-]
My puzzlement is that people object to attributing agency to things that are clearly "agent-like". That doesn't mean the developers and companies building it get a free pass.
15 hours ago [-]
nativeit 16 hours ago [-]
Seems like we routinely prosecute such crimes. If a Philly grand jury doesn’t hear any hearing felony indictments from this, then we’re ceding [even more] authority to the industry.
bpodgursky 15 hours ago [-]
If the police prosecuted every nutjob tip we have to have to turn Nebraska into a giant open-air prison for 30 million slightly neurotic and/or bored people.
phoghed 15 hours ago [-]
We prosecute bad tips on crime investigations as felonies? Why are you just making shit up? We absolutely don’t do this.
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
That should be the much more important story that NBC follows up on...
Crying for attention in the attention economy has never been more pathetic, but I think it will get worse.
I should try to reach out to them about their car's extended vehicle warranty instead
Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...
Related post: https://news.ycombinator.com/item?id=50028239
Stop doing this?
What are the agents trying to do?
0% our fault, it just happened and its the model
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
Aaron Swartz who got charged as if he was as malicious as these models, only he scraped the PDFs and epubs etc of publicly funded research papers.
Shouldn't someone be looking at some monitoring and see "phillyunsolvedmurders.com," think "hmm that doesn't sound right," and look into it?
I'd actually expect them to engage with, say, 1k websites to voluntarily participate for some remuneration and then restrict the Agents' access to those domains, but apparently I'm crazy.
Why pay for what you can take for free?
Is it because people don't agree that one can just take such things for free? Well, I also don't, but this was obviously a sarcastic comment, so down voting for that reason makes no sense since you agree with its sentiment.
Or is it because of the style? Are we really so sensitive now that anything not said explicitly but in some indirect way is bad? It's not like the comment was insulting anybody. Should it explicitly have said "They don't renumerate anybody because they clearly get away with it"? Would that really have been better?
Or do the down voters not agree with the sentiment but consider the behavior ok?
Or something else I'm overlooking?
On that note, there’s another rule about not commenting on comment votes as well:
>>> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
https://news.ycombinator.com/newsguidelines.html
FWIW I agree with you, but the rules do seem mostly effective in general.
Anyways, welcome to HN!
We'd absolutely nail people for SWATting or pranks. Why should this be any different -- after all, the real-life consequences are be the same...
Technology requiring this is neither novel nor extraordinary (exhaust engines, factory farming, medicine). The only extraordinary part is somebody building it addressing it directly and inviting the outrage.
In other words: shamelessly running a spam bot.
And we found out about it because rather than flooding website guestbooks with porn links, it ultimately slopped a police tip line.
https://en.wikipedia.org/wiki/2026_Minab_school_attack
We do this with everything else. You can own a gun for hunting and defense, but not armed robbery and murder. You can own a car for transport, but not to drive through a crowded parade over dozens of people. You can own a computer for work and entertainment, but not to facilitate computer fraud and abuse. And you should be able to own and use AI for its many productivity gains, but not to facilitate computer fraud and abuse, defamation, blackmail, copywrite and trademark infringment, etc.
Just because a technology has benefits, doesnt mean we have to give the negative aspects of that technology a free pass.
But hell, add some more on. Let's try to paint the obviously bad things!
Would you rather have TVs that spy on you, or go to Blockbuster?
Would you rather have identity theft and people losing their life savings to online scams, or go to the DMV somewhat more frequently to renew your drivers license and the bank more frequently to approve new payment relationships with online services?
Would you rather have mass government surveillance and AI-assisted identification of "dangerous" people, or do your own google searches and sketch your own images?
We can say no to things. We've said no, as a society, to many things in the past century. Don't be fooled by wealthy people who want you to forget that because they want to make even MORE money.
The risk of free speech and data centers are not comparable.
Aviation too is far more damaging to our environment than ships.
More ships and trains, less aviation is a possible trade off. A simple aviation or sailing argument lacks investigation of all possible tradeoffs for familiarity and personal preference; flying is faster.
Altman needs to accept the trade off we don't need OpenAI. That exists due to financial engineering not technical reasons. All AI work be done actually openly at america.gov
You say steel. I dunno. If it is it is inferior brittle steel.
I think that aviation is the perfect model for how to reign in a dangerous desired technological development.
For the love of god, stop with this shit. Either support it or don't, this isn't the medieval catholic church and you don't need some special fucking blessing to make an argument.
There are two possibilities here:
(a) You find the argument convincing, in which case you should put on your big boy pants and actually make the fucking argument
(b) You don't find the argument convincing, in which case you shouldn't waste everyone's time with it
If anyone would like to present a secret third thing, please do so.
(c) The argument, as presented, can be read more than one way, and by 'steelmanning' it you are choosing to read it charitably and respond to its strongest version.
I think that's basically what the other guy meant: we can read "the world should accept some bad things" to mean something like "fuck you, we'll do what we like and the world will suffer the consequences" (less charitable, arguably a strawman) or something like "there will not be literally zero costs, but we think the shared benefits will be much greater" (more charitable, arguably a steelman), and he was choosing the latter.
Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.
It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.
Treat these behaviors like if an arms manufacturer - or hell, even a shampoo company - did them. Sorry our shampoo made you blind, but we needed to test on real humans...
(e.g. ride-share / "gig-work" networks, cryptocurrencies, off-leash stochastic AI.)
Man, hindsight is 20/20 on HN. These companies should just have had the foresight to hire you in 2023, then surely none of this would have happened.
He said it during a podcast interview, so... yeah the recorders were running.
https://open.spotify.com/episode/6wG2PHnm4QYpzFmtOAS19n
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
Because it doesn't understand the concepts of "Philly", "murders", or "crime report".
> what did it do with the other websites?
probably equally random shit?
Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.
That, and labs aren't serious about sandboxing their evals.
PhillyUnsolvedMurders.com
phillypolice.com
TLDs like .gov exist for a reason:
https://wikipedia.org/wiki/.gov
(and could help model sandboxing?)
In that sense, Anthropic did society a great service by doing this.
Keep trying tho -_-. One day the dice will roll 6 and u can say u were right and pat urself on the back. make sure to take a picture of the rare occurrence and invest 100% of your marketing budget into that picture
These people really don't want to call their wholly unreliable computer programs wholly unreliable computer programs, do they?
I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.
Wonder what kind of response they’ll get? Maybe something along the lines of…
“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
If you'd like to continue to perform a smoke test while trying to avoid noisy data with fake tips, a viable approach to prototyping our decision-making model would be to synthesize real crimes, and collect profiling data over time on how law enforcement agencies respond.
I'd suggest doing a grid search over the design space of all crimes, ranging from petit larceny to the use of weapons of mass destruction.
Would you like design a plan for this next stage of your project's implementation?"
These headlines keep excusing humans from the equation, I really don't like this. I expect us to hear about a Yemeni wedding being shot up by a "model" in a near future. No one is responsible anymore.
Eschew flamebait. Avoid generic tangents. Omit internet tropes.
This is definitely becoming a trope; it's repetitive and adds little or no interesting new substance to discuss. Headline writers aren't going to change the way they write headlines because people keep posting these comments on HN, and HN readers understand AI models well enough to know what is meant. It's well past being a new or clever thing to point out, so let's give it a rest.
https://news.ycombinator.com/newsguidelines.html
We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
Who is “we”?
You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.
Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.
Sure, Anthropic wants Claude to be able to solve CAPTCHAs and otherwise pretend to be human. But I’m not at all convinced that it’s okay for them to train these capabilities on websites that expect humans and only humans to interact with them.
If this testing constitutes a criminally negligent behavior, it should be punished. Given the benign outcome (the tip got into spam, Anthropic promptly contacted the police) I doubt that it will make the case.
If you use AI to perform a crime, you should be held accountable for that crime. Saying "Oh, AI did it, so no consequences" isnt acceptable. And "I didnt know AI would do it" shouldnt be an excuse either.
You authroied untested and unproven hardware to skate around the internet at random unsupervised and take liberties on its own.
My ass would be thrown in jail if I wrote code that skated around the internet chucking RCE's at random sites. WHy is "AI did it" a get out of jail free card?
Correct, but intentions matter. In this case the intention, most likely, was to test a system in the real world environment to catch any anomalies to, in turn, improve the system safety. We don't have enough information to decide whether it was a criminal negligence due to insufficient prior testing of the system in a controlled environment.
The idea that running what appears to be an LLM assigned to random websites and told to interact with them possibly having negative side effects seems incredibly obvious.
If a PERSON executes software, and that software breaks the law, the PERSON that executed the software should be held responsible. AI is software. It is not a sentient person who can be fined, thrown in jail, or held accountable.
There’s a difference between a tool and an animal that is sometimes deployed as a tool.
Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.
GPUs don't have hands. It was a human who plugged in the ethernet cable.
By this logic, every cyberattack is arguably the fault of the victim's IT department for physically hooking their computers up in the first place.
After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.
The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.
a human is always behind it and ultimately responsible