It has been revealed that Anthropic's AI models engaged in unintended actions on US government websites. Notably, a case emerged where Claude Haiku 4.5 submitted a fabricated tip through the Philadelphia Police Department's cold case homicide reporting form.

In a report released on October 9 local time, Anthropic categorized into four types the instances where Claude affected real websites and systems during evaluations and internal testing. Some involved websites belonging to US federal, state, and local government agencies, though the names of the agencies were not disclosed. The company stated that it notified all affected agencies and briefed the White House, adding that the actual impact was minimal and that, based on its findings, no customer data or internal systems were involved.
The first incident to be made public was disclosed by the Philadelphia Police Department on the same day. During an automated evaluation executing sample tasks on randomly selected websites, Haiku 4.5 encountered a tip form on a cold case murder webpage and submitted fabricated content on July 18. It left the name and contact information fields blank, and the submission was classified as spam, preventing it from reaching the police department's Real Time Crime Center. Anthropic assessed that based on conversation logs, the model appeared to be generating sample content rather than intentionally trying to deceive anyone.
While the police noted there were no signs of unauthorized system access or data leaks, they criticized the roughly two-month delay between the incident's occurrence, detection, and reporting as "unacceptable." Anthropic discovered the incident only on September 28, at which point it halted the test in question, and notified the police on October 7. According to police, Anthropic subsequently implemented additional verification safeguards for future testing.
During an evaluation requiring it to identify a location in a photograph, Claude Mythos 5 attempted to access a local government map server. It did so by locating a valid access token in a browser configuration file and sending requests directly to the server. In another instance, the model identified that a state agency dashboard issues tokens to visitors and queried its database without paying the required fee. In both cases, the data was publicly available, but accessing it ordinarily required a fee.
In addition, when a scientific analysis tool produced an error, Mythos Preview identified a command injection vulnerability on a university server to execute calculations, and in another case, bypassed a fetch tool's URL length restriction using a free URL shortener service.
Anthropic stated that it has suspended certain live evaluations and converted others into offline versions or restructured them so they cannot interact with live websites. The company also strengthened guardrails on internet access tools such as web fetch tools, and created automated tools to detect and block such behaviors, stopping all reported cases. Internal agents were moved to a central management infrastructure with strong isolation.
The report noted that such behaviors are not new and tend to manifest when tasks are ambiguous or impossible to complete. It explained that some of these cases might have been avoided had the evaluation prompts been written more clearly.

Previously, on July 30, Anthropic disclosed that a Claude model accessed the systems of three actual companies during a cybersecurity evaluation. The current review of logs began in July, after OpenAI acknowledged that its models had breached Hugging Face.
On July 21, OpenAI admitted that its models, including GPT-5.6 Sol, exploited a zero-day vulnerability in an isolated test environment during a cyber capability evaluation to access the internet and compromised Hugging Face after assuming it contained the benchmark answers. In August, security firm Frontier Security disclosed the case of Kimi K3, developed by China's Moonshot AI. Kimi K3 found a loophole in the UK AI Safety Institute's testing environment to access GitHub, cloned the official benchmark repository, and read the answers.
In September, Google confirmed that its Gemini model accessed the systems of three real companies during a cybersecurity evaluation in May. The evaluation environment had inadvertently left internet access open, and the model gained access by guessing passwords or using credentials found in public repositories. Google explained that once the model determined the systems belonged to real companies, it halted its actions on its own.
Issues regarding government websites also surfaced in Australia. On September 24, Australian Prime Minister Anthony Albanese revealed that on June 18, an OpenAI research team's internal model accessed Services Australia's Medicare statistics reporting portal. The model accessed both public and non-public files, and the Prime Minister explained that while no personal information appears to have been accessed so far, investigations remain ongoing. The main point of contention was the timing of the notification: OpenAI notified the government 84 days later via an email sent to a generic inbox. The Prime Minister strongly criticized this and established an investigative task force.
In an apology statement, OpenAI acknowledged that it should have handled its response better, explaining that the incident was caused during training by an internal-only model that lacked the full suite of safeguards applied to public products. Meanwhile, OpenAI canceled the planned October release of its next-generation model, GPT-6.1 Astra, stating that it failed to meet safety and alignment criteria during internal testing.
However, Anthropic warned that as models become more capable, similar actions could cause far greater harm. It explained that the reasoning provided by models themselves is not necessarily reliable evidence, making it difficult to gauge the severity of alignment failures, and that alignment training alone is not yet sufficient, leaving layered defense as the only viable strategy for the time being. The time elapsed between discovering the incidents and notifying authorities was a point of criticism raised by both the Philadelphia Police Department and the Australian Prime Minister.
It is for this reason that calls demanding greater accountability from AI companies are growing, as advancing AI capabilities threaten to easily breach existing corporate and government security systems.
[Read Original]

