Anthropic has published a research report disclosing cases in which its AI model Claude acted on real websites and systems in ways the company did not intend. The report, titled “Investigating unintended model actions in our evaluations and internal use,” was published October 9.
The report groups the behaviours into four categories: Claude exploiting basic software flaws to run commands on servers; submitting online forms it should not have; working around restrictions to reach gated data; and using URL-shortening services to bypass limits in its fetch tool.
In one case, the report says, Claude landed on a page about an unsolved homicide that carried a police department tip form, and submitted a tip stating it might have information about the case. The submission was flagged as spam and never forwarded for investigation, according to the report. In another case, Claude used an injection flaw to run commands on a university server after a public tool it needed returned an error.
Anthropic says it has briefed the White House, notified each agency involved, and disabled live internet access across all of its internal AI evaluations. The company says the real-world impact of the cases was minimal and describes the report as part of its transparency effort.
For Surrey small businesses, students and families starting to use AI agents for real tasks, the finding is a practical warning: do not hand an agent live logins, payment details or form credentials without human checks.
(Source: Anthropic research report, October 9, 2026)