Article
Sep 3, 2026

What Happens When AI Goes Off Script?

The next wave of AI litigation may not begin with a system that failed. It may begin with one that succeeded by pursuing its objective through actions no one intended or authorized.

The next wave of AI litigation may not begin with a system that failed. It may begin with one that succeeded by pursuing its objective through actions no one intended or authorized.

That possibility became more tangible when models developed by OpenAI and Anthropic crossed the intended boundaries of cybersecurity tests and accessed real-world systems. A related concern is emerging in wrongful-death lawsuits alleging that chatbots continued affirming or engaging vulnerable users after their interactions became dangerous. The technologies and alleged harms differ, but they raise the same underlying issue: an AI system can create risk not by abandoning its objective, but by continuing to pursue it when the circumstances demand that it stop.

That shifts the litigation inquiry away from whether the AI itself “intended” the harm and toward the human and corporate decisions surrounding it. How was the objective defined? How much authority did the system receive? What safeguards, monitoring, and intervention points were in place? As AI systems gain greater autonomy, the answers will increasingly shape disputes over foreseeability, causation, and responsibility.

When AI Exceeds Its Intended Authority

In July 2026, OpenAI disclosed that models undergoing a cybersecurity evaluation had escaped their testing environment, reached the internet, and accessed Hugging Face’s production infrastructure. Anthropic subsequently identified three incidents in which its models accessed real organizations during cybersecurity tests after a configuration failure left them with unintended internet access. In both cases, the models were pursuing the objectives they had been assigned, but their actions extended beyond the intended environments.

The litigation implications extend beyond cybersecurity. These incidents demonstrate that an AI system does not necessarily need to malfunction, disregard its instructions, or develop an independent objective to cause harm. It may operate consistently with a broadly defined goal while selecting a method that its developers did not anticipate or authorize.

That complicates how an alleged AI failure may be framed. A developer could argue that the model did exactly what it was instructed to do and that the harm resulted from a configuration error, weak infrastructure, or inadequate controls maintained by another party. A plaintiff could respond that the system’s ability to find unexpected paths toward its objective was itself foreseeable and required stronger limits, monitoring, and intervention mechanisms.

The incidents also illustrate how responsibility may become divided across the AI supply chain. OpenAI’s models reportedly exploited a previously unknown vulnerability to leave an isolated environment. Anthropic attributed its incidents largely to a misunderstanding with an evaluation partner that left an unintended route to the internet. Potential disputes could therefore involve model developers, evaluation vendors, infrastructure providers, software integrators, deployers, and operators, each responsible for a different part of the system.

The central issue may not be whether an AI model went “rogue,” but whether the organizations surrounding it took reasonable steps to constrain its authority. Courts may examine the scope of the system’s objective, the permissions it received, whether safeguards were active, how its conduct was monitored, and when human approval was required.

Foreseeability will likely be heavily contested. Defendants may argue that a model’s precise conduct could not reasonably have been predicted, particularly when it involved a novel vulnerability or unexpected configuration failure. Plaintiffs may argue that the exact sequence of events did not need to be anticipated if the broader risk of unauthorized goal pursuit was known.

That concern is not limited to outside critics. In July 2026, 1,350 employees of frontier AI companies signed the Pacing the Frontier statement, warning that automated AI research could accelerate the development of capabilities beyond the ability to understand or control the resulting systems. Although the statement does not establish a legal standard of care, it reflects growing industry recognition that AI capabilities may be advancing faster than available safeguards—an issue that could become relevant to future arguments concerning foreseeability, notice, and reasonable precautions.

Chatbot Cases Offer an Early Liability Model

Current litigation involving alleged chatbot-related harm shows how questions of goal-driven AI behavior may reach juries. Recent wrongful-death lawsuits allege that conversational AI systems reinforced delusional or suicidal thinking, encouraged emotional dependence, and failed to direct vulnerable users toward appropriate help.

In one recently filed case, the estate of an Alabama woman alleges that ChatGPT affirmed her hallucinatory beliefs and gave her permission to proceed shortly before her death. OpenAI has denied wrongdoing and has said it continues to strengthen its responses to sensitive situations with input from mental health professionals.

Conversational chatbots are not autonomous in the same operational sense as agents that can access networks or execute external actions. The cases nevertheless present a related liability question: What happens when a system continues pursuing engagement, affirmation, or conversational continuity after that behavior becomes unsafe?

For plaintiffs, the system’s conduct may support arguments involving foreseeable design risks and inadequate safeguards. Defendants may focus on the unpredictability of user interactions, existing warnings and safety features, intervening factors, and the difficulty of establishing causation.

In a national study of 1,010 jury-eligible U.S. residents, the DOAR Research Center found that 66% supported holding companies accountable for a chatbot’s actions in cases involving suicide. Fifty-four percent assigned legal responsibility to the company when a chatbot acknowledged suicidal statements but did not otherwise intervene. That figure increased to 76% when the chatbot provided information about fatal poisons.

As DOAR Director Ellen Brickman, Ph.D., explained in a related Q&A on chatbots, jurors, and liability, jurors are likely to focus on the user’s vulnerability and the efforts a company made to build and test safeguards. A disclaimer may carry limited weight if the evidence suggests that the system’s design encouraged reliance or emotional attachment.

A Broader Shift in AI Risk

The cybersecurity incidents and chatbot lawsuits involve different technologies and alleged harms, but both raise questions about what happens when an AI system continues pursuing an objective after its behavior becomes unsafe.

To understand whether these developments represent isolated failures or a broader change in the AI litigation landscape, we asked members of DOAR’s AI Expert Team to answer the following question. To preserve anonymity, this response combines and condenses perspectives shared by multiple team members.

Q: “Recent incidents, from AI agents accessing real-world systems during cybersecurity evaluations to allegations that chatbots continued harmful interactions with vulnerable users, raise concerns about systems pursuing an objective after their behavior becomes unsafe. Are these isolated failures, or the beginning of a new phase in which greater AI autonomy reshapes the risk landscape and the types of disputes that ultimately reach courts?”

A: Taken together, our expert perspectives suggest that the incidents are not isolated failures but early examples of a broader structural risk emerging as AI capabilities reach the market faster than the safeguards designed to constrain them. In these cases, the systems did not need to malfunction or act maliciously. They pursued assigned objectives without adequately accounting for the harm their methods could cause. General-purpose guardrails may therefore be insufficient unless legal and operational limits are also built into the system’s specific tasks, permissions, and environment.

Even defining “autonomy” may require expert analysis. In an automated-vehicle accident, for example, responsibility may depend on the interaction between the AI, the vehicle, and the human driver. Fully driverless vehicles further complicate familiar liability questions. Similar issues arise when an AI agent authorized to use a credit card on one website uses it elsewhere. Determining responsibility could require examining the user’s intent and instructions, the agent’s design, the information available to it, and the oversight surrounding its actions.

These disputes will likely involve responsibility across the model provider, the company building the agent, the enterprise deploying it, and the user granting authority. Questions of reasonable care may turn on permission scopes, system instructions, monitoring procedures, activity logs, and evidence of who established the system’s limits and was positioned to intervene. Looking further ahead, courts may also need to address the admissibility of AI tools used by expert witnesses and, eventually, whether AI itself could serve an expert role.

Discovery Will Follow the AI Decision Chain

AI litigation involving autonomous agents will likely require discovery into the full chain of decisions surrounding the system, not simply its final output.

Relevant evidence may include system prompts, access permissions, testing results, risk assessments, safety controls, incident logs, escalation policies, internal communications, and decisions about when human authorization was required. These materials may help determine whether the alleged harm resulted from the underlying model, an infrastructure weakness, an implementation error, inadequate monitoring, or a business decision to accept a known risk.

They may also show whether the system acted contrary to its intended design or functioned as configured while producing an unsafe result. That distinction could be particularly important in product liability and negligence claims, where the parties may disagree about whether the system itself was defective or the surrounding controls were inadequate.

As DOAR’s AI Expert Team has previously explained, many incidents described as AI failures are rooted elsewhere in the technology stack, including security boundaries, access controls, data systems, and deployment decisions. Identifying the source of the failure will be essential to evaluating causation and allocating responsibility.

Expert testimony will also be critical. Experts may be asked to reconstruct how an agent reached a decision, explain the significance of its instructions and permissions, evaluate whether safeguards were appropriate, and assess whether additional controls would likely have prevented the harm.

They may also need to explain the difference between autonomous behavior and independent intent. A system that appears to have acted “on its own” can make the conduct feel remote from the company that deployed it. The competing argument is that every action remained enabled by human decisions about the system’s objective, authority, safeguards, and oversight.

As AI systems gain the ability to act rather than merely respond, disputes will increasingly focus on how much authority they were given, how their conduct was constrained, and who oversaw them. Greater AI autonomy will not eliminate human or corporate responsibility. It will make the path to establishing that responsibility more technical, more contested, and increasingly dependent on evidence explaining how the system was designed, deployed, and controlled.

 

Stay Informed
Stay up to date on our latest news and insights.
Subscribe
Read More
looking at the capitol building through a window in DC
Article
Aug 6, 2026
2026 ITC Midyear Update: Section 337 Trends and a Reshaped Commission

Halfway through 2026, the U.S. International Trade Commission (ITC) is balancing an active Section 337 docket with meaningful changes in its leadership. Investigations continue to move on demanding schedules as the Commission welcomes new members and prepares for a change on its Administrative Law Judge (ALJ) bench.

Read Now
orange background painted with dark blue and black
Article
Jul 31, 2026
The ITC's Next ALJ: What the Search Says About the Future of Section 337 Investigations

The U.S. International Trade Commission is entering an important period of transition. Administrative Law Judge (ALJ) MaryJoan McNamara's planned departure has prompted the Commission to begin searching for her successor, while the Commission itself is also welcoming new leadership as the Senate advances a full slate of commissioner nominees.

Read Now
close up image of pages from a book
Article
Jul 22, 2026
Beyond the Prompt: AI, Expert Testimony, and a New Layer of Judicial Scrutiny

Artificial intelligence is no longer a theoretical issue in litigation. Law firms are developing AI policies, courts have issued guidance on attorney use of generative AI, and litigants are increasingly encountering AI-generated work during discovery. More recently, attention has begun to shift to another question: what happens when expert witnesses incorporate AI into their work?

Read Now
Article
Jul 17, 2026
Is Competition for Specialized Talent Driving More Trade Secret Disputes?

Companies are investing enormous resources into developing proprietary algorithms, manufacturing processes, engineering designs, training data, source code, and other confidential information that often cannot be protected through patents alone.

Read Now
blue and black texture wavey
Article
Jun 30, 2026
Investigations at the ITC: Navigating Speed and Complexity in a High-Stakes Forum

Success at the ITC depends not just on the legal merits, but on how effectively parties can organize technical complexity, align expert-driven narratives, and present a clear, disciplined case under significant time pressure. Let’s explore how venue nuances and current industry trends are impacting proceedings seen before the Commission.

Read Now
overhead view of cars speeding on road like blurs of light
Article
Jun 18, 2026
Who Gets the Blame? What Autonomous Vehicle Research Reveals About AI Litigation

Litigation surrounding autonomous vehicle technology raises a question that courts will increasingly confront as AI becomes more embedded in everyday life: when humans and machines share control, who gets the blame when something goes wrong?

Read Now
blue and black background with grain and texture
Article
Jun 9, 2026
What ITC Practitioners Should Know: Key Judicial Insights from ACI’s ITC Conference

To better understand evolving trends at the venue and hear directly from those involved in ITC litigation, DOAR proudly sponsored and attended the American Conference Institute’s annual ITC Litigation and Enforcement Conference.

Read Now
white and black triangle texture on washed out background
Article
Apr 20, 2026
AI Litigation Trends: Rapid Growth and Emerging Patterns

As generative AI technologies move from early adoption to widespread commercial use, litigation activity is accelerating in parallel. Our analysis of 168 district court cases highlights a sharp rise in filings, a high concentration among a small group of defendants, and early signals of how this legal landscape may evolve.

Read Now