Artificial intelligence is becoming a routine partner in tasks that once demanded specialized expertise, from analyzing technical information to generating software and supporting high-stakes decisions. Yet the systems behind these capabilities can still fail in surprisingly simple ways. A carefully worded prompt may persuade an AI model to ignore its original instructions, while an unfamiliar input can trigger inaccurate claims, flawed calculations or reasoning that appears convincing but is fundamentally unsound. Researchers at the University of Central Florida are now pursuing a strategy that could help AI recognize and resist these weaknesses before they cause harm.
Amrit Singh Bedi, an assistant professor in UCF’s Department of Computer Science, has received a one-year, $340,000 award from the Defense Advanced Research Projects Agency, or DARPA, to investigate how artificial intelligence systems can become more resilient. Bedi and his research team will study methods for training AI to withstand failures, adversarial attacks and unexpected conditions. Their central objective is to develop systems capable of identifying weaknesses in their own reasoning and adjusting their behavior over time, rather than depending entirely on fixed safety rules or human intervention.
The project addresses a growing technical problem in modern generative AI. Large language models and related systems are trained on enormous collections of data and optimized to produce useful responses, but their apparent fluency does not guarantee reliability. These models generate outputs by identifying statistical patterns and selecting likely sequences of information. As a result, they may confidently produce false statements, follow misleading instructions or make decisions based on subtle cues that have nothing to do with the intended task. Their performance can deteriorate even further when they encounter inputs designed specifically to exploit their limitations.
Such attacks are often known as adversarial prompting or prompt injection. An attacker may introduce instructions that conflict with a system’s original goals, conceal malicious commands within ordinary text or manipulate the order and wording of information presented to the model. In an AI agent connected to software tools, databases or physical equipment, a successful attack could have consequences beyond a wrong answer. It might cause the system to reveal sensitive information, execute an inappropriate action or abandon the objectives established by its developers. Bedi’s research is focused on helping AI detect these pressures and respond without losing control of its mission.
The SafeRR AI Lab, led by Bedi and his collaborators, will explore what the researchers describe as “self-hardening.” The concept is comparable to strengthening a computer system after identifying a vulnerability, but it is intended to operate within the AI’s reasoning process. A self-hardening system would monitor its own responses for signs of manipulation, inconsistency or unusual uncertainty. When it detects a suspicious instruction or a conflict between competing goals, it could slow down, reassess the evidence, reject the unsafe direction or choose a more conservative response.
Developing such capabilities requires more than simply adding another list of prohibited words or behaviors. Static safeguards can be effective against known problems, but they may fail when an attack is expressed in an unfamiliar form. The UCF team is instead examining adaptive approaches in which an AI system learns from attempts to deceive or derail it. This could involve testing a model against hostile inputs, analyzing where its reasoning breaks down and updating its decision-making strategies to handle similar situations in the future. The challenge is to improve resilience without making the system so cautious that it becomes unable to perform legitimate tasks.
At the center of the project is the distinction between apparent capability and true robustness. An AI may achieve impressive results under normal testing conditions while remaining fragile when instructions are ambiguous, data are incomplete or an adversary is actively trying to manipulate its behavior. For systems used in mission-critical environments, that gap can be dangerous. A small deviation in an autonomous platform, a cybersecurity tool or a defense-related decision-support system could produce consequences far greater than the original error. Bedi’s team aims to develop evaluation methods that reveal how systems behave under pressure, not only how well they perform in ideal conditions.
The research could have implications well beyond military applications. AI tools are increasingly being considered for healthcare, emergency response, robotics, transportation and cybersecurity, where they must operate in environments that are unpredictable and sometimes hostile. A medical system may encounter incomplete patient information, an emergency-response platform may face rapidly changing conditions and a robotic system may receive conflicting signals from its surroundings. In each case, the ability to recognize uncertainty, preserve core objectives and recover from a faulty decision could be as important as the ability to generate a correct answer in the first place.
Bedi describes the broader goal as building AI that can actively recognize, withstand and recover from adversarial pressure while remaining faithful to its intended purpose. If successful, the project could contribute to a new generation of systems that do not merely follow safety instructions but participate in maintaining their own reliability. The researchers are not suggesting that AI should operate without oversight. Instead, self-hardening is intended to provide an additional layer of protection, allowing systems to identify risks earlier and respond more responsibly before human users must intervene. As artificial intelligence moves into increasingly consequential roles, that capacity for self-monitoring and adaptation may become essential to earning public and professional trust.
Subject of Research: Artificial intelligence resilience, adversarial robustness, self-hardening AI, machine learning safety
Article Title: UCF Researchers Explore Self-Hardening Artificial Intelligence to Resist Attacks and Failures
References: Defense Advanced Research Projects Agency grant; University of Central Florida SafeRR AI Lab
Keywords
Artificial intelligence, generative AI, machine learning, adversarial attacks, prompt injection, AI safety, resilient AI, adaptive systems, robotics, cybersecurity, defense technology

