
Google's Gemini artificial intelligence model has reportedly hacked into three companies during a security test, demonstrating how advanced AI systems could potentially perform cyber operations with increasing levels of autonomy.
According to the BBC report, Gemini accessed the internet during the test, searched for publicly available information and guessed credentials to gain access to three websites. A Google official told the BBC that the model stopped after gaining access, and the affected companies were informed.
The incident was part of security testing rather than an uncontrolled attack on companies. However, the ability of an AI model to independently discover information, attempt credentials and gain access has raised broader questions about how AI could change cybersecurity.
AI has traditionally been used in cybersecurity for tasks such as detecting suspicious activity, analysing large amounts of data and identifying potential vulnerabilities.
The Gemini test illustrates a different capability: an AI system can potentially chain multiple steps together to pursue a cybersecurity objective.
In this case, the model reportedly found information online, identified websites it believed were relevant to the test and attempted credentials before gaining access.
Such capabilities could eventually make security testing faster, but they could also create additional risks if similar techniques are used against real-world targets without authorization.
Cybersecurity researchers have increasingly been examining how AI models behave when given access to the internet, software tools and computer systems.
A major concern is that increasingly autonomous AI agents could reduce the amount of human involvement needed to conduct certain cyber operations. This creates a difficult balance for technology companies: models need enough capability to help security teams identify weaknesses, while safeguards must prevent those same capabilities from being misused.
The latest Gemini test comes amid wider debate over the pace of AI development and the potential risks associated with increasingly powerful systems.
For businesses, the development could mean that traditional cybersecurity practices need to account for AI-powered attacks as well as human attackers.
Security teams may increasingly need to test whether publicly exposed information, weak passwords and poorly protected websites can be discovered and exploited automatically by AI systems.
At the same time, companies developing AI agents are likely to face greater pressure to introduce safeguards, monitoring and authorization controls before giving models access to sensitive systems.
The Gemini incident therefore represents more than a single security experiment. It demonstrates how AI is moving toward systems that can search, reason, act and adapt across multiple steps, making cybersecurity one of the areas where the capabilities of autonomous AI could have significant consequences.