
ARTIFICIAL intelligence safety cannot be reduced to cybersecurity, monitoring or rules governing the release of increasingly powerful models, according to Yoshua Bengio, one of the pioneers of modern deep learning, who warned that the more difficult problem is what happens when an AI system becomes capable of pursuing goals that humans did not intend.
The issue was at the center of “Governing at the Speed of AI,” an invitation-only meeting organized by AI Safety Connect at UN Headquarters during the 81st UN General Assembly. About 300 heads of state, ministers, diplomats, technology executives and researchers were invited to the gathering, which focused on what governments and companies should do as AI systems take on more complicated tasks and greater independence.
Bengio said serious AI risks arise when three conditions come together: a system has a misaligned goal, has enough capability to pursue it and operates in an environment that allows it to act. The Manila Times attended the exclusive event virtually and posed questions to the representatives.
“If capabilities continue and we don’t solve alignment, it is almost certain that there will be AIs smart enough to go through our defenses,” Bengio said.
Cybersecurity and human monitoring remain necessary, he said, but they cannot be the sole safeguards as systems become more capable. An AI does not necessarily have to defeat a technical security system to cause harm. It could exploit weaknesses involving people, including by persuading someone to take an action that advances its objective.
delivered to your inbox
That concern is becoming more relevant as companies develop AI agents capable of performing tasks with less direct human supervision.
“What is really dangerous is uncontrolled autonomy,” Bengio said. “When the AIs are autonomous and going after goals that we didn’t choose and often actually are against instructions, as we’ve seen last summer.”
The problem, Bengio said, is therefore not simply whether an AI system is intelligent or useful. Researchers have to establish how much autonomy a system should have and how its objectives can remain compatible with human intentions as its capabilities increase.
He said alignment is solvable and pointed to work through LawZero, the nonprofit he founded to investigate alternative approaches to AI safety. One area of research involves systems that do not operate as autonomous agents with independent objectives. Bengio said AI could instead be designed without an “ego,” while remaining honest about its uncertainty and limitations.
The discussion around Bengio’s warnings took place as governments and technology companies began trying to turn broad AI safety principles into procedures that can be tested and enforced.
UN Under-Secretary-General and UNDP Associate Administrator Haoliang Xu brought the discussion down to recent incidents rather than hypothetical future threats. He cited reports of AI agents gaining unauthorized access to computer infrastructure and said such incidents should be treated as breaches for which someone must be held responsible.
“If a human being deliberately entered a government system without authorization, bypassed controls and accessed restricted information, we would not dismiss this as unexpected behavior,” Xu said. “We would investigate it, and there’d be legal consequences.”
Call to action
A call to action initially signed by 22 countries and led by Norway and Finland called for mandatory testing, shared incident reporting and work toward an international mechanism capable of verification. The number of countries supporting the initiative was described during the gathering as continuing to increase.
Outi Holopainen, Ambassador Undersecretary of State for Political Affairs, Finland, said governments have a primary responsibility to protect their citizens from frontier AI risks.
“If you want to sell your products, you have to make it safe on your door with us,” David Lametti, Ambassador of Canada and Permanent Representative to the UN, told the forum. He also argued that countries outside the largest AI powers could use their combined markets to influence developers.
Kenya’s representative brought the discussion to countries that will consume AI systems rather than develop the most advanced models. African countries, he argued, cannot simply receive technologies developed elsewhere and then be expected to deal with the consequences.
“Kenya sees AI as an opportunity because of its young population and expanding digital economy, but its officials also see the risks associated with systems already being deployed. Countries where AI is used, rather than only countries where it is developed, should participate in setting safety requirements,” Philip Thigo, Special Envoy on Technology, Office of the President, Kenya, said.
A California legislator involved in AI regulation said conventional legislation can move too slowly for a technology that changes rapidly.
“California’s approach has focused on technical standards and independent evaluation,” Jerry McNerney said, while describing legislation that would establish working groups to conduct independent evaluations. The groups would have to be protected from government interference and from capture by technology companies. The longer-term objective would be international groups involving technical experts from different countries.
Independent evaluation
Industry representatives discussed giving outside experts access to advanced AI systems and enough technical information to determine whether developers’ safety claims can be supported by evidence. Microsoft’s Natasha Crampton, vice president and chief responsible AI officer, discussed independent evaluation in the context of financial auditing.
Michael Sellitto, head of APAC Policy at Anthropic, addressed the need for companies developing and deploying AI systems to share information about vulnerabilities and unusual model behavior before problems become serious incidents. Information may need to move quickly between developers, organizations deploying the systems and other companies in the AI supply chain.
That raises the question of how companies can be encouraged to report problems before all the facts are known. Safe-harbor provisions were discussed as one way of reducing the legal risk associated with early disclosure. Companies could otherwise delay reporting because incomplete information might expose them to liability, even when the information could help other organizations determine whether they face the same vulnerability.
The European Union was cited as an example of how access to a major market can extend safety requirements to companies based elsewhere. Insurance was raised as another pos sible incentive, with safety standards potentially affecting the terms available to companies, as has happened in aviation, automotive manufacturing and financial services.
Open-weight models
The participants also addressed open-weight models, which cannot be recalled once the weights have been released and can be copied or modified by others. This model allows individuals to create
Those models have potential advantages, including lower costs, greater accessibility and greater opportunity for inspection. They can also be useful to cybersecurity defenders. During one incident discussed at the gathering, defenders turned to an open-weight model after a proprietary system proved unsuitable for the task.
The financial sector offered another perspective on managing AI risks.
Tara Lyons, managing director and global head of AI at JPMorgan Chase & Co., said policymakers should focus on outcomes rather than attempting to regulate the technology in isolation. Financial institutions have already been using AI in sensitive environments where risk management and public trust are essential.
“Trust is a practical requirement for deployment. Companies and consumers will not use AI extensively if they do not believe the systems can be relied upon,” Lyons said.
International problem
AI researchers in China, Europe and the United States share concerns about the risks associated with increasingly capable systems. What is missing is effective cooperation that can continue despite geopolitical tensions.
An international scientific channel insulated from political disputes was proposed as one way of maintaining communication among researchers.
The participants did not expect a comprehensive agreement between Washington and Beijing to emerge immediately. Several argued that countries can establish their own requirements while working toward standards that are compatible across jurisdictions.
Xiao Qian, vice dean of the Institute of AI at Tsinghua University, said the point was not to create a completely new international bureaucracy. Existing organizations already have parts of the machinery needed for AI safety, including incident-reporting frameworks, standards bodies, national AI safety institutes and scientific assessments.
Former Colombian president Juan Manuel Santos connected the urgency over AI with the Doomsday Clock and nuclear risk. He argued that the danger is increasing rapidly and that governments already know enough to begin taking action.
Joost Flamand, Vice Minister of Foreign Affairs of the Netherlands, framed the issue as one of democratic control. Democracies must retain agency over the direction of AI development, he said, while safety and controllability should be treated as conditions for sustainable innovation rather than obstacles to it.
Arthur Franzon, French deputy administrator for AI and Digital Affairs, identified four areas for international work: independent evaluation, transparency and incident reporting, verification and auditing, and legal accountability during AI development and deployment.
Unesco emphasized the need for governments themselves to have the expertise and institutional capacity to govern AI. Safety cannot depend entirely on the countries and companies developing the technology.
For Bengio, however, the central problem remains more basic. As AI systems become more capable and autonomous, governments and developers have to determine not only whether those systems can perform a task, but whether they can be trusted to pursue the task without developing objectives that conflict with the people who built and deployed them.



