Normal view

Can the finance sector oversee AI innovation while maintaining its rapid progress?

3 September 2026 at 07:32

Disclaimer: The opinions expressed and arguments employed herein are solely those of the authors and do not necessarily reflect the official views of the OECD, the GPAI or their member countries.

Artificial intelligence (AI) is no longer an emerging feature of financial systems. From credit scoring and fraud detection to financial advice and customer support, AI is reshaping how finance operates. Yet, as AI innovation accelerates, so too does a critical policy question: How can regulators keep pace with AI to ensure positive outcomes for consumers and markets? And how can they identify and assess emerging risks to market integrity and stability?

By focusing on supervisory practice – how existing financial services rules are interpreted, implemented and enforced – policymakers can create an environment where innovation in financial markets thrives without compromising trust, stability and the pace of innovation.

Strong regulatory foundations, complex supervisory realities

Under the broadly technology-neutral principle that guides financial regulation in OECD economies, existing requirements apply regardless of the technology used to deliver a financial service or product. Financial supervision serves as the practical enforcement mechanism for financial regulation, ensuring that policies translate into effective oversight and resilient financial markets.

Practical implementation of AI policies in finance may face challenges due to the intrinsic characteristics of AI innovation, especially advanced AI such as agentic AI systems, the growing volume and speed of transactions and the opacity and complexity of some advanced models that challenge human oversight, as well as their potentially autonomous nature. Limited data on AI adoption by financial services firms makes it difficult to evaluate its use and may hinder monitoring of associated vulnerabilities and their impact on markets and consumers more generally.

The OECD report Supervision of Artificial Intelligence in Finance offers a timely analysis of supervisory approaches and practices that support policy objectives while promoting safe and responsible AI adoption in finance, ensuring that innovation and oversight evolve together.

Supervisory challenges and barriers to wider AI adoption in finance

In some jurisdictions, evaluating bodies and firms seeking compliance face similar difficulties, which could directly impede the wider AI uptake in financial services. Many of the challenges are predictable and include model risk management and validation, testing the outcomes of complex AI systems and limited AI explainability. Challenges also include the practical assessment of the robustness and fairness of model outputs, and data governance and management – many elements highlighted in the OECD AI Principles.

For example, many find it challenging to articulate how the “human in the loop” concept should work in practice. Others struggle to determine appropriate fairness thresholds (including what industry benchmarks should look like) and assess whether they are adequate.

Effective oversight of AI in finance is more about better supervision than new regulation

While existing requirements still apply and supervised entities are expected to adapt their risk-management frameworks to AI-specific matters, some jurisdictions could benefit from guidance and clarification on how these existing regulatory requirements apply to advanced AI models.

In some instances, guidance might clarify ambiguities in interpretation and implementation, for example, in model risk management frameworks, especially regarding issues such as model validation, testing and monitoring, as well as regulatory requirements, such as governance.

Rather than imposing rigid, overly prescriptive requirements that could hinder AI adoption, providing interpretative guidance and practical clarifications on applying existing risk-management and governance frameworks in AI contexts may be preferable, where necessary.

From static compliance to dynamic understanding

Closer engagement between supervisors and industry stakeholders, beyond standard supervisory activities, could foster mutual understanding. Such engagement can yield significant benefits for supervised entities by helping to clarify ambiguity, while also improving authorities’ understanding of the challenges they face in their supervisory efforts.

One way financial supervisors can engage with the industry is through AI-specific testing, which can provide confidence and clarity, encouraging innovation while protecting markets, their participants and stability.

Initiatives such as the UK’s Financial Conduct Authority (FCA) AI Live Testing programme could help promote constructive dialogue between firms developing or using AI and regulatory agencies, encouraging mutual understanding and learning.

Novel approaches to AI supervision: the FCA’s AI Live Testing

The FCA’s AI Live Testing programme allows financial services firms to work directly with its regulatory and technical teams when deploying AI systems in controlled live market environments.  The programme focuses on how an AI model interacts with its environment. It is not a simulated lab environment, but rather aims to understand the behaviour of deployed AI systems, taking relevant rules and regulations into account. The objective is to support safe, responsible AI adoption without creating new AI‑specific rules, while giving firms regulatory confidence and helping the FCA understand how AI affects UK markets and consumers.

How AI Live Testing works: From discovery to live testing

The AI Live Testing programme comprises two phases: one for discovery and one for live testing. The programme is designed to be proportionate, with a strong emphasis on in‑person engagement, structured workshops and iterative feedback.

The discovery phase is intended to build a clear understanding of the AI system used by a given firm, including its testing and monitoring approach, governance arrangements, controls and key areas of potential risk. Through a series of continued engagement (including in-person workshops) and voluntary evidence‑gathering, the AI Live Testing team assesses:

  • AI system architecture and design
  • Model development and performance
  • Data integrity and pipeline management
  • AI robustness and reliability
  • Explainability and transparency
  • Monitoring, logging and version control
  • Security and technical resilience

The live testing phase includes a practical assessment of the firm’s testing and monitoring approaches. Building on the discovery phase, it uses structured workshops and feedback sessions to explore:

  • How to test the AI use case most effectively
  • What evidence is required to demonstrate that an AI system is safe and responsible, including delivering the right outcomes for consumers and markets
  • How testing can demonstrate alignment with FCA regulations and requirements

It should be noted that the FCA does not test the AI system on the firm’s behalf. Firms are responsible for testing, pilots and evidence generation, including demonstrating that a given AI system is safe and responsible. AI Live Testing’s role is to analyse the approach taken by firms, review evidence and identify areas where risk management, governance or the interpretation of policies and test results may be unclear or require adjustment to secure the best outcomes for consumers and markets. This ensures that responsibility for the technology, including risk identification and mitigation, remains with the firm.

Observed benefits of FCA’s AI Live Testing

The aim of AI Live Testing is to help firms design, deploy and monitor safe and responsible AI, not to provide regulatory approval, audit or sign-off. To achieve this, it uses the following approaches:

  • AI flaws are always understood in the context of the firm’s AI use case that is within the scope of AI Live Testing. This avoids an esoteric or overly technical view of harms, keeping the focus on real-world consumer impact.
  • The programme is designed to encourage two-way learning between the FCA and participating firms through feedback loops at all key stages of discovery and testing. It is intended to be exploratory rather than supervisory, focused on insights rather than regulatory judgements, approvals, AI audits or endorsements.
  • No “hands-on” testing:  The FCA does not test the system for the firm.  It analyses firms’ testing, monitoring and piloting approaches and provides feedback on potential gaps in the firm-level testing regime.

AI Live Testing helps build an evidence-based understanding of what works in safe and responsible AI deployment, especially where ambiguities exist, and what may need to change in testing, governance, risk management and continuous improvement

Issues brought to the surface by advances in AI innovation

To date, AI Live Testing has identified two critical challenges:

  1. The first issue is ensuring that AI governance keeps pace with increased complexity as a result of the pace of technological innovation combined with the pace of adoption and deployment of these more advanced AI systems. In practice, this means that effective policies and frameworks are key in designing and deploying safe and responsible AI. This also requires confidence in testing, risk mitigation strategies and the ability of individual firms to monitor.
  2. The second issue concerns agentic AI systems. They have multiple possible execution paths, but all relevant eventualities cannot be tested prior to deployment. This presents a fundamental challenge in pre-deployment testing because, by definition, this process is insufficient. Often, the response is to fall back on human control. However, this is a limited means of effectively verifying all the potential actions of an AI model.

These developments raise further questions: What do they mean for the overall effectiveness of AI system governance, including the role of human oversight? Where can regulatory requirements on governance, accountability, and systems and controls make a positive difference? And how can regulatory requirements adapt over time to stay relevant and deliver the right outcomes for consumers and markets?

Flexible, agile and adaptive AI supervision

Maintaining a flexible, agile, and adaptive stance towards supervision is one way financial supervisors can keep pace with technological advances, such as agentic AI. Novel methodologies and techniques could enrich the supervisory toolkit and help to ensure beneficial innovation, stability and trust. They also provide an opportunity for supervisors to ask the right questions of firms in a timely manner, based on insights from these toolkits. Public-private dialogue and proactive engagement between supervisors and regulated entities should support the adoption of new supervisory methods and tools. This collaboration will serve as the foundation for deploying safe and responsible AI in financial markets. At the same time, it enables institutional capacity building, which is key to more effective monitoring of AI activity.

By remaining open to enriching and adapting frameworks to reflect the realities and specific characteristics of AI innovation, financial supervisors can foster responsible AI innovation in finance while mitigating associated risks. This shift will not happen overnight and will require investment, experimentation, co-operation with industry and, above all, a willingness to rethink and enrich long-standing institutional practices.

The post Can the finance sector oversee AI innovation while maintaining its rapid progress? appeared first on OECD.AI.

A five-step roadmap to closing the AI evaluation gap

31 July 2026 at 10:19
two men facing each other

Policy debates continue over how best to regulate artificial intelligence (AI) and harness its economic and other benefits while safeguarding society from its harms and risks.  For example, the European Union and China have opted to govern AI through different regulatory approaches, while the United States and some other countries have adopted different policy tools that they believe will better promote AI innovation. While countries and regions take different approaches, consensus is emerging across jurisdictions and among leading experts that there is an AI evaluation gap that must be promptly closed. 

Today’s AI evaluations fall short

The 2026 International AI Safety Report, prepared by more than one hundred experts and supported by more than 30 countries and multilateral organisations, explains that today’s AI evaluation techniques often fail to anticipate real-world performance.  This can occur when AI models produce overinflated test results or when AI testing environments materially differ from the real world.  Compounding the challenge, the “AI evidence dilemma” arises from the difficulties of assessing the risks of this rapidly evolving technology.

Boosting AI trust, diffusion, security and investment returns

Closing the AI evaluation gap will bring many benefits.  Sound AI evaluations would help increase understanding of AI’s performance and reliability and better inform decisions about its use.  This is critical since AI’s performance remains “jagged,” with some AI applications performing better than others.  

Helping people better understand AI’s reliability across contexts would enhance their trust in its appropriate use.  Similarly, reliable evaluations can also help buttress the security of AI systems.  All this, in turn, would help to support greater adoption and diffusion of secure and trusted AI applications and help organisations more fully reap the benefits of their AI investments.  

Reliable evaluation helps to reduce uncertainty for policymakers 

Better evaluation provides a greater evidence-based to inform AI policy choices. As explained in the 2026 International AI Safety Report, this could help reduce some of the uncertainty policymakers currently face.  Equipping policymakers with this information could lead to swifter, more confident policy decisions and help regulatory frameworks keep pace with AI’s rapid technological progress.  

Better evaluation could decrease operational costs and expand competition and access

Closing the AI evaluation gap could have the added benefit of reducing AI operating costs, potentially making trusted AI cheaper and more accessible.  Even with mature AI evaluations, there could be variation across jurisdictions on whether they are applied in a voluntary or mandatory manner.  However, if the evaluations underpinning different regulatory approaches are standardised, with little variation across regions, this could help companies and organisations reduce costs and increase ease of operating across borders.  In other words, reliable and standardised evaluation could help to foster both regulatory interoperability and innovation.  This, in turn, could lower barriers for new AI market entrants and support a healthy competition ecosystem.

Governments want better AI evaluations

Already, several governments are investing in closing the AI evaluation gap.  The White House AI Action Plan calls for building an evaluation ecosystem. In June 2026, US President Trump signed a new AI Executive Order establishing a voluntary framework for the government to test covered frontier models to improve secure innovation and cybersecurity.

These actions build on other important government-led AI evaluation efforts.  Following the establishment in 2023 of AI Safety Institutes by the UK and the US[SR1] [SR2]  (both renamed in 2025), a total of 11 jurisdictions, including the EU, India and Singapore, have now followed a similar model and established AI safety institutes or similar organisations that conduct testing and evaluation of foundation models.  Avenues have been paved for international co-ordination among these organisations.

Prior to the new US AI Executive Order, the US Center for AI Standards and Innovation (CAISI), mentioned above, announced voluntary agreements with xAI, Google, and Microsoft for pre-deployment frontier model testing.  These add to CAISI’s voluntary testing arrangements with Anthropic and OpenAI, as well as its collaborations with these companies to boost AI security and related measurement techniques.  The US National Institute of Standards and Technology (NIST) has evaluation programmes for generative AI.  Similarly, Korea and Singapore recently concluded joint tests of AI agents to evaluate data leakage.    The UK AI Security Institute also continues its cutting-edge frontier AI model evaluations.

The private sector is also expanding evaluation initiatives

After discovering that Mythos could detect severe vulnerabilities in all major web browsers and operating systems, Anthropic launched Project Glasswing to make the unreleased frontier model available to several organisations to help secure their systems.  Anthropic went a step further and committed to sharing its learnings with the broader community.  Additionally, several major AI developers launched the Frontier Model Forum (FMF), a collective effort to advance AI safety and security, and have released several publications, including a recent report, Managing Advanced Cyber Risks in Frontier AI Models.

A 5-step roadmap for closing the AI evaluation gap

To successfully close the AI evaluation gap, AI actors can take several steps, building on today’s existing efforts.  

 Step 1: Balance standardisation and customisation

First, to account for different languages, cultures, use cases, and norms, the evaluations should strive to balance standardisation and customisation.  At the 2026 AI Impact Summit in India, several companies pledged to improve multilingual and contextual evaluations to help achieve this balance.  

Step 2: Test throughout the AI lifecycle

Second, to address AI performance differences between controlled environments and the real world, evaluations should be conducted throughout the AI system lifecycle.  This approach is already embraced in several key publications and leading frameworks, including the International AI Safety Report and NIST’s AI Risk Management Framework.  The next step is to develop and implement evaluations for these different contexts.  

Step 3:  Build the right ecosystem 

Third, evaluations must be supported by a robust ecosystem, including qualified examiners.  This should be accompanied by methodologies for effectively communicating AI evaluation results to diverse audiences, including business users and affected individuals. At the same time, evaluations must preserve proprietary information about the AI systems.  Existing assurance practices used in the financial services and other sectors could further inform this work.  

Step 4:  Consider the AI value chain, technology and context.

Fourth, AI evaluations should be tailored to the needs of different actors in the AI value chain and different types of AI deployments.  For example, testing conducted upstream by large language model (LLM) developers may vary from the evaluations performed by companies deploying LLMs downstream.  In other words, the role an organisation plays in the AI ecosystem, including whether it enhances models downstream, should help determine the types of evaluations it implements.  

The rise of AI agents and Agentic AI capable of acting autonomously presents new evaluation challenges that governments and other stakeholders are working to address.  This reinforces the need to continuously assess the suitability of evaluation methodologies for different AI technologies and deployment settings.  

Step 5:  Create an efficient and trusted process

Finally, the process for developing AI evaluations also merits careful consideration.  To help ensure that evaluations address the appropriate factors, the process should capture global inputs from diverse stakeholders, including industry, government, academia and civil society.  To help keep pace with AI’s rapid development, the process should also leverage and co-ordinate the good work already being done. This includes the efforts of safety institutes and standards organisations, such as the “Zero Draft” project launched by NIST to expedite the standards process.  It should also consider the outputs of the OECD Hiroshima AI Process (HAIP) Reporting Framework.  A key purpose of the evaluation process is to  instill trust in everyone affected by a given AI system. 

The time is now

In sum, while much work remains to close the AI evaluation gap, as discussed above, there is already an emerging consensus, a solid foundation, and momentum to do so.  Closing the evaluation gap holds great promise of increasing AI trust, adoption, security, diffusion and investment returns.  Furthermore, it can help reduce policy uncertainty, increase regulatory interoperability, reduce costs for AI companies and organisations, and expand the availability of AI services and competition.  Simply put, the prize is worth the effort. 

The post A five-step roadmap to closing the AI evaluation gap appeared first on OECD.AI.

Why AI Sandboxes matter for responsible innovation and public trust

18 March 2026 at 21:36

Among the various tools available to policymakers, regulatory sandboxes have gained considerable prominence in the AI governance landscape because they enable supervised innovation testing under controlled conditions and within limited timeframes. This can help to identify risks early, foster regulatory learning and refine regulatory requirements before they are applied at scale.

As AI regulatory sandboxes expand across jurisdictions and sectors, common design principles, recurring challenges and opportunities for greater effectiveness and policy coherence are emerging. As this happens, institutional co-operation and knowledge sharing are more important for ensuring coherent and effective regulatory experimentation both nationally and across borders.

In November 2025, the OECD webinar “AI Sandboxes: Sharing knowledge for success” brought together government officials, regulators and policy experts from seven countries to discuss the design and implementation of AI regulatory sandboxes. Here are the event’s key takeaways.

What is an ‘AI regulatory sandbox’?

Although there is no universally accepted definition, a regulatory sandbox generally offers temporary regulatory flexibility or waivers, allowing innovative products, services, or business models to be tested under controlled conditions and regulatory oversight. This approach promotes responsible experimentation and innovation while protecting the public interest.

To cite a few examples, Singapore’s AI healthcare sandbox offers guidelines for synthetic data to minimise privacy risks while allowing realistic testing. In the UK and other countries, AI-powered innovations in financial services are being tested under supervision that helps to prevent consumer harms such as biased scoring and automated decision-making. 

In AI, this approach aligns with the OECD AI Principles – specifically Principle 2.3, which encourages governments to promote experimentation to enable the safe testing and scaling of AI systems. Similarly, the Recommendation of the Council for Agile Regulatory Governance to Harness Innovation urges governments to facilitate greater experimentation, testing, and trialling to stimulate innovation under regulatory supervision.

AI regulatory sandboxes are valuable for testing new technologies and rules in a safe, controlled way before full rollout. They offer less benefit if risks are low or if rules are already well established, but can be useful for compliance and learning in more complex regulatory environments. Decisions to utilise sandboxes should follow clear criteria to ensure efforts are appropriately targeted. Generally, initiatives with high innovation potential, substantial risks, and opportunities for regulatory discovery and improvement (including by removing barriers to beneficial innovation) should be prioritised. Key regulators, industry actors and other relevant stakeholders should be involved in this process.

In July 2023, the OECD published the policy paper Regulatory Sandboxes in artificial intelligence. Building on lessons from fintech, the report highlights the benefits of AI sandboxes, including accelerating market entry, improving regulatory understanding, and stimulating investment. It also explains why adapting the traditional sandbox model to AI presents unique technical and governance challenges. As a cross-sectoral technology, AI covers multiple legal, ethical, and technical domains, requiring strong coordination among several regulatory authorities.

Since the paper’s release, the use of AI sandboxes has accelerated. By February 2025, the Datasphere Initiative identified over 60 sandboxes worldwide related to AI, data, and technology. Furthermore, key regulatory and policy frameworks, including the European Union’s AI Act and America’s AI Action Plan, view regulatory sandboxes as essential tools for fostering AI innovation and ensuring the safe development and adoption of AI.

Insights shared during the webinar by experts from Spain, Thailand, Luxembourg, Brazil, Korea, Israel and Singapore offer valuable lessons on how different jurisdictions design and operate AI sandboxes, highlighting what works, where challenges arise, and how approaches vary across contexts. For example, Spain provides appropriate, tailored guidance to ensure effectiveness and facilitate regulatory compliance further down the line. Brazil’s sequenced approach includes capacity-building to enable participants to contribute to effective experimentation and evaluation.

>> REVISIT THE WEBINAR AND RELATED PUBLICATIONS <<

Six insights about AI regulatory sandboxes from around the globe

1. AI sandboxes are not uniform

  • According to the Datasphere Initiative, three primary types of sandboxes are emerging worldwide, especially within the context of AI. Regulatory sandboxes: Collaborative processes where regulators work with innovators to test innovations under regulatory supervision.
  • Operational sandboxes: Testing environments and infrastructure where data can be hosted and accessed in controlled conditions.
  • Hybrid models: Combining regulatory oversight with operational capabilities, sometimes offering infrastructure and operational spaces for testing and experimentation (e.g., “supercharged sandbox” in the UK).

These models intervene at different phases of the policy and regulatory lifecycle. Some are employed before formal regulation to identify gaps and suggest necessary updates. Others operate during the development process, supporting iterative regulatory design. Some focus on helping understand legal obligations and ensure regulatory compliance, such as under the EU AI Act. Sector-specific sandboxes are also common, with countries adopting different approaches depending on regulatory priorities and institutional settings. Across these models, regulatory waivers are frequently used to enable experimentation under regulatory supervision. 

Several experimentation-related initiatives, such as regulatory testbeds, living labs, or policy prototyping, share certain features and objectives with regulatory sandboxes. What truly distinguishes sandboxes is that they are the most institutionalised form of regulatory experimentation, usually led by regulators and integrated with regulatory supervision.

Follow us on LinkedIn

2. Coordination is essential

AI does not always fit neatly within existing sectoral, jurisdictional, or administrative boundaries. Its development and deployment span multiple regulatory domains, making effective coordination crucial. Luxembourg’s approach demonstrates this well, showing that AI sandboxes are more than testing spaces—they are platforms for regulatory collaboration and coordination. Luxembourg’s model brings together 11 authorities and innovation actors to align priorities and prevent fragmentation, emphasising the need for skilled project management alongside legal and technical expertise. 

Specific stakeholders within the AI ecosystem pursue different objectives: data protection authorities concentrate on privacy, cybersecurity authorities on resilience, and innovators on efficiency and speed. They also offer different kinds of expertise. Sandboxes can offer a neutral space to reconcile these priorities, fostering trust and mutual understanding. To achieve this, managing the expectations of involved parties and clearly defining the objectives of a sandbox are particularly important steps. 

In Thailand, a multi-faceted approach to AI regulatory sandboxing shows how balancing safety and flexibility depends on agile cooperation between sectoral regulators and industry. This approach integrates three complementary pathways: in the short term, fostering AI deployment where existing rules already allow it; in the medium term, establishing sector-specific sandboxes to manage domain-specific risks and opportunities; and eventually, developing system-wide sandboxes, including for government use, to test cross-cutting applications. Together, these mechanisms help promote AI-driven innovation within current legal frameworks while leveraging testing and experimentation to better understand the implications of emerging AI applications. Effective coordination is essential to prevent duplication of effort, regulatory gaps or conflicting rules.

Sandboxes can also play a valuable role in involving expert and academic communities in the development of AI regulation, with countries such as Spain, Luxembourg, and Brazil benefiting from such expertise at multiple stages of sandbox design and operation.

3. From policy to practice, and back again

AI sandboxes are increasingly used to bridge the gap between regulatory frameworks and real-world implementation. For example, Spain’s regulatory sandbox pilot translates the EU AI Act’s requirements for high-risk AI applications into practical compliance steps, enabling early identification of gaps and clarifying obligations for deployers. In December 2025, the Spanish AI Supervision Agency (AESIA) published a series of introductory and technical resources, developed from insights gathered during the regulatory sandbox pilot, that demonstrate how sandboxes can support evidence-based compliance guidance. Luxembourg’s AI sandbox, in turn, acts as a coordination platform to ensure lessons learned on overlapping obligations feed into the domestic operationalisation of the EU AI Act and related future guidance.

In July 2025, Singapore launched its Global AI Assurance Sandbox, building on insights from a previous pilot phase, to create a testing environment where creators or deployers of GenAI applications can have their applications evaluated by expert technical testers. Key risk aspects examined during testing include hallucination, undesirable content, data leakage, and vulnerability to adversarial prompts, with the findings informing policy guidance. Brazil’s Regulatory Sandbox on AI and Data Protection also exemplifies this trend. It aims to promote transparency, privacy by design and responsible innovation in AI systems that handle personal data, using structured experimentation to help innovators achieve regulatory compliance and assist regulators in understanding how rules work in practice and where adjustments may be necessary.

Simultaneously, AI sandboxes continue to shape future regulatory frameworks. In Thailand, sector-specific sandboxes for digital payments, digital assets, banking and insurance are expected to help regulators understand real-world AI applications and prepare for system-wide governance. As sandboxes move regulation from theory to practice and back to policy, they create an iterative loop that can strengthen trust and adaptability in AI governance. To achieve this, sandboxes should generate insights to inform better regulation. This, in turn, requires consistent reporting, sharing of results, and the establishment of feedback loops across sectors and countries to boost compliance and policy development.

This is one example of a knowledge-sharing process in AI regulatory sandboxes.

Nevertheless, translating sandbox results into regulatory improvements remains challenging, even in countries like Korea, which has considerable experience conducting regulatory experiments across sectors.

4. Incentives matter

Participation in AI sandboxes is not automatic. Clear and well-designed incentives are essential for both innovators and regulators. Israel, for instance, has introduced a government fund that provides financial support, legal counselling and mentorship for regulators launching AI sandboxes, while also offering grants to participating firms. Similarly, Singapore reduces testing-related barriers to GenAI adoption through practical guidance and access to specialised testing partners.

These models recognise a fundamental challenge: AI experimentation is resource-intensive and needs to focus on areas where it matters most. Without targeted support, regulators may struggle to operate sandboxes, and companies might be hesitant to participate. Furthermore, when offering incentives, authorities should encourage a diverse mix of participants—small firms, big players, different sectors, and different AI applications—to ensure that sandbox insights are both representative and robust. The complex nature of regulatory sandboxes themselves may also pose challenges for some applicants or even participants. In this context, Brazil outlined a three-stage execution framework, starting with capacity building for selected participants undertaken by a partner university, before advancing to the experimentation and evaluation phases. 

5. The growing need for interoperability and cross-border collaboration

As AI systems operate across borders, there is a growing need for AI sandboxes to extend beyond national borders. Without international coordination, firms may engage in ‘jurisdiction hopping’, seeking the most permissive regulatory environments. Interoperability between sandboxes is thus becoming a governance necessity. Cross-border collaboration is also crucial for international regulatory cooperation. In Brazil’s case, preparatory work to develop the sandbox included international consultations on the experimental methodology. This approach enables the benefit from international practices and standards and facilitates the sharing of experiences in later stages of the project. 

Cross-border sandboxes have already proven their worth. Singapore’s Global AI Assurance Pilot, for example, involved 17 AI deployers from nine countries collaborating with 16 specialised testers from the US, UK and Europe. Use cases include summarisation, chatbots to AI applications in healthcare, finance and human resource management. These cross-border tests allowed regulators and companies to understand how AI performs in different legal, cultural and technical environments. For example, a chatbot that safely managed English queries inadvertently leaked confidential information when prompted in Mandarin, demonstrating the importance of multilingual testing. 

6. AI sandboxes come with challenges of their own

AI sandboxes face several challenges. Regulators often encounter capacity limitations and lack the technical expertise or project management skills necessary to supervise complex AI systems. Fragmentation and coordination issues also arise, as AI spans multiple sectors and necessitates collaboration among numerous authorities, including across borders. 

Designing suitable requirements and safeguards for sandbox frameworks can be challenging, especially when multiple regulatory regimes are involved, as demonstrated by Brazil. Deciding the appropriate level of transparency, human oversight, and data governance can be particularly difficult when firms seek waivers to speed up testing.

At the same time, Korea’s experience demonstrates that although safety and consumer protection rules are vital, excessively strict requirements may deter participation, especially among SMEs, and hinder experimentation. Safeguards should therefore be proportionate to and aligned with the risks posed by the technology, as overly complex procedures can undermine the agility required in regulatory sandboxes in a rapidly evolving AI landscape.

Measuring the impact of AI sandboxes is also difficult. Without clear metrics, sandboxes risk becoming isolated experiments rather than influential policy tools. Ideally, impact should be monitored across various areas, such as faster time-to-market for compliant AI systems, increased regulatory clarity and coherence (including through less fragmentation), and tangible updates to laws and standards shaped by sandbox insights. 

Furthermore, as highlighted in a 2024 OECD policy paper, there are potential limitations concerning legality, feasibility, resources, and equity. Regulatory experiments should adhere to constitutional norms, including those concerning equal treatment.

A vital element for responsible innovation and public trust?

As a relatively new regulatory tool designed to address a rapidly evolving general-purpose technology, AI sandboxes raise significant questions about their role. Some of the questions raised during the online workshop include:

  • What role might civil society play in AI sandboxing?
  • How can public institutions build the expertise required to supervise regulatory sandboxes involving frontier-level AI systems?
  • How can sandboxes balance flexibility with protecting long-term societal values (e.g., what types of safeguards should be in place regarding regulatory exemptions)?
  • What measures, such as reporting and documentation requirements, talent management, and capacity building, are necessary to ensure transparency and build trust in AI regulatory sandboxes? 

As AI governance develops and these questions are addressed, AI sandboxes hold the potential to become key tools for promoting responsible innovation, enhancing governance and building public trust in AI systems across the globe. 

The OECD is well placed to advance these objectives by facilitating the systematic exchange of knowledge and expertise, and by developing standardised, comparable frameworks for measuring outcomes. It can also use its convening role to support alignment on guidance for the targeting, design and implementation of AI regulatory sandboxes. By grounding this work in empirical evidence and practical experience, the OECD can help strengthen the overall evidence base and inform more effective policy approaches.

The authors would like to thank Natalie Cohen, Lucia Russo, Guillermo Hernandez, Xavier Pearson, Viktor Samek and John Leo Tarver for their contributions to this piece.

The post Why AI Sandboxes matter for responsible innovation and public trust appeared first on OECD.AI.

Can we create a clear understanding of what agentic AI is and does?

3 March 2026 at 08:38
chalk drawing of two heads with messy string

AI agents and agentic AI based on large language models are becoming more autonomous and capable of interacting with both physical and virtual environments. As the capabilities of these AI systems grow, they are gaining visibility, and with reason. It is reaching a point where they could become the driving force behind innovation, investment and improved productivity across sectors by streamlining processes and enabling more efficient operations.

While ideas related to agency have long been explored in academic research in fields such as philosophy, economics and computer science, recent advances in AI are stretching conceptual boundaries. As AI’s capabilities evolve, so do our shared understanding of what qualifies as AI agent and agentic AI.

The OECD report, The agentic AI landscape and its conceptual foundations, developed by the OECD.AI Expert Group on Agentic AI, helps clarify what AI agents and agentic AI are and how they differ. Grounded in the OECD AI system definition, the analysis examines how these terms are defined and used across the literature. By analysing key features, overlaps and distinctions and mapping them to the core elements of the OECD definition of an AI system, the report helps to establish more precise and consistent terminology. And in a rapidly evolving field, conceptual precision is essential for effective, well-informed governance.

Three key messages stand out in the report:

  • AI agents and agentic AI are closely related, but not interchangeable.
  • Agentic AI ought to be seen as a socio-technical paradigm.
  • Despite technological gaps and varying levels of maturity in areas such as digital security and privacy, uptake is growing.

The common foundations and meaningful distinctions of AI agents and agentic AI

Our analysis shows that AI agents and agentic AI share foundational characteristics. Both involve systems with a degree of autonomy that pursue goals and can perceive and act within physical and virtual environments.

However, there are differences that mean these terms are not interchangeable.

  1. AI agents can be understood as systems that perceive and act on their environment with a degree of autonomy, using tools as needed to achieve specific goals and adapt to changing inputs and contexts.
  2. By contrast, agentic AI generally refers to systems composed of multiple co-ordinated AI agents that can break down tasks, collaborate and pursue complex objectives autonomously over extended periods. Agentic AI systems are designed to operate in more open-ended, less predictable physical and virtual environments, and to function with minimal human supervision.

In short, agentic AI is more complex, as it can co-ordinate multiple agents, perform task decomposition and delegation, and sustain operations over longer periods. It can also operate in more complex, less predictable environments with limited human oversight.

Agentic AI as a socio-technical paradigm

Agentic AI systems are not isolated technical artefacts. They are frequently embedded in social contexts and interactions and operate within a socio-technical paradigm.

Their value lies not only in autonomous action, but in interaction with other AI agents, humans and institutional processes. Co-ordination and negotiation across these actors require advanced reasoning capabilities, robust infrastructure and reliable communication protocols.

This relational perspective is an essential part of what agentic AI is. This means that understanding how they interact within broader ecosystems is essential to designing agentic AI systems that function responsibly and effectively, particularly in open or high-stakes environments.

Uptake is accelerating, but maturity is uneven

The report also presents descriptive evidence on trends in AI agent adoption. Many developers have already integrated them into their toolkits, and survey data indicate that nearly half of respondents on Stack Overflow use them or plan to do so.

To be clear, adoption should not be confused with maturity. Developers highlight opportunities to further strengthen the security, privacy and accuracy of AI agents. These concerns underscore an important point: as the capabilities of agentic AI advance rapidly, progress in robust, trustworthy AI systems must keep pace.

A foundation for further analysis

Overall, the report provides a descriptive overview of the agentic AI landscape, clarifying key concepts and characteristics and establishing a shared analytical foundation. By anchoring the discussion in the OECD AI system definition, it aims to promote coherence across technical and policy communities.

Looking ahead, an improved understanding of real-world use will be essential to identify where safeguards, standards, and governance mechanisms will be most effective. Policy-relevant typologies that build upon this work could help guide governance efforts to distinguish systems by level of autonomy, degree of adaptiveness, domain of operation and scale of impact. Evidence-based policymaking will require more empirical data on how AI agents and agentic AI are being adopted and used across sectors, as well as clearer evidence of their broader implications and impacts.

This report contributes to a clearer, shared understanding of agentic AI and provides a basis for thoughtful, forward-looking policy grounded in conceptual clarity. As agentic AI systems become more capable of coordinating multiple AI agents, taking action and operating over longer periods, governance conversations have to keep pace.

The post Can we create a clear understanding of what agentic AI is and does? appeared first on OECD.AI.

The OECD’s new responsible AI guidance: A compass for businesses in a complex terrain

19 February 2026 at 09:30
people talking in a server room

Companies hoping to take advantage of AI’s opportunities need to be trustworthy. Whether investing in, developing, or using AI, the OECD’s new Due Diligence Guidance for Responsible AI provides businesses with an internationally agreed, government-backed tool to demonstrate that markets and societies can trust their AI systems.   

Recent international reporting underscores a growing consensus: AI is not just a technological shift. It is a major geopolitical, economic, and societal phenomenon that demands coordinated action amongst all actors, including companies. 

AI has the potential to transform society through productivity, economic value and solutions to complex challenges, but for these benefits to materialise, AI needs trust.  So far, the technology seems to advance faster than its guardrails. The gap between AI systems and appropriate safeguards is now one of the defining challenges for policymakers and global businesses alike. Both are under pressure to balance AI innovation and diffusion with safety and risk management. Success depends on getting the balance right.

Risks throughout the AI value chain are continually evolving

Risks to people and the environment can manifest at any point along the AI value chain. The OECD actively tracks and categorises risks through its AI Incidents and Hazards Monitor.

Here are a few examples. At one end of the AI value chain, there are the people who label, clean, and moderate the vast datasets required to train AI models. They can face low wages, long hours, and suffer psychological distress from exposure to harmful content. Companies need to ensure decent work for data enrichment workers.

The environmental costs of running AI systems can also be significant, particularly for energy and water consumption by data centres that power AI development and deployment, which may lead to higher energy prices.

Data privacy is another critical concern. AI models are trained on massive datasets that may include personal or sensitive information. If these datasets are not properly anonymised and secured, it can lead to data breaches. If AI models “memorise” and reproduce sensitive data in their outputs, they can expose confidential details, creating legal and ethical dilemmas.

At the other end of the AI value chain, the potential for AI misuse poses risks such as reputational harm and the spread of misinformation. AI-generated deepfakes, for instance, can be used to create realistic but fabricated content, damaging reputations or manipulating public opinion. Similarly, AI can be used to generate and disseminate mis and dis-information at speed and scale, eroding trust in institutions and potentially influencing events.

Worldwide, governments, consumers, and markets are calling for responsible and trustworthy AI. This is one of the reasons for the surge in mandatory and voluntary AI risk management frameworks, responsible AI initiatives, global agreements, academic research and statements from industry leaders and investors. However, this surge in frameworks is also increasing complexity for companies, as risk management is defined differently across jurisdictions and understanding of AI-related risks is evolving.

OECD Due Diligence Guidance for Responsible AI: A flexible, whole-of-value-chain approach to support businesses in navigating evolving risks and rules

This is why the OECD has now developed the first internationally agreed, government-backed Due Diligence Guidance for Responsible AI. Backed by all the OECD’s member countries, plus 17 partner governments and the EU, this Guidance helps enterprises navigate the complex terrain of AI risk management. It is designed to help businesses ensure that the AI systems they develop are trustworthy, used and developed safely and responsibly, and aligned with broad societal values.

Concretely, this Guidance offers:

  • A step-by-step framework for enterprises to set up internal management systems capable of proactively identifying and responding to risks related to human rights, labour standards, and environmental impacts.
  • Comprehensive coverage of all risk areas from the leading international standards that it is built on and reflects, notably, the OECD Guidelines for Multinational Enterprises on Responsible Business Conduct (MNE Guidelines) and the OECD Recommendation on Artificial Intelligence (AI Principles);
  • Recommendations and implementation examples for everyone in the AI value chain, from data suppliers and infrastructure providers to financiers and end-users – including enterprises. The guidance emphasises a “whole-of-value-chain” approach to support secure and resilient AI value chains more resistant to supply chain shocks and interference.
  • A roadmap of related provisions in existing frameworks, indicating how each step complements and relates to relevant provisions from AI risk management frameworks. This feature helps enterprises understand how implementing this guidance can help them meet expectations from multiple sources and navigate the current landscape of AI risk management frameworks.

Responsibility and trust can give a competitive edge

Responsibility and innovation not only coexist but also reinforce each other. Companies that show a commitment to responsible AI and actively address potential risks can gain trust from investors, customers, regulators, and policymakers. This trust leads to a competitive edge. Instead of hindering innovation, responsible AI practices can speed up growth by reducing obstacles and preventing costly damage to reputation, legal issues, and society.

Responsible and trustworthy AI is becoming increasingly crucial for accessing global markets as international regulatory and voluntary risk management frameworks evolve. Companies in the AI value chain that meaningfully implement the Guidance’s recommendations can position themselves advantageously for cross-border expansion, potentially avoiding the substantial costs of retrofitting systems to meet various regional requirements.

As AI continues to develop rapidly, frameworks and best practices for responsible AI are likely to evolve as well. To help stakeholders keep pace, the OECD will launch an online navigation tool later this year with updates on new frameworks and use cases.

The post The OECD’s new responsible AI guidance: A compass for businesses in a complex terrain appeared first on OECD.AI.

The Global South can shape AI in practical terms: Why the India AI Impact Summit Matters

15 February 2026 at 14:13
india gate new delhi

Artificial intelligence is changing fast, and the world is feeling both excited and uneasy about it. People use AI tools every day in hospitals, classrooms, companies and public services, yet the rules that guide these tools are still developing. Many governments are trying to find a balance between innovation and safety. Others are trying to make sure that AI actually improves people’s lives without widening gaps.

This is the backdrop against which the India AI Impact Summit 2026 will take place in New Delhi in February. Earlier global AI meetings, including the 2023 gathering at Bletchley Park and subsequent summits in Asia and Europe, helped define the risks and push for action.

These summits did not occur in isolation but are part of broader global efforts to coordinate responsible approaches to AI. The G7 Hiroshima Process in 2023–24 established a shared commitment to trustworthy, human-centric AI, leading to the adoption of the Hiroshima Declaration, which calls for international cooperation on safety, transparency, and risk mitigation.

Building on that, the Paris AI Summit in 2025 moved the conversation toward implementation, with an early agreement on safety evaluations, incident-reporting mechanisms and commitments to support countries with limited technical capacity. The India AI Action Summit represents the next step in this progression: translating these collective principles into measurable on-the-ground outcomes.

In recent months, people have repeatedly asked me two questions. Why should India host such a major global meeting? And is this summit actually useful for India and the world?

The simple answer here is that the next phase of AI will not be decided by a small number of companies or countries. It will depend on whether billions of people, especially in the Global South, can use AI safely, affordably and accountably. India, with its linguistic diversity, strong digital public infrastructure and experience deploying technology at a population scale, is well positioned to help shape this practical phase.

However, AI comes with challenges related to privacy, digital exclusion and the balance between innovation and oversight. But these very tensions make India’s experience pertinent to other countries facing the same trade-offs.

This blog post explains why that matters, what the international community can expect in Delhi, and how we should measure progress at the end of the summit.

Why India, and why now

AI deployment in the Global South will shape global outcomes

Much of the world’s discussion on AI has focused on frontier models, international competition and long-term safety. These debates are important, but AI’s greatest impact will be felt in how it reaches ordinary people. From farmers and students to small businesses, frontline health workers and local governments.

More than half of the world’s population lives in countries categorised as the Global South — a term first popularised in the late 1960s to describe post-colonial economies, and one I don’t fully agree with, as it often flattens diverse countries into a single broad category.

If AI is to be truly global, it must work for multilingual, resource-constrained and diverse environments. This includes reliable translations, culturally grounded datasets, accessible interfaces and low-cost deployments. It also means designing systems that respect human rights and democratic norms even in places with limited regulatory capacity.

India sits at the intersection of these challenges. With over a billion people, 22 official languages and thousands of dialects, any technology deployed at scale must be inclusive by design. India’s experience offers lessons for many other countries navigating the same realities.

India has a strong track record in large-scale digital public infrastructure

India’s digital public infrastructure, or DPI, is one of the most widely referenced success stories of how technology can enable access and accountability. Systems like Aadhaar, UPI and DigiLocker have helped millions access identification, financial services and digital records. These platforms were built with interoperability and openness in mind, which has led to a wave of public and private innovations.

At the same time, these systems have also raised important questions about privacy, data security, and exclusion of marginalised communities who lack documentation or digital access. India’s ongoing work to address these concerns—through data protection legislation, improved grievance mechanisms, and efforts to reach the digitally excluded—provides practical lessons about implementation challenges that other countries will inevitably face.

The India AI Impact Summit is expected to draw on this experience, including both successes and areas for improvement. The global community is watching to see how India will frame the link between AI and digital public goods, and how these tools can be used responsibly in sectors such as education, healthcare and social protection.

International expectations are focusing on implementation leadership

The earlier global AI safety and governance summits created momentum. They helped identify risks, promote transparency and encourage cooperation. But now, many countries and organisations want clarity on what should happen next.

The India summit is an opportunity to shift the conversation from what AI might do to what it should deliver. This includes measurable improvements in public services, clearer accountability mechanisms and more inclusive access to AI tools. By focusing on implementation, India can complement the work of the OECD-GPAI, UNESCO and other international bodies.

 What the India AI Impact Summit should prioritise

A conversation about measurable, real-world outcomes

The summit should begin by asking a straightforward question: What changes on the ground when AI is deployed responsibly at scale? To answer it, discussions need to move beyond broad aspirations and focus on concrete domains like public healthcare triage, classroom support tools, agricultural advisory systems, and other public-sector applications where impact can be seen and measured.

Government delegates should be encouraged to present evidence, not statements of intent. That means clear baselines, transparent evaluation methods, and metrics that reflect real improvements: higher diagnostic accuracy, increased crop yields and shorter benefit-processing times all achieved without compromising fairness or human oversight.

If the summit succeeds, it will shift the global conversation toward what works, for whom, and under what conditions.

Three ways the Global South can shape the international agenda

A meaningful summit requires a wide range of voices—especially from regions where AI deployment will shape social and economic outcomes for decades to come. Countries across Asia, Africa, Latin America and the Middle East bring their lived experiences of linguistic diversity, data scarcity, affordability constraints and uneven digital access.

The summit should create space for these countries to set priorities rather than simply respond to frameworks developed elsewhere. Their perspectives are vital for building governance models that reflect the realities of low-resource contexts, rather than idealised assumptions from high-income environments.

A more pluralistic conversation would reinforce a simple principle: responsible AI cannot be universal if it is not also contextual.

Rebalancing the narrative with the immediate societal, environmental and institutional challenges

One of the most important roles the summit can play is to broaden the global AI discourse. Today, existential risk narratives dominate many international forums, often overshadowing more immediate and systemic issues. The India AI Impact Summit should refocus attention on the present: AI’s energy footprint, labour displacement, rising misinformation, digital exclusion and the growing pressure on public institutions to oversee algorithmic systems they are not adequately equipped to oversee.

The environmental dimension deserves particular attention. Training and deploying large AI models require significant energy resources, disproportionately affecting the Global South. Many of these countries face climate vulnerability, fragile grids and competing development priorities. For regions already grappling with heatwaves, droughts and energy shortages, the cost of “AI at scale” cannot be separated from broader planetary concerns. If AI is to be deployed responsibly, discussions must also consider energy and natural resource efficiency and equitable access to compute.

These issues determine how people experience AI today and whether they trust it tomorrow. Giving them equal weight would help correct the imbalance in global discussions and lead to governance that addresses risks people actually face, not only those imagined at the far horizon.

Potential wins for the India AI Impact Summit

Practical pathways for responsible public-sector deployment

Across sectors, governments are eager to use AI to strengthen healthcare, expand access to education and streamline welfare delivery. Yet many lack clarity on how to procure, evaluate or oversee these systems responsibly. A meaningful outcome of the summit would be simple, actionable pathways that public agencies can adopt without specialised expertise. These might take the form of evaluation checklists with acceptable error and bias thresholds, procurement templates with human oversight requirements, or clear guidance on when and how officials should override an AI recommendation. Transparent case studies and training for civil servants would also help countries move from hesitation to informed, confident experimentation.

Strengthened mechanisms for trust and accountability

Concerns about misinformation, bias, privacy and security continue to rise, and many countries, particularly those with limited technical capacity, need practical tools to manage these risks. The summit could make a real contribution by advancing shared approaches to incident reporting, auditing and assurance, as well as safety testing methods that work across varied deployment contexts. Small pilot frameworks would help establish a common baseline of accountability. Such efforts would not only support global cooperation but also build public trust at a time when many citizens and policymakers remain uncertain about the reliability of AI systems.

Broader cooperation on multilingual and inclusive AI with robust safety infrastructure

Many countries struggle with adapting AI to their unique linguistic profiles. India’s long-standing work in language technologies positions it to convene collaborations on multilingual and inclusive AI. New partnerships on datasets, dialect-specific models, local-first interfaces and research on linguistic bias could meaningfully expand access for millions of people worldwide.

But inclusion must be matched with safeguards. As AI tools become more widely available, countries will need parallel investments in risk-assessment expertise, regional coordination on harmful content and support for the development of first-generation regulatory frameworks. Striking the right balance between openness and safety would reinforce core OECD AI Principles and help ensure expanded access does not bring greater vulnerability.

Broader implications for global AI policy

Everyday impact before frontier risks

Research on frontier AI risks must continue, but the India AI Impact Summit signals an important rebalancing of global attention, as mentioned before. It asks policymakers to look beyond hypothetical future scenarios to acknowledge how AI is already shaping critical aspects of our daily lives, from healthcare triage and classroom instruction to welfare delivery, agricultural advice and urban mobility.

For most people, the urgent question is not whether AI poses an existential threat, but whether the systems they encounter today are reliable, safe and genuinely useful. The summit’s focus on practical impact aligns global governance with lived reality.

Ensuring AI benefits for everyone

A second implication is the reaffirmation that inclusion is not a downstream concern but a prerequisite for responsible AI. Global conversations often gravitate toward powerful models built in highly resourced environments, yet billions of people rely on limited connectivity, low digital literacy and minority-language interfaces.

India’s leadership places these conditions at the centre of the global agenda. It broadens the imagination of what “good AI” must account for, reminding the world that both equitable deployment and cutting-edge capability are essential to whether AI helps or harms societies.

Building a more open and collaborative ecosystem

The summit also nudges the world toward a more open and cooperative model of AI development. Some countries can share tools, datasets, and governance mechanisms to help each other build their own capabilities rather than remain passive consumers. Openness here is not about lowering standards; it is about raising the global floor and ensuring that safety capacity grows alongside access. Many countries want to participate meaningfully in the AI economy, and the summit offers a platform to explore practical pathways for doing so.

What the world should take away from Delhi

The India AI Impact Summit 2026 is more than just another international meeting because it represents a shift from abstract debates to concrete action. Its core question — how to make AI useful, safe and inclusive at scale — goes to the heart of global governance. If the summit delivers practical tools, clearer deployment pathways and stronger cross-regional collaboration, it will set a new benchmark for what international coordination on AI can achieve.

The world is watching India, not because it claims to have all the answers, but because it has repeatedly demonstrated the ability to turn large-scale ideas into real-world outcomes.  And it has done so while openly confronting the tensions and trade-offs that accompany such efforts. In a period of rapid technological change, this experience is invaluable.

As AI evolves, the global community will increasingly need countries that can translate principles into practice at a population scale. The India AI Impact Summit is a chance to advance that work. If successful, its influence will extend far beyond India, shaping how the world understands and pursues responsible AI in the years ahead.

The post The Global South can shape AI in practical terms: Why the India AI Impact Summit Matters appeared first on OECD.AI.

❌