
Tldr
The document argues that current AI dialogue is split between “alarmist” existential warnings and immediate empirical harms, but lacks a robust definition of risk itself. It proposes adapting natural risk frameworks (Hazard, Exposure, Vulnerability) to AI, while concluding that General Purpose AI (GPAIS) operates under “Severe Uncertainty” rather than calculable risk, necessitating adaptive, control-based regulation rather than simple probability assessments.
The document categorizes AI risks into two distinct camps:
- Existential & Societal (Long-termism): Focuses on catastrophic scenarios like Loss of Control, where super-intelligent entities might render humans obsolete, and the risk of extinction is placed alongside nuclear war and pandemics.
- Concrete & Immediate (Empirical): Focuses on currently observable harms such as Algorithmic Discrimination (e.g., facial recognition bias), Privacy Violations (intrusive surveillance), and Environmental Impact (carbon footprint of training LLMs).
The EU has adopted a risk-based approach visualized as a pyramid, classifying systems by threat level:
- Unacceptable Risk: Banned (e.g., social scoring, real-time biometric ID).
- High Risk: Heavily regulated (e.g., critical infrastructure, employment tools) requiring strict conformity assessments.
- Limited & Minimal Risk: Subject to transparency (e.g., chatbots) or codes of conduct (e.g., spam filters).
- Critique: The document argues this framework lists “hazards” but lacks a deep epistemological conceptualization of risk, creating a gap in addressing unforeseen dangers.
To better analyze risk, the document adapts the “Natural Risk” methodology to technology:
- Hazard (The Source): The intrinsic potential for harm (e.g., a bug in code, a bias in a dataset). Mitigation involves debugging or prohibiting specific algorithms.
- Exposure (The Scale): The proximity or digital reach of the hazard (e.g., viral spread of Deep Fakes). Mitigation involves limiting access, such as age-gating or restricting high-stakes use cases.
- Vulnerability (The User): The susceptibility of the target (e.g., children, the elderly, the digitally illiterate). Mitigation focuses on protecting these specific groups from manipulation or emotional dependence (e.g., with “carebots”).
A critical epistemological finding is that General Purpose AI (like LLMs) defies standard risk probability calculations (
) because their use cases are unbounded and unknown.
- The Tuxedo Fallacy: The error of treating the real world (an unknown “Jungle”) as if it were a casino with known odds (a “Tuxedo” environment). We cannot calculate the probability of specific AI hazards because we do not know what all the hazards are.
- Conclusion: AI operates under Severe (Knightian) Uncertainty, not Risk. Therefore, policy must shift from Ex-Ante prediction (which is impossible) to incremental introduction, continuous monitoring, and adaptive regulation.
The contemporary dialogue concerning Artificial Intelligence (AI) is characterized by a significant tension between alarmist warnings of hypothetical, long-term existential threats and the growing body of evidence regarding immediate, concrete harms. An epistemological analysis—that is, a critical examination of how AI risks are defined, structured, and understood—is fundamentally necessary to inform and ultimately create robust, effective regulatory and ethical frameworks. This approach allows for a precise categorization of risks, moving beyond a simple dichotomy to a nuanced spectrum of potential impact.
The Spectrum of AI Risks
The broad spectrum of AI-related risks can be formally segmented into two primary categories: speculative, high-impact societal or existential threats and currently observable, tangible risks. Both categories, while differing in temporal and empirical immediacy, demand serious consideration in the academic and policy arenas.
Existential and Societal-Scale Risks (Long-Termism)
A substantial portion of the high-level public discourse is dominated by “long-termism” and scenarios involving catastrophic global consequences. Prominent organizations and influential figures have issued strident warnings, emphasizing the need for extreme caution regarding the unrestrained trajectory of current AI development.
One of the central concerns is the Loss of Control over advanced artificial systems, a theme critically examined in initiatives such as the Future of Life Institute’s “Pause Giant AI Experiments” letter.
Cite
“Should we let machines flood our information channels with propaganda and untruth? Should we automate away all the jobs, including the fulfilling ones? Should we develop nonhuman minds that might eventually outnumber, outsmart, obsolete and replace us? Should we risk loss of control of our civili/action?”
Future of Life Institute’s open letter Pause Giant AI Experiments
This perspective posits foundational questions about maintaining civilization-level control over super-intelligent entities. Specific warnings include the potential for AI to autonomously flood information channels with sophisticated, untraceable propaganda, thereby destabilizing democratic institutions; the automation of fulfilling human labor, which risks profound societal and economic dislocation; and, most fundamentally, the development of nonhuman minds whose goals may be orthogonal to—or directly conflict with—human survival and flourishing, potentially leading to human obsolescence or replacement.
Further solidifying the catastrophic viewpoint, the Center for AI Safety has formally categorized the risk posed by AI in the same gravity class as pandemics and nuclear warfare.
Cite
“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”
Center for AI Safety’s Statement on AI Risk
Their assertion is that the mitigation of the risk of extinction from AI must be elevated to a paramount global priority, necessitating coordinated international action and regulatory safeguards against runaway technological progress.
Concrete and Immediate Risks (Empirical Harms)
In sharp contrast to the speculative nature of existential risks, a growing body of academic literature and empirical studies has identified tangible harms that are actively manifesting in society today. These risks are not theoretical but represent demonstrable failures in system design, data governance, and deployment ethics.
A critical and pervasive immediate risk is Algorithmic Discrimination, which involves AI systems producing systematically unfair or biased outcomes against specific demographic groups. The seminal “Gender Shades” project by Buolamwini and Gebru (2018) provided compelling evidence that commercial facial recognition systems exhibited significantly higher error rates when identifying darker-skinned individuals and women, compared to white men. This differential performance is directly traceable to non-representative or poorly curated training data, which lacks diversity and embeds historical societal biases into the resulting algorithmic model. This systemic bias translates into immediate real-world harms in areas like law enforcement and hiring.
Furthermore, the operational requirements of AI—particularly the voracious appetite for data—pose significant risks to personal autonomy and liberty, culminating in widespread Privacy Violations. The relentless drive for higher performance and accuracy often leads to intrusive surveillance and the potential for data misuse across various applications. The mere existence of pervasive AI-powered monitoring infrastructure fundamentally alters public behavior and erodes the expectation of privacy.
Finally, the substantial computational requirements necessary to train and operate state-of-the-art large language models (LLMs) and other complex systems constitute a significant Environmental Impact. The massive energy consumption associated with the training process results in a large carbon footprint, contributing materially to climate change concerns and presenting an ethical dilemma regarding the sustainability of current AI research paradigms.
Regulatory Frameworks: European Union’s AI Act
In response to the growing complexity of AI-related challenges, the European Union (EU) has taken a pioneering legislative step by proposing and developing the EU AI Act. This landmark legislation is designed to codify AI safety and ethical deployment by adopting a risk-based regulatory approach. This methodology dictates that the stringency of legal obligations placed upon an AI system is directly proportional to the severity and nature of the risk the system poses to fundamental rights and safety.
The central tenet of the EU AI Act is its hierarchical classification system, which organizes AI systems into four distinct risk categories. This framework is conceptually illustrated as a pyramid, with the most severe risks occupying the prohibited apex and the least severe forming the broad base.

| Risk Level | Regulatory Consequence | Examples of System Use |
|---|---|---|
| Unacceptable Risk | Strict Prohibition and outright banning of deployment. These systems are deemed to pose a clear, egregious threat to human safety, livelihoods, and fundamental rights. | Systems used for social scoring by governmental or public authorities; the deployment of real-time remote biometric identification in public spaces (with narrowly defined exceptions for law enforcement); the use of subliminal manipulative techniques; and the enforcement of dark-pattern AI interfaces. |
| High Risk | Mandated Conformity Assessment before any system can be legally placed on the market. These systems trigger strict, specific obligations related to the quality of training data, comprehensive technical documentation, robust record-keeping, and the necessity of adequate human oversight measures throughout the system’s lifecycle. | AI utilized in critical infrastructure (e.g., energy, transport); systems governing access to and grading within education and vocational training; technologies involved in employment (e.g., automated CV sorting and candidate evaluation); applications within law enforcement and criminal justice; and systems managing migration, asylum, and border control. |
| Limited Risk | Primary focus on Transparency Obligations. Users of these systems must be explicitly and clearly informed that they are interacting with or subject to a machine-driven process. | General-purpose AI chatbots and conversational agents; emotion recognition systems operating in non-high-risk contexts; and deep fakes or synthetically generated media, which must be clearly and unambiguously labeled as artificial. |
| Minimal Risk | Encouragement of voluntary compliance through Codes of Conduct. These systems do not impose new mandatory legal obligations, reflecting their low potential for harm. | Routine technologies such as email spam filters and AI used for video games. |
Despite the detailed and structured nature of this regulatory framework, we can identify a fundamental epistemological gap in its foundation.
While the EU AI Act effectively regulates different levels of risk, it fundamentally lacks a proper, comprehensive conceptualization of the notion of risk itself.
The current approach tends to operate as a functional list of prohibited or regulated applications (i.e., specific hazards or verticals such as discrimination or privacy breaches) rather than establishing a general, comprehensive philosophical and theoretical approach to AI risk.
To advance effective and future-proof policy, academic analysis must transcend a simple tabulation of hazards. It necessitates a deep, epistemological analysis—a process of systematically breaking down the complex notion of risk into its fundamental, constituent components. This theoretical work is crucial for moving beyond ad hoc regulation of specific applications toward a robust, generally applicable framework capable of addressing emerging and unforeseen AI risks.
Adapting Natural Risk Frameworks
To establish a robust theoretical foundation for AI safety and policy, it is constructive to adapt the established methodological principles of Natural Risk Analysis—traditionally applied to phenomena like earthquakes or floods—and transpose them into the context of complex technological artifacts. This transference allows for a systematic and rigorous deconstruction of AI risk into measurable and manageable components.
The Quantitative Definition of Risk
The concept of risk inherently lacks a single, univocal definition, varying significantly between the precise, quantitative language used in technical engineering domains and the broader, more qualitative interpretations prevalent in non-technical public discourse. However, the cornerstone of risk quantification is the classical approach, popularized by the Royal Society in 1983, which defines
Definition
Risk as a mathematical expectation:
where
represents the Probability of a specific harmful event occurring, and denotes the Magnitude of the adverse consequences resulting from that event.
The Multi-Dimensional Analysis
For a more granular and actionable approach to technological risk mitigation, a multi-dimensional analysis is preferred. This methodology breaks down the holistic concept of risk into three distinct, interconnected components: Hazard, Exposure, and Vulnerability. Analyzing risk across these three axes enables the design of targeted and context-specific intervention strategies, moving beyond simple probability calculations.

By systematically decomposing risk into its three constituent elements—Hazard, Exposure, and Vulnerability—we gain distinct mitigation levers that can be strategically manipulated to reduce overall risk. This analytical separation is critical because the most effective intervention strategies often differ fundamentally depending on whether the risk originates from an immutable natural phenomenon or a malleable technological artifact like an AI system.
Hazard
Definition
The Hazard component is defined as the intrinsic source of potential harm, whether it originates from a natural phenomenon or is embedded within a technological artifact or system. It is the originating event or condition with the potential to cause loss, damage, or harm.
In the context of technology, examples include catastrophic structural failure, such as a bridge collapse; an uncontrolled event like a toxic chemical leak; or, critically for AI systems, a massive data leakage event where sensitive or private information is unintentionally disclosed, representing an intrinsic systemic flaw. Reducing the hazard component involves directly addressing and minimizing the intrinsic source of potential harm.
- Natural Risk Context: In the domain of natural risk, reducing the intrinsic hazard is rarely feasible. Humanity possesses limited ability to stop or significantly diminish the force of major events like volcanic eruptions or earthquakes. Mitigation is therefore generally focused on the other two components.
- Technological (AI) Risk Context: Conversely, in the context of technology, hazard reduction is often highly viable. Because AI systems are human-made artifacts, the technology itself can be modified and controlled.
- Intervention Strategies: These include the prohibition of specific algorithms deemed too dangerous or ethically compromising (as seen in the EU AI Act’s unacceptable risk category); the withdrawal of demonstrably dangerous or defective products from the market; and, most commonly, the process of debugging code and refining models to actively identify and remove biases or functional flaws that constitute the hazard. This is a direct intervention on the artifact’s capacity to cause harm.
Exposure
Definition
Exposure quantifies the extent to which people, critical infrastructure, economic assets, or environmental elements are geographically or digitally situated in an area where they are susceptible to be affected by the identified hazard. It is a spatial or locational concept. In a natural context, this involves physical proximity, such as constructing residential housing at the base of a volcanically active mountain.
In the technological domain, exposure is defined by the number of users utilizing a specific online service that harbors a data vulnerability, or the connection of critical national infrastructure to an inadequately secured network that is a known target for cyber intrusion. Reducing exposure involves limiting the spatial or digital proximity of potential victims or assets to the source of the hazard. This strategy is highly effective and applicable across both natural and technological domains.
- Natural Risk: Common strategies include the designation of mandatory evacuation zones or denying construction permits in geologically unstable or flood-prone areas, thereby physically limiting who is exposed.
- Technological (AI) Risk: In the digital realm, exposure reduction is achieved through controlled access and segmentation. Examples include implementing strict age restrictions for social networks to limit the exposure of minors to potentially manipulative content; denying internet or AI access to mission-critical systems through air gapping (physical isolation); or implementing stringent access control mechanisms to critical datasets.
Vulnerability
Definition
Vulnerability refers to the physical, social, economic, or environmental processes and factors that increase the susceptibility of an individual, group, or community to the harmful impacts of a hazard. Unlike hazard (the source of harm) and exposure (being in the line of harm), vulnerability describes the ability to cope with or recover from the harm.
General examples include populations with inherent physical fragility, such as older adults or children. In the specialized context of AI, vulnerability manifests in scenarios where populations with low digital literacy interact with complex, opaque automated decision-making systems, making them less able to detect or contest algorithmic error or manipulation. Similarly, the exploitation of minors through manipulative social media algorithms designed for engagement highlights a specific social vulnerability being targeted by technological hazards.
Reducing vulnerability focuses on strengthening the inherent resilience, capacity, and defense mechanisms of the individuals or systems that might be affected by the hazard. This approach aims to ensure that even if a hazard occurs and exposure exists, the resulting impact will be minimized.
- Natural Risk: Strategies involve bolstering physical resilience, such as mandating the construction of earthquake-resistant shelters and creating robust, widely practiced evacuation plans and community support systems.
- Technological (AI) Risk: In the digital context, this involves increasing the capacity of the end-user. Key interventions include promoting universal digital literacy education to help individuals identify and avoid manipulative or deceptive AI interactions; deployment of robust defensive measures like antivirus and anti-malware software; and enacting extra legal protections for marginalized or susceptible groups to grant them enhanced recourse or protection when interacting with automated systems.
The Necessity of a Multi-Dimensional Approach to AI Risk
While traditional risk management frameworks often treat risk as a monolithic entity—a single likelihood-consequence product—the application of these methods to the complex and heterogeneous domain of Artificial Intelligence necessitates a multi-dimensional analysis. The utility of the Hazard/Exposure/Vulnerability (HEV) framework is affirmed by the observation that the riskiness of different AI applications stems not from a single source, but predominantly from one of these three distinct dimensions. This distinction is crucial for regulatory bodies, as effective mitigation strategies must target the specific dimension that drives the risk.
Hazard-Critical AI Systems
These AI systems operate in critical sectors where a decision delegation to the machine has profound, direct, and often irreversible impacts on human life or fundamental rights. When society delegates decisions with high stakes—such as life, liberty, or physical integrity—the intrinsic hazard level of the technology is automatically classified as high.
Example
Medical AI: Diagnostic tools, such as those employing machine learning to detect malignancies from radiological imagery. A false negative error in this system constitutes a direct, high-level hazard, as it can result in missed diagnosis and ultimately, patient mortality (Panayies et al. 2020).
Judicial/Legal AI: Algorithms utilized in complex decision-making processes, such as determining sentencing recommendations or bail risk assessments. An algorithmic error or bias here directly translates into the severe legal harm of wrongful imprisonment or disproportionate legal consequences (Queudot & Meurs 2018).
Military AI: Systems deployed for autonomous weapons or strategic command analysis in high-stakes combat scenarios. The hazard here is clear: an algorithmic malfunction could lead to unintended escalation, civilian casualties, or breaches of international humanitarian law (Amoroso & Tamburrini 2020).
Exposure-Critical AI Systems
In contrast to hazard-critical systems, some AI applications possess a relatively low immediate hazard—a single instance of failure may not be individually catastrophic—but their risk profile becomes elevated due to the sheer scale of their deployment. This is the domain governed by the Exposure dimension.
- The “Recommender System” Paradox: Historically, systems like recommender algorithms (which drive content suggestion on major platforms) were initially overlooked in high-risk categories. Their function—suggesting the next video, article, or product—seemed trivial.
- Regulatory and Epistemological Evolution: This perception has demonstrably shifted, as evidenced by amendments in the EU AI Act (e.g., in June 2023) that recognized certain recommender systems as being critical due to their potential for “systemic risk.”
The risk is not derived from the depth of harm in a single instance, but from the cumulative, systemic effect across a massive user base. The risk materializes through “very large scale influence.” The core identified risks are therefore systemic: pervasive privacy erosion from mass data consumption; the cultivation of behavioral addiction and passive consumption; and the potential for large-scale political manipulation and propaganda through algorithmic amplification and filter bubbles. The risk is thus a societal one, driven by the breadth of the system’s deployment.
Vulnerability-Critical AI Systems
The third dimension, Vulnerability, focuses less on the AI system itself and more on the specific characteristics of the user base with whom the system interacts. Risk in this context is heightened when the user lacks the necessary cognitive, emotional, or social resilience to withstand the system’s mechanisms. This is particularly relevant in areas like Affective Computing and Social Robotics, which are explicitly designed for emotionally or socially salient interaction. Examples include carebots deployed for elderly individuals and AI tutors utilized in educational settings.
The inherent vulnerability of the target demographics—such as children, individuals with cognitive decline, or the elderly—is exploited when a machine mimics empathy, companionship, or authority. The risk is not physical but psychological and emotional, stemming from the potential for algorithmic deception, creation of unhealthy dependence, or manipulative social engagement by an artifact that does not genuinely possess the traits it simulates.
Consequently, an AI system that is inherently Low Hazard (e.g., a robot with no capacity for physical harm) can still be classified as “Highly Risky” because the user base is psychologically susceptible to its influence, proving the multi-dimensional framework essential for comprehensive risk assessment.
The Integrated Risk Model
Sole reliance on the Hazard component to define AI risk creates significant regulatory gaps, potentially leading to the under-regulation of systems whose danger is driven by scale or user context. A multi-dimensional analysis provides a far more robust and nuanced lens for policy-making, particularly in the context of the EU AI Act. By identifying the dominant risk vector, policies can be meticulously tailored to address the specific source of danger—be it the intrinsic code, the scale of distribution, or the susceptibility of the user.
The effectiveness of mitigation policies is maximized when they precisely correspond to the specific dimension generating the risk:
| System Type | Dominant Risk Dimension | Mechanism of Risk | Policy Implication and Mitigation Strategy |
|---|---|---|---|
| Deep Fakes | Exposure | The hazard (the synthetic media) is dangerous primarily due to its potential for viral spread and mass deception across a large population. | Mitigation focuses on limiting large-scale distribution (exposure) or mandates transparency obligations (e.g., clear labeling of synthetic content), rather than an outright prohibition of the underlying generation technology. |
| Emotion Recognition | Vulnerability | The system’s primary danger lies in its application to vulnerable populations (e.g., fatigued truck drivers being monitored, students subjected to automated emotional assessments). | Policy must focus on protecting these specific subjects by restricting the use of such systems in sensitive environments or by requiring heightened oversight and user recourse mechanisms. |
| Medical/Military AI | Hazard | The danger is intrinsic to the mission-critical task itself, where even a single error has catastrophic consequences (e.g., a fatal diagnostic mistake or an unintended military strike). | Mitigation demands rigorous ex-ante testing, validation, and certification of accuracy and reliability, ensuring the core technology meets stringent safety standards before deployment. |
By moving away from a “one-size-fits-all” approach, policymakers can effectively target the precise origin of the danger: modifying the code (Hazard), restricting the deployment footprint (Exposure), or strengthening the protective mechanisms around the users (Vulnerability).
The Epistemological Distinctions of AI
To fully grasp why AI necessitates this specific epistemological analysis and a dedicated regulatory framework like the EU AI Act, one must define and analyze what fundamentally distinguishes AI from previous technological paradigms.
The field of AI can be broadly characterized by two principal types of definitions: the behaviorist and the operational.
- The Behaviorist Definition (Nilsson, 1998): This classic approach focuses on the capabilities of the artifact, defining AI as being “concerned with intelligent behavior in artifacts.” This intelligent behavior is further characterized by a suite of complex functions, including perception, reasoning, learning, communication, and acting effectively within complex and dynamic environments.
Cite
“AI, broadly (and somewhat circularly) defined, is concerned with intelligent behavior in artifacts. Intelligent behavior involves perception, reasoning, learning, communication and acting in complex environments.”
Nils Nilsson 1998
- The Operational Definition (OECD, 2023): This modern definition, often adopted by regulatory bodies, emphasizes the function and autonomy of the system. An AI system is defined as a machine-based system capable of influencing the environment by producing specific outputs (such as predictions, recommendations, or decisions) to fulfill a given set of objectives. Crucially, these systems are designed to operate with varying levels of autonomy, signifying a departure from pre-programmed deterministic software.
Cite
“An AI system is a machine-based system that is capable of influencing the environment by producing an output (predictions, recommendations or decisions) for a given set of objectives. It uses machine and/or human-based data and inputs to (i) perceive real and/or virtual environments; (ii) abstract these perceptions into models through analysis in an automated manner (e.g., with machine learning), or manually; and (iii) use model inference to formulate options for outcomes. AI systems are designed to operate with varying levels of autonomy.”
OECD AI Principles overview 2023
A critical epistemological distinction is that AI systems maintain their status as experimental technologies even after their initial market deployment. Unlike mature technologies such as a traditional bridge or a simple appliance like a toaster, which possess predictable and well-understood failure modes, AI operates within a persistent state of uncertainty and novelty.
Cite
“Risks and benefits of experimental technologies may not only be hard to estimate and quantify, sometimes they are unknown.”
van de Poel 2016
As van de Poel (2016) highlights, the risks and benefits of experimental technologies are not only challenging to estimate and quantify, but are sometimes entirely unknown prior to wide-scale interaction with the real world. This inherent uncertainty complicates traditional risk assessment models which rely on historical data and probabilistic predictability.
The unique and challenging nature of AI risk stems from 5 core technical and operational features that fundamentally differentiate it from conventional software and engineering:
- Complexity: AI systems are engineered to achieve sophisticated goals, often requiring them to adapt and learn from new, often unbounded, data inputs and environmental changes, leading to intricate and sometimes unforeseen interactions.
- Machine Execution: The tasks are performed solely by non-human agents, removing the instantaneous human intuition or judgment that might intercept errors in conventional systems.
- Autonomy: AI systems possess the ability to operate without constant or direct human intervention, making real-time oversight challenging and expanding the scope for unmanaged errors or unforeseen actions.
- Prediction: Unlike deterministic software that yields binary or fixed results, AI systems often produce probabilistic outputs (predictions). This reliance on statistical inference rather than hard logic introduces an inherent, non-zero risk of error that is difficult to eliminate.
- Opacity: The use of advanced techniques like Deep Learning often results in “Black Box” issues, where the internal reasoning, weights, and decision-making processes of the system are opaque and largely incomprehensible to human operators or auditors. This lack of interpretability makes the a priori prediction and a posteriori investigation of failures profoundly difficult, hindering conventional risk analysis.
General Purpose AI and the Implementation Gap
Despite the theoretical elegance and utility of the multi-dimensional risk analysis framework (Hazard, Exposure, and Vulnerability), its practical application to modern Artificial Intelligence systems encounters substantial limitations. This implementation gap arises directly from the peculiar technical characteristics of AI, namely its opacity and high degree of autonomy, which fundamentally challenge the applicability of standard risk analysis and assessment tools. The inherent complexity and unpredictability of these systems often render the precise quantification and enumeration of risks difficult, if not impossible in a traditional engineering sense.
The most formidable challenge currently disrupting established AI risk methodologies is the rapid emergence and deployment of General Purpose AI Systems (GPAIS), such as advanced Large Language Models (LLMs). These systems are defined by their remarkable and innate ability to be flexibly utilized, adapted, and fine-tuned for an exceptionally wide variety of disparate purposes. The lack of a predefined, narrow scope of use in GPAIS fundamentally breaks the traditional, linear timeline of risk assessment relied upon in conventional engineering.
The Ex-Ante Problem: Risk Assessment Before Deployment
In traditional product development (e.g., structural engineering), the assessment of risk is predominantly Ex-Ante—performed before the product’s release—because the context of use is known and bounded (e.g., a bridge is designed specifically for vehicular traffic across a defined geographical span).
- The Context Vacuum: With GPAIS, a critical issue arises from the difficulty in identifying and enumerating specific risks before the system is inserted into a specific use context. The inherent adaptability of a single Large Language Model means it can be immediately appropriated for wildly divergent high-stakes tasks, ranging from medical diagnosis and scientific research to complex financial advice or computer programming.
- Consequence for Hazard Enumeration: Because the intended “context of use” is indeterminate or virtually unbounded at the design and development phase, it becomes impossible for the initial developer to fully enumerate all the specific hazards that may emerge when the technology is later applied in an unforeseen critical domain. This prevents a complete Ex-Ante Hazard analysis.
The Ex-Post Problem: Difficulties in Incident Tracing and Evaluation
Even after a GPAIS has been deployed and a harmful incident or failure has occurred, the evaluation and tracing of the root cause—the Ex-Post analysis—remains significantly complicated by the system’s inherent design features.
- Opacity and Traceability: The “black box” nature inherent to modern deep learning techniques makes it exceptionally difficult to retrospectively trace the root cause of an error. Unlike traditional software where a bug can be isolated to a line of code, the specific weights, parameters, and training data instances that led to a flawed prediction or decision remain frustratingly opaque.
- Lack of Informational Transparency: The auditing process is further hindered by a frequent lack of transparency from developers regarding the model’s essential inputs, including the composition of the proprietary training data and the specifics of the complex model architecture. This essential information is often unavailable to external auditors or regulatory bodies investigating a harmful incident.
- Multi-Hazard Scenarios and Prediction Complexity: GPAIS introduce a massive confluence of potential hazards. The risks stem not only from internal “errors” (unforeseen system malfunctions) but also from “possible abuses” (malicious or unintended user intent, such as prompting the model for harmful instructions). The resulting interaction creates a vast, open-ended number of potential hazards, making their prediction and mitigation a combinatorial challenge that far exceeds traditional risk management capabilities.
Redefining the Role of Risk Analysis
Given the demonstrated infeasibility of predicting every specific hazard associated with a General Purpose AI System (GPAIS), the fundamental objective of technological risk analysis must undergo a paradigm shift. The goal is no longer the unattainable ideal of eliminating all hazards; rather, it transitions to the practical necessity of identifying broad areas of concern and implementing effective measures to limit the scope of use and mitigate the system’s potential for harm.
Mitigation Strategies for Unspecified Hazards
Since the specific mode of error cannot be accurately predicted for complex, opaque, and adaptable GPAIS, mitigation strategies must pivot to interventions on the other two dimensions of risk: Exposure and Vulnerability.
-
Limiting Contexts (Exposure Reduction): Instead of making the technically impossible attempt to perfect an LLM for every high-stakes application, regulatory intervention can be implemented to limit the context of use. For example, due to the potentially fatal consequences of misdiagnosis, regulators may restrict or outright prohibit the deployment of general-purpose LLMs in sensitive medical diagnostic contexts. This shifts the focus from fixing the hazard to controlling the system’s opportunity to cause high-consequence harm.
-
Technical Guardrails and Accountability: Mitigation efforts must integrate legislative requirements with technological design. This includes mandating the implementation of reliable detection tools, such as watermarking and digital fingerprinting for AI-generated content. These technical guardrails are crucial for preventing the large-scale, untraceable spread of deep fakes and disinformation by ensuring a chain of accountability and the ability to verify content authenticity, thereby reducing exposure to manipulation.
-
Incremental Introduction and Continuous Monitoring: Rather than an abrupt, full market release, AI systems should be introduced through a strategy of incremental deployment. This approach allows developers and society to manage and limit exposure levels to a controlled subset of the population or environment while simultaneously implementing a regime of continuous monitoring to detect and respond to emerging, unforeseen risks and hazards as they manifest in real-world conditions.
-
Focus on Vulnerability: A non-negotiable strategy involves mandating special attention and heightened protections for demonstrably vulnerable subjects (such as children and the elderly) and in vulnerable social contexts (like healthcare, educational assessment, and access to social services). This means building in extra layers of human oversight, robust recourse mechanisms, and explicit warnings whenever AI interacts with these susceptible populations.
From Risk to Severe Uncertainty
The overarching epistemological analysis of AI risk concludes by revealing a crucial conceptual shift in how safety must be approached, moving away from a traditional, quantitative engineering model.
The core findings of this analysis are summarized as follows:
-
Applicability of Frameworks: The multi-dimensional analysis (Hazard, Exposure, and Vulnerability) retains its theoretical utility and should be applied to AI, even in the face of experimental and general-purpose systems, as it isolates the necessary mitigation levers.
-
New Mitigation Paradigm: The focus of safety analysis must shift beyond simple ex-ante (probabilistic prediction) and ex-post (retrospective audit) to a dynamic approach centered on identifying and regulating broad areas of concern and potential harm vectors.
-
The Probability Problem: The central limitation for GPAIS is the inability to perform a reliable probabilistic assessment of risk. Because the full spectrum of potential hazards is not known, calculating the likelihood of a specific event—the
term in the classical definition—becomes difficult or impossible. -
The Domain Shift: This difficulty forces a fundamental conceptual transition in the governance of AI. We are moving away from the well-defined domain of “Risk” (where outcomes are known and probabilities can be calculated or estimated) and into the domain of “Severe Uncertainty” (where both the potential outcomes/hazards and their associated probabilities are either unknown or unquantifiable).
AI and Severe Uncertainty
The final stage of the epistemological analysis of AI risk necessitates acknowledging that the domain of AI safety often transcends the traditional realm of measurable risk and enters the more challenging domain of severe uncertainty. The peculiar nature of autonomous, general-purpose AI systems (GPAIS) renders the standard toolkit of probabilistic risk assessment fundamentally inadequate.
The Tuxedo Fallacy
A significant methodological error prevalent in AI discourse is the mistaken assumption that all potential outcomes are known and that the probability of these outcomes can be precisely calculated. This error is formalized as the “Tuxedo Fallacy,” a concept articulated by Sven Ove Hansson (2009).
The fallacy is illuminated by a powerful metaphor that contrasts two environments requiring decision-making:
-
The Casino (The Tuxedo): This represents an environment of calculable risk. When engaging in a game like roulette, the financial stakes are high, but the probabilities are perfectly known (e.g., the chance of landing on red is
). The system is closed, the rules are fixed, and all possible outcomes are enumerated. Decision theory is highly effective here. -
The Jungle: This represents an environment of severe uncertainty. An expedition into an unknown jungle presents dangers that are not only unquantifiable but may be entirely unknown (e.g., we don’t know the probability of encountering a specific undiscovered pathogen, a hidden sinkhole, or a hostile tribe). The system is open, dynamic, and its rules are undefined.
The “Tuxedo Fallacy” is the epistemological mistake of treating decisions made under conditions of real-world uncertainty (The Jungle) as if they were taking place under the controlled, predictable epistemic conditions of a game of chance (The Casino).
Cite
“Life is more like an expedition into an unknown jungle than a visit to the casino. Most of the time we have to deal with dangers without knowing their probabilities, and often we do not even know what dangers we have ahead of us. […] I propose to call this the tuxedo fallacy. It consists in treating all decisions as if they took place under epistemic conditions analogous to gambling at the roulette table.”
— Hansson, 2009
Applying this to AI, the error is made when analysts attempt to assign precise, quantitative probabilities to AI risks (e.g., stating a “5% chance of the Large Language Model hallucinating” in a novel context). Given the opacity, autonomy, and adaptability of GPAIS, AI operates in an open-world “jungle” where the precise probabilities and the very nature of future emergent hazards are often fundamentally unknowable.
The Nature of Uncertainty
To properly construct a regulatory framework, one must move past the simplistic definition of “risk” and recognize the qualitative difference presented by uncertainty.
Drawing on the seminal work of economist Frank Knight (1921), the lecture emphasizes the concept of Severe Uncertainty (often termed Knightian Uncertainty). Knight clearly distinguished between Risk (measurable uncertainty, where outcomes are known and probabilities can be calculated, often insurable) and Uncertainty (unmeasurable uncertainty, where probabilities cannot be calculated a priori or empirically).
Severe Uncertainty is not simply a deficit of data; it represents a qualitative difference in our lack of knowledge that resists probabilistic quantification. Hansson (2022) further notes that uncertainty is a multifaceted phenomenon, stressing that there is no single, unified formal model capable of representing all types of epistemic uncertainty (e.g., factual, structural, agential, or possibilistic). This suggests that a purely mathematical solution to the problem of AI safety is inherently insufficient.
General Conclusions and Policy Recommendations
The lecture concludes by synthesizing the epistemological argument, confirming the need to fundamentally overhaul the approach to AI safety.
While much of the contemporary debate focuses intensely on specific risks (e.g., the symptoms, like algorithmic bias, privacy violations, or existential threats), a general epistemological framework capable of addressing the core nature of AI’s unpredictable risk generation has been notably absent. Effective policymaking requires shifting the focus from the symptoms to the underlying epistemological condition.
-
Multi-Dimensional Utility: The analytical decomposition of risk into Hazard, Exposure, and Vulnerability remains a fruitful and necessary method. It provides the distinct levers needed to intervene, focusing not only on fixing the intrinsic code (Hazard) but critically on controlling the scale of deployment (Exposure) and protecting the affected people (Vulnerability).
-
Intrinsic Uncertainty: The generation and operation of AI systems inherently occur in contexts of severe uncertainty. This is not a temporary technical flaw but an intrinsic feature of autonomous, complex, and experimental technologies.
-
Beyond the Illusion of Control: Policymakers and technologists must abandon the illusion that AI outcomes can be fully predicted, managed, or controlled through traditional, probabilistic risk assessment techniques, thereby explicitly avoiding the Tuxedo Fallacy.
Recommendations for Policy and Ethics
To cope with the reality of severe uncertainty, ethical and regulatory frameworks must systematically adopt a control-based approach rather than relying solely on a probability-based one:
-
Gradual Introduction: Systems should be released incrementally with controlled exposure, rather than full-scale deployment, enabling real-time observation of emergent hazards that could not be predicted ex-ante.
-
Continuous Monitoring: Regulation must move away from a fixed, “one-off” Ex-Ante approval process toward a mandatory model of continuous assessment and audit, adapting to the system’s performance and observed harms over its entire lifecycle.
-
Adaptive Regulation: Legal frameworks must be designed with the necessary flexibility and adaptivity to swiftly address novel dangers that were entirely unknown at the time of the system’s initial design, thereby embedding a permanent mechanism for managing Knightian Uncertainty.