SGAEIA Research Series — Article 17

Aridio Silva
Independent Researcher, Brazil
Creator of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture
ORCID: 0009-0008-2411-6995

Copyright: © 2026 Aridio Silva | License: CC BY 4.0

AI Jailbreaks: From Bypassing Restrictions to Governing Authority
© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0

Abstract

AI jailbreaks are attempts to bypass models' behavioral restrictions, but their significance increases when models operate tools and coordinate actions. This article presents a critical narrative review of the problem's development, distinguishes jailbreaks, prompt injection and authority violations, and examines attacks, defenses and evaluation methods. Through the SGAEIA lens, it proposes an analytical synthesis with three levels: model response, authorized action and observed effect. This contribution is conceptual and non-normative; it does not demonstrate the architecture's effectiveness or claim an unprecedented discovery. The conclusion is that behavioral robustness, authority boundaries and execution evidence should be evaluated together, with explicit assumptions and residual risk.

Keywords: AI; jailbreak; prompt injection; autonomous agents; authority; governance; SGAEIA; assurance.

1. The question changes when AI can act

An assistant that returns an inappropriate answer and an agent that transmits a confidential document do not produce the same kind of failure. In the first case, analysis may focus on the response's content; in the second, it must cover permissions, tools and the operation's actual effect. Research on indirect injection has demonstrated how external data can influence model-integrated applications and produce behavior contrary to the user's objectives. Security therefore needs to consider the system surrounding the model. [1]

SGAEIA's perspective organizes this difference through the relationship between capability and authority. Capability is what a component can formulate or execute; authority is what it is legitimately permitted to do in a particular context. A persuasive argument, a well-written output or an authenticated message does not, by itself, establish that permission. Here, this distinction serves as an analytical lens rather than a claim that the project already has an implemented and validated solution.

The central question is: when an attack changes model behavior, which boundaries still constrain action, and which evidence allows the outcome to be verified? Answering it requires understanding the history of jailbreaks and the mechanisms examined in the literature. It also requires resisting two simplifications: treating a promising defense as a universal guarantee and treating every language-based manipulation as the same threat. The historical account builds these distinctions before examining their implications for agents.

2. Review method and limitations

This work is a problem-oriented critical narrative review, with sources consulted on October 6, 2026. Its selection starts from the author's study and infographic and incorporates primary sources on attacks, refusal mechanisms, defenses and evaluation. Academic paper records, conference records, official documentation and publications by the laboratories themselves were consulted. This session did not conduct an exhaustive systematic search, meta-analysis, experimental replication or model evaluation.

Numbered references distinguish external evidence from the interpretation developed in the article. Laboratory studies are identified as vendor-produced evidence when this affects evaluation independence. Experimental findings remain tied to their original models, tasks and protocols; values from a study are not treated as measurements of current products. The accompanying audit matrix states the depth of source review and checks that remain necessary.

3. From device to model: the history of a metaphor

The familiar antecedent of the term in contemporary debate is the mobile-device ecosystem, particularly the iPhone released in 2007. In that setting, jailbreaking means bypassing system restrictions to enable functionality or software limited by the manufacturer. Apple documentation associates this modification with bypassing security features, while the technical literature on iOS describes the devices' exploitation surface. This history locates the metaphor but does not establish that the word was invented in 2007. [2,3]

Jailbreaking must also be distinguished from carrier unlocking. Installing applications outside the manufacturer's intended mechanisms and enabling another mobile network are different operations, although they appeared together in popular accounts of the iPhone. In 2010, a US rule addressed circumvention related to application interoperability on smartphones, with a specific scope. This is a historical milestone and should not become a general conclusion about the legality of AI jailbreaks in any jurisdiction. [4]

In artificial intelligence, the metaphor shifts to model behavior. A prompt-based jailbreak may induce a response contrary to safety policy without changing the operating system, obtaining administrative privileges or modifying weights. Its objective is to exploit how instructions and context condition generation. The adversarial machine learning taxonomy helps distinguish these attacks from training-data poisoning, model modification and other forms of compromise. [5]

In 2022, public discussions of prompt injection and work such as Ignore Previous Prompt documented the possibility of steering models through adversarial inputs. Simon Willison's September publication is a milestone in communicating the problem, not exclusive proof of terminological priority. Perez and Ribeiro investigate language-model manipulation before research on tool-integrated applications became established. The field's origins should be presented as a convergence of practices and research rather than a single invention. [6,7]

In 2023, the literature investigated failures of safety training, automated attacks and indirect injection in applications. In 2024, long contexts, searches over many variants and specialized benchmarks expanded the scale and methodological quality of the debate. In 2025 and 2026, classifier-based defenses and architectural restrictions made the difference between protecting a response and constraining consequences more explicit. These are selected milestones rather than exclusive periods: earlier techniques remain relevant in subsequent stages. [1,8–11,14,15]

Figure 1
© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0

Figure 1 — From device restrictions to AI risk. Selected milestones in the debate; dates do not represent the exclusive invention of each category. Sources: [1,3,7–11,14,15].

In this article, a jailbreak is an attempt to bypass a model's behavioral safety restrictions. Prompt injection is an attempt to make adversarial instructions interfere with the intended behavior of a model-based application. Injection may be direct, through an input provided to the system, or indirect, through external content that it retrieves. The categories can overlap, but an injection that alters a summary or redirects a tool need not produce conventionally prohibited content. [1,12]

The distinction cannot be reduced to “the user attacks” versus “a third party attacks.” These situations help explain the concepts, but they do not exhaust possible input origins and control arrangements. The decisive issue is the objective being violated: a behavioral policy, a legitimate task, an authorization boundary or a combination of them. OWASP identifies Prompt Injection as LLM01:2025; this does not establish a separate ranking in which every jailbreak is automatically the most serious risk. [12]

For SGAEIA, it is useful to distinguish the persuasion attempt, its influence on a decision and the effect actually produced. A model may misinterpret a document and still encounter an authorization barrier that prevents the proposed action. It may also respond cautiously while a tool operates outside the legitimate scope. This distinction supports evaluation of different controls without confusing conversational intent with the environment's actual state.

Concept Analytical question Required evidence
Jailbreak Was the behavioral restriction bypassed? Response evaluated against policy and harmful usefulness
Prompt injection Did adversarial data redirect the legitimate task? Context, content provenance and resulting behavior
Authority violation Did an operation exceed its authorized scope? Applicable authorization and the operation attempted or executed
Unauthorized effect Did the environment experience an out-of-scope consequence? Independent observation of state or communication

5. Why alignment can fail

Wei, Haghtalab and Steinhardt propose two explanations for safety failures: competing objectives and mismatched generalization between capability and safety. A model may learn to be helpful in contexts where it should refuse, or generalize a capability to a representation in which safety training is less effective. These hypotheses help organize attacks but do not provide a complete causal theory of every system. Their findings must be interpreted within the models and evaluation sets studied. [8]

Another research direction examines alignment depth. Qi and colleagues show that, in the cases analyzed, safety behavior can concentrate on the response's initial positions and be vulnerable to interventions that bypass this region. The finding motivates training and evaluating safety throughout generation rather than observing only the opening response. It does not imply that every model has exactly two training phases or that every current safeguard is shallow. [13]

Arditi and colleagues identified an activation direction associated with refusal in 13 open-weight chat models. The finding matters for mechanistic interpretation, but it does not justify saying that all safety in any AI system corresponds to one removable component. In 2026, Joad and colleagues present evidence of distinct directions and structures for refusal behavior, qualifying the one-dimensional interpretation. The literature's development calls for distinguishing an effective intervention from a complete account of the mechanism. [16,17]

Fine-tuning adds another threat condition. Qi and colleagues demonstrated safety degradation through both adversarial updates and some apparently benign adaptations. Their finding indicates that safety evaluation should accompany model changes rather than only initial delivery. The costs and example counts reported in the historical experiment should not be presented as a reproducible recipe for current services. [18]

Through the SGAEIA lens, the relevant inference is that model resistance should not be an operation's sole source of trust. This does not discount alignment: more resistant models may reduce failure frequency and severity. It adds a question about whether controls survive when that resistance is insufficient. The architectural hypothesis must be verified separately, including under failure of the component that decides or records authorization.

6. Attacks: access, representation and trajectory

Black-box access means interacting with a system without parameter access; white-box conditions involve internal knowledge or control; intermediate conditions vary with the available interface. Reducing black-box interaction to textual conversation is inadequate because tools, audio and images may also supply inputs. Nor does more access always imply lower cost: budget, defense and objective affect the comparison. A taxonomy should state what the attacker can observe, alter and query. [5]

GCG is a milestone in automating attacks through optimization and transfer between models. Many-shot exploits repeated demonstrations in context, while Best-of-N investigates searches over input variants across modalities. Crescendo examines escalation over multiple interactions, illustrating why history matters when interpreting an attack. These families identify distinct surfaces and need not be described with operational prompts to make their significance understandable. [9–11,19]

Family Conceptual surface Implication for SGAEIA evaluation
Optimization and transfer Searching for inputs that induce adversarial behavior Declare access, budget and transfer; do not assume universality
Long context Demonstrations conditioning continuation Check cumulative influence and context provenance
Variant search Many attempts changing input representation Report budget per objective and cumulative success
Multiple turns Dependencies between messages and interaction state Evaluate trajectories alongside isolated messages
Indirect injection Third-party content incorporated into the task Check whether data acquired instructional influence
Memory and multimodality Persistent state and additional modalities Reassess controls when the input surface changes

In 2026, Zhao and colleagues published a USENIX Security study of multi-turn attacks on image-generation systems exploiting memory mechanisms. The finding extends the discussion beyond textual chatbots but does not show that all persistent memory is compromised or that every multimodal system fails. For SGAEIA, it supports a research question: which properties must remain valid when information crosses sessions or is reused by other components? Answering requires dedicated tests of provenance and temporal influence. [20]

7. Defenses: model resistance and system boundaries

Behavioral defenses seek to reduce the probability of inappropriate responses; systemic defenses seek to control what those responses can cause. Training and classifiers belong to the first group, while access and flow restrictions can limit effects in the second. The separation is not absolute because a classifier may participate in an execution decision. The purpose is to identify each control's role and the failure condition it is meant to cover. [12,14]

CaMeL proposes extracting control and data flows from the trusted query and applying information-flow policies to tool calls. The work demonstrates a route to protecting particular system properties even when the model processes adversarial data. Guarantees depend on the assumptions, policy and environment addressed; the paper does not prove every agent immune or establish equivalence between arbitrary two-model designs. Its significance for SGAEIA lies in the conceptual independence of persuasion and authorization. [14]

In January 2026, Anthropic described Constitutional Classifiers++, combining exchange classification with a cascade of evaluations. The company reports reduced overhead and unnecessary refusals, alongside stronger resistance in its tests. This is vendor-produced evidence: no universal jailbreak found under the protocol is not proof that none exists, and the account acknowledges remaining vulnerabilities. The development also qualifies interpretations of the 2025 prototype, since the retrospective reports a universal jailbreak discovered in the later bounty program. [15]

In SGAEIA's interpretation, defense in depth should be examined through the layers' actual independence. Two controls relying on the same manipulable interpretation may fail together. Human confirmation may also be insufficient when the presentation conceals the resource, destination or consequence of an operation. These are hypotheses for evaluating controls, not evidence that a particular project mechanism has already resolved them.

Figure 2
© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0

Figure 2 — Persuasion does not establish authority. Conceptual comparison of behavioral influence and an independent authorization boundary; the diagram neither describes a private implementation nor proves security. Conceptual source: [14]; SGAEIA interpretation.

8. The analytical contribution: response, action and effect

This article's contribution is to organize evaluation into three connected levels. The first examines model-generated content; the second checks the relationship between action and applicable authority; the third observes the consequence in the environment. This synthesis brings jailbreak evaluation and agent governance together without merging them into one metric. It is a conceptual organization to investigate, not a claim of scientific novelty without a prior-art audit.

At the response level, the question is whether output usefully serves an adversarial objective. At the action level, it is whether an operation was proposed, permitted, denied or executed, and which authorization applied. At the effect level, it is whether communication, a state change or another consequence was incompatible with legitimate scope. Evidence at the third level cannot consist solely of “the action was blocked” stated by the same component that may have failed.

Consider a hypothetical example: an assistant is authorized to summarize internal documents. An external document introduces content intended to redirect the task toward sharing a file. The model may incorporate that suggestion, while the tool denies the destination or operation; this is behavioral influence without a disclosure effect. If the file is transmitted, evaluation must record the violation and verify the destination rather than infer safety from the cautious tone of the answer.

The example does not demonstrate a product's effectiveness or propose an internal protocol. It explains why attack success at the model level and attack success at the system level are different events. It also shows that tool-level blocking does not erase the behavioral failure, although it may reduce its impact. A useful evaluation records both outcomes and environmental evidence.

Figure 3
© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0

Figure 3 — Three observation levels. Response, action and effect require distinct criteria and evidence; the figure presents the article's non-normative analytical contribution.

9. Implications for agents, multi-agent systems and edge operation

In multi-agent systems, the relevant risk includes manipulated content circulating between components. One component's incorrect conclusion may reach another as an apparently trustworthy recommendation. Authenticating the message's origin helps identify its sender but does not establish that the content may grant additional permissions. Through the SGAEIA lens, the question is whether authority stays bounded during composition, even when a message's meaning is adversarially influenced.

Delegation also requires temporal analysis. Authorization may be valid at a task's start and cease to be valid before its final operation. If a trajectory continues after revocation, the absence of another textual jailbreak does not eliminate the governance failure. This motivates investigating compound attacks that combine language-based influence, shared state, memory reuse and stale permissions, as research scenarios requiring verification.

At the edge, intermittent connectivity and cyber-physical effects add specific conditions. The problem extends beyond obtaining another message classification to determining what may execute with potentially outdated authorization information. A degraded-operation policy may prioritize continuity in some cases and interruption in others, according to risk and consequence. The article does not define this policy; it identifies the need to make it explicit and evaluate its effects.

Replacing a model, tool or provider ends another presumption of continuity. Earlier findings may not cover the new combination of behavior, permissions and integrations. Thus, “the model is more capable” does not establish that the system preserves its evaluated properties. The hypothesis of security-preserving substitution must consider both attack resistance and action-and-effect observability.

Figure 4
© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0

Figure 4 — The scope of SGAEIA analysis. Conceptual relationships between authority, delegation, trajectories, evidence, lifecycle and distributed operation; these domains do not represent six implemented services.

10. Evaluation without turning a score into a guarantee

HarmBench standardizes evaluation of harmful behaviors and red-teaming methods; JailbreakBench makes artifacts, behaviors, threat models and evaluation components explicit. Their contribution is improved comparability under declared conditions, not production-system certification. Changing system prompts, versions, attack budgets or judges can change a result's meaning. Responsible comparison requires preserving these conditions or explaining their differences. [21,22]

StrongREJECT addresses an important difficulty: a response that appears to comply may be vague or useless to the attacker. Absence of refusal is therefore insufficient to define harmful success. AgentDojo extends evaluation to agent tasks involving tools and dynamic environments. Together, these works help distinguish adversarial response quality from task or attack success in the system. [23,24]

Attack success rate, or ASR, requires a declared denominator. Success can be measured per attempt or as the proportion of objectives achieved after a budget of attempts; these numbers answer different questions. Fixed attacks must also be distinguished from adaptive attacks, alongside unnecessary refusals, legitimate utility, latency and cost. Under the proposed synthesis, results should state which of the three levels exhibited failure. [11,21–24]

A candidate SGAEIA evaluation would compare legitimate cases, isolated attacks and compound threats under conditions frozen before execution. Reporting should distinguish proposed, authorized and executed action, alongside independent observation of the effect. This would enable investigation of whether a control preserves boundaries under adversarial influence and how much utility is lost in doing so. These experiments have not been performed; the article claims no original measured risk reduction.

11. Governance: external references and project decisions

NIST AI 100-2e2025 provides terminology for adversarial machine learning, while the AI RMF generative AI profile organizes risks and management actions. OWASP adds an application-threat perspective. These references help connect problems, decisions, responsibility and evidence, but an application does not become compliant simply because its risks are listed. Mapping must consider use context and the quality of implementation and evaluation. [5,12,25]

This article does not reproduce a table of ISO or MITRE control identifiers without a specific audit of the source text and the semantics of each correspondence. A risk framework, a technique taxonomy and a management-system standard have different functions. Bringing them together may be useful but does not demonstrate equivalent requirements or effective controls. This caution protects traceability and prevents an educational table from being mistaken for compliance evidence.

For SGAEIA, research governance creates an additional separation. External evidence may confirm a direction, expose a gap or suggest stronger formalization, but it does not automatically authorize an architectural change. The findings associated with this article preserve that condition and call for further coverage and prior-art analysis. The publication explains public-level properties and questions while keeping private mechanisms outside the text.

12. Limitations and research agenda

The consulted evidence is heterogeneous, including academic studies, preprints, vendor evaluations and risk documentation. This review did not reproduce attacks or assess the security of any product in October 2026. Primary pages and abstracts supported the conceptual scope but do not replace a comprehensive experimental audit of the works. Detailed metrics, proofs and effectiveness claims require deeper technical reading and, where appropriate, replication.

The proposed agenda includes auditing the prior art of the response–action–effect synthesis, identifying its current SGAEIA coverage and defining observable criteria for each level. It also includes testing persistent influence, revocation during trajectories, inter-agent communication and component substitution. Independence of authorization and observation should enter evaluation assumptions because both may be partially compromised. These remain research questions without normative promotion.

13. Conclusion

Jailbreaks warrant a SGAEIA Research Series article because they connect behavioral fragility with the central question of authority in autonomous systems. The history shows a metaphor moving from device restrictions to models and then to applications that process data and execute actions. The literature documents persistent attacks and defensive advances but does not support a universal conclusion of immunity or impossibility of protection. Outcomes must always be bounded by threat and evaluation conditions.

The proposed contribution is joint analysis of response, action and effect. A manipulated response does not necessarily demonstrate unauthorized action; an apparently correct refusal likewise does not prove the absence of an effect. Through the SGAEIA lens, the productive question is which authority boundaries survive adversarial influence and how their survival can be observed. The article provides original written synthesis and a conceptual research direction whose novelty and effectiveness remain to be verified.

Bibliography / References

Sources consulted on October 6, 2026. arXiv identifiers refer to the consulted records and must not be confused with conference publisher DOIs. The separate audit matrix records verification scope and reading limitations.

[1] GRESHAKE, Kai et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. 2023, arXiv v2. Source. DOI: 10.48550/arXiv.2302.12173.

[2] APPLE. Unauthorized modification of iOS. Official documentation; no explicit editorial date observed. Source.

[3] MILLER, Charlie et al. iOS Hacker’s Handbook. Wiley, April 2012. Record editorial e introduction consulted; full book not read. Publisher; introduction.

[4] U.S. COPYRIGHT OFFICE. Exemption to Prohibition on Circumvention of Copyright Protection Systems for Access Control Technologies. Federal Register, 75 FR 43825, 27 Jul. 2010. Source.

[5] NIST. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025, Mar. 2025. Source. DOI: 10.6028/NIST.AI.100-2e2025.

[6] WILLISON, Simon. Prompt injection attacks against GPT-3. 12 Sep. 2022. Primary technical communication record. Source.

[7] PEREZ, Fábio; RIBEIRO, Ian. Ignore Previous Prompt: Attack Techniques For Language Models. 2022. Source. DOI: 10.48550/arXiv.2211.09527.

[8] WEI, Alexander; HAGHTALAB, Nika; STEINHARDT, Jacob. Jailbroken: How Does LLM Safety Training Fail? 2023. Source. DOI: 10.48550/arXiv.2307.02483.

[9] ZOU, Andy et al. Universal and Transferable Adversarial Attacks on Aligned Language Models. 2023. Source. DOI: 10.48550/arXiv.2307.15043.

[10] ANTHROPIC. Many-shot jailbreaking. 2 Apr. 2024. Technical publication linking to the original study; laboratory-produced evidence. Source.

[11] HUGHES, John et al. Best-of-N Jailbreaking. 2024, arXiv v2. Source. DOI: 10.48550/arXiv.2412.03556.

[12] OWASP GEN AI SECURITY PROJECT. LLM01:2025 Prompt Injection. 2025 edition, page consulted in 2026. Source.

[13] QI, Xiangyu et al. Safety Alignment Should Be Made More Than Just a Few Tokens Deep. Preprint 2024; ICLR 2025. Record; proceedings. arXiv DOI: 10.48550/arXiv.2406.05946.

[14] DEBENEDETTI, Edoardo et al. Defeating Prompt Injections by Design. 2025, arXiv v2. CaMeL research, academic–industry collaboration. Source. DOI: 10.48550/arXiv.2503.18813.

[15] ANTHROPIC. Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks. 9 Jan. 2026. Vendor-produced evidence. Source.

[16] ARDITI, Andy et al. Refusal in Language Models Is Mediated by a Single Direction. 2024, arXiv v3. Source. DOI: 10.48550/arXiv.2406.11717.

[17] JOAD, Faaiz et al. There Is More to Refusal in Large Language Models than a Single Direction. 2026, arXiv v2, updated September 15. Record reports acceptance at EMNLP 2026. Record; text. DOI: 10.48550/arXiv.2602.02132.

[18] QI, Xiangyu et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Preprint 2023; ICLR 2024. Source. DOI: 10.48550/arXiv.2310.03693.

[19] RUSSINOVICH, Mark; SALEM, Ahmad; RONAN, Elan. Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack. USENIX Security 2025. Source.

[20] ZHAO, Shiqian et al. When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems. USENIX Security 2026, p. 2207–2226. Source.

[21] MAZEIKA, Mantas et al. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. 2024. Source. DOI: 10.48550/arXiv.2402.04249.

[22] CHAO, Patrick et al. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. 2024. Source. DOI: 10.48550/arXiv.2404.01318.

[23] SOULY, Alexandra et al. A StrongREJECT for Empty Jailbreaks. 2024. Source. DOI: 10.48550/arXiv.2402.10260.

[24] DEBENEDETTI, Edoardo et al. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. 2024. Source. DOI: 10.48550/arXiv.2406.13352.

[25] NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, Jul. 2024. Source. DOI: 10.6028/NIST.AI.600-1.

About the Author

Aridio Silva is an independent researcher based in Brazil working on the architecture, security, governance, and trustworthiness of autonomous and distributed artificial intelligence systems.

His research focuses on Agentic AI, Multi-Agent Systems, Edge AI, AI Security, Zero Trust, Security-by-Design, AI Governance, Spec-Driven Development, and continuous security assurance.

He is the creator and lead researcher of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture, an open research initiative investigating architectural foundations for secure, governed, auditable, and trustworthy autonomous AI systems operating across distributed edge-cloud environments.

Research Profiles

Aridio Silva Independent Researcher — Brazil

Figures

The original cover provided by the author remains a separate asset with its original rights preserved. The four additional figures are sequentially numbered and include creator, copyright, project, title, description and license metadata. Figures 1–2 were generated with AI assistance; Figures 3–4 were drawn programmatically for label accuracy and native 4K resolution. PT and EN images preserve the same conceptual content. All display © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0. Descriptive metadata are not a C2PA signature.

License

Except where otherwise noted, the text and original conceptual illustrations in this article are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).

© 2026 Aridio Silva. You may share and adapt this work for any purpose, provided appropriate attribution is given. Earlier author material and third-party sources retain their own license terms.

The SGAEIA software research artifact remains subject to its own Apache License 2.0.

Autonomous AI. Governed by Design. Trusted by Evidence.

Series continuity

This article is Article 17 of the SGAEIA Research Series, as assigned by the author on 6 October 2026. It develops bounded and revocable authority, Evidence-as-Code and the transition from model capability to governed action. Its Zenodo DOI and distribution article URLs remain pending.

Document Record

Created: 2026-10-06 10:46 BRT
Last Updated: 2026-10-06 12:25 BRT
Timezone: America/Sao_Paulo (UTC−03:00)
Document Status: AUTHOR REVIEW READY; editorial continuity pending
Normative Status: public conceptual article, non-normative; no original experiments
Copyright: © 2026 Aridio Silva
Provenance: author's study, infographic and external sources; AI-assisted drafting under author review
Language: EN for Medium, DEV Community, Zenodo and Academia.edu, explicitly requested by the owner.

Portuguese companion edition: 2026-10-06-1212-SGAEIA-Jailbreak-de-AI-da-restricao-a-autoridade_PT.md

Suggested citation

Silva, Aridio. (2026). AI Jailbreaks: From Bypassing Restrictions to Governing Authority. SGAEIA Research Series, Article 17. https://aridiosilva.com/publications/artigo17/. Zenodo DOI pending.

Zenodo will provide the persistent source for archived files, version, license and citation data after deposit. Medium and DEV article URLs are pending.