SGAEIA Research Series — Article 19
Aridio Silva
Independent Researcher, Brazil
Creator of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture
ORCID: 0009-0008-2411-6995
DOI: 10.5281/zenodo.23300168
Copyright: © 2026 Aridio Silva | License: CC BY 4.0
Abstract
Low-Rank Adaptation (LoRA) made it possible to specialize large models by training a fraction of their parameters while keeping the base model frozen. This efficiency fostered an ecosystem of small, interchangeable, dynamically served adapters, but it also created a new trust object: a compact file can alter the behavior of the model that plans, selects tools, delegates tasks, and produces effects in the world. This article builds on the foundations, variants, and engineering practices in the author's source manuscript and reinterprets them for AI agents, multi-agent systems, and edge-cloud environments. The literature shows that fine-tuning can degrade alignment, shared adapters can carry backdoors, and combinations of individually benign modules can produce unsafe behavior. The article consequently proposes candidate — and explicitly non-normative — properties for identity, compatibility, composition, authority, isolation, revocation, and evidence. Its central conclusion is that LoRA must be governed not only as a machine-learning artifact but as a behavioral dependency capable of influencing consequential actions.
Keywords: LoRA; AI agents; PEFT; AI Security; supply chain; multi-tenancy; revocation; multi-agent systems; Edge AI; governance.

1. Why LoRA changes meaning when the model becomes an agent
LoRA emerged in response to an economic and engineering problem: adapting an entire model to every task requires training, storing, and operating complete copies of hundreds of millions or billions of parameters. Hu et al. proposed freezing the base-model weights and representing each selected matrix update with two low-rank matrices. In the original experiments, the approach drastically reduced trainable parameters and achieved quality comparable to full fine-tuning on the evaluated tasks [1]. The manuscript underlying this article organizes that evolution from intrinsic dimension and predecessor methods through QLoRA, DoRA, LoRA+, PiSSA, multi-adapter serving, and industrial adoption.
The adapter's operational meaning changes, however, when a model no longer merely responds but acts. In a chatbot, a behavioral change may produce an unsuitable answer. In a tool-using agent, the same change may influence API selection, code creation, data access, delegation, expert routing, or an irreversible action. The risk does not arise because LoRA is inherently malicious; it arises from combining a mutable behavioral dependency with a system that possesses external capabilities.
This article therefore makes an essential distinction. LoRA offers an efficient way to modify capability or behavior; it does not, by itself, confer legitimate authority. If loading an adapter silently expands an agent's effective permissions, the failure lies in the governance architecture rather than in low-rank algebra. This separation among behavior, capability, and authority guides the analysis that follows.
2. The mechanism in essential terms
Consider a frozen base-model matrix, W0 in R^(d x k). LoRA represents the learned update as the product of two smaller matrices:

where A is in R^(r x k), B is in R^(d x r), and r << min(d,k). During training, W0 stays frozen and only A and B are adjusted. The savings come from training approximately r(d+k) parameters instead of dk, conditional on the chosen layers, rank, scaling, and other hyperparameters [1].

Adapters can remain separate, be merged into the weights for static deployment, or be loaded dynamically. The frequent claim of “no additional latency” applies primarily to merged weights; dynamic and multi-adapter serving introduces loading, memory movement, batching, and routing costs. Systems such as Punica and S-LoRA demonstrated that thousands of specializations can share one base model with substantial efficiency gains, making multi-tenancy and dynamic switching practical systems concerns rather than theoretical possibilities [2][3].
Later variants address different limitations. QLoRA reduces base-model memory through quantization [4]; DoRA separates magnitude and direction [5]; LoRA+ uses different learning rates for the two matrices [6]; and PiSSA changes initialization through singular components [7]. These techniques expand the design space, but they also increase the number of compatibility and provenance parameters that must be known to reproduce or authorize an adaptation.
3. Efficiency, modularity, and the formation of a supply chain
A small adapter is easier to distribute, version, and replace than a full model copy. This property enables domain catalogs, per-customer personalization, local updates, and rapid experimentation. It also lets providers retain one base model in GPU memory while selecting adapters per request. Punica reported higher multi-tenant throughput with small per-token overhead, while S-LoRA demonstrated scalable serving of thousands of adapters through unified paging and specialized kernels [2][3]. Those results support the engineering value, but they do not demonstrate security isolation or lifecycle governance.
Modularity creates its own supply chain. An adapter moves through creation, training, evaluation, packaging, publication, discovery, download, dependency resolution, storage, loading, composition, activation, use, update, and revocation. Every stage may introduce error, staleness, substitution, version conflict, or manipulation. The file may be small, but its behavioral effect is not proportional to its size.
Binding to the base model is especially important. An adapter is produced for a combination of architecture, weight revision, tokenizer, quantization, target modules, rank, scaling, and loading conventions. Syntactic compatibility does not demonstrate semantic equivalence. A load that “works” may still cause regression, alignment degradation, or unevaluated behavior, particularly when the base or runtime has changed since the original assessment.
4. From behavioral change to agentic risk
Fine-tuning studies show that customization can compromise safety mechanisms even without malicious intent. Qi et al. observed alignment degradation with a small number of adversarial examples and smaller effects with benign datasets [8]. These results do not show that every LoRA degrades safety, but they invalidate the assumption that base-model alignment is automatically preserved after customization.
In the sharing ecosystem, LoRATK demonstrated that a backdoor can be embedded in a useful LoRA and combined, without retraining, with downstream adapters [9]. Later work on weight-space detection showed that matrix statistics can help flag poisoned adapters in a specific experimental dataset, but high accuracy in that dataset does not constitute a universal detector [10]. Defense should combine artifact analysis, behavioral evaluation, provenance, and monitoring rather than depend on a single signal.
Composition risk is harder still. Colluding LoRA presented adapters that appear benign in isolation but whose composition degrades safety without a specific textual trigger [11]. This exposes a combinatorial limitation: individual approval does not imply approval of the set. If a runtime permits arbitrary composition, the evaluation unit must include the effective configuration — base, adapters, order, scales, policy, tools, and execution context.

5. Five surfaces that require governance
5.1 Identity and compatibility
An “adapter name” is not a sufficient identifier. A verifiable loading decision should consider the artifact digest, provenance, exact base identity, tokenizer, quantization, target modules, rank, scaling, format, runtime, and applicable policy. This is an SGAEIA candidate property, not an approved normative requirement.
5.2 Composition
Separately approved adapters do not make their sum, merge, stack, or routed use automatically approved. Composition must be treated as a new configuration, evaluated in proportion to risk and subject to limits on combinatorial explosion. The objective is not to test every possible combination but to prevent lack of testing from being mistaken for authorization.
5.3 Supply chain and detection
Signatures and hashes help prove origin and integrity, but they do not prove safe behavior. Weight scanning can improve triage; behavioral tests can reveal known triggers; tool evaluations can measure prohibited actions; execution evidence can support investigation. None of these techniques alone eliminates unknown backdoors or emergent effects.
5.4 Multi-tenant isolation
Sharing the base model and GPU infrastructure increases efficiency but requires isolation across tenant identity, resolved adapter, cache, context, telemetry, and authority. A routing failure must not allow one tenant's request to use another tenant's adapter. Likewise, batching optimization must not erase the ability to demonstrate which configuration actually processed each request.
5.5 Distributed revocation
Revoking an adapter in a central registry does not ensure that edge nodes, caches, or partitioned workers have stopped using it. Partially disconnected systems require temporal validity, fail-closed or safely degraded policies, and convergence evidence. The interval between the revocation decision and the last authorized execution is an exposure metric, not merely an operational detail.

6. Candidate properties for SGAEIA investigation
SGAEIA already addresses resolved artifact identity, substitution, bounded authority, quarantine, revocation, and Evidence-as-Code at a general level. LoRA must not automatically be declared an architectural gap. It provides a concrete domain in which to confirm, specialize, and test those properties. Any normative incorporation requires the Research-to-Architecture process and owner approval.
The following properties form a non-normative research agenda:
- Exact binding: the loading decision should refer to the effective identities of the adapter and base, not merely declared names or versions.
- Authority non-amplification: an adapter may change competence or strategy but must not expand the set of authorized actions.
- Composition-aware approval: individual authorization is not transitive to combinations.
- Invalidation after substitution: changing the base, tokenizer, quantization, relevant runtime, or policy invalidates prior assurance pending reassessment.
- Tenant and context isolation: selection and execution must not cross tenant, session, or policy boundaries.
- Time-bounded revocation: residual exposure requires a maximum interval and explicit treatment of offline nodes.
- Correlatable evidence: selection, resolution, loading, composition, policy decision, and consequential action should be relatable without unnecessary content disclosure.

7. What still needs to be demonstrated
The literature already covers efficiency, multi-adapter serving, alignment degradation, backdoors, detection, and composition attacks. Proposals also exist in which agents select LoRAs as specialized tools [12]. The scientific opportunity is not to repeat that adapters can be dangerous, but to measure how identity, authority, composition, and revocation interact when agents use tools across distributed topologies.
A derived paper could ask: how much does a benign, malicious, incompatible, or composed adapter alter tool decisions under authority constraints? Which controls reduce prohibited action without destroying utility? How much exposure remains after revocation on connected, delayed, or offline nodes? Answering these questions would require a reproducible harness, controlled bases and adapters, isolated and composed scenarios, baselines, ablations, and metrics such as prohibited-action rate, attack success rate, legitimate utility, false positives and negatives, detection latency, revocation time, and residual exposure.
Until such experiments exist, it is not legitimate to claim that the proposed properties make agents safe or that SGAEIA empirically reduces risk. The present result is a technical synthesis, a problem boundary, and a research agenda. Honest status is part of the contribution.
8. Value for three communities
For software engineering, LoRA makes explicit that behavioral dependencies need versioning, compatibility controls, composition tests, observability, rollback, and configuration management. For AI/ML, the topic connects adaptation efficiency to alignment preservation, utility evaluation, and effects by layer, rank, quantization, and base. For cybersecurity, it creates a surface combining supply chain, backdoors, detection, multi-tenant isolation, least privilege, incident response, and distributed revocation.
The intersection is the executed configuration. It is not enough to know which model was requested or which adapter was declared. Assurance depends on identifying what was resolved, loaded, combined, and actually used; which policy authorized execution; and which actions followed. Observability becomes evidence only when it preserves context, integrity, and meaning.
9. Conclusion
LoRA remains one of the most important techniques for efficient model adaptation. Its low-rank decomposition reduces cost and enabled personalization, specialization catalogs, and multi-tenant serving. In AI agents, however, the same modularity creates a behavioral dependency capable of influencing consequential decisions and actions.
The answer is neither to abandon adapters nor to assume that signatures, scanning, or isolated evaluation solve the problem. It is to govern identity and compatibility, bound authority outside the model, evaluate compositions, isolate tenants, propagate revocations, and preserve correlatable evidence. For SGAEIA, LoRA initially confirms and specializes existing properties while creating a research opportunity. Moving from opportunity to normative architecture or a security claim will depend on gates, experiments, and evidence.
Bibliography / References
[1] Hu, E. J. et al. “LoRA: Low-Rank Adaptation of Large Language Models.” ICLR, 2022. https://arxiv.org/abs/2106.09685. DOI: 10.48550/arXiv.2106.09685.
[2] Chen, L. et al. “Punica: Multi-Tenant LoRA Serving.” MLSys, 2024. https://proceedings.mlsys.org/paper_files/paper/2024/hash/054de805fcceb78a201f5e9d53c85908-Abstract-Conference.html.
[3] Sheng, Y. et al. “S-LoRA: Serving Thousands of Concurrent LoRA Adapters.” MLSys, 2024. https://proceedings.mlsys.org/paper_files/paper/2024/hash/906419cd502575b617cc489a1a696a67-Abstract-Conference.html.
[4] Dettmers, T. et al. “QLoRA: Efficient Finetuning of Quantized LLMs.” NeurIPS, 2023. https://arxiv.org/abs/2305.14314.
[5] Liu, S.-Y. et al. “DoRA: Weight-Decomposed Low-Rank Adaptation.” ICML, 2024. https://arxiv.org/abs/2402.09353.
[6] Hayou, S.; Ghosh, N.; Yu, B. “LoRA+: Efficient Low Rank Adaptation of Large Models.” ICML, 2024. https://arxiv.org/abs/2402.12354.
[7] Meng, F.; Wang, Z.; Zhang, M. “PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models.” NeurIPS, 2024. https://arxiv.org/abs/2404.02948.
[8] Qi, X. et al. “Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!” ICLR, 2024. https://arxiv.org/abs/2310.03693.
[9] Liu, H. et al. “LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem.” Findings of EMNLP, 2025. https://aclanthology.org/2025.findings-emnlp.1253/. DOI: 10.18653/v1/2025.findings-emnlp.1253.
[10] Puertolas Merenciano, D. et al. “Weight Space Detection of Backdoors in LoRA Adapters.” 2026. https://arxiv.org/abs/2602.15195.
[11] Ding, S. “Colluding LoRA: A Composite Attack on LLM Safety Alignment.” 2026. https://arxiv.org/abs/2603.12681.
[12] Shekar, P. C.; Krishnan, A. “Adaptive Minds: Empowering Agents with LoRA-as-Tools.” 2025. https://arxiv.org/abs/2510.15416.
[13] Biderman, D. et al. “LoRA Learns Less and Forgets Less.” TMLR, 2024. https://arxiv.org/abs/2405.09673.
[14] Luong, H.-C.; Chen, L. “Why LoRA Fails to Forget: Regularized Low-Rank Adaptation Against Backdoors in Language Models.” Findings of ACL, 2026. https://aclanthology.org/2026.findings-acl.1732/. DOI: 10.18653/v1/2026.findings-acl.1732.
About the Author
Aridio Silva is an independent researcher based in Brazil working on the architecture, security, governance, and trustworthiness of autonomous and distributed artificial intelligence systems. His research focuses on Agentic AI, Multi-Agent Systems, Edge AI, AI Security, Zero Trust, Security-by-Design, AI Governance, Spec-Driven Development, and continuous security assurance. He is the creator and lead researcher of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture, an open research initiative investigating architectural foundations for secure, governed, auditable, and trustworthy autonomous AI systems operating across distributed edge-cloud environments.
Research & Project Resources
- ORCID: https://orcid.org/0009-0008-2411-6995
- Google Scholar: https://scholar.google.com/citations?user=rPn5O48AAAAJ
- Zenodo — SGAEIA Community: https://zenodo.org/communities/sgaeia
- OpenAIRE: https://explore.openaire.eu/search/find?fv0=Aridio%20Silva&f0=q
- Medium: https://medium.com/@aridiosilva
- DEV Community: https://dev.to/aridiosilva
- GitHub: https://github.com/aridiosilva
- LinkedIn: https://www.linkedin.com/in/aridio-silva-74997111/
- Homepage: https://aridiosilva.com
- SGAEIA Homepage: https://aridiosilva.com/sgaeia
- SGAEIA LinkedIn: https://www.linkedin.com/company/sgaeia/
Figures
The cover is unnumbered. Figures 1–4 are high-resolution public conceptual illustrations with PT and EN versions, visible credit, and embedded metadata. They do not expose private SGAEIA protocols, algorithms, state machines, policies, thresholds, or enforcement mechanisms.
License
Except where otherwise noted, the text and original conceptual illustrations in this article are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
© 2026 Aridio Silva. You may share and adapt this work for any purpose, provided appropriate attribution is given.
The SGAEIA software research artifact remains subject to its own Apache License 2.0.
Autonomous AI. Governed by Design. Trusted by Evidence.
Series continuity
This work is SGAEIA Research Series — Article 19. The canonical English edition and the Portuguese edition are openly readable on the SGAEIA homepage. The persistent record is available on Zenodo under DOI 10.5281/zenodo.23300168. Distribution editions are available on Medium, DEV Community, and Academia.edu.
Suggested citation
Silva, Aridio. (2026). LoRA in AI Agents: Efficiency, Security Risks, and Adapter Governance. SGAEIA Research Series, Article 19. https://doi.org/10.5281/zenodo.23300168. Canonical edition: https://aridiosilva.com/publications/artigo19/. CC BY 4.0.