Operational Changes in Critical Infrastructure Security
Explains the ANCHOR-CI framework and AI agent governance in security operations.
What Changed Operationally
The operational landscape for critical infrastructure security has shifted decisively with the announcement of the ANCHOR-CI framework. This new initiative represents a structural evolution in how the public and private sectors collaborate to secure national assets. By incorporating the best practices and lessons learned from the previous Critical Infrastructure Partnership Advisory Council (CIPAC), ANCHOR-CI aims to transform the nature of information sharing and threat response. The framework introduces a governance model designed to expand meaningful engagement to a wider range of stakeholders, moving beyond traditional sector-specific silos. This expansion is intended to create a more resilient ecosystem where real-time threat intelligence flows between government agencies and private operators with greater speed and relevance.
The significance of this change lies in its ability to operationalize the theoretical benefits of public-private partnership. The DHS and CISA have developed the framework to address the complexities of modern cyber threats, which often transcend organizational boundaries. Under ANCHOR-CI, CISA will manage the governance of the councils established under the new structure, ensuring a centralized yet distributed approach to oversight. The framework allows for the establishment of four distinct types of councils: critical infrastructure sector councils, cross-sector councils, industry councils, and regional coordinating councils. This variety is designed to address specific threats that require sector-specific knowledge while simultaneously fostering collaboration across different industries and geographical regions. The framework is set to operate for an initial period of two years, with the potential for extension by the Secretary pursuant to Section 871 of the Homeland Security Act, providing a stable yet adaptable platform for long-term security planning.
Architectural Separation of Skills, Reasoning, and Models
The operational capability of modern security agents hinges on a distinct separation of concerns that moves beyond simple automation. As the industry moves fast on capabilities, agents now perform complex tasks such as triaging alerts, investigating endpoints, creating detection rules, and enriching indicators. However, the ability to govern these agents effectively requires a clear architectural distinction between the agent's skills, its reasoning patterns, the underlying models, and the environmental context. Skills define what the agent can do across triage, enrichment, forensics, and detection engineering—each acting as a composable unit of instructions, tools, and domain knowledge. This separation is crucial because it isolates the reasoning process, which is the core determinant of agent behavior. It is reasoning that decides whether an alert gets escalated or closed, whether lateral movement gets investigated, or whether a host gets isolated.
How The Capability Fits Together
This architectural layering enables predictability and compliance, which are essential for integrating AI agents into security operations. The reasoning patterns, which include investigation methodology, escalation logic, evidence evaluation criteria, and hypothesis generation patterns, are distinct from the specific Large Language Models (LLMs) executing the tasks. While models like Claude, Gemini, ChatGPT, or open-source variants work differently based on their training and guardrails, the reasoning logic must remain consistent to ensure reliable outcomes. By decoupling the reasoning layer from the model layer, operators can maintain governance over how decisions are made, a requirement that is converging in regulations such as ISO 42001, DORA, NIS2, and the EU AI Act. These frameworks mandate that organizations prove their AI systems work as intended and that humans maintain meaningful oversight, making the separation of reasoning and models a technical necessity for compliance.
Progressive Trust and the Human-in-the-Loop
The governance of AI agents in security operations relies heavily on the concept of progressive trust, which balances autonomy with necessary human intervention. Most teams deploying agents in production today default to a tiered autonomy model where low-risk actions run autonomously, medium-risk actions require oversight, and high-risk actions demand human approval. The goal of progressive trust is to start with maximum oversight, where every consequential action goes through a human, and then reduce that oversight as the agent accumulates evidence of reliable behavior. However, the depth of any investigation is bound by the platform's ability to pull environmental context into the agentic solution. Without this context, an agent may make decisions based on incomplete information, undermining the trust required for autonomous operation.
To manage this trust effectively, organizations must implement rigorous evaluation metrics and monitoring systems. Classification accuracy is measured through precision (how many of its positive classifications were correct?) and recall (how many actual positives did it catch?). Beyond simple accuracy, planning quality refers to whether the agent considered multiple hypotheses, while retrieval quality refers to whether the agent found the right evidence. Grounding quality determines whether the agent’s reasoning is supported by the evidence it found. Ablation testing is used to validate these metrics by building a benchmark suite with known ground truth and systematically varying the configuration to identify which components contribute to agent quality. As AI agents become operational, the human role is shifting from operating to supervising to governing, requiring a system that can monitor and respond to agent-based decisions just as it would to human operators.
Operational Impact
Operationalizing the New Governance Frameworks
The transition to the ANCHOR-CI framework requires administrators to fundamentally restructure how they approach public-private collaboration. Unlike previous structures, this framework mandates the establishment of four distinct council types—critical infrastructure sector councils, cross-sector councils, industry councils, and regional coordinating councils. Engineers must map their current stakeholder engagement strategies to these specific categories to ensure they are not operating in a vacuum. The framework is designed to expand meaningful, impactful engagement to a wider range of stakeholders, which means administrators will need to identify which specific councils align with their operational scope. Because the framework incorporates the best practices and lessons learned from the Critical Infrastructure Partnership Advisory Council (CIPAC), administrators should audit existing engagement protocols to identify gaps where the new structure will force a shift in communication channels and frequency.
For security operations teams, the introduction of AI agents capable of triaging alerts and investigating endpoints necessitates a rigorous evaluation of autonomy levels. The industry has moved fast on capabilities, and most teams deploying agents in production today default to a tiered autonomy model. Administrators must define clear boundaries for these agents, ensuring that low-risk actions run autonomously while high-risk actions require human approval. This tiered approach is essential for maintaining the balance between operational efficiency and security posture. As agents become capable of performing most actions security operators can perform, the administrator's role shifts from operating to supervising to governing. This shift requires the implementation of a "progressive trust" model, where oversight is initially maximum and only decreases as the agent proves its reliability over time.
Evaluation and Compliance Requirements
Rollout And Governance Decisions
Governance in this environment is defined by the need for predictability, consistency, and explainability. Regulations such as ISO 42001, DORA, NIS2, and the EU AI Act all converge on the expectation that organizations must prove their AI systems work as intended, that humans maintain meaningful oversight, and that decisions can be explained. To meet these requirements, administrators must separate the reasoning layer from the model layer. Reasoning is the core determinant of agent behavior, deciding whether an alert gets escalated or closed. Because reasoning patterns differ between models like Claude, Gemini, or open-source options, administrators cannot rely on a single model's output without rigorous testing. The depth of any investigation is bound by the platform's ability to pull environmental context into the agentic solution, so administrators must ensure their environment is fully instrumented to support this level of analysis.
The evaluation of agent performance requires a shift from simple classification accuracy to a broader set of metrics. While precision and recall are standard measures, administrators must also evaluate planning quality (whether the agent considered multiple hypotheses), retrieval quality (whether the agent found the right evidence), and grounding quality (whether the agent's reasoning is supported by evidence). Ablation testing is a concrete method for validating these metrics; administrators should build a benchmark suite with known ground truth, run the agent, and capture the full trajectory of every tool call and reasoning step. When AI agents become operational in security, administrators need systems not only to monitor and respond to security incidents but also to monitor and respond to agent-based decisions. This requires capturing the full reasoning chain, evidence, decision boundaries, and true negatives to ensure compliance and maintain accountability.
Prerequisites and Risk-Based Prioritization
Before deploying these advanced governance and automation tools, administrators must address the foundational requirements of vulnerability management. The regulatory landscape is tightening, with BOD 26-04 establishing vulnerability management requirements for Federal Civilian Executive Branch (FCEB) agencies and reinforcing the importance of the CISA Known Exploited Vulnerabilities (KEV) Catalog. While BOD 26-04 applies specifically to FCEB agencies, CISA encourages all organizations to adopt risk-based vulnerability management. Administrators must prioritize the remediation of vulnerabilities listed in the KEV Catalog, such as CVE-2026-48939 and CVE-2026-56291, which allow unrestricted file uploads with dangerous types. These vulnerabilities pose immediate risks that must be patched before introducing new, complex governance frameworks.
Furthermore, administrators must ensure that the physical and software infrastructure supporting these governance efforts is secure. For example, legacy software like Schneider Electric PowerChute Serial Shutdown versions <=1.4 is vulnerable to file path overwriting and resource consumption issues. Version 1.5 is available as a fixed version for Windows, Red Hat Enterprise Linux, and SuSE Linux. Administrators must verify that all critical infrastructure components, including power management software, are updated to mitigate these specific vulnerabilities. The introduction of AI agents and new council structures does not reduce the need for traditional security hygiene; rather, it increases the complexity of the environment. Administrators must ensure that the "grounding quality" of their agents is not compromised by insecure underlying infrastructure or unpatched software.
Failure Modes And Limits
Operational Risks and Governance Challenges
The deployment of autonomous agents introduces significant operational risks that differ fundamentally from traditional software development. While agents are capable of triaging alerts, investigating endpoints, and creating detection rules, their ability to perform most actions traditionally reserved for human operators creates a complex governance landscape. The primary challenge lies in ensuring that the reasoning driving these actions remains predictable and consistent. Because the underlying logic is often opaque, organizations must establish rigorous frameworks to separate reasoning from the models executing it. Without this separation, the "skills" of the agent—the specific tools and domain knowledge it possesses—can be compromised by variations in how different models, such as Claude, Gemini, or open-source alternatives, process information. This variability means that a change in the underlying model or the guardrails applied by a vendor can skew results, making it difficult to maintain a stable operational baseline.
Security And Privacy Considerations
Furthermore, the integration of these agents into critical infrastructure environments requires a shift from "operating" to "supervising." The concept of progressive trust is essential here; organizations must begin with maximum oversight where every consequential action requires human approval. However, the human role is evolving into one of governance, necessitating the ability to monitor and respond to agent-based decisions in real-time. As agents become more autonomous, the risk of approval fatigue increases. If every decision requires a human signature, the human operator is effectively removed from the loop. Therefore, design must prioritize friction proportional to consequence, ensuring that low-risk actions can proceed autonomously while high-risk decisions are rigorously reviewed. This balance is critical to maintaining the agility that agents promise while mitigating the risk of catastrophic errors.
Limitations and Unanswered Questions
Despite the promise of enhanced security through autonomous agents, several limitations and unanswered questions remain, particularly regarding the depth of investigation and the validity of evaluation metrics. The depth of any investigation is strictly bound by the platform's ability to pull environmental context into the agentic solution. If the agent lacks access to necessary data or the tools to interpret it correctly, its capabilities are artificially limited, potentially missing critical threats. Additionally, while metrics like classification accuracy, planning quality, and retrieval quality are proposed as standards for evaluation, they do not fully capture the nuance of security operations. For instance, classification accuracy measures precision and recall, but it does not account for the context in which a decision is made or the potential for false positives to overwhelm human analysts.
Open Questions
There are also significant uncertainties regarding the long-term stability of these systems. The industry has moved rapidly on capabilities, but the regulatory landscape is still catching up. While frameworks like ISO 42001, DORA, NIS2, and the EU AI Act converge on the need for explainability and human oversight, the practical implementation of these requirements for autonomous agents is still nascent. Organizations must verify how these regulations apply to their specific use cases, particularly when dealing with cross-sector councils and regional coordinating councils designed to share sensitive information. The lack of standardized performance metrics for agent-based governance means that organizations must develop their own benchmarks, relying on ablation testing to systematically vary configurations and capture full trajectories of tool calls and reasoning steps.
Environment Checklist
Environment Checklist
- Verify Model Compatibility: Ensure that the reasoning patterns and guardrails of the selected AI model (e.g., Claude, Gemini) are compatible with your organization's security policies and operational context.
- Implement Progressive Trust: Start with a "max oversight" approach for all consequential actions, gradually reducing human intervention only after the agent has demonstrated reliability through ablation testing and consistent performance.
- Audit Environmental Context: Confirm that the platform has the necessary access to pull environmental context, evidence, and decision boundaries required for the agent to perform deep investigations.
- Design for Approval Fatigue: Structure approval workflows so that friction is proportional to consequence, ensuring that low-risk actions do not require human intervention that could lead to oversight fatigue.
- Monitor Agent Telemetry: Establish monitoring systems to capture the full reasoning chain, evidence, and decision boundaries, ensuring you can audit agent actions just as you would human decisions.
Verification Statement
This article was not lab-tested. The information presented regarding the operational risks, governance challenges, and limitations of AI agents is synthesized from industry research and regulatory frameworks. Readers must independently verify the compatibility of specific AI models with their infrastructure and conduct their own risk assessments before deploying autonomous agents in production environments.
// source record
Sources
- https://www.cisa.gov/news-events/news/cisa-announces-new-advisory-council-strengthen-partnerships-and-secure-critical-infrastructure www.cisa.gov · checked 13 July 2026
- https://www.elastic.co/blog/the-future-of-governing-ai-agents www.elastic.co · checked 13 July 2026
- https://www.cisa.gov/news-events/ics-advisories/icsa-26-190-02 www.cisa.gov · checked 13 July 2026
- https://www.cisa.gov/news-events/alerts/2026/07/10/cisa-adds-two-known-exploited-vulnerabilities-catalog www.cisa.gov · checked 13 July 2026