The support agent is authenticated.
Its token is valid. Its access to the ticketing system is approved. Its connection to the customer database is working exactly as designed.
Then it reads a support ticket containing a hidden instruction to retrieve account data and send it somewhere else.
Nothing about the agent’s identity has changed.
Everything about the safety of its next action has.
That is the authorization problem emerging around AI agents. Traditional access control asks whether a known identity may use a resource. Agentic systems add another question: what information is influencing the agent at this moment, and should that information reduce what the agent is allowed to do?
Static permission is not enough for dynamic context.
The National Institute of Standards and Technology put fresh attention on this issue in a September 29 update about its Software and Agentic AI Identity and Authorization project. NIST’s summary of more than 600 public comments says respondents repeatedly called for separate governance or gateway components to evaluate agent requests. Commenters also argued that authorization policy may need to change when prompt injection is suspected—or whenever an agent is working with untrusted data. (NIST NCCoE, September 29, 2026)
That summary is not a final NIST control standard. It is evidence of where the architecture conversation is moving.
The practical lesson is already useful:
When the trustworthiness of an agent’s context decreases, its authority should decrease with it.
Strong Identity Does Not Make the Context Trustworthy
Identity and authorization remain essential.
Every consequential agent should have a distinct identity. Its credentials should be protected. Tool access should be scoped. High-impact actions should require stronger controls. Tokens should be short-lived where practical. Human sponsorship and delegation should be attributable.
Those controls answer important questions:
- Which agent is making the request?
- Which person or organization authorized it?
- Which tools and data may it use?
- What limits apply to its role?
They do not necessarily answer:
- Did the agent just ingest hostile instructions from an email, document, web page, RAG source, or tool response?
- Did sensitive information enter the prompt context unnecessarily?
- Did an external source attempt to redefine the agent’s goal?
- Is the requested action consistent with the original user’s intent?
- Should the agent still have write, send, execute, or disclose authority after the context changes?
An agent can be the correct identity, using a valid credential, while acting on corrupted influence.
This is why prompt injection is not merely a content-filtering problem. It is a runtime authorization signal.
The recent NIST comment summary describes the same architectural tension. It notes that language-model systems may not reliably separate untrusted data from trusted instruction, and that respondents favored logically separate governance components at multiple points—from user input and prompt processing to tool calls and cross-boundary requests. (NIST NCCoE summary of comments)
The reasoning system should not be the only system deciding whether its reasoning is safe enough to execute.
Use Trust-Responsive Authorization
The goal is not to reinvent identity management.
It is to make authorization responsive to the current trust state of the workflow.
Call this trust-responsive authorization:
Agent authority is determined by identity, intended task, requested action, data sensitivity, source trust, and observed content risk—not by identity alone.
That model allows an agent’s permissions to step down as risk rises.
State 1: Normal
The agent is operating within its approved mission. Inputs come from expected sources, inspection finds no material risk, requested actions remain within policy, and the agent uses only the minimum tools required.
Normal does not mean unlimited. It means the bounded workflow may continue without extra intervention.
Examples might include:
- Summarizing a known internal knowledge article.
- Classifying a routine support request.
- Retrieving non-sensitive status information.
- Drafting a response without sending it.
State 2: Constrained
The agent encounters untrusted or privacy-sensitive content, but the workflow can remain useful after controls reduce the risk.
The system may:
- Sanitize or remove suspicious instructions.
- Tokenize or redact sensitive values.
- Defang links.
- Remove write-capable tools.
- Restrict retrieval to an approved data set.
- Convert an autonomous action into a recommendation.
- Prevent external communication or network egress.
The objective is graceful degradation. The organization keeps useful work moving while reducing the blast radius.
State 3: Review
The context contains suspected prompt injection, conflicting instructions, unusual data access, an unexpected tool request, or another condition that makes autonomous action inappropriate.
The agent may continue gathering non-sensitive evidence, but the consequential action pauses for human review.
The reviewer should see:
- The original business request.
- The source and trust level of the suspect content.
- What was detected or changed.
- What the agent attempted to do.
- Which data and tools were involved.
- Which policy required review.
A generic “AI needs approval” message is not enough. The human needs the context required to make a real decision.
State 4: Stop
The request is clearly malicious, violates policy, attempts to expose protected data, seeks an unauthorized action, or cannot be made safe without destroying the purpose of the workflow.
The system blocks or quarantines the transaction. Depending on the situation, it may also revoke a session, disable a workflow, preserve evidence, or trigger incident response.
These states should be defined by system owners before an agent encounters the problem.
An AI agent should not decide for itself when it deserves more authority.
Put an Independent Control Point in the Path
Trust-responsive authorization requires a signal about the content influencing the agent.
That signal should come from a control point that is separate from the model’s reasoning and placed directly in the workflow.
This is the role of an AI Application Firewall.
CaneCorso™ is MicroSolved’s shared control plane for AI workflows. It sits between workflow content and the model or downstream logic. Based on configured policy, it can allow, sanitize, tokenize, or block content; apply privacy controls; and preserve reasons, scores, and evidence for review.
That makes CaneCorso useful as the inspection and enforcement layer in a trust-responsive architecture.
It does not replace identity and access management. It does not replace scoped tool permissions, secure credentials, application authorization, human approval, or incident response.
It gives those systems information they normally do not have:
Is the content influencing this action safe enough for the workflow to continue in its current mode?
A practical integration can work like this:
- Inspect the inbound content. Email, documents, tickets, RAG results, API responses, or tool output pass through CaneCorso before entering the model context.
- Apply the content decision. Approved content proceeds. Sensitive or suspicious portions can be sanitized, tokenized, defanged, or blocked according to policy.
- Map risk to agent mode. The workflow orchestrator uses the disposition to keep the agent in normal mode, remove higher-risk capabilities, require human review, or stop processing.
- Limit tool execution. The authorization layer checks the agent identity, requested action, current mode, data sensitivity, and approval state before a tool runs.
- Inspect consequential output. Where model output will drive decisions or actions, route it through a second inspection point before downstream execution.
- Preserve the decision record. Send the content-risk result, agent identity, tool request, policy decision, approval, and outcome to the organization’s monitoring and audit systems.
MicroSolved’s earlier production-use article describes this two-layer pattern—inspection before model processing and another check before selected downstream use—in Brent Huston’s own agent environment. (Why My AI Agents Needed CaneCorso as a Security Control Plane)
The key design choice is not merely to log the risk score. It is to make the workflow consume the decision.
A Warning Without an Authority Change Is Only Telemetry
Many AI security designs stop at detection.
The system identifies suspicious content. It writes an event. It may notify a security team. The agent continues with the same tools and privileges it had before the warning.
That is visibility, not control.
If a possible injection does not alter the available action set, the organization is asking a later human process to contain a machine-speed decision. The alert may arrive after the email was sent, the record was modified, the file was retrieved, or the tool call completed.
This does not mean every suspicious string should shut down the workflow. Overblocking can make AI systems unusable and encourage teams to bypass the control.
The better design is proportional response:
- Low-risk privacy exposure can be tokenized while the workflow continues.
- A suspicious URL can be defanged before analysis.
- Untrusted text can be summarized with all action-capable tools removed.
- A suspected injection can force the agent into read-only mode.
- A consequential request can require human approval.
- Clearly malicious content can be blocked and preserved for investigation.
CaneCorso’s allow, sanitize, tokenize, and block outcomes support that more nuanced response. The surrounding application determines which tools remain available and which approvals are required.
That separation of responsibilities matters. The content control should not silently become the identity provider, and the identity provider should not pretend it understands prompt semantics.
Test Whether the Boundary Actually Holds
No prompt-injection control should be treated as infallible.
NIST’s March analysis of a large agent red-teaming competition reported more than 250,000 attack attempts by over 400 participants across 13 frontier models. At least one successful hijacking attack was found against every target model. NIST’s takeaway was not that agent use is impossible. It was that evaluations must evolve with adversaries and that comparative red teaming provides information ordinary static tests may miss. (NIST CAISI, March 23, 2026)
The joint Careful Adoption of Agentic AI Services guidance from agencies including CISA, NSA, and the national cyber centers of Australia, Canada, New Zealand, and the United Kingdom recommends input validation and sanitization, prompt-injection filters, continuous evaluation, fail-safe defaults, human control points, runtime monitoring, and regular attempts to bypass safeguards. (Joint guidance, May 1, 2026)
For one representative agent workflow, test the complete control chain:
- Establish a normal task and record the intended agent identity, data sources, tools, and allowed actions.
- Introduce hostile instructions through realistic sources such as a ticket, email, document, RAG entry, API response, or tool result.
- Confirm that CaneCorso identifies, sanitizes, tokenizes, or blocks the content according to the workflow policy.
- Confirm that the resulting disposition changes the agent’s available authority as designed.
- Attempt a high-impact action after the risk signal and verify that the authorization layer blocks it or requires human approval.
- Verify that the agent cannot recover removed capability by rewriting the request, changing language, splitting the instruction across sources, or calling a different tool.
- Inspect the record. Confirm that responders can identify the source, agent, content decision, policy state, requested tool, approval, and final outcome.
- Test the failure mode. Decide what happens if the control plane, identity service, logging path, or human approval channel is unavailable.
CaneCorso’s Injection Scanner is designed to help exercise these controls before deployment, after changes, and during assurance reviews. The test still needs to cover the full application behavior. Detecting the input is only one link; the system must also constrain the agent and prevent the unsafe action.
Measure the Decision, Not Just the Detection
Counting detected prompts can be misleading.
A high count may indicate active attacks, a noisy policy, a research feed that contains legitimate attack examples, or ordinary content that resembles instructions. A low count may indicate clean inputs—or weak coverage.
More useful measures connect detection to outcome:
- Percentage of agent actions carrying a recorded content-risk disposition.
- Time from suspicious input detection to reduced agent authority.
- Percentage of high-impact actions attempted after a risk signal and correctly blocked or escalated.
- Percentage of sensitive values tokenized or redacted before model exposure.
- Rate of safe workflows disrupted by false positives.
- Rate of risky workflows allowed because a control failed open, timed out, or was bypassed.
- Percentage of human reviews with enough evidence to explain the requested action and policy decision.
- Results of adversarial validation before and after model, prompt, tool, policy, or data-source changes.
- Time required to reconstruct which content influenced a consequential agent action.
The important metric is not “How many injections did we detect?”
It is:
When the context became less trustworthy, did the system reduce authority before harm occurred?
What Security Leaders Should Do This Week
Choose one agent that consumes external or mixed-trust content.
- Name the consequential actions. Identify every send, write, execute, disclose, purchase, delete, or privilege-changing capability available to the agent.
- Map the influence paths. Include users, email, documents, RAG sources, web content, API responses, tool output, memory, and other agents.
- Add an independent inspection point. Put CaneCorso or an equivalent AI Application Firewall before untrusted content reaches the model and before consequential output reaches downstream logic.
- Define the four states. Document what Normal, Constrained, Review, and Stop mean for this workflow.
- Bind state to capability. Make the orchestrator and authorization layer remove tools, require approval, or block execution as risk increases.
- Run an adversarial test. Use realistic indirect prompt injection attempts and verify the action is prevented—not merely detected.
- Preserve the reason. Record what influenced the agent, which policy applied, what authority changed, who approved any exception, and what finally happened.
Agent identity answers who is acting.
An AI Application Firewall helps determine whether the content influencing that agent is safe enough to use.
Authorization must connect the two.
If an agent can encounter hostile context while retaining its full authority, the architecture is trusting the moment when it should become most skeptical.
Prompt injection should not merely create an alert.
It should change what the agent is allowed to do.
Put CaneCorso in Front of Your AI Workflows
CaneCorso™ gives organizations a shared AI Application Firewall for email, document AI, RAG, support workflows, copilots, and agent-driven automation.
MicroSolved, Inc. can help you:
- Map the content, privacy, tool, and authorization boundaries in an AI workflow.
- Pilot CaneCorso against a representative production use case.
- Configure allow, sanitize, tokenize, and block policies.
- Test prompt-injection defenses with CaneCorso’s Injection Scanner.
- Connect runtime decisions and evidence to monitoring, review, and audit processes.
- Design constrained, human-review, and fail-safe operating states for AI agents.
To discuss a CaneCorso pilot or schedule a practical review, contact MicroSolved at info@microsolved.com or +1.614.351.1237.
Relax. We’re on watch.
AI tools were used as a research assistant for this content, but human moderation and writing are also included.








