Prompt
How do I secure APIs for AI agents?
Latest observation
Securing APIs for AI agents is mostly about controlling who can act, what they can access, and how much damage they can do if something goes wrong. Because agents can plan, chain calls, and sometimes be influenced by untrusted input, you need stronger controls than a normal app.
Core principles
1) Use least privilege
Give the agent only the permissions it absolutely needs.
- Prefer narrow, task-specific scopes
- Separate read vs write permissions
- Separate environments and tenants
- Don’t give agents broad admin keys
Example:
- Good:
calendar.read,calendar.create_event - Bad: full user account access
2) Treat the agent as untrusted
Even if the agent is “yours,” its prompts, retrieved data, and tool outputs may be manipulated.
- Assume prompt injection is possible
- Assume retrieved web content may be malicious
- Assume tool outputs may contain unsafe instructions
- Never let model output directly become an authorization decision
3) Put authorization outside the model
The model should not decide whether it can call an API.
Use a separate enforcement layer for:
- Authentication
- Authorization
- Policy checks
- Rate limits
- Data access controls
The agent can request an action, but your backend must approve it.
Authentication patterns
4) Use short-lived credentials
Prefer ephemeral tokens over long-lived API keys.
- OAuth 2.0 access tokens
- Signed, short-lived service tokens
- Rotating credentials
- Per-session tokens for agents
Avoid:
- Hardcoded API keys
- Shared static secrets
- Keys with broad privileges
5) Bind tokens to context
Limit token reuse if stolen.
Examples:
- Bind token to user session
- Bind to device, workload, or mTLS identity
- Use audience restrictions
- Use nonce or proof-of-possession where possible
6) Separate human and agent identity
An agent should act either:
- as itself, with its own service identity, or
- on behalf of a user, with explicit delegated permissions
Don’t blur the two.
Use:
- OAuth delegated consent for user-scoped actions
- Service accounts for autonomous background tasks
- Clear audit trails for impersonation/delegation
Authorization patterns
7) Use policy-based access control
Implement rules in a policy engine or backend guardrails.
Examples of policy checks:
- Is this agent allowed to use this tool?
- Is this action allowed for this user?
- Is the data within this tenant?
- Is the request within time, budget, or risk limits?
Good approaches:
- RBAC for coarse roles
- ABAC for context-aware decisions
- Policy engines like OPA, Cedar, or custom policy middleware
8) Constrain tools, not just endpoints
If the agent can call tools, each tool should have:
- A strict schema
- Validated arguments
- Explicit side effects
- Well-defined scopes
For example:
search_orders(query)is safer than raw database accesssend_email(to, subject, body)should be stricter than generic SMTP access
9) Require step-up approval for high-risk actions
For sensitive actions, add human confirmation or extra checks.
Examples:
- Sending money
- Deleting data
- Sharing private records
- Changing security settings
- Accessing regulated data
Use:
- Human-in-the-loop approval
- Two-person approval
- Transaction signing
- Re-authentication
Input and output safety
10) Validate every tool call
Never trust the model’s arguments.
- Enforce JSON schema validation
- Check required fields, types, ranges, and enums
- Reject unexpected parameters
- Canonicalize inputs
- Protect against injection in SQL, shell, URLs, headers, and templates
11) Sanitize untrusted content before feeding it to the agent
Retrieved text, emails, tickets, pages, and docs may contain malicious instructions.
Mitigations:
- Label untrusted content clearly
- Strip or segregate instructions from data
- Use retrieval filters
- Don’t allow retrieved content to override system policy
- Consider content classification before presenting it to the model
12) Don’t let the model construct raw privileged requests
Avoid patterns where the model generates:
- SQL directly
- Shell commands directly
- IAM policies directly
- Production config changes directly
Instead:
- Have the model choose from safe, pre-approved actions
- Translate intent into constrained operations server-side
Network and infrastructure controls
13) Isolate agent execution
Run agents in restricted environments.
- Containers or sandboxes
- Minimal filesystem access
- No unnecessary network egress
- Separate prod from non-prod
- Limit access to secrets and metadata services
14) Restrict egress
An agent should not be able to call arbitrary internet destinations unless needed.
- Allowlist domains
- Block internal metadata IPs
- Filter outbound requests
- Monitor DNS and HTTP activity
15) Segment data access
Don’t give agents direct access to your whole database or object store.
- Use per-tenant views
- Use scoped search indices
- Use read replicas where possible
- Mask or tokenize sensitive fields
Secrets management
16) Never expose secrets to the model if you can avoid it
If the model does not need a secret, don’t provide it.
- Use backend tool execution
- Keep secrets in vaults
- Inject secrets only at the point of use
- Rotate and revoke regularly
17) Prevent secret exfiltration
If an agent can read sensitive data, it may leak it.
Mitigations:
- DLP checks
- Response filtering
- Redaction of secrets, tokens, and PII
- Output classifiers for sensitive content
- Least-privilege access to source data
Abuse prevention
18) Add rate limits and quotas
Prevent runaway cost and abuse.
- Per user
- Per agent
- Per tool
- Per tenant
- Per time window
Also consider:
- Budget limits
- Maximum number of tool calls per task
- Circuit breakers on repeated failures
19) Log everything important
You need auditable traces of:
- Prompt context used for authorization-sensitive actions
- Tool calls
- Input parameters
- Authorization decisions
- Human approvals
- Token issuance and revocation
Make logs tamper-resistant where possible.
20) Monitor anomalous behavior
Watch for:
- Unusual tool call patterns
- Sudden spikes in access
- Repeated denied requests
- Attempts to access unrelated data
- Large exports
- Off-hours high-risk actions
Common mistakes to avoid
- Giving the model a master API key
- Trusting the model to enforce access control
- Allowing direct database or shell access
- Letting retrieved web content override system instructions
- Using long-lived tokens with broad scope
- Failing to validate tool arguments
- No approval flow for destructive actions
- No audit logs or alerting
- Mixing user identity and agent identity
- Not isolating tenants
A practical secure architecture
A solid pattern is:
- Agent receives a user request
- Agent proposes an action
- Backend validates intent and arguments
- Policy engine checks permissions
- Sensitive actions require approval
- Tool executes with short-lived scoped credentials
- Results are filtered/redacted
- Audit logs are stored
In this setup, the model is a planner, not the authority.
Quick checklist
- Short-lived credentials
- Least privilege scopes
- Separate agent identity from user identity
- Backend-enforced authorization
- Strict tool schemas
- Input validation and output filtering
- Sandbox and network egress controls
- Human approval for high-risk actions
- Comprehensive logging and monitoring
- Secret management and rotation
If you want, I can turn this into:
- a reference architecture diagram,
- a security checklist for production, or
- a sample implementation using OAuth + policy engine + tool sandboxing.