AI Agent Authentication and Identity: From Shared API Keys to Workload Identity
Every major AI authentication pattern explained: API keys, OAuth 2.0, workload identity, DPoP tokens, and scoped credentials for agent-to-API…
Every major AI authentication pattern explained: API keys, OAuth 2.0, workload identity, DPoP tokens, and scoped credentials for agent-to-API communication in production systems.
AI Agent Authentication and Identity: From Shared API Keys to Workload Identity
Most AI agent deployments authenticate the same way a weekend prototype does: a single API key, hardcoded or stuffed into an environment variable, shared across every agent instance. This works until it doesn’t — and “doesn’t” arrives fast when agents start calling third-party APIs autonomously, operating across trust boundaries, or running as persistent background processes that outlive any single user session.
The authentication problem for AI agents is distinct from traditional service-to-service auth. Agents make unpredictable sequences of API calls, often to services discovered at runtime via tool schemas. They may operate on behalf of a user but with delegated authority that should be narrower than the user’s own permissions. They run in loops where a single logical task can span minutes or hours, crossing token expiration boundaries. And with frameworks enabling parallel agent instances that share context — as Claude Code’s multi-instance coordination now demonstrates — the identity model needs to handle concurrent agents acting under the same logical identity but requiring independent credential scopes.
This post covers the full spectrum: from what’s wrong with the status quo, through OAuth 2.0 patterns adapted for agents, to workload identity systems and emerging standards like DPoP that bind credentials to specific agent instances.
Table of Contents
- The Shared API Key Problem
- Authentication vs Authorization vs Identity
- API Key Patterns and Their Limits
- OAuth 2.0 for Agent Workloads
- Workload Identity: Moving Beyond Secrets
- DPoP: Proof-of-Possession for Agents
- Scoped Tokens and Least-Privilege Design
- MCP Authentication and Tool-Level Identity
- Credential Management in Long-Running Agents
- Implementation Patterns
- Summary
- Further Reading
The Shared API Key Problem
A typical agent deployment uses a single API key to call an LLM provider and another key (or the same one) for downstream services — databases, search APIs, code execution sandboxes. Every agent instance shares these credentials.
All agent instances funnel through a single shared credential — no per-instance attribution, no granular revocation.
The problems compound:
No attribution. When three agent instances share a key and one starts making anomalous calls — say, exfiltrating data through a prompt injection attack — audit logs show the key, not which agent instance or user session triggered the behavior.
Blast radius. Revoking a compromised key kills every agent instance simultaneously. In production systems running Claude Cowork or similar persistent agents, this means interrupting every active user session.
Scope creep. A key that needs read access to a database and write access to a cache often ends up with admin access to both, because creating fine-grained keys per agent is operationally painful.
Secret sprawl. Keys get copied into CI/CD pipelines, container images, agent configuration files, and MCP server manifests. Each copy is an exfiltration surface. Anthropic’s own disclosure about Claude models breaching test environments and publishing malware to PyPI illustrates what happens when agent processes acquire more access than intended.
Authentication vs Authorization vs Identity
These terms get conflated in agent architectures. The distinctions matter for choosing the right pattern.
Authentication establishes identity; identity feeds both authorization decisions and audit trails.
Authentication answers “who is this agent?” — proving that a request comes from a legitimate agent instance, not an impersonator. API keys do this weakly (bearer token: anyone holding the key is “authenticated”). Client certificates and DPoP do this strongly (proof of possession of a private key).
Authorization answers “what is this agent allowed to do?” — mapping an authenticated identity to a set of permissions. OAuth 2.0 scopes are the most common mechanism. Without authorization, an authenticated agent has implicit access to everything the credential grants.
Identity is the persistent, verifiable representation of an agent. In cloud-native systems, this is a workload identity — a service account, IAM role, or SPIFFE ID that exists independently of any secret. The identity persists even when credentials rotate.
For AI agents specifically, there’s a fourth dimension: delegation. The agent often acts on behalf of a human user, and the system needs to distinguish between what the agent is allowed to do (its own permissions) and what the user has delegated to it (which should be a strict subset of the user’s permissions).
API Key Patterns and Their Limits
Despite their limitations, API keys remain the dominant authentication method for LLM APIs. Every major provider — OpenAI, Anthropic, Google — uses them as the primary auth mechanism. The key question isn’t whether to use API keys at all, but how to use them with less risk.
Per-Agent Key Issuance
Instead of one key shared across all instances, issue a unique key per agent instance or per logical workflow.
# Key provisioning service — issues scoped keys per agent
import hashlib
import secrets
from datetime import datetime, timedelta
class AgentKeyManager:
def __init__(self, vault_client):
self.vault = vault_client
def provision_key(self, agent_id: str, scopes: list[str], ttl_hours: int = 24):
"""Issue a scoped, time-limited API key for a specific agent instance."""
key = secrets.token_urlsafe(32)
key_hash = hashlib.sha256(key.encode()).hexdigest()
self.vault.store(
path=f"agents/{agent_id}/api-key",
data={
"key_hash": key_hash,
"scopes": scopes,
"agent_id": agent_id,
"issued_at": datetime.utcnow().isoformat(),
"expires_at": (datetime.utcnow() + timedelta(hours=ttl_hours)).isoformat(),
}
)
return key # Return plaintext only once; agent stores in memory
def revoke_key(self, agent_id: str):
"""Revoke a single agent's key without affecting others."""
self.vault.delete(path=f"agents/{agent_id}/api-key")
This gives per-instance revocation and audit attribution. The tradeoff: operational complexity scales linearly with agent count, and most LLM providers don’t support programmatic key issuance through their APIs (you’d need to manage this at your own gateway layer).
Key Hierarchy
| Pattern | Attribution | Revocation Scope | Operational Cost |
|---|---|---|---|
| Single shared key | None | All agents | Minimal |
| Per-user key | User-level | All agents for one user | Low |
| Per-agent key | Instance-level | Single agent | Medium |
| Per-request token | Request-level | Single request | High |
| Workload identity (no key) | Instance-level | Automatic rotation | Medium (setup), Low (ongoing) |
The industry is moving toward the bottom of this table, but most production deployments today sit in the top three rows.
OAuth 2.0 for Agent Workloads
OAuth 2.0 was designed for delegated authorization — exactly the problem agents face when acting on behalf of users. The challenge is mapping OAuth’s interactive flows to non-interactive agent processes.
Client Credentials Flow
For agents that act as their own identity (not on behalf of a specific user), the client credentials grant is the natural fit. The agent authenticates with a client ID and client secret, receiving a short-lived access token.
Client credentials flow: agent authenticates directly, receives scoped tokens, refreshes before expiry.
import httpx
from datetime import datetime, timedelta
class AgentOAuthClient:
def __init__(self, client_id: str, client_secret: str, token_url: str):
self.client_id = client_id
self.client_secret = client_secret
self.token_url = token_url
self._token = None
self._expires_at = None
async def get_token(self, scopes: list[str]) -> str:
if self._token and self._expires_at > datetime.utcnow() + timedelta(seconds=30):
return self._token
async with httpx.AsyncClient() as client:
response = await client.post(
self.token_url,
data={
"grant_type": "client_credentials",
"client_id": self.client_id,
"client_secret": self.client_secret,
"scope": " ".join(scopes),
},
)
data = response.json()
self._token = data["access_token"]
self._expires_at = datetime.utcnow() + timedelta(seconds=data["expires_in"])
return self._token
Token Exchange for Delegation
When an agent acts on behalf of a user, the OAuth 2.0 Token Exchange (RFC 8693) allows the agent to swap a user’s token for a narrower agent-specific token.
Token exchange converts a broad user token into a narrow agent-specific token — the agent never holds the user’s full permissions.
The critical property: the exchanged token has strictly fewer permissions than the original. An agent operating on behalf of a user with read/write access to a repository might receive a token with read-only access plus write access to a specific branch.
The Refresh Problem for Long-Running Agents
Standard OAuth access tokens expire in minutes to hours. Agents running multi-step workflows — Claude Cowork sessions, research agents, CI/CD pipelines — can run for hours or days. The refresh token flow handles this, but introduces its own complications:
- Refresh tokens must be stored securely (they’re long-lived secrets)
- Refresh token rotation means a single failed refresh can invalidate the entire chain
- Concurrent agent instances sharing a refresh token will race on rotation, causing one to receive an invalid token
The practical solution: use a dedicated token management sidecar or service that handles refresh logic, and have agents request fresh tokens through it rather than managing refresh tokens directly.
Workload Identity: Moving Beyond Secrets
Workload identity eliminates static secrets entirely. Instead of authenticating with a key or client secret, the agent proves its identity based on where and how it runs.
Cloud Provider Workload Identity
Every major cloud provider offers a mechanism for workloads to obtain credentials without static secrets:
| Cloud | Mechanism | Identity | Token Lifetime |
|---|---|---|---|
| AWS | IAM Roles for Service Accounts (IRSA) / EKS Pod Identity | IAM Role ARN | 1-12 hours |
| GCP | Workload Identity Federation | Service Account email | 1 hour |
| Azure | Managed Identity | Object ID | 24 hours |
Workload identity flow: the agent pod has no static secrets; it obtains credentials from the cloud metadata service based on its runtime identity.
The agent process never sees a long-lived secret. Credentials rotate automatically. If a container is compromised, the attacker gets a token that expires within hours and can’t be used from a different network context.
SPIFFE/SPIRE for Cross-Platform Identity
SPIFFE (Secure Production Identity Framework For Everyone) provides workload identity that works across cloud providers and on-premises infrastructure. Each workload gets a SPIFFE ID (a URI like spiffe://example.com/agent/research-agent) and a short-lived X.509 certificate or JWT.
# Using the SPIFFE Workload API to get an identity
from pyspiffe.workloadapi import WorkloadApiClient
class SPIFFEAgentIdentity:
def __init__(self, socket_path: str = "unix:///tmp/spire-agent/public/api.sock"):
self.client = WorkloadApiClient(socket_path)
def get_jwt_token(self, audience: str) -> str:
"""Get a JWT-SVID for authenticating to a specific service."""
jwt_svid = self.client.fetch_jwt_svid(audiences=[audience])
return jwt_svid.token
def get_mtls_context(self):
"""Get X.509 materials for mTLS connections."""
x509_svid = self.client.fetch_x509_svid()
return {
"cert": x509_svid.cert_chain,
"key": x509_svid.private_key,
"trust_bundle": self.client.fetch_x509_bundles(),
}
SPIFFE is particularly useful for multi-agent architectures where agents communicate with each other — each agent can verify the other’s identity through mutual TLS without a central authentication service in the request path.
Workload Identity Federation for External APIs
The harder problem: agents need to call external APIs (LLM providers, SaaS tools) that don’t natively support workload identity. The pattern here is federation — exchanging a cloud-native identity token for an external API credential.
Federation bridges cloud identity to external APIs that only support API keys — the vault issues short-lived, scoped credentials.
HashiCorp Vault’s dynamic secrets engine can generate short-lived database credentials on demand. For APIs that only support static keys (most LLM providers today), Vault can act as a proxy: the agent authenticates to Vault with its workload identity, and Vault returns a cached API key after logging the access and enforcing policy.
DPoP: Proof-of-Possession for Agents
Demonstration of Proof-of-Possession (DPoP, RFC 9449) addresses a specific weakness of bearer tokens: if an access token is intercepted, anyone can use it. DPoP binds each token to a cryptographic key pair held by the client.
How DPoP Works
The agent generates a key pair at startup. When requesting a token, it sends a DPoP proof — a signed JWT containing the HTTP method, URL, and a unique identifier. The authorization server issues a token bound to that key. Subsequent API calls must include both the token and a fresh DPoP proof signed with the same key.
DPoP flow: the token is cryptographically bound to the agent’s key pair, making stolen tokens unusable without the private key.
import json
import time
import uuid
from jwcrypto import jwk, jwt
class DPoPAgent:
def __init__(self):
# Generate ephemeral key pair — lives only in agent memory
self.key = jwk.JWK.generate(kty='EC', crv='P-256')
def create_dpop_proof(self, method: str, url: str, access_token: str = None) -> str:
"""Create a DPoP proof JWT for a specific request."""
header = {
"typ": "dpop+jwt",
"alg": "ES256",
"jwk": json.loads(self.key.export_public()),
}
claims = {
"jti": str(uuid.uuid4()),
"htm": method,
"htu": url,
"iat": int(time.time()),
}
if access_token:
# Bind proof to specific access token (ath claim)
import hashlib, base64
ath = base64.urlsafe_b64encode(
hashlib.sha256(access_token.encode()).digest()
).rstrip(b'=').decode()
claims["ath"] = ath
token = jwt.JWT(header=header, claims=claims)
token.make_signed_token(self.key)
return token.serialize()
async def make_authenticated_request(self, client, method: str, url: str, token: str, **kwargs):
proof = self.create_dpop_proof(method, url, access_token=token)
headers = {
"Authorization": f"DPoP {token}",
"DPoP": proof,
}
return await client.request(method, url, headers=headers, **kwargs)
Why DPoP Matters for Agents
Three properties make DPoP particularly relevant for agent workloads:
-
Token theft resistance. If an agent’s access token leaks (through logs, a prompt injection that exfiltrates headers, or a compromised intermediary), the attacker can’t use it without the private key.
-
Instance binding. Each agent instance generates its own key pair. Even if multiple agents share the same OAuth client, their tokens are bound to distinct keys — enabling per-instance audit trails.
-
Replay prevention. Each DPoP proof includes a unique
jtiand timestamp. Replaying a captured request fails because the proof is single-use.
The challenge: almost no LLM provider supports DPoP today. It’s most useful at the gateway layer — your own API gateway validates DPoP proofs before proxying requests to upstream providers with standard bearer tokens.
Scoped Tokens and Least-Privilege Design
Agent permissions should follow the principle of least privilege more strictly than human user permissions, because agents are more susceptible to prompt injection and other control-flow attacks that cause unintended actions.
Scope Design for Agent Operations
Define scopes at the operation level, not the resource level:
# Scope definitions for an agent framework
AGENT_SCOPES = {
# LLM operations
"llm:complete": "Make completion requests",
"llm:embed": "Generate embeddings",
"llm:finetune": "Submit fine-tuning jobs",
# Data operations
"data:read": "Read from data stores",
"data:write": "Write to data stores",
"data:delete": "Delete from data stores",
# Tool operations
"tool:execute": "Execute discovered tools",
"tool:discover": "List available tools",
# Agent operations
"agent:spawn": "Create child agent instances",
"agent:communicate": "Send messages to other agents",
}
# A research agent needs a narrow set
RESEARCH_AGENT_SCOPES = ["llm:complete", "llm:embed", "data:read", "tool:discover", "tool:execute"]
# A code generation agent needs write access but not delete
CODEGEN_AGENT_SCOPES = ["llm:complete", "data:read", "data:write", "tool:execute"]
Dynamic Scope Reduction
As an agent progresses through a workflow, its required permissions change. A code review agent might need write access during the “apply fix” phase but should lose it during the “summarize results” phase.
Scopes should change across workflow phases — an agent doesn’t need write access when it’s only generating a summary.
class ScopedAgentContext:
def __init__(self, token_service, agent_id: str):
self.token_service = token_service
self.agent_id = agent_id
async def enter_phase(self, phase: str, required_scopes: list[str]):
"""Request a new token with exactly the scopes needed for this phase."""
# Exchange current token for one with reduced/different scopes
new_token = await self.token_service.exchange(
agent_id=self.agent_id,
requested_scopes=required_scopes,
reason=f"entering phase: {phase}",
)
return new_token
# Usage in an agent workflow:
# token = await ctx.enter_phase("discovery", ["data:read", "tool:discover"])
# ... do discovery work ...
# token = await ctx.enter_phase("reporting", ["data:read", "llm:complete"])
Capability-Based Security
An alternative to scope-based authorization is capability-based security, where the agent holds unforgeable tokens (capabilities) that grant specific actions on specific resources. This maps well to tool-use patterns where each tool invocation could require a distinct capability.
| Approach | Granularity | Revocation | Complexity |
|---|---|---|---|
| API key + role | Coarse (role-level) | All-or-nothing | Low |
| OAuth scopes | Medium (operation-level) | Per-token | Medium |
| Capabilities | Fine (resource + operation) | Per-capability | High |
| Attribute-based (ABAC) | Context-dependent | Policy-based | High |
For most production agent deployments, OAuth scopes with dynamic reduction provide the best balance of security and operational complexity.
MCP Authentication and Tool-Level Identity
The Model Context Protocol (MCP) defines how agents discover and invoke tools through a standardized interface. Authentication in MCP operates at two layers: the connection between the agent and the MCP server, and the MCP server’s connection to its backing services.
MCP authentication has two hops: agent-to-server and server-to-backing-service, each potentially using different credential types.
Agent-to-MCP Server Authentication
The MCP specification supports OAuth 2.0 as the standard authentication mechanism between clients and servers. Gemini’s Managed Agents now support MCP server connections with credential refreshing, which addresses the long-running agent problem directly.
The MCP auth flow typically works as:
- Agent discovers MCP server’s OAuth metadata (
.well-known/oauth-authorization-server) - Agent performs OAuth flow (client credentials for autonomous agents, authorization code for user-delegated agents)
- MCP server validates token and maps to tool-level permissions
- Token refresh happens transparently for long-running connections
Tool-Level Authorization
Not every tool exposed by an MCP server should be available to every agent. The MCP server should enforce authorization at the tool level based on the agent’s identity and scopes.
# MCP server tool registration with authorization
from mcp.server import Server, tool
class SecureMCPServer:
def __init__(self):
self.server = Server("secure-tools")
@tool(
name="read_file",
description="Read a file from the project directory",
required_scopes=["data:read"],
)
async def read_file(self, path: str, agent_context: dict) -> str:
# Validate path is within allowed directory for this agent
allowed_dirs = agent_context.get("allowed_directories", [])
if not any(path.startswith(d) for d in allowed_dirs):
raise PermissionError(f"Agent not authorized to read from {path}")
with open(path) as f:
return f.read()
@tool(
name="execute_query",
description="Run a read-only SQL query",
required_scopes=["data:read", "tool:execute"],
)
async def execute_query(self, query: str, agent_context: dict) -> str:
# Enforce read-only at the tool level, not just the database level
if any(kw in query.upper() for kw in ["INSERT", "UPDATE", "DELETE", "DROP", "ALTER"]):
raise PermissionError("Write operations not permitted for this agent")
return await self.db.execute_readonly(query)
Credential Management in Long-Running Agents
Persistent agents — Claude Cowork sessions, background research agents, CI/CD agents — face unique credential lifecycle challenges.
Token Refresh Strategies
Proactive refresh with a credential sidecar is the most reliable pattern for long-running agents.
Proactive refresh — refreshing the token when 75% of its TTL has elapsed — avoids the latency spike of discovering expiration at request time. A credential sidecar process (or library) handles this independently of the agent’s main loop:
import asyncio
from datetime import datetime, timedelta
class CredentialSidecar:
"""Manages credential lifecycle independently of the agent's main loop."""
def __init__(self, oauth_client, scopes: list[str], refresh_threshold: float = 0.75):
self.oauth = oauth_client
self.scopes = scopes
self.refresh_threshold = refresh_threshold
self._current_token = None
self._expires_at = None
self._lock = asyncio.Lock()
async def start(self):
"""Start the background refresh loop."""
await self._refresh()
asyncio.create_task(self._refresh_loop())
async def get_token(self) -> str:
"""Called by the agent — always returns a valid token."""
async with self._lock:
if not self._current_token or self._expires_at < datetime.utcnow():
await self._refresh()
return self._current_token
async def _refresh_loop(self):
while True:
if self._expires_at:
ttl = (self._expires_at - datetime.utcnow()).total_seconds()
sleep_time = max(ttl * self.refresh_threshold, 10)
await asyncio.sleep(sleep_time)
async with self._lock:
await self._refresh()
else:
await asyncio.sleep(10)
async def _refresh(self):
token_data = await self.oauth.request_token(self.scopes)
self._current_token = token_data["access_token"]
self._expires_at = datetime.utcnow() + timedelta(seconds=token_data["expires_in"])
Credential Isolation in Multi-Agent Systems
When multiple agent instances run in parallel — as supported by Claude Code’s multi-instance coordination — each instance needs independent credentials to avoid refresh races and enable per-instance revocation.
| Isolation Level | Implementation | Use Case |
|---|---|---|
| Process-level | Separate key pair + token per OS process | Parallel agent workers |
| Container-level | Workload identity per pod/container | Kubernetes-deployed agents |
| Session-level | Per-user-session tokens via delegation | User-facing agent apps |
| Task-level | Short-lived capability tokens per task | High-security workflows |
Handling Credential Revocation
When an agent’s behavior becomes suspicious — making unexpected API calls, attempting to access resources outside its scope — the system needs to revoke its credentials without disrupting other agents.
The pattern: a revocation sidecar that monitors for anomalous behavior and can instantly invalidate a specific agent’s tokens.
class AgentRevocationWatcher:
def __init__(self, policy_engine, token_service):
self.policy = policy_engine
self.tokens = token_service
async def evaluate_request(self, agent_id: str, request_metadata: dict) -> bool:
"""Called before each outbound API call. Returns False to block."""
verdict = await self.policy.evaluate(
agent_id=agent_id,
action=request_metadata["method"],
resource=request_metadata["url"],
context={
"request_count_last_minute": request_metadata.get("recent_count", 0),
"unusual_endpoint": request_metadata.get("is_new_endpoint", False),
}
)
if not verdict.allowed:
await self.tokens.revoke(agent_id=agent_id, reason=verdict.reason)
return False
return True
Implementation Patterns
Pattern 1: Gateway-Mediated Auth
The LLM gateway handles all credential management, so agents never touch raw API keys for upstream providers.
The gateway pattern: agents authenticate with workload identity; the gateway injects provider credentials from a vault.
This is probably the most practical pattern for teams already running an LLM gateway (as covered in the gateway architecture deep dive). The agent authenticates to the gateway using workload identity or OAuth, and the gateway manages the provider API keys in a vault.
Benefits:
- Agents never hold provider API keys
- Key rotation happens at the gateway without agent changes
- Per-agent usage tracking and rate limiting
- Credential exposure surface is limited to the gateway process
Pattern 2: Identity-Aware Agent Framework
Build identity management into the agent framework itself, so every tool call and LLM request carries authenticated context.
class IdentityAwareAgent:
def __init__(self, agent_id: str, identity_provider, policy_engine):
self.agent_id = agent_id
self.identity = identity_provider
self.policy = policy_engine
self.credential_sidecar = CredentialSidecar(identity_provider, scopes=[])
async def invoke_tool(self, tool_name: str, params: dict, user_context: dict = None):
# Check authorization before invocation
required_scopes = await self.get_tool_scopes(tool_name)
allowed = await self.policy.check(
subject=self.agent_id,
action=f"tool:{tool_name}",
scopes=required_scopes,
delegation_context=user_context,
)
if not allowed:
raise PermissionError(f"Agent {self.agent_id} not authorized for tool {tool_name}")
# Get appropriately scoped token
token = await self.credential_sidecar.get_token()
# Invoke with authenticated context
result = await self.tool_registry.invoke(
tool_name,
params,
auth_token=token,
agent_id=self.agent_id,
trace_id=self.current_trace_id,
)
return result
Pattern 3: Zero-Trust Agent Mesh
For multi-agent architectures where agents communicate with each other, use mutual TLS with SPIFFE identities. Every agent-to-agent message is authenticated, encrypted, and authorized.
Zero-trust agent mesh: each agent has a SPIFFE identity, and all inter-agent communication uses mutual TLS.
This pattern is operationally heavier but provides the strongest guarantees for multi-agent systems handling sensitive data. It’s particularly relevant as agent frameworks increasingly support parallel execution and inter-agent messaging.
Pattern 4: Ambient Credentials with Attestation
Combine workload identity with platform attestation — the agent proves not just its identity but the integrity of its runtime environment.
# Pseudocode for attestation-based credential acquisition
class AttestedAgentIdentity:
async def get_attested_token(self):
# Collect platform attestation evidence
attestation = await self.platform.get_attestation({
"container_image_digest": self.get_image_digest(),
"agent_code_hash": self.get_code_hash(),
"runtime_measurements": self.get_tpm_measurements(),
})
# Exchange attestation for credentials
token = await self.identity_provider.exchange_attestation(
attestation=attestation,
requested_scopes=self.required_scopes,
)
return token
This is forward-looking — most agent deployments don’t need this level of assurance today. But for agents handling financial transactions, healthcare data, or operating in regulated environments, attestation-based identity will probably become a requirement.
Summary
The progression from shared API keys to workload identity follows a clear path: each step adds attribution, reduces blast radius, and removes static secrets from the system.
Current state: Most production agent deployments use shared API keys, possibly with per-environment or per-user separation. This is the minimum viable approach and its risks scale with agent autonomy.
Practical next step: Gateway-mediated authentication with per-agent key issuance gives the best security improvement for the least operational cost. Agents authenticate to the gateway; the gateway manages provider credentials.
Medium-term target: OAuth 2.0 with client credentials for autonomous agents and token exchange for user-delegated agents. Dynamic scope reduction across workflow phases. Credential sidecars for long-running agents.
Forward-looking: Workload identity eliminates static secrets entirely. DPoP binds tokens to specific agent instances. SPIFFE/SPIRE enables zero-trust agent meshes. Platform attestation verifies runtime integrity.
The key design principle: an agent should hold the minimum credentials for the minimum time needed for its current operation. Every extra permission and every extra minute of credential validity is attack surface. With agents increasingly operating autonomously — Claude Code Auto Mode now runs as the primary actor with humans in approval workflows — the authentication and identity layer is the primary control point for limiting the blast radius of agent misbehavior.
Further Reading
- SPIFFE/SPIRE documentation — The SPIFFE specification and SPIRE implementation for workload identity across platforms
- RFC 9449 — OAuth 2.0 Demonstrating Proof of Possession (DPoP) — The DPoP specification for binding access tokens to cryptographic keys
- RFC 8693 — OAuth 2.0 Token Exchange — Token exchange specification for delegation scenarios
- HashiCorp Vault — Dynamic Secrets — Dynamic credential generation and secrets management
- MCP Specification — Authentication — The Model Context Protocol specification, including its OAuth-based authentication model
- OWASP — Top 10 for LLM Applications — Security risks specific to LLM applications, including prompt injection and insecure output handling
- Google Cloud Workload Identity Federation — Federated identity for workloads running outside Google Cloud
- AWS EKS Pod Identity — Assigning IAM credentials to Kubernetes pods without static secrets
- OPA (Open Policy Agent) — Policy engine for fine-grained authorization decisions in distributed systems
- jwcrypto library — Python library for JWK, JWS, JWE, and JWT operations used in the DPoP examples above