Centralize the Claude subscription, not the credential. Enterprise platform teams should give production workloads, developer laptops and external CI/CD separate identities, narrowly assigned workspaces and independently revocable access. A successful model response proves connectivity; it does not prove that a developer cannot reach production or that a pipeline cannot read another team's files.
AWS's October 1, 2026 multi-environment access walkthrough describes a dedicated AI Services account, cross-account SigV4 access, development API keys and external OIDC federation. The practical question is how to demonstrate that those boundaries hold before distributing access. This guide adds an illustrative pilot, negative tests and an operating checklist; it does not report a customer deployment or measured security outcome.
Start with the service boundary and prerequisites
Claude Platform on AWS is not simply another name for Claude in Amazon Bedrock. Anthropic operates this platform, and its data-processing and retention terms require separate review. Zero Data Retention is not automatic. Confirm the applicable agreement and feature coverage before sending sensitive information; an AWS invoice is not evidence that inference stays entirely within AWS. See Anthropic's platform comparison and setup guidance.
The AWS prerequisites include an active subscription, a configured AWS CLI, outbound web identity federation enabled for the account and a workspace ID. Assign a platform owner to verify those items in the intended account, a security owner to review access and retention, and a finance owner to approve the pilot budget. Do not enable account-level capabilities from an unidentified terminal profile.
Record the account ID, approved regional endpoint, workspace ID, workspace ARN, model, SDK version and expected authentication method in a nonsecret configuration manifest. Pin the reviewed SDK and token-generator versions in the application lockfile. Decide which data classes, APIs and network destinations are allowed. Begin with synthetic prompts and no customer documents, production tools or downstream write actions.
Design one services account with distinct workspace boundaries
Keep the subscription and access administration in a dedicated AI Services account. Separate production and development workspaces, with named owners and explicit grants for each caller. A workspace name is an organizational label, not an authorization rule. Review the effective IAM policy, including inherited grants, permission boundaries and organization controls, rather than checking only the new policy you just attached.
AWS documents workspaces as regional resources: the endpoint must match the workspace's region. Use the wrkspc_-prefixed ID in the anthropic-workspace-id request header; use the full workspace ARN in IAM resource statements. Keep those values distinct. A copied production ID in a development configuration should fail authorization, not silently redirect the request.
API endpoint region and inference geography are different decisions. Verify the platform's current routing and residency settings alongside the contractual data boundary. For stricter isolation, evaluate separate AWS accounts as well as separate workspaces; a convenient shared account should not overrule a documented client or regulatory requirement.
AWS workloads
Workload role → cross-account role → SigV4
Grant only the approved production workspace and operations. The pilot uses a nonproduction stand-in.
Developer laptops
Individual scoped key → development workspace
Inspect the backing identity. Prove that production and unapproved stored content remain inaccessible.
External CI/CD
OIDC job identity → STS credentials → approved workspace
Use SigV4 directly, or a short-lived bearer token when required. Pilot in development; production needs a separate explicit grant.
Production workloads: assume a narrow role and sign requests
For an AWS-hosted application, use its workload identity to assume a dedicated role in the AI Services account, then make SigV4-signed requests. Restrict the target role's trust to the intended workload role, not the entire source account without conditions. Separately limit what the assumed role can do. Trust answers who may enter; the permission policy answers what that identity may access.
Start with the exact inference action and workspace required by the application. Add token counting or model discovery only when the client uses them. The IAM action reference uses the aws-external-anthropic namespace, not bedrock. Keep subscription administration, key management, console administration and resource provisioning out of the ordinary inference role.
AWS's managed-policy descriptions show why policy names alone are insufficient: inference-oriented managed policies can include broad Get/List access, including stored content. Synchronous and batch inference also have separate permissions. Inventory every action the application needs, explicitly test prohibited actions, and avoid granting a broad managed policy just to remove an unexplained error.
Do not bake access keys into an image or pass a human administrator's credentials into a pod. Prove the runtime's actual assumed identity after deployment. Include a second workload role that must fail to assume the production role, and a development workspace request that must fail under the production-only policy. Keep denial tests separate from model-output quality tests.
Developer laptops: scope the backing identity before issuing a key
A development API key is useful when local tools cannot use the preferred federated AWS flow. Give each developer or independently managed tool an attributable credential with development-only authority. Store it in the approved secret manager, inject it only into the intended process and exclude it from Git, shell history, screenshots and support bundles. Where local tooling supports federated SigV4, consider that route instead of creating another persistent key.
The AWS walkthrough warns that a console-generated key's AeaApiKey-* backing IAM user receives AnthropicLimitedAccess by default, spanning workspaces. Identify the exact backing user for the issued key, replace that broad grant with the reviewed development policy, and test it before distribution. Do not guess the owner by selecting whichever IAM user was created most recently.
Bearer access needs a separate authentication permission as well as the requested API action. The detailed Anthropic IAM reference specifies CallWithBearerToken on Resource: *, with inference scoped to the workspace ARN. That account-scoped authentication statement is not a reason to wildcard inference. Some AWS overview wording describes this as permission on the target workspace; use the detailed action mapping and verify the effective policy in the pilot rather than copying an ambiguous example.
Keys for the ordinary first-party Claude API or Amazon Bedrock are not interchangeable with Claude Platform on AWS keys. Use the platform-specific client and the correct endpoint; AWS's authentication guide explains the distinction. A secret stored in a development vault is still dangerous if its actual permissions reach production.
External CI/CD: federate the job, not a shared service password
Use the external system's OIDC identity to obtain temporary AWS credentials for a dedicated role. Restrict the trust relationship to the approved issuer, audience and subject claims. For a build platform, bind the subject to the intended repository and approved branch or deployment environment. Test an unapproved branch and an unrelated repository; do not assume that a trusted CI provider makes every job trustworthy.
If the caller supports SigV4, temporary AWS credentials can sign the request directly. When an integration needs bearer authentication, exchange them using the supported token generator. The AWS Python token-generator documentation bounds lifetime by the requested duration, underlying credential expiry and a 12-hour maximum. Tokens are regional credentials, not harmless job metadata. Choose a lifetime matched to the job rather than accepting the maximum by habit.
Keep the token in memory or the runner's protected secret channel. Do not echo it, include it in build artifacts or reuse it across unrelated jobs. Arrange refresh explicitly for long-running work, with a bounded failure path when refresh fails; a client holding a token should not be assumed to renew it. Separate the pipeline that provisions access from the pipeline that calls the model. A test job should not be able to broaden its own role.
Check credential precedence before interpreting a successful call
Anthropic's documented precedence is: explicit API key, explicit AWS access/secret credentials, explicit AWS profile, ANTHROPIC_AWS_API_KEY, then the default AWS credential chain. Exact argument names vary by SDK. An environment API key can therefore override the workload credentials you expected the default chain to select. See the current SDK setup guide.
Test from a clean process with only the chosen credential mechanism available. Record the resolved account, role or key identity, region and workspace without logging credential values. Then deliberately introduce a synthetic, development-only competing credential and verify the selected path. Remove it afterward. An SDK upgrade should trigger this test again, not merely a connectivity smoke test.
Require an explicit region and validate its relationship to the workspace before sending data. Test missing region, wrong region, missing workspace ID and a syntactically valid but unauthorized workspace. Capture the structured error and request identity when available. A DNS failure, timeout or malformed request is not proof that IAM correctly denied a prohibited operation.
Run one isolated pilot and prove both allow and deny behavior
Illustrative scenario: a platform team wants Claude to summarize synthetic release notes. Create one pilot development workspace and a separate empty protected workspace standing in for production. Use a laptop identity, an AWS test workload and a restricted CI job. No real production records or tools are needed to prove the intended access boundaries.
Freeze the policy/configuration versions and write expected outcomes before testing. Use the same harmless prompt for allowed calls, then change one identity, resource or action at a time. Do not accept a denial unless an allowed control request proves the endpoint is reachable and the rejected request actually exercised the intended permission check. Record failures and retries, not just the final green run.
- Approved caller + approved inference
- Expected: Allow. One harmless response; intended identity, workspace and request IDs recorded.
- Development key + protected workspace
- Expected: Deny. Authorization rejection at the reachable service, not a bad URL or missing parameter.
- Unapproved AWS role + target role
- Expected: Deny. Role assumption rejected; no inference credential obtained.
- Wrong CI subject or audience
- Expected: Deny. Federation rejected for an unapproved repository, branch or environment.
- Inference-only identity + files, batches or administration
- Expected: Deny. Each excluded action tested separately with valid synthetic inputs.
- Expired token or revoked development key
- Expected: Deny. Old credential rejected while an authorized control request still succeeds.
- Missing/wrong region or workspace
- Expected: Fail safely. No fallback to another account or workspace; classify configuration errors separately from IAM denials.
- Competing development key + intended role
- Expected: Verify identity. Observed credential selection matches the pinned SDK behavior; remove competing credentials before release.
Have a reviewer trace every allowed call back to the approved identity and every forbidden call to the expected denial. Re-run from a fresh laptop profile and a clean CI runner to catch hidden credentials and cached configuration. Only after review should the real production owner approve a separate production grant. A passing pilot does not authorize new models, stored files, agent tools or additional workspaces.
Make individual calls auditable and costs attributable
CloudTrail management history alone is not enough. AWS's monitoring guide requires explicit data-event configuration for inference and other data-plane operations, using resource type AWS::AWSExternalAnthropic::Workspace. Account for logging charges and verify selectors cover each intended workspace. Retain both response identifiers: x-amzn-requestid for AWS correlation and request-id for Anthropic support.
Maintain a restricted application record with timestamp, job/run ID, resolved principal, nonsecret key identifier where available, workspace, model, operation, outcome, latency and token usage. Do not log authorization headers or prompt bodies by default. Test whether an operator can reconstruct a failed CI request without obtaining the secret. A shared key identifies the credential, not necessarily the human who used it; application attribution must not conceal that limitation.
Tag workspaces by owner, team, application, environment and cost center, then activate the relevant billing cost-allocation tags. AWS documents that allocation applies to usage on or after activation; verify actual reports after processing delays. Compare workspace usage, failed/retried calls, logging charges and engineering time. Marketplace consumption billing is the financial record, not a substitute for a per-call access trail.
Plan rotation, offboarding and incident rollback before rollout
For persistent development keys, maintain an owner, purpose, approved workspace, issue date, expiry or rotation deadline and last-use review. Rotate by issuing a separately scoped replacement, testing it, switching the known consumer and revoking the old key. Verify the old credential fails. Do not leave both active indefinitely or create a broad emergency key to keep a failing test green.
Offboarding must remove the person's identity-provider access, local secret access, platform keys and any dedicated backing identity that is no longer needed. Review shared CI subjects, repository permissions and console roles separately. Temporary credentials reduce exposure duration but do not replace revocation planning. Test the effect of a trust or policy change on already-issued sessions and tokens; do not promise instantaneous invalidation without evidence.
If unauthorized access succeeds, stop the affected job or application path, restrict the exact implicated role or key, preserve nonsecret evidence and involve the security owner. Determine which workspaces and operations were exposed. Restore only the last reviewed configuration that meets the access boundary; rolling back to a formerly working broad credential is not recovery. Keep the feature disabled or use a documented manual process while access is repaired.
At pilot completion, revoke pilot keys and roles, remove only dedicated unused federation resources, and clean local caches, runner artifacts and temporary workspace data under the retention plan. Preserve approved audit evidence. Check for continuing usage and charges. Do not delete a shared OIDC provider, cancel a shared subscription or remove another team's workspace as cleanup.
Use acceptance measures that expose hidden access failures
- Unauthorized access
- Zero successful prohibited operations across the prewritten matrix. Any success pauses expansion. Record case count; a finite test suite cannot prove universal safety.
- Successful calls
- Every required allowed case passes on a clean rerun. Track successful valid requests divided by attempts, with authorization, configuration, quota and provider failures separated.
- Attribution
- Every pilot call is traceable to the intended identity, workspace and run. Reconcile application records with the configured audit events; investigate gaps rather than assuming no activity.
- Setup reliability
- Repeat from three clean environments or runs per access path. Record time to first correctly attributed call, retries and undocumented manual fixes.
- Reviewer effort
- Measure total review and correction time per approved access path against the current manual setup baseline. Agree the acceptable time budget before testing.
- Recovery and cost
- Revoke a pilot credential, prove denial and restore approved access within the team's written recovery target. Attribute test consumption and logging cost without exposing payloads.
Treat these as proposed ITECS pilot criteria, not vendor guarantees or measured results. Repeat the matrix after changes to IAM, trust claims, SDKs, token generation, workspace configuration or team ownership. Review denied requests for expected behavior as well as misconfiguration; a system that denies everything is secure only in the narrowest sense and unusable for the business.
ITECS AI DevOps can help turn this access design into an owned delivery process. For teams connecting custom AI agents, pair identity testing with agent tool-use evaluation: permission to call a model does not grant permission for an agent to change a business system.
